2026-06-21 11:02:35 +00:00
|
|
|
|
import json
|
|
|
|
|
|
import os
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
import re
|
2026-06-21 11:02:35 +00:00
|
|
|
|
import shlex
|
|
|
|
|
|
import subprocess
|
2026-05-31 23:58:26 +09:00
|
|
|
|
from copy import deepcopy
|
|
|
|
|
|
|
2026-06-11 01:43:49 +03:00
|
|
|
|
from fastapi import APIRouter, HTTPException
|
|
|
|
|
|
|
2026-06-21 11:02:35 +00:00
|
|
|
|
from core.platform_compat import run_ssh_command
|
2026-06-11 01:43:49 +03:00
|
|
|
|
from routes._validators import validate_remote_host, validate_ssh_port
|
2026-05-31 23:58:26 +09:00
|
|
|
|
|
|
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
# Backends the manual hardware simulator accepts. Must stay a subset of what
|
|
|
|
|
|
# services.hwfit.fit understands so a simulated box ranks like a real one:
|
|
|
|
|
|
# "metal" routes through the Apple-Silicon path (GGUF-only, llama.cpp/Ollama),
|
|
|
|
|
|
# the CPU backends through the RAM/offload path, cuda/rocm through vLLM.
|
|
|
|
|
|
_MANUAL_BACKENDS = {"cuda", "rocm", "metal", "cpu_x86", "cpu_arm"}
|
2026-05-31 23:58:26 +09:00
|
|
|
|
|
|
|
|
|
|
|
2026-06-11 01:43:49 +03:00
|
|
|
|
def _validate_detection_target(host: str = "", ssh_port: str = "") -> tuple[str, str]:
|
|
|
|
|
|
host_value = validate_remote_host(host) or ""
|
|
|
|
|
|
port_value = validate_ssh_port(ssh_port) or ""
|
|
|
|
|
|
if port_value and not host_value:
|
|
|
|
|
|
raise HTTPException(400, "ssh_port requires host")
|
|
|
|
|
|
return host_value, port_value
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
def _apply_manual_hardware(system, manual_mode="", manual_gpu_count="", manual_vram_gb="", manual_ram_gb="", manual_backend=""):
|
|
|
|
|
|
"""Manual hardware is a "what if I had this setup" simulator —
|
|
|
|
|
|
REPLACES the detected hardware entirely instead of adding to it.
|
2026-05-31 23:58:26 +09:00
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
The previous additive behavior averaged the manual VRAM across
|
|
|
|
|
|
all GPUs (base + manual), which meant adding "1× 400 GB" on top
|
|
|
|
|
|
of "2× 70 GB" only nudged the per-GPU cap from 70 to 180 GB
|
|
|
|
|
|
(= 540 / 3), so GGUF models bigger than that still didn't surface
|
|
|
|
|
|
— exactly the "cap stuck at detected level" bug the user hit.
|
|
|
|
|
|
"""
|
|
|
|
|
|
manual_mode = (manual_mode or "").lower()
|
|
|
|
|
|
if manual_mode not in {"gpu", "ram"}:
|
|
|
|
|
|
return system
|
2026-05-31 23:58:26 +09:00
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
try:
|
|
|
|
|
|
override_ram_gb = float(manual_ram_gb) if manual_ram_gb else 0
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
override_ram_gb = 0
|
|
|
|
|
|
override_ram_gb = max(0.0, override_ram_gb)
|
|
|
|
|
|
if override_ram_gb:
|
|
|
|
|
|
# Replace RAM, don't add. The number in the field is the
|
|
|
|
|
|
# TOTAL system memory the user wants to simulate.
|
|
|
|
|
|
system["available_ram_gb"] = round(override_ram_gb, 1)
|
|
|
|
|
|
system["total_ram_gb"] = round(override_ram_gb, 1)
|
|
|
|
|
|
system["manual_hardware"] = True
|
2026-05-31 23:58:26 +09:00
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
if manual_mode == "ram":
|
|
|
|
|
|
# RAM-only simulation — wipe GPU entirely so the ranker uses
|
|
|
|
|
|
# CPU/RAM paths.
|
|
|
|
|
|
system["has_gpu"] = False
|
|
|
|
|
|
system["gpu_name"] = None
|
|
|
|
|
|
system["gpu_vram_gb"] = 0
|
|
|
|
|
|
system["gpu_count"] = 0
|
|
|
|
|
|
system["gpus"] = []
|
|
|
|
|
|
system["gpu_groups"] = []
|
|
|
|
|
|
system["backend"] = "cpu_x86"
|
|
|
|
|
|
system.pop("unified_memory", None)
|
2026-05-31 23:58:26 +09:00
|
|
|
|
return system
|
|
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
try:
|
|
|
|
|
|
count = int(manual_gpu_count) if manual_gpu_count else 1
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
count = 1
|
|
|
|
|
|
try:
|
|
|
|
|
|
vram_each = float(manual_vram_gb) if manual_vram_gb else 8.0
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
vram_each = 8.0
|
|
|
|
|
|
count = max(1, min(count, 16))
|
|
|
|
|
|
vram_each = max(1.0, vram_each)
|
|
|
|
|
|
backend = (manual_backend or system.get("backend") or "cuda").lower()
|
|
|
|
|
|
if backend not in _MANUAL_BACKENDS:
|
|
|
|
|
|
backend = "cuda"
|
|
|
|
|
|
total_vram = round(vram_each * count, 1)
|
|
|
|
|
|
gpu_name = f"Simulated {backend.upper()} GPU" + (f" × {count}" if count > 1 else "")
|
|
|
|
|
|
system["has_gpu"] = True
|
|
|
|
|
|
system["gpu_name"] = gpu_name
|
|
|
|
|
|
system["gpu_vram_gb"] = total_vram
|
|
|
|
|
|
system["gpu_count"] = count
|
|
|
|
|
|
system["gpus"] = [
|
|
|
|
|
|
{"index": i, "name": gpu_name, "vram_gb": vram_each}
|
|
|
|
|
|
for i in range(count)
|
|
|
|
|
|
]
|
|
|
|
|
|
# Single homogeneous pool — vram_each here is the ACTUAL per-GPU
|
|
|
|
|
|
# VRAM the user entered, not an average. That's the whole point:
|
|
|
|
|
|
# raising vram_each lifts the per-GPU cap (GGUF, tensor-parallel
|
|
|
|
|
|
# math) all the way up, not just by a small fraction.
|
|
|
|
|
|
system["gpu_groups"] = [{
|
|
|
|
|
|
"name": gpu_name,
|
|
|
|
|
|
"vram_each": vram_each,
|
|
|
|
|
|
"count": count,
|
|
|
|
|
|
"indices": list(range(count)),
|
|
|
|
|
|
"vram_total": total_vram,
|
|
|
|
|
|
}]
|
|
|
|
|
|
system["homogeneous"] = True
|
|
|
|
|
|
system["backend"] = backend
|
|
|
|
|
|
# Apple Silicon shares one unified memory pool with the GPU; flag it so
|
|
|
|
|
|
# the API/UI report it the way real Metal detection does. Discrete GPUs
|
|
|
|
|
|
# (cuda/rocm) and the CPU backends carry separate VRAM, so clear any
|
|
|
|
|
|
# stale flag a previous detection left on the dict.
|
|
|
|
|
|
if backend == "metal":
|
|
|
|
|
|
system["unified_memory"] = True
|
|
|
|
|
|
else:
|
|
|
|
|
|
system.pop("unified_memory", None)
|
|
|
|
|
|
return system
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-21 11:02:35 +00:00
|
|
|
|
def _run_model_probe(host: str, ssh_port: str, cmd: str) -> str:
|
|
|
|
|
|
try:
|
|
|
|
|
|
if host:
|
|
|
|
|
|
r = run_ssh_command(
|
|
|
|
|
|
host,
|
|
|
|
|
|
ssh_port or None,
|
|
|
|
|
|
cmd,
|
|
|
|
|
|
timeout=15,
|
|
|
|
|
|
connect_timeout=5,
|
|
|
|
|
|
strict_host_key_checking=False,
|
|
|
|
|
|
text=True,
|
|
|
|
|
|
)
|
|
|
|
|
|
else:
|
|
|
|
|
|
r = subprocess.run(["bash", "-lc", cmd], capture_output=True, text=True, timeout=15)
|
|
|
|
|
|
if r.returncode == 0:
|
|
|
|
|
|
return (r.stdout or "").strip()
|
|
|
|
|
|
except Exception:
|
|
|
|
|
|
return ""
|
|
|
|
|
|
return ""
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _inspect_model_path(model_path: str, host: str = "", ssh_port: str = "") -> dict:
|
|
|
|
|
|
"""Read lightweight metadata from a local or SSH-visible HF model folder."""
|
|
|
|
|
|
path = (model_path or "").strip()
|
|
|
|
|
|
if not path or path.startswith(("http://", "https://")):
|
|
|
|
|
|
return {}
|
|
|
|
|
|
if not (path.startswith("/") or path.startswith("~")):
|
|
|
|
|
|
return {}
|
|
|
|
|
|
|
|
|
|
|
|
qpath = shlex.quote(path)
|
|
|
|
|
|
qconfig = shlex.quote(os.path.join(path, "config.json"))
|
|
|
|
|
|
out = {}
|
|
|
|
|
|
exists = _run_model_probe(host, ssh_port, f"test -d {qpath} && printf found || printf missing")
|
|
|
|
|
|
if exists != "found":
|
|
|
|
|
|
target = host or "local container"
|
|
|
|
|
|
out["model_probe_error"] = f"Model path is not visible on {target}: {path}"
|
|
|
|
|
|
return out
|
|
|
|
|
|
raw_config = _run_model_probe(host, ssh_port, f"test -f {qconfig} && sed -n '1,240p' {qconfig}")
|
|
|
|
|
|
if raw_config:
|
|
|
|
|
|
try:
|
|
|
|
|
|
cfg = json.loads(raw_config)
|
|
|
|
|
|
except Exception:
|
|
|
|
|
|
cfg = {}
|
|
|
|
|
|
for key in ("context_length", "max_position_embeddings", "n_ctx_train", "model_max_length", "max_seq_len"):
|
|
|
|
|
|
value = cfg.get(key)
|
|
|
|
|
|
if isinstance(value, (int, float)) and value > 0:
|
|
|
|
|
|
out["model_ctx_max"] = int(value)
|
|
|
|
|
|
break
|
|
|
|
|
|
else:
|
|
|
|
|
|
out["model_probe_error"] = f"config.json not found in model path: {path}"
|
|
|
|
|
|
|
|
|
|
|
|
size_cmd = (
|
|
|
|
|
|
f"find {qpath} -type f \\( -name '*.safetensors' -o -name '*.bin' -o -name '*.gguf' \\) "
|
|
|
|
|
|
"-printf '%s\\n' 2>/dev/null | awk '{s+=$1} END {if (s>0) printf \"%.6f\", s/1073741824}'"
|
|
|
|
|
|
)
|
|
|
|
|
|
weights = _run_model_probe(host, ssh_port, size_cmd)
|
|
|
|
|
|
try:
|
|
|
|
|
|
weights_gb = float(weights)
|
|
|
|
|
|
except Exception:
|
|
|
|
|
|
weights_gb = 0.0
|
|
|
|
|
|
if weights_gb > 0:
|
|
|
|
|
|
out["model_weights_gb"] = round(weights_gb, 3)
|
|
|
|
|
|
elif "model_probe_error" not in out:
|
|
|
|
|
|
out["model_probe_error"] = f"No model weight files found in: {path}"
|
|
|
|
|
|
return out
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-02 10:12:34 -04:00
|
|
|
|
def setup_hwfit_routes():
|
|
|
|
|
|
router = APIRouter(prefix="/api/hwfit", tags=["hwfit"])
|
|
|
|
|
|
|
2026-05-31 23:58:26 +09:00
|
|
|
|
@router.get("/system")
|
|
|
|
|
|
def get_system(host: str = "", ssh_port: str = "", platform: str = "", fresh: bool = False):
|
|
|
|
|
|
"""Detect and return current system hardware info. Pass host=user@server for remote.
|
|
|
|
|
|
fresh=true bypasses the per-host cache (the Rescan button)."""
|
|
|
|
|
|
from services.hwfit.hardware import detect_system
|
2026-06-11 01:43:49 +03:00
|
|
|
|
host, ssh_port = _validate_detection_target(host, ssh_port)
|
2026-05-31 23:58:26 +09:00
|
|
|
|
return detect_system(host=host, ssh_port=ssh_port, platform=platform, fresh=fresh)
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/models")
|
Open email context for agent, email search across All Mail, cookbook serve polish
- Agent: pass the open email reader (uid/folder/account/from/subject/body
preview) on every chat submit so 'reply to this' / 'write email saying
hi' route to ui_control open_email_reply with the right UID instead of
inventing a new .md draft. Code-level enforcement (chat_routes strips
create_document + send_email when active_email is set); cross-session
active_doc_id is now trusted instead of being silently dropped.
set_active_email/clear_active_email tool-layer helpers in
tool_implementations.
- ui_control open_email_reply: optional body argument so the agent can
open-and-write in one call; envelope now forwards uid/folder/account/
body/panel through tool_output. Tool description sharpened and the
parser rejects empty bodies on reply/reply-all (forces the agent to
write rather than open an empty draft).
- Email library: search now runs against [Gmail]/All Mail when the
current folder is INBOX (archived emails surface). Whirlpool spinner
+ 'Searching…' placeholder while in flight. Each search result is
stamped with its source folder so clicks open the right email instead
of whatever shares its UID in INBOX. Search no longer re-applies the
same text pill locally (which only checks subject/from/snippet, never
body) so body-only matches don't get dropped after IMAP returns them.
Initial inbox load bumped 100→500.
- Email favorites: 'Favorite (pin to top)' / 'Unfavorite' in both the
card menu and the open-reader more menu, backed by a new
/api/email/flag/{uid}?on=true|false endpoint. Flagged emails always
bubble to the top of the grid regardless of active sort.
- AI reply in doc editor: never overwrites existing draft text or the
quoted history. AI suggestion is prepended; AI-generated 'On …
wrote:' re-quotes are stripped so the original quote isn't visually
edited.
- Cookbook serve: pre-launch GPU driver / has_gpu / install / version-
floor checks (vllm minimax_m2 needs 0.10.0+, deepseek_r1 needs 0.7.0
etc.) before the launch chain starts. Detect 'another model already
running on this host' and offer Stop & launch (with graceful then
force tmux kill helpers, port release wait). Per-vendor deep-link
buttons (vLLM recipe / SGLang cookbook) with hardware hash. Backend
picker is now a custom dropdown with accent-coloured logos for vLLM,
SGLang, llama.cpp, Ollama, Diffusers; same glyphs added next to
package names in Dependencies. Runtime-readiness note moved inside
the panel (green when ready, red when missing) with an × dismiss.
Esc collapses the expanded card; expanded card scrolls when it
overflows; Trust Remote / Auto Tool / Reasoning Parser / Enforce
Eager / Prefix Caching / Expert Parallel / Speculative / MoE Env on
one row (Reasoning Parser auto-detected per model family).
Dtype→Row 1, GPUs→Row 2 (rightmost). Removed redundant GPU 'auto'
input — command builders read from the GPU button strip. Default
cookbook open is Download tab.
- Cookbook hwfit: 'Model (latest)' / 'Model (oldest)' header sorts by
release_date; release dates can be backfilled with the new
scripts/backfill_model_release_dates.py and recipe metadata pulled
with scripts/import_from_vllm_recipes.py against the upstream
vllm-project/recipes catalog (vllm_recipe + min_vllm_version stamped
on entries).
- Calendar: Quick add hint cycles a random Odysseus-themed example per
open (wooden horse Friday, crew muster 10am daily, council on
Ithaca, …). Typing a time like '11pm' in the event title updates
the hero clock live.
- Doc editor: email-mode Reply button (sparkle icon, accent) opens the
same Fast/Full + context popover the email reader uses; Ctrl+Alt+M
toggles markdown preview.
- Memories panel: custom sort picker with per-option icons, default
'Latest', visible Enabled/Disabled toggle text matching the section
description style.
2026-06-15 20:47:51 +09:00
|
|
|
|
def get_models(use_case: str = "", sort: str = "newest", limit: int = 50, search: str = "", host: str = "", quant: str = "", ctx: str = "", gpu_count: str = "", gpu_group: str = "", ssh_port: str = "", platform: str = "", fresh: bool = False, manual_mode: str = "", manual_gpu_count: str = "", manual_vram_gb: str = "", manual_ram_gb: str = "", manual_backend: str = "", ignore_detected_gpu: bool = False, ignore_detected_ram: bool = False, fit_only: bool = False):
|
2026-05-31 23:58:26 +09:00
|
|
|
|
"""Rank LLM models against detected hardware and return scored results.
|
|
|
|
|
|
gpu_count: override GPU count (0 = CPU only, 1-N = simulate N GPUs of the
|
|
|
|
|
|
active group). gpu_group: index into system.gpu_groups (the homogeneous
|
|
|
|
|
|
pools) to target — empty/auto = the largest pool. vLLM can only
|
|
|
|
|
|
tensor-parallel across identical GPUs, so we never mix pools.
|
|
|
|
|
|
fresh=true bypasses the hardware-detection cache."""
|
|
|
|
|
|
from services.hwfit.hardware import detect_system
|
|
|
|
|
|
from services.hwfit.fit import rank_models
|
|
|
|
|
|
from services.hwfit.models import get_models, model_catalog_path
|
2026-06-11 01:43:49 +03:00
|
|
|
|
host, ssh_port = _validate_detection_target(host, ssh_port)
|
2026-05-31 23:58:26 +09:00
|
|
|
|
system = deepcopy(detect_system(host=host, ssh_port=ssh_port, platform=platform, fresh=fresh))
|
|
|
|
|
|
if system.get("error"):
|
|
|
|
|
|
return {"system": system, "models": [], "error": system["error"]}
|
|
|
|
|
|
if not get_models():
|
|
|
|
|
|
return {
|
|
|
|
|
|
"system": system,
|
|
|
|
|
|
"models": [],
|
|
|
|
|
|
"error": f"Model catalog missing or empty: {model_catalog_path()}",
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
if ignore_detected_gpu:
|
|
|
|
|
|
system["has_gpu"] = False
|
|
|
|
|
|
system["gpu_name"] = None
|
|
|
|
|
|
system["gpu_vram_gb"] = 0
|
|
|
|
|
|
system["gpu_count"] = 0
|
|
|
|
|
|
system["gpus"] = []
|
|
|
|
|
|
system["gpu_groups"] = []
|
|
|
|
|
|
if ignore_detected_ram:
|
|
|
|
|
|
system["available_ram_gb"] = 0
|
|
|
|
|
|
system["total_ram_gb"] = 0
|
|
|
|
|
|
|
|
|
|
|
|
system = _apply_manual_hardware(system, manual_mode, manual_gpu_count, manual_vram_gb, manual_ram_gb, manual_backend)
|
|
|
|
|
|
|
|
|
|
|
|
# Keep the raw detection around so the UI can still show the box's full
|
|
|
|
|
|
# GPU complement even while we rank against one homogeneous pool.
|
|
|
|
|
|
system["detected_gpu_vram_gb"] = system.get("gpu_vram_gb")
|
|
|
|
|
|
system["detected_gpu_count"] = system.get("gpu_count")
|
|
|
|
|
|
|
|
|
|
|
|
groups = system.get("gpu_groups") or []
|
|
|
|
|
|
# Resolve the target homogeneous pool. Default (auto) = the largest pool,
|
|
|
|
|
|
# which for a uniform box is simply "all the GPUs" — no behaviour change.
|
|
|
|
|
|
grp = None
|
|
|
|
|
|
if groups:
|
|
|
|
|
|
try:
|
|
|
|
|
|
gidx = int(gpu_group) if gpu_group != "" else 0
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
gidx = 0
|
|
|
|
|
|
if 0 <= gidx < len(groups):
|
|
|
|
|
|
grp = groups[gidx]
|
|
|
|
|
|
|
|
|
|
|
|
def _apply_group(g, n):
|
|
|
|
|
|
n = max(1, min(n, g["count"]))
|
|
|
|
|
|
system["gpu_count"] = n
|
|
|
|
|
|
system["gpu_vram_gb"] = round(g["vram_each"] * n, 1)
|
|
|
|
|
|
system["gpu_name"] = g["name"]
|
|
|
|
|
|
system["active_group"] = {**g, "use_count": n}
|
|
|
|
|
|
|
2026-06-11 02:01:58 +03:00
|
|
|
|
# Parse the optional count defensively (matches the gpu_group guard
|
|
|
|
|
|
# above): a non-numeric query param previously raised ValueError ->
|
|
|
|
|
|
# HTTP 500. A malformed value is ignored, same as omitting it.
|
|
|
|
|
|
try:
|
|
|
|
|
|
n = int(gpu_count) if gpu_count != "" else None
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
n = None
|
|
|
|
|
|
if n is not None:
|
2026-05-31 23:58:26 +09:00
|
|
|
|
if n == 0:
|
|
|
|
|
|
# RAM-only mode: rank against system memory, offload allowed.
|
|
|
|
|
|
system["has_gpu"] = False
|
|
|
|
|
|
system["gpu_vram_gb"] = 0
|
|
|
|
|
|
system["gpu_count"] = 0
|
|
|
|
|
|
system["gpu_only"] = False
|
|
|
|
|
|
system.pop("active_group", None)
|
|
|
|
|
|
elif grp:
|
|
|
|
|
|
_apply_group(grp, n)
|
|
|
|
|
|
system["gpu_only"] = True
|
|
|
|
|
|
else:
|
|
|
|
|
|
# No per-GPU detail (older detection) — assume uniform split.
|
|
|
|
|
|
single_vram = (system.get("gpu_vram_gb") or 0) / (system.get("gpu_count") or 1)
|
|
|
|
|
|
system["gpu_count"] = max(1, n)
|
|
|
|
|
|
system["gpu_vram_gb"] = round(single_vram * max(1, n), 1)
|
|
|
|
|
|
system["gpu_only"] = True
|
|
|
|
|
|
elif grp:
|
|
|
|
|
|
# No explicit count, but we still pin to one pool so heterogeneous
|
|
|
|
|
|
# boxes rank against a real mixable group, not a fictional VRAM sum.
|
|
|
|
|
|
# gpu_only stays off here so the default view still surfaces offload.
|
|
|
|
|
|
_apply_group(grp, grp["count"])
|
|
|
|
|
|
|
Cookbook: scoring fixes, UI polish, false-finished + stale-state bug fixes
Backend (services/hwfit + routes):
- rank_models picks visible set by REQUESTED column, not always score —
sorting by Param now shows highest-param models PERIOD (incl. too_tight).
- New fit_only param. Multi-GPU rigs filter GGUF Q*/IQ quants (vLLM/SGLang
cannot serve them); default non-prequantized to BF16 on 2+ GPUs.
- AWQ / GPTQ-8bit get a -1.0 quality penalty (was 0.0, tied with FP8), so
FP8 wins when both fit.
- Version-aware tiebreaker (parse Mn.n / Vn) — MiniMax-M2.7 ranks above
M2.5 on equal composite score; >=100B integers not misread as versions.
- /api/cookbook/hf-latest no longer drops models without an "NB" pattern in
the repo id (MiniMax-M2.7, DeepSeek-V4-Pro etc. were silently filtered).
- Cached-model scan: atexit flushes models JSON even if the script is
killed mid-walk; each scan_dir wrapped in try/except; timeout 60s -> 180s.
- KB granularity for sub-MB sizes (was "0 MB" for 12 KB shells). New
"stalled" status for shells <1 MB with no .incomplete files.
- /api/cookbook/state POST guard: rejects "done" download tasks lacking
DOWNLOAD_OK / DOWNLOAD_FAILED / /snapshots/ when the last-mentioned
shard is N<total — stops stale tabs from poisoning persisted state.
- hf_models.json: add zai-org/GLM-5.1; flip zai-org/GLM-5 quantization
Q4_K_M -> BF16 (it is the native base, not a quant).
Frontend (static/js):
- Scan/Download toolbar: quant defaults to All; ctx slider (8k/16k/32k/
50k/128k/Max) ported from origin/main with sort=fit on drag, sort=score
on Max. GPU toggle commits _activeCount to maxGpu on initial render. Fit
column header tagged with active budget (RAM / GPU / N GPU).
- Foldable Download admin-card: the Download h2 is the chevron trigger;
state persists in localStorage.
- Download card surfaces destination dir (Dir: <path>). Same dir on running
task row, font/color matched to uptime (9px Fira Code muted, opacity .4).
- Serve panel ctx text input always resets to model max on open. Sub-MB
cached models show with red "download stalled" badge.
- Bulk-select Cancel + Delete reset the Select button label on exit.
- Cookbook running: false-finished bug fixed — DOWNLOAD_OK or /snapshots/
required; bare "Download complete" no longer marks the task done after
the first config file. Clear button now sends tmux kill-session too.
True overall % for multi-shard downloads: ((N-1)+frac)/total instead of
hf_transfer per-shard aggregate.
- Diagnosis card simplified: removed fold toggle, copy button, dismiss X.
Suggestion font matches message body (12px).
- HF token field flashes green check + "Saved" on save.
- Cached scan no longer counts stalled rows as downloaded in Scan/Download.
CSS:
- dep Install button width pinned to 76px to match Installed split.
- task-sub row +1px; task-status badge gets margin-right 8px.
- Ctx slider styled like gallery editor sliders (thin pill rail, red thumb).
- Bulk-select cancel button top -3px -> -5px.
2026-06-03 16:32:20 +09:00
|
|
|
|
try:
|
|
|
|
|
|
target_context = int(ctx) if ctx else None
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
target_context = None
|
|
|
|
|
|
if target_context is not None:
|
|
|
|
|
|
target_context = max(1024, min(target_context, 1000000))
|
|
|
|
|
|
|
Cookbook UI: Ollama browser, advanced serve fold, API tokens form, diagnosis toolbar, polish
Surface a lot of accumulated cookbook + UI work as a single non-agent
commit so the agent rework lands cleanly.
Highlights:
- Ollama as a first-class backend in the Cookbook:
* Download input accepts ollama-style names (name:tag) → backend=ollama
* /api/cookbook/ollama/library (cached scrape of ollama.com + curated
fallback so classic models like qwen2.5 stay reachable)
* "Browse Ollama library" toggle below Download with size chips
* Engine=Ollama in hwfit toolbar merges the Ollama library into the
main scan list as per-tag rows with the same Fit/Param/Quant/VRAM
columns; click → fills Download input
- API Tokens form added to Integrations panel (matching wired
loadTokens()/initTokenForm() that had no HTML)
- Serve panel polish: Advanced fold tightening (-8px nudges on vLLM
checks, Extra args, Spec row), n_cpu_moe + Split Mode controls
pulled up 8px to align with the row's checkboxes, GGUF File dropdown
exposed for Ollama backend, GPU re-render on Edit serve restore,
_forceBackend flag so saved serveState wins over backend detection,
cookbook:servers-changed CustomEvent so panels don't need refresh
- Models page redesign: Add Models row (URL + hidden API key reveal +
Type select + Scan/Ollama/Key/Test/Add icon buttons), Probe All +
Clear-offline buttons in Added Models toolbar, offline-pill removed
(opacity already conveys state), Engine dropdown gains Ollama option
- _ping_endpoint probes /v1/models then base, accepts 4xx as
reachable (vLLM returns 404 on bare /v1, fully working endpoints
were showing offline)
- Diagnosis card: × dismiss + Copy bundle buttons restored on the
serve error feedback card
- Orphan tmux sweep re-enabled behind a 60s rate-limit + background
Thread (off the main event loop) so dead serves get discovered
- cookbook_routes auto-register watchdog: drops the endpoint if the
serve session exits non-zero within the first ~3min
- ollama-rocm sidecar awareness in download wrapper (`docker exec
ollama-rocm ollama pull` when host ollama isn't installed)
- Skill extractor sets initial_status="published" when
auto_approve_skills pref is on (audit demotes later)
- Skill list / model list / cookbook scan misc polish
2026-06-08 22:38:49 +09:00
|
|
|
|
rank_kwargs = {
|
|
|
|
|
|
"use_case": use_case or None,
|
|
|
|
|
|
"limit": limit,
|
|
|
|
|
|
"search": search or None,
|
|
|
|
|
|
"sort": sort,
|
|
|
|
|
|
"quant": quant or None,
|
|
|
|
|
|
"fit_only": fit_only,
|
|
|
|
|
|
}
|
|
|
|
|
|
if target_context is not None:
|
|
|
|
|
|
rank_kwargs["target_context"] = target_context
|
|
|
|
|
|
try:
|
|
|
|
|
|
import inspect
|
|
|
|
|
|
supported = set(inspect.signature(rank_models).parameters)
|
|
|
|
|
|
rank_kwargs = {k: v for k, v in rank_kwargs.items() if k in supported}
|
|
|
|
|
|
except Exception:
|
|
|
|
|
|
rank_kwargs.pop("target_context", None)
|
|
|
|
|
|
rank_kwargs.pop("fit_only", None)
|
|
|
|
|
|
results = rank_models(system, **rank_kwargs)
|
2026-05-31 23:58:26 +09:00
|
|
|
|
return {"system": system, "models": results}
|
|
|
|
|
|
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
@router.get("/profiles")
|
2026-06-21 11:02:35 +00:00
|
|
|
|
def get_serve_profiles(model: str = "", model_path: str = "", host: str = "", ssh_port: str = "", platform: str = "", fresh: bool = False, serve_weights_gb: float = 0.0, serve_quant: str = ""):
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
"""Compute llama.cpp serve profiles (Quality/Balanced/Speed) for `model`
|
|
|
|
|
|
against the detected hardware on `host` (or local). Returns concrete
|
|
|
|
|
|
flags (n_gpu_layers, n_cpu_moe, cache_type, ctx) the serve UI can apply.
|
|
|
|
|
|
|
|
|
|
|
|
`model` is matched against the catalog by name; if it's not in the
|
|
|
|
|
|
catalog (e.g. an ad-hoc HF repo), pass enough hints via a minimal synthetic
|
|
|
|
|
|
entry isn't possible here, so we return [] and the UI keeps manual flags.
|
|
|
|
|
|
"""
|
|
|
|
|
|
from services.hwfit.hardware import detect_system
|
|
|
|
|
|
from services.hwfit.models import get_models
|
|
|
|
|
|
from services.hwfit.profiles import compute_serve_profiles
|
2026-06-11 01:43:49 +03:00
|
|
|
|
host, ssh_port = _validate_detection_target(host, ssh_port)
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
system = detect_system(host=host, ssh_port=ssh_port, platform=platform, fresh=fresh)
|
|
|
|
|
|
if system.get("error"):
|
|
|
|
|
|
return {"system": system, "profiles": [], "error": system["error"]}
|
|
|
|
|
|
catalog = {m.get("name"): m for m in (get_models() or [])}
|
|
|
|
|
|
|
|
|
|
|
|
def _norm(s):
|
|
|
|
|
|
# Normalize for matching: drop org/ prefix, a trailing -GGUF/-gguf
|
|
|
|
|
|
# marker, and any quant tag, lowercase. So "DeepSeek-Coder-V2-Lite-
|
|
|
|
|
|
# Instruct-GGUF" (a local folder name) matches catalog entry
|
|
|
|
|
|
# "deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct".
|
|
|
|
|
|
s = (s or "").lower().strip()
|
|
|
|
|
|
s = s.split("/")[-1] # drop org prefix
|
2026-06-22 02:39:18 +00:00
|
|
|
|
for suffix in ("-gguf", "_gguf", ".gguf", "gguf"):
|
|
|
|
|
|
if s.endswith(suffix):
|
|
|
|
|
|
s = s[: -len(suffix)]
|
|
|
|
|
|
break
|
|
|
|
|
|
cut_at = None
|
|
|
|
|
|
for idx, ch in enumerate(s):
|
|
|
|
|
|
if ch not in "-_." or idx + 1 >= len(s):
|
|
|
|
|
|
continue
|
|
|
|
|
|
suffix = s[idx + 1:]
|
|
|
|
|
|
if (
|
|
|
|
|
|
suffix in {"fp8", "bf16", "f16"}
|
|
|
|
|
|
or suffix.startswith(("awq", "gptq", "iq"))
|
|
|
|
|
|
or (suffix.startswith("q") and len(suffix) > 1 and suffix[1].isdigit())
|
|
|
|
|
|
):
|
|
|
|
|
|
cut_at = idx
|
|
|
|
|
|
if cut_at is not None:
|
|
|
|
|
|
s = s[:cut_at]
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
return s
|
|
|
|
|
|
|
|
|
|
|
|
m = catalog.get(model)
|
|
|
|
|
|
if m is None and model:
|
|
|
|
|
|
want = _norm(model)
|
|
|
|
|
|
for name, entry in catalog.items():
|
|
|
|
|
|
nn = _norm(name)
|
|
|
|
|
|
if nn and (nn == want or want.endswith(nn) or nn.endswith(want)):
|
|
|
|
|
|
m = entry
|
|
|
|
|
|
break
|
2026-06-21 11:02:35 +00:00
|
|
|
|
path_meta = _inspect_model_path(model_path or model, host=host, ssh_port=ssh_port)
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
if m is None:
|
2026-06-21 11:02:35 +00:00
|
|
|
|
return {
|
|
|
|
|
|
"system": system,
|
|
|
|
|
|
"profiles": [],
|
|
|
|
|
|
"error": "model not in catalog",
|
|
|
|
|
|
"model_ctx_max": int(path_meta.get("model_ctx_max") or 0),
|
|
|
|
|
|
"model_weights_gb": float(path_meta.get("model_weights_gb") or 0),
|
|
|
|
|
|
"model_probe_error": path_meta.get("model_probe_error") or "",
|
|
|
|
|
|
}
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
# Surface the model's trained context limit so the serve UI can clamp a
|
|
|
|
|
|
# user-typed context down to it (asking for ctx > n_ctx_train overflows
|
|
|
|
|
|
# and, with a quantized KV cache, can crash the GPU).
|
|
|
|
|
|
model_ctx_max = 0
|
|
|
|
|
|
for k in ("context_length", "max_position_embeddings", "n_ctx_train", "context"):
|
|
|
|
|
|
v = m.get(k)
|
|
|
|
|
|
if isinstance(v, (int, float)) and v > 0:
|
|
|
|
|
|
model_ctx_max = int(v)
|
|
|
|
|
|
break
|
2026-06-21 11:02:35 +00:00
|
|
|
|
path_ctx_max = int(path_meta.get("model_ctx_max") or 0)
|
|
|
|
|
|
if path_ctx_max > 0:
|
|
|
|
|
|
model_ctx_max = max(model_ctx_max, path_ctx_max)
|
|
|
|
|
|
model_weights_gb = float(path_meta.get("model_weights_gb") or 0)
|
|
|
|
|
|
if model_weights_gb <= 0:
|
|
|
|
|
|
for k in ("min_vram_gb", "required_gb", "size_gb", "recommended_ram_gb", "min_ram_gb"):
|
|
|
|
|
|
v = m.get(k)
|
|
|
|
|
|
if isinstance(v, (int, float)) and v > 0:
|
|
|
|
|
|
model_weights_gb = float(v)
|
|
|
|
|
|
break
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
return {
|
|
|
|
|
|
"system": system,
|
|
|
|
|
|
"profiles": compute_serve_profiles(
|
|
|
|
|
|
system, m,
|
|
|
|
|
|
serve_weights_gb=(serve_weights_gb or None),
|
|
|
|
|
|
serve_quant=(serve_quant or None),
|
|
|
|
|
|
),
|
|
|
|
|
|
"model_ctx_max": model_ctx_max,
|
2026-06-21 11:02:35 +00:00
|
|
|
|
"model_weights_gb": model_weights_gb,
|
|
|
|
|
|
"model_probe_error": path_meta.get("model_probe_error") or "",
|
Cookbook serve profiles and engine filter
* Cookbook: Engine filter + intelligent hardware-computed serve profiles
Two related Cookbook serving improvements for accurate, hardware-aware model
serving (especially on consumer GPUs that can only run GGUF/llama.cpp).
Engine filter
- New "Engine" dropdown (All / llama.cpp / vLLM / SGLang) beside the quant
picker. Pure client-side view filter over the fetched list via the same
_detectBackend() the serve commands use, so what you filter to is exactly what
would launch. Re-renders from cache (no refetch). Empty-state message + the
instant-cache-paint path account for it too.
Intelligent serve profiles (Quality / Balanced / Speed)
- services/hwfit/profiles.py: compute_serve_profiles() turns detected VRAM +
model size into concrete llama.cpp flags (n_gpu_layers, n_cpu_moe, cache-type,
context). Encodes the by-hand tuning: a too-big MoE offloads experts to CPU
instead of failing; a model that fits stays fully on GPU; quant tracks profile
intent; vision models keep image-encoder headroom. Reuses models.py VRAM math
so filtering and serving agree on what fits. Pure/deterministic (no t/s claims
— partial-offload speed isn't reliably predictable; fit is what's computed).
- /api/hwfit/profiles endpoint returns the profiles + the model's trained
context limit, with loose name matching (strips org/ prefix, -GGUF suffix,
quant tag) so a local GGUF folder name resolves to its catalog entry.
- _buildServeCmd (llama.cpp) now emits --n-cpu-moe / --flash-attn /
--cache-type-k/v when set, with llama-cpp-python fallback equivalents. It
previously only set -ngl/-c, which is why it OOM'd or ran slow.
- Serve panel: profile chips that fill the fields on click, plus CPU-MoE / KV
Cache / Flash Attn fields. Context is clamped to the model's trained limit
(and an absolute 1M sanity ceiling) on type/blur/profile-load and at launch —
fixes a crash where a stale 256k/16M preset + quantized KV cache caused an
amdgpu ErrorDeviceLost.
Tests: tests/test_serve_profiles.py (7) — offload vs full-GPU fit, never exceed
VRAM, context cap, launchable flags, vision headroom, no-GPU empty.
Checks: py_compile + node --check pass; pytest test_serve_profiles + test_hwfit_amd
green; verified live on an RDNA4 box (gfx1200) — Balanced lands ~ncm18 q4 128k,
matching hand-tuning.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook: make column-header sorting discoverable (incl. Newest)
Sorting in Cookbook is via clickable column headers (pewds' design), but the
headers had no visual cue that they're interactive — so sorting in general, and
the Newest sort on the Model header specifically, was undiscoverable.
- Style sortable headers as interactive: pointer cursor, hover underline, and
the active sort column bolded/highlighted. There was no CSS for
.hwfit-sortable / .hwfit-sort-active at all; this helps every existing sort,
not just Newest.
- The Model column header sorts by release_date (newest first), reusing the
existing header-click sort wiring and the "newest" SORT_KEY.
No new sort control — uses the existing column-header paradigm.
Checks: node --check passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve profiles: keep the on-disk file's quant fixed (don't propose Q6/Q2)
In the Serve tab the model is a specific GGUF file already on disk, so its quant
can't change — but the profiles were suggesting "Quality · Q6_K" / "Speed · Q2_K"
as if you could re-quantize it. That's meaningless when serving a fixed file.
- compute_serve_profiles gains serve_weights_gb / serve_quant. When set (SERVE
mode), the quant is locked to the file's and profiles differ only in the real
serving knobs — n_cpu_moe, KV-cache type, context. _weights_gb / _cpu_moe_for_budget
use the file's actual size instead of a quant-derived estimate. DOWNLOAD mode
(no override) still varies the quant to show download options.
- /api/hwfit/profiles accepts serve_weights_gb & serve_quant.
- The Serve panel parses the file's size (from m.size "20.6 GB") and quant (from
the repo/file name) and passes them, so profiles match what's actually served.
Result for a 20.6 GB Q4_K_M file: all three profiles stay Q4_K_M and differ by
KV/ctx/offload (Quality q8 KV 128k ncm21, Balanced q4 128k ncm17, Speed q4 32k
ncm15) — no nonsensical quant changes.
Tests: test_serve_mode_keeps_fixed_quant. Full serve-profile suite green (9).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Cookbook serve: Vision toggle (auto-find mmproj) + live VRAM/RAM-spillover monitor
Two serve-panel additions:
1. **Vision toggle.** A "Vision" checkbox that serves the model with its
multimodal projector so it can read images. The mmproj path is resolved at
runtime (find mmproj-*.gguf next to the model), so dropping an mmproj file in
the model folder makes the toggle just work; `--mmproj … --image-max-tokens
1024` (native) / `--clip_model_path` (llama-cpp-python) only when on + found.
2. **Live GPU-memory monitor.** A readout that polls /api/cookbook/gpus every 4s
while the panel is open and shows VRAM used/total/%, free, and — crucially on
a discrete card — **RAM spillover** (AMD gtt_used_mb), with a plain-language
health hint: green/healthy, amber/tight, red/"spilled to RAM — slow (raise
CPU MoE or lower context)". Surfaces gtt_used_mb from the gpus endpoint
(previously read for total only and discarded for 'used').
Lets you see at a glance whether a config fits VRAM (fast) or is paging to system
RAM over PCIe (slow) instead of guessing.
Checks: node --check + py_compile pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 05:34:42 +02:00
|
|
|
|
}
|
|
|
|
|
|
|
2026-05-31 23:58:26 +09:00
|
|
|
|
@router.get("/image-models")
|
|
|
|
|
|
def get_image_models(sort: str = "fit", search: str = "", host: str = "", gpu_count: str = "", ssh_port: str = "", platform: str = "", fresh: bool = False, manual_mode: str = "", manual_gpu_count: str = "", manual_vram_gb: str = "", manual_ram_gb: str = "", manual_backend: str = "", ignore_detected_gpu: bool = False, ignore_detected_ram: bool = False):
|
|
|
|
|
|
"""Rank image generation models against detected hardware."""
|
|
|
|
|
|
from services.hwfit.hardware import detect_system
|
|
|
|
|
|
from services.hwfit.image_models import rank_image_models
|
2026-06-11 01:43:49 +03:00
|
|
|
|
host, ssh_port = _validate_detection_target(host, ssh_port)
|
2026-05-31 23:58:26 +09:00
|
|
|
|
system = deepcopy(detect_system(host=host, ssh_port=ssh_port, platform=platform, fresh=fresh))
|
|
|
|
|
|
if system.get("error"):
|
|
|
|
|
|
return {"system": system, "models": [], "error": system["error"]}
|
|
|
|
|
|
if ignore_detected_gpu:
|
|
|
|
|
|
system["has_gpu"] = False
|
|
|
|
|
|
system["gpu_name"] = None
|
|
|
|
|
|
system["gpu_vram_gb"] = 0
|
|
|
|
|
|
system["gpu_count"] = 0
|
|
|
|
|
|
system["gpus"] = []
|
|
|
|
|
|
system["gpu_groups"] = []
|
|
|
|
|
|
if ignore_detected_ram:
|
|
|
|
|
|
system["available_ram_gb"] = 0
|
|
|
|
|
|
system["total_ram_gb"] = 0
|
|
|
|
|
|
system = _apply_manual_hardware(system, manual_mode, manual_gpu_count, manual_vram_gb, manual_ram_gb, manual_backend)
|
|
|
|
|
|
# Image models use a single GPU — always use per-GPU VRAM
|
|
|
|
|
|
gpu_vrams = [float(g.get("vram_gb") or 0) for g in (system.get("gpus") or []) if isinstance(g, dict)]
|
|
|
|
|
|
single_vram = max(gpu_vrams) if gpu_vrams else ((system.get("gpu_vram_gb") or 0) / max(system.get("gpu_count") or 1, 1))
|
|
|
|
|
|
system["gpu_vram_gb"] = single_vram
|
|
|
|
|
|
system["gpu_count"] = 1 if single_vram > 0 else 0
|
|
|
|
|
|
results = rank_image_models(system, search=search or None, sort=sort)
|
|
|
|
|
|
return {"system": system, "models": results}
|
|
|
|
|
|
|
|
|
|
|
|
return router
|