對話摘要 2026-08-11

2026-08-11

Auto Router 測試(深夜完成 ✅)

用戶要求:今晚測試 LiteLLM Auto Router (beta),聽日俾報告。

測試配置:新增 test-smart-router(auto_router/complexity_router,weighted_scores)

  • SIMPLE→deepseek-first / MEDIUM→free-first / COMPLEX→deepseek-first / REASONING→azure-gpt-4o

結果:9 個 request 全部成功,router 有真正生效(4 種 model 路由)

  • ✅ REASONING 強制規則生效:多重 reasoning markers → azure/gpt-4o
  • ✅ SIMPLE 簡單查詢 → deepseek-v4-flash
  • ⚠️ Code 任務傾向去 MEDIUM(free-first CF 模型,質素參差)
  • ⚠️ decision log 冇輸出,要靠 response header x-litellm-model-name 驗證

完整報告:auto-router-report-2026-08-11.md(workspace) 測試 script:auto_router_test.py + auto_router_test2.py 建議:MEDIUM 唔好放 code;REASONING 可考慮 deepseek-v4-pro 代替 azure-gpt-4o;可加 keyword_tier_rules / adaptive / session_affinity

安排:cron(明早 08:00)自動發送報告去 Telegram。model test-smart-router(ebc969ee-…)留低等用戶決定。

磁碟清理(✅ 完成)

用戶要求立即清理磁碟。執行前磁碟 91%(剩 18G)。清理完成:刪除 9 個 OpenRAG 殘留鏡像(~40GB,2026-08-10 移除 stack 時漏咗鏡像)+ build cache 3.2GB,磁碟 91%→72%,可用空間 18G→52G(釋放 ~34GB)。保留 caddy:2.9-alpine(sandbox-gateway 用緊)、/home/opc/backup(933M)、/opt/rclone/data(898M 備份 staging)。

Hermes 轉返 gpt-4o(✅ 完成)

用戶要求幫 Hermes 轉做 gpt-4o。用 azure-switch-to-deepseek.sh --back 執行:config default 由 deepseek-first → azure-gpt-4o,key 換返 Azure 專用 key,max_tokens 16384,restart 完成,API healthy(status 200)。lock 已移除——若 Azure credit 再觸發自動切換(n8n 每日檢查)Hermes 會跟住切返 deepseek-first。

LiteLLM Auto-Routers 詳解

官方 2026-07 新功能 v1.94.x+(用戶機行 main-latest 可直接用):4 tier(SIMPLE→MEDIUM→COMPLEX→REASONING)、3 種分類方式(heuristic scorer / LLM classifier / keyword rules)、session affinity 預設已改 off、Decision log;Adaptive Router(bandit,需 Postgres,用戶有 pg-main);用戶現用 simple-shuffle+fallbacks 未有 auto router,約定下次 session 起測試 router smart-router(SIMPLE: deepseek-first、MEDIUM: free-first、COMPLEX: deepseek-first、REASONING: azure-gpt-4o),注意 fallback 鏈唔會自動套入 tier。

n8n MCP token 分析

實測對比 n8n MCP(34 tools / 8,290 tok / avg 244)vs fastmcp(171 tools / 11,171 tok / avg 65):n8n 平均每個工具係 fastmcp 嘅 3.75 倍,34 個已佔 fastmcp 171 個嘅 74%。最肥係 update_workflow 單一 tool 1,727 tok(佔 n8n 21%),Top 5(update_workflow 1,727 + create_workflow_from_code 561 + explore_node_resources 483 + execute_workflow 424 + validate_node_config 420)已佔 44%。n8n 肥因內嵌完整 node definition 及大量 JSON Schema 約束噪音;fastmcp 瘦因 description 精簡、參數少、冇 validation 噪音(150/171 個 <100 tok,零個超過 300)。疑點:n8n.yaml whitelist tools: [search_workflows, get_workflow_details, execute_workflow, search_executions] 寫咗 4 個但 system prompt 顯示 34 個全 expose,疑似只影響 config_service 嘅 enabled flag 唔影響 runtime builder。慳 token 兩個方向:① 查清 n8n whitelist 真 filter 機制(慳 ~6.3K);② 開 QP_PROGRESSIVE_TOOLS=on 令 DeepSeek 用動態載入,240 tools 全部 defer,每次 request 慳 ~20K schema。

QwenPaw patches

QwenPaw 兩項 framework patch 均已完成並生效:①MCP whitelist filter(方案 B,mcp.py list_capabilities() 根因 patch 兩處 src+site-packages,n8n 縮至 6 個常用 tools 慳 ~6.5K tokens/request,fastmcp 移除 whitelist 保持全量防漏 cf_browser/grafana/openhands);②Secret Redaction 密鑰遮罩(secret_redaction.py + react_agent.py 兩處 redact_agent_context call,每次 model call 前 in-place scrub tool_result 已知 secret/credential 格式,QP_SECRET_REDACTION 預設 on + TTL 60s cache,實測 sk-/ghp_/cfut_/ak2_/st./JWT 全遮罩、unprefixed hex token 防誤傷保留、太短 key 保留)。對比:QwenPaw = 儲存層加密(Fernet)+顯示層遮罩,Hermes = 資料流 scrubbing。代價:模型見唔到明文 secret,Infisical→Dockhand 搬 secret 須工具封裝或 n8n 零 LLM(Vikunja token 輪換已係咁)。QwenPaw 升級會覆蓋全部 patch 要 re-apply。app 進程 PID 25231 21:41 起動已 restart 生效,勿睇容器 init 誤判。