Dropstone 如何選擇其模型
每個月 Blankline 都會針對 Dropstone 各層級可用的開放權重模型,以三項測試重新評估:完成工作的表現、提供服務的成本,以及使用工具時的安全行為。新模型必須擊敗目前佔據該位置的模型。
Fast、Pro 和 Heavy 是透過每月評估來決定,而非永久性的選擇。本頁說明評估的內容、由誰執行,以及要贏得一個位置需要具備什麼條件。
誰來決定?
評估由 Blankline(Dropstone 的開發公司)內部執行。這是一項持續性的承諾,而非一次性活動:模型陣容每個月都會重新評估,包括那些結論是「無需變更」的月份。
評估什麼?
每個候選模型都會接受三項測試。
| 評估項目 | 評估內容 |
|---|---|
| 完成工作的表現 | 程式碼生成、編輯現有儲存庫、工具使用、從長文件中檢索資訊,以及跨多個步驟的規劃 |
| 提供服務的成本 | 典型重度編碼回合的實際量測成本,而非模型的標價 |
| 使用工具時的安全行為 | 在長對話中遵循指令、拒絕應拒絕的請求,以及在允許行動時抵禦提示注入攻擊 |
模型如何取得位置
候選模型必須擊敗或至少持平於目前佔據該位置的模型。僅憑「更新」本身並不構成資格。
三個層級是分開評判的,因為工作內容不同,所以一個模型可能贏得某個位置,卻輸掉另一個位置。失去位置的模型也可能降級而非被淘汰:在 Kimi K3 接手 Heavy 之前,原本佔據 Heavy 的模型是降級到 Pro,而不是被直接移除。
選定之後會發生什麼
一旦某個層級的選擇變更,版本就會更新,新的模型陣容隨之上線。模型版本與每月陣容說明了版本編號規則,而發行說明則記錄了變更內容。
除此之外,其他一切都不受影響。你的配額、你的記憶,以及你的對話都不會受到影響。
評估理由發布在哪裡
Blankline 會發布模型陣容背後的評估理由,而不只是結果。詳細說明與基準測試資料位於 blankline.org/research,而 Dropstone 部落格則以更平易近人的語言介紹模型陣容。
如果某個週期沒有任何變更,就沒有什麼需要宣布的,版本也會維持原狀。
Related articles
- Which AI models does Dropstone use?Dropstone runs three tiers. On the 1.8 lineup, Fast runs DeepSeek V4.1 Flash, Pro runs GLM-5.3 Flash, and Heavy runs Kimi K3, re-evaluated every month.
- Model versions and the monthly lineupDropstone numbers each lineup with the year and the month it was cut, and a new version ships only when a tier's pick changes. How to read the number, and which lineup you get.
- Why Dropstone runs open-weight modelsThe models behind Dropstone Fast, Pro, and Heavy are open-weight, so the weights are published and anyone can check which model is behind a tier. Here is what that gives you, and what it does not mean.
- Which model am I using?The button at the foot of the composer names the tier, the lineup version and the effort level, and your usage page records which model handled each request.
- Choose a modelHow to pick between Dropstone Fast, Pro, and Heavy, where the selector lives in chat and in the CLI, and why the name Fast describes what a tier costs rather than how fast the model inside it runs.
Ctrl+I