Site icon Kimlud.co

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling






Underdog, the on-device assistant from Conway Research, has released Saluki 27B under Apache 2.0. Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent. For developers, that means a 27B-class agent model that runs in stock llama.cpp.

TL;DR

  • Size: 27B dense parameters. 7.89 GB GGUF versus 54 GB for BF16.
  • Runs on: stock llama.cpp and apps built on it, with full GPU offload. Optional 629 MB or 928 MB vision add-on.
  • Performance: 96% average retention across 9 benchmarks versus full Qwen3.8-27B.
  • Best: Parallel tool calls, 42 versus 35 for the full model (120% retention).
  • Worst: AIME 2025, 79.2 versus 96.7 (about 82% retention).
  • Bottom line:
    • Best: beats the 54 GB original at tool calling in a sub-8 GB file.
    • Worst: competition math and multi-step reasoning drop 12 to 18 points.

What is Underdog Saluki 27B?

Saluki 27B is a 2-bit, mixed-precision GGUF of Qwen3.8-27B built for local agents. It stacks 3 layers of work:

  • The base is Qwen3.8-27B, a dense 27B model from the Qwen team. It has 64 layers, mixes Gated DeltaNet linear attention with gated attention, and supports 262,144 tokens natively.
  • The second layer is ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF. GSQ learns accurate low-bit scalar grids per tensor. RCO assigns a quantization type to each tensor under a fixed size budget. ISTA’s smallest file, IQ2_XS, is 8.4 GB at 2.50 bits per weight.
  • The third layer is Underdog’s own pass. It shrank the file to 7.89 GB and targeted tool calling. The file is named IQ2-mix and carries an imatrix tag. Underdog has not published the full recipe for this pass.

How does Saluki perform on benchmarks?

Underdog splits its results into 2 groups.

The first group ran both models in the same harness:

  • Underdog Bench: 120 tasks from BFCL v4, frozen before testing. Thinking off, temperature 0. Saluki scores 88, the full model 84, and PrismML’s Bonsai 2 scores 70.
  • Parallel tool calls: 100 BFCL v4 parallel tasks with the official checker. Saluki 42, full model 35.
  • SWE-bench Verified: 50 issues. Saluki fixes 30, the full model 33.

The second group compares Saluki with public full-size scores:

Benchmark Saluki 27B Qwen3.8-27B (public)
IFEval (prompt-loose) 93.5 91.5
IFBench (prompt-loose) 72.7 71.0
MBPP+ 78.0 83.9
MuSR 67.5 79.6
AIME 2025 (avg@4) 79.2 96.7
AIME 2026 (avg@4) 80.0 94.6

‘+r[0]+’
‘+r[1]+’ vs ‘+r[2]+(r[4]?’ ‘+r[4]:”)+’
‘+ret+’%

‘});
BM.innerHTML=h;BN.textContent=N[g]+’ Right column: Saluki as a share of the full model.’;grow(BM);setTimeout(post,80)}
q(‘.seg button’).forEach(function(b){b.onclick=function(){q(‘.seg button’).forEach(function(x){x.classList.toggle(‘on’,x===b)});draw(b.getAttribute(‘data-g’))}});draw(‘a’);
var RUN=[‘1. Download the 7.89 GB GGUF

huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir .

‘,
‘2. Serve with stock llama.cpp

llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768

–jinja turns on the Qwen3.8 chat template. OpenAI chat API on port 8080.‘,
‘3. Pick settings

General use: thinking on, temperature 0.6, top_p 0.95, top_k 20.
Fast tool calls: thinking off, temperature 0:

"chat_template_kwargs": {"enable_thinking": false}

Add –mmproj with the vision file for image input.‘];
var ri=0,RB=document.getElementById(‘mtps-run’),dots=q(‘.dot’),pv=document.getElementById(‘mtps-prev’),nx=document.getElementById(‘mtps-next’);
function setR(i){ri=i;RB.innerHTML=RUN[i];dots.forEach(function(d,j){d.classList.toggle(‘on’,j===i)});pv.disabled=i===0;nx.disabled=i===2;post()}
dots.forEach(function(d,i){d.onclick=function(){setR(i)}});pv.onclick=function(){if(ri>0)setR(ri-1)};nx.onclick=function(){if(ri<2)setR(ri+1)};setR(0);
q(‘.lc’).forEach(function(c){c.onclick=function(){c.classList.toggle(‘open’);c.querySelector(‘.h span’).textContent=c.classList.contains(‘open’)?’-‘:’+’;setTimeout(post,320)}});
grow(R.querySelector(‘.pn.on’));window.addEventListener(‘load’,post);window.addEventListener(‘resize’,post);setTimeout(post,300);
})();



Source link

Exit mobile version