ALIGNED BY HARMAN SANDHU
DPO · QLoRA PREFERENCE OPTIMIZATION

Gemma 2 2B · DPO

Our Gemma QA fine-tune, aligned with Direct Preference Optimization on 500 AI-judged preference triplets — no reward model.

HuggingFace weights →
2.6B
Parameters
256K
Vocab
8,192
Context
4-bit
NF4 quantized
92.4%
Judged correct
46%
Win-rate vs SFT
90.6%
Groundedness
Model lineage
interactiveconnecting…

Ask a question

Connecting to the inference endpoint…
The model's answer will appear here.
preference-aligned (QLoRA)

What this is

The Gemma QA adapter, further trained with DPO against a frozen reference. Preferences were written by Gemini and verified by a blind LLM judge that the chosen answer is genuinely better.

750 steps, ~7 min on an L4. Judged correctness 0.924 [0.906–0.942]; the paired win-rate against the SFT checkpoint is 46% (ties split), so DPO did not raise judged accuracy on the already-strong Gemma — it mostly shifted style rather than correctness.

Architecture
ClassGemma2ForCausalLM
Layers26
Hidden size2,304
Attention8 heads / 4 KV · dim 256 · GQA
Feed-forwardGeGLU · inner 9,216
Attention windowsliding 4,096 (alternating) · logit soft-cap
NormRMSNorm
Context8,192 tokens
Vocabulary256,128
Embeddingstied input/output
Training
Init fromgemma-2-2b QA-SFT adapter
MethodQLoRA-DPO · rank-16
Trainable params20.8M LoRA (0.8%)
Training data500 preference pairs
Training tokens335K (chosen + rejected)
Optimizer steps750
What this model cost to build

$11.75 total Modal usage

our cost begins at fine-tuning — the base is Google's, imported free.

StageDetailCost
QLoRA QA fine-tunethe adapter DPO starts from$9.84
QLoRA-DPO alignmentdirect preference optimization on L4$0.43
Evaluation (shared)13 versions on 500 held-out questions, this model's share$1.48
Total$11.75

Figures are Modal GPU usage (time × rate) across this model's lineage; shared datasets are charged at this model's share. Whether base pretraining is included is stated above — it is for the models pretrained here, and excluded for imported bases. Evaluation-derived metrics come from an independent blind-judge harness on a frozen, decontaminated held-out set. Serving is billed separately and scales to zero.