🏥 Clinical LLM — QLoRA Fine-Tune on MedQA
Base model: Llama 3.2 3B Instruct | Fine-tune: QLoRA r=32 | Dataset: MedQA (2,500 samples) | Accuracy: Base 10% → Fine-tuned 59.2%
These 6 questions were hand-selected from the validation set to give an honest picture of what fine-tuning achieves. The fine-tuned model is ~6x more accurate than the base model (59.2% vs 10%), largely because it learns to respond in the correct concise format rather than generating verbose explanations. That said, 59.2% accuracy means the model still gets a significant portion of questions wrong — the samples below reflect both the strengths and the honest limitations of a 3B parameter model fine-tuned on 2,500 examples. Try a few and see for yourself.
📊 Ablation Results
| Rank | Trainable Params | Train Loss | Accuracy | ROUGE-L |
|---|---|---|---|---|
| Base | — | — | 0.100 | 0.128 |
| r=4 | 6M (0.19%) | 1.1689 | 0.560 | 0.319 |
| r=8 | 12M (0.38%) | 1.1389 | 0.576 | 0.356 |
| r=16 | 24M (0.75%) | 1.1013 | 0.560 | 0.342 |
| r=32 | 48M (1.50%) | 1.0496 | 0.592 | 0.355 |
⚡ Quantization Benchmark (r=32, L4 GPU)
| Quantization | Latency | Accuracy |
|---|---|---|
| 4-bit | 0.458s | 0.588 |
| fp16 | 0.583s | 0.596 |