🏥 Clinical LLM — QLoRA Fine-Tune on MedQA

Base model: Llama 3.2 3B Instruct  |  Fine-tune: QLoRA r=32  |  Dataset: MedQA (2,500 samples)  |  Accuracy: Base 10% → Fine-tuned 59.2%

These 6 questions were hand-selected from the validation set to give an honest picture of what fine-tuning achieves. The fine-tuned model is ~6x more accurate than the base model (59.2% vs 10%), largely because it learns to respond in the correct concise format rather than generating verbose explanations. That said, 59.2% accuracy means the model still gets a significant portion of questions wrong — the samples below reflect both the strengths and the honest limitations of a 3B parameter model fine-tuned on 2,500 examples. Try a few and see for yourself.


📊 Ablation Results

Rank Trainable Params Train Loss Accuracy ROUGE-L
Base — — 0.100 0.128
r=4 6M (0.19%) 1.1689 0.560 0.319
r=8 12M (0.38%) 1.1389 0.576 0.356
r=16 24M (0.75%) 1.1013 0.560 0.342
r=32 48M (1.50%) 1.0496 0.592 0.355

⚡ Quantization Benchmark (r=32, L4 GPU)

Quantization Latency Accuracy
4-bit 0.458s 0.588
fp16 0.583s 0.596