Posted by:

Ali

Category:

LLM

Posted on:

December 17, 2025

Parameter Efficient Fine Tuning of LLMs Using Unsloth

Parameter efficient fine tuning adapts large language models to specific tasks without retraining all weights. By training only a small set of parameters, tools like Unsloth make LoRA and DoRA practical on a single GPU or Google Colab.

Why Not Full Fine Tuning

Full fine tuning updates every parameter in a model, which is expensive and often unnecessary. PEFT focuses training on a small number of added parameters while keeping the base model frozen.

This is especially useful when you want to specialize a model for a narrow domain, iterate quickly, or deploy multiple task specific adapters without storing multiple full model copies.

Where Unsloth Fits

Unsloth is designed to make fine tuning faster and more memory efficient. It supports quantized model loading, efficient gradient checkpointing, and optimized training paths for common PEFT methods.

In practice, it helps you run experiments that would otherwise require larger GPUs, while keeping the developer experience straightforward and compatible with the Hugging Face ecosystem.

LoRA and DoRA

LoRA, or Low Rank Adaptation, injects small low rank matrices into specific layers, typically attention and feedforward projections. During training, only these new matrices are updated and the base model stays unchanged.

DoRA, or Weight Decomposed Low Rank Adaptation, builds on LoRA by separating the update into magnitude and direction components, which can improve optimization behavior and stability in some settings.

  • LoRA — one low rank update per targeted layer, base weights frozen
  • DoRA — the same update split into magnitude and direction, often steadier to train
  • Both leave the base model untouched, so adapters stay small and swappable

Practical Tips

Before training, it helps to confirm your task format, label constraints, and sequence length distribution. During training, start with conservative batch sizes, use gradient accumulation for stability, and monitor loss for divergence.

After training, save adapters and keep the base model unchanged so you can reuse it across multiple tasks.

  • Freeze the base model and train only adapters
  • Target attention and MLP projection layers first
  • Use 4 bit loading when GPU memory is limited
  • Use gradient accumulation to increase effective batch size
  • Save adapters per task instead of saving full model copies

LoRA vs DoRA in Code

In Unsloth the two methods are the same call. You load the base model once, then attach the adapters — DoRA is a single extra flag, which makes it cheap to run both and compare.

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name     = "unsloth/llama-3-8b-bnb-4bit",
    max_seq_length = 2048,
    load_in_4bit   = True,          # 4 bit base keeps memory low
)

TARGETS = ["q_proj", "k_proj", "v_proj", "o_proj",
           "gate_proj", "up_proj", "down_proj"]

# LoRA — a single low rank update per targeted layer
model = FastLanguageModel.get_peft_model(
    model,
    r              = 16,
    lora_alpha     = 16,
    lora_dropout   = 0,
    target_modules = TARGETS,
    use_gradient_checkpointing = "unsloth",
)

# DoRA — same call, one flag: the update is decomposed
# into magnitude and direction instead of one low rank term
model = FastLanguageModel.get_peft_model(
    model,
    r              = 16,
    lora_alpha     = 16,
    target_modules = TARGETS,
    use_dora       = True,
)
                

Closing Thoughts

PEFT techniques such as LoRA and DoRA are now standard tools for adapting large language models efficiently. Unsloth makes these workflows faster and more accessible by reducing memory pressure and simplifying the training stack.

LOADING