LoRA and PEFT: Make Model Adaptation Modular, Not Free

Full finetuning says: make every parameter in the model eligible to change. That is a powerful default and an expensive one.

Parameter-efficient finetuning, or PEFT, starts from a different premise: the base model already contains most of the capability the task needs. Freeze those weights and learn a small, task-specific update instead.

PEFT does not make adaptation free. It changes the thing you pay to train, store, evaluate, and serve. That distinction explains why LoRA has become so useful for specialized features and why it can still be the wrong answer for a particular task.

Full, Partial, and Parameter-Efficient Updates

In full finetuning, every model weight is trainable. Gradients and optimizer state therefore scale with the full model. A 7B model stored at 16 bits needs around 14 GB for weights alone; using Adam to update every weight adds roughly 42 GB for gradients and optimizer state before activation memory enters the calculation.

Partial finetuning reduces that bill by freezing some layers, often keeping early layers unchanged and updating later ones. It is a reasonable idea, but it is not automatically parameter-efficient. Updating a quarter of a large model is still a large training and serving decision.

PEFT methods aim for a stronger result: performance close to full finetuning while changing orders of magnitude fewer parameters. They typically belong to one of two families:

  • adapter-based methods add small trainable modules or updates while leaving the base model untouched;
  • soft-prompt methods learn continuous input-like vectors that guide the model but are not human-readable prompt text.

The latter are useful in some settings, but adapters are often easier to reason about operationally because the resulting artifact is explicit and can be versioned, evaluated, and attached to a known base model.

LoRA Learns a Small Update, Not a Small Model

LoRA, short for Low-Rank Adaptation, is the common adapter-based approach. Consider a base weight matrix $W$ with dimensions $n \times m$. Rather than training a complete replacement, LoRA learns two smaller matrices:

$$ A: n \times r, \quad B: r \times m $$

Their product has the shape of the original matrix, allowing the serving model to use an update such as:

$$ W' = W + \alpha AB $$

The rank $r$ is deliberately small. A $4096 \times 4096$ matrix contains nearly 16.8 million values. With rank 8, the two LoRA factors contain $4096 \times 8 \times 2 = 65,536$ trainable values. The base matrix stays frozen.

This is not a claim that LoRA compresses the full base model into two tiny matrices. The base model remains necessary. LoRA compresses the adaptation needed for a task into a small update.

That works because a well-pretrained model often needs a constrained behavioral shift for a downstream task rather than a new world model. The practical implication is more modest than the theory: start from a capable base model, and do not assume a low-rank update can rescue a model that lacks the underlying ability.

The Important Knobs Are Architectural

"Use LoRA" is not yet a configuration. A team must choose which matrices receive updates, how much adapter capacity to allocate, and how strongly those updates contribute.

For transformers, LoRA is commonly applied to attention projections: query, key, value, and output. With a fixed parameter budget, spreading a low rank across more of those matrices can outperform spending all of it on a single matrix. Attention is not the only option: feedforward layers can add quality, especially when paired with attention adapters, but they consume more of the budget.

Rank is a capacity control, not a quality dial that should always be turned up. Small ranks such as 4 through 64 often work well. Larger ranks increase trainable state and can overfit without improving the target metric. The scaling value $\alpha$ adds another interaction: the same rank can behave differently when the learned update is scaled differently.

Treat these as experiments, not folk wisdom. Define a budget and compare configurations against a holdout set that includes the normal workflow, boundary cases, and tasks that must not regress.

QuestionWhy it matters
Which base model and revision?An adapter is meaningful only with its compatible base weights.
Which modules receive LoRA?Architecture choices affect quality, memory, and the tasks that change.
What rank and scaling are used?They set adaptation capacity and can change overfitting behavior.
Which data and evaluation version?The adapter must be reproducible and its claims auditable.
Is the adapter merged or loaded dynamically?This determines latency, storage, switching, and rollback behavior.

Merged and Dynamic Adapters Optimize Different Workloads

A LoRA update can be merged into the base weights before serving. The resulting model adds no adapter computation at inference time, which is attractive for a single specialized deployment. The tradeoff is storage: every specialization becomes a full model artifact.

Alternatively, keep the base weights and adapter factors separate and combine them during inference. This introduces some overhead, but many adapters can share one loaded base model.

Imagine a product with one 16.8M-parameter base matrix and 100 customer-specific rank-8 adapters. Storing 100 merged versions means carrying that full matrix 100 times. Storing one base plus 100 small adapter pairs is dramatically smaller. It also makes an adapter switch faster than loading a different full model.

That makes dynamic adapters appealing for multi-tenant or multi-feature systems, but it creates a new control plane. An adapter selection is now security-sensitive routing. The service must determine which adapter a request may use from authenticated server-side state, not from a client-provided model name. Cache keys, traces, rate limits, and evaluation records need the base-model and adapter identifiers together.

type ModelProfile = {
  baseModel: string
  adapterId: string
  adapterVersion: string
  maxInputTokens: number
}

function modelProfileForAccount(account: Account): ModelProfile {
  return account.aiProfile
}

The frontend should receive the result of that decision, not the ability to select an arbitrary adapter. A mismatch can leak a tenant's specialized behavior or quietly produce an answer from the wrong policy configuration.

Quantized LoRA Solves a Different Memory Problem

Once the adapter is small, reducing its size further does little for the total footprint. The large term is still the frozen base model.

QLoRA attacks that term by storing base weights at low precision, commonly 4 bits, while dequantizing them for the forward and backward calculations where needed. Combined with techniques such as paging data between CPU and GPU, this can make a much larger model finetunable on a single GPU than conventional 16-bit storage would allow.

The tradeoff is time and numerical complexity. Quantization and dequantization can slow training, and the model, hardware, and framework need a compatible path. "Fits in memory" must still be measured against training throughput, quality, sequence length, and operational cost.

PEFT Still Needs an Exit Criterion

PEFT is usually a sensible first finetuning experiment because it lowers the cost of learning whether adaptation is valuable. It can work with fewer examples than full finetuning and makes multiple variants practical to store and compare.

It does not guarantee full-finetuning quality. It can fail when the requested capability requires a broader model change, when the base model is weak, when data is unrepresentative, or when the adapter configuration targets the wrong modules. It also cannot make changing business facts current; that remains a retrieval and data-governance problem.

Choose LoRA when the product needs a stable specialization and benefits from modular variants that share a base model. Keep the base, adapter, dataset, configuration, and evaluation result as one versioned release. That turns a small set of matrices into an engineered artifact rather than a mysterious file that happened to improve a demo.