Why Finetuning Needs More Memory Than Inference
A practical memory model for finetuning: why gradients, optimizer state, activations, precision, and batch size turn an inference-sized model into a training capacity problem, and which tradeoffs actually reduce the bill.





