ASKA DIGITAL
Explore
b.log

Post-Training Your Own Models

How to take an open model and teach it a custom task, a workflow, or your own writing voice.
In this article
  1. 1. Do you even need to fine-tune?
  2. 2. The method map
  3. 3. Supervised fine-tuning (SFT)
  4. 4. LoRA and QLoRA: the default for everyone else
  5. 5. Preference alignment: RLHF, DPO, and friends
  6. 6. RL post-training: GRPO and RLVR
  7. 7. Knowledge distillation
  8. 8. Choosing a base model
  9. 9. Building the dataset
  10. 10. Mimicking a writing style
  11. 11. Compute and cost
  12. 12. The tooling stack
  13. 13. Evaluation
  14. 14. Serving the finished model
  15. 15. Pitfalls and honest caveats
  16. 16. Three playbooks
  17. Sources and further reading

1. Do you even need to fine-tune?

Fine-tuning is the heavy answer. Most problems have lighter ones. Run through this list before spending a dollar on training.

  • Prompt engineering first. A good system prompt with a few examples (few-shot prompting) solves many style and format problems. It costs nothing and takes minutes. If the model can do the task when shown how, it does not need training.
  • Retrieval (RAG) for knowledge. Fine-tuning is a bad way to teach facts. Models trained on new facts hallucinate them back with confidence. If your task needs private documents, wire retrieval into the prompt instead.
  • Fine-tune for behavior, not information. The cases where training earns its keep: a consistent voice across thousands of outputs, a rigid output format (JSON schemas, tool calls), a domain workflow the model keeps getting wrong, or latency and cost (a small tuned model replacing a large general one).

Rule of thumb

If you can describe the desired behavior in under a page of instructions and the model follows it, stop there. Fine-tune when you need the behavior baked in: always-on, no prompt to maintain, cheaper to serve, and robust to prompt injection drift.

2. The method map

Post-training means changing a model's weights after pretraining. The methods below stack. A common pipeline is: base model → SFT on demonstrations → preference alignment (DPO) → quantize and serve.

MethodWhat it doesWhen to use itCompute
Full SFTUpdates every weight on input→output examplesLarge domain shift; you have serious data and GPUsHigh
LoRA / QLoRATrains small adapter matrices; base weights stay frozenDefault choice for tasks and style; single-GPU friendlyLow
DPO / ORPO / KTOAligns on chosen vs rejected response pairs, no reward modelPolishing tone, format, and preference after SFTLow–medium
RLHF (PPO)Trains a reward model, then optimizes against itLarge labs; rarely worth it for individuals nowHigh
GRPO / RLVRReinforcement learning with verifiable rewards (math, code, tools)Reasoning tasks with checkable answersMedium–high
DistillationA small student learns from a large teacher's outputsCompressing capability into a cheap servable modelMedium
Model mergingCombines weights of two tuned models, no trainingBlending skills (e.g. style adapter + task adapter)Near zero

3. Supervised fine-tuning (SFT)

SFT is the foundation. You show the model thousands of input→output examples and train it to predict the outputs. Everything else in this guide either replaces SFT's full-weight update with something cheaper or polishes the result afterward.

How it works, in one paragraph

Each training example is a prompt plus the desired response. The model reads the prompt, predicts the response token by token, and the loss measures how far its predictions were from the actual tokens. Gradient descent nudges the weights to close the gap. Repeat for 1–3 passes over the data (epochs). That is the whole trick.

Full SFT vs adapters

Full SFT updates every parameter. For a 7B model in 16-bit precision that means ~14 GB just for the weights, plus gradients and optimizer state. Adam keeps a master copy plus two running statistics per parameter, so full fine-tuning needs roughly 16 bytes per parameter at 16-bit: about 112 GB for 7B. That is why individuals rarely do full SFT on anything above 3B without serious hardware. Section 4 covers the standard workaround.

Hyperparameters that actually matter

  • Learning rate: 5e-6 to 2e-5 for full SFT, far below LoRA rates. Too high and the model collapses into repeating itself. Too low and nothing changes.
  • Epochs: 1–3 for large datasets; up to 5–10 for tiny ones (under 1,000 examples). More epochs on small data = memorization, not learning.
  • Batch size: effective batch of 32–128 via gradient accumulation. Larger batches stabilize training but need more memory.
  • Warmup: ramp the learning rate over the first ~5% of steps, then cosine decay. Skipping warmup is the most common cause of early-training collapse.
  • Weight decay: 0.01–0.1. Mild regularization against overfitting.
  • Loss masking: compute loss on the assistant's tokens only, not the prompt. Pack short examples together to stop wasting compute on padding.

The one failure mode to respect

Catastrophic forgetting. SFT on narrow data erodes the model's general ability. A model trained only on your support tickets will get worse at everything else. Mix in 10–30% general instruction data to keep the base capabilities intact, and check both directions: your task metrics and a general benchmark.

4. LoRA and QLoRA: the default for everyone else

LoRA (Low-Rank Adaptation) freezes the base model and trains two small matrices per layer whose product approximates the weight change. Instead of updating a 4096×4096 weight matrix (16.8M numbers), you train a 4096×16 and a 16×4096 pair (131K numbers). Same architecture, 0.1–1% of the parameters.

Why LoRA usually wins for individuals

  • Memory: the frozen base can be quantized to 4-bit. A 7B model fits in ~6 GB; the adapters and optimizer state for the adapters add a few GB more. Total: trainable on a single 24 GB consumer GPU (RTX 4090), often on 16 GB cards.
  • Speed: fewer parameters to update means faster steps and cheaper experiments. You can try ten ideas in the time one full run takes.
  • Modularity: adapters are small files (tens to hundreds of MB). Swap them at load time: one adapter for style, one for a task, one per client. The base model stays untouched.
  • Quality: on instruction-following and style tasks, LoRA matches full fine-tuning. It falls behind only when the task needs deep knowledge changes (new languages, heavy domain shift).

QLoRA: the 4-bit variant

QLoRA combines LoRA with a 4-bit quantized base (usually NF4) plus paged optimizers. It is the reason a single GPU can fine-tune 7B–13B models at all. The finding that launched it: 4-bit quantization barely dents fine-tuning quality. Use QLoRA unless you have a reason not to.

Settings that work

  • Rank (r): 16 is the default. Start at r=8 for style-only shifts; go 32–64 for harder tasks. Higher rank = more capacity, more memory, slower.
  • Alpha: set to 2× rank (the classic recipe). It just scales the adapter's influence.
  • Target modules: start with attention (query, key, value, output). If the model underfits, add the MLP projections (gate, up, down) or target all linear layers.
  • Learning rate: 2e-4 with a cosine schedule, higher than full SFT because you are moving fewer parameters.
  • Dropout: 0.05 on the adapter, cheap insurance on small datasets.

When overfitting hits

Overfitting is the number one LoRA failure mode: training loss falls while held-out scores stall. Fix in this order: reduce rank, raise dropout, cut epochs, add data. Do not reach for a bigger rank first.

Variants worth knowing

DoRA splits magnitude and direction of the update and slightly improves quality at small extra cost. rsLoRA fixes a scaling quirk so higher ranks actually help. LoRA+ uses different learning rates for the two adapter matrices. None of these change the workflow; they are drop-in upgrades in most libraries.

5. Preference alignment: RLHF, DPO, and friends

SFT teaches the model what to say. Preference methods teach it which answer is better. You give pairs: prompt, chosen response, rejected response. The model learns to prefer the chosen one.

DPO killed the RLHF pipeline for most people

Classic RLHF trains a separate reward model on human preferences, then optimizes the policy against it with PPO. Three models, finicky hyperparameters, frequent reward hacking. DPO (Direct Preference Optimization) skips all of it: a single loss function that pushes probability toward the chosen response and away from the rejected one. Same data format, one training run, stable optimization. For style polish and format compliance, DPO after SFT is the standard stack.

The family

  • DPO: the default. Key knob is beta (0.1), which controls how far the model may drift from the reference. Low beta = stronger preference following, higher risk of weirdness. Never reuse the policy as its own reference; the distribution collapses within an epoch or two.
  • ORPO: folds SFT and preference learning into one step, no separate reference model, at roughly a quarter of PPO's cost. Good when you want a single-stage run.
  • KTO: needs only good/bad labels, not pairs. Useful when your feedback is thumbs-up/thumbs-down rather than A-vs-B comparisons.
  • SimPO: drops the reference model and normalizes for length, which reduces the classic failure where the model learns that longer = better.

What DPO cannot do

DPO is a gentle nudge, not a capability builder. It re-ranks behavior the model already has; it does not teach new skills. If the base model cannot do the task at all, fix that with SFT or distillation first.

Length bias

Preference data is full of an accidental signal: chosen responses tend to be longer. Models learn to ramble. Counter it by keeping chosen and rejected responses similar in length, or use a length-normalized method like SimPO.

Where do preference pairs come from? Generate two responses per prompt from your SFT model, then rank them: by hand for a few hundred (gold standard), with a stronger model as judge (cheap, slightly noisy), or from implicit signals like user edits and thumbs votes. A few hundred to a few thousand pairs is enough for a noticeable polish pass.

6. RL post-training: GRPO and RLVR

The biggest post-training story of the last two years is reinforcement learning with verifiable rewards. Instead of a learned reward model guessing what humans like, you use rewards you can check: the math answer is right, the code passes its tests, the tool call returned the right schema. DeepSeek-R1-Zero showed what this unlocks: pure RL on a base model took AIME math scores from 15.6% to 77.9%, with no human reasoning examples at all.

Why it matters

RLVR (reinforcement learning with verifiable rewards) is what made small models reason. DeepSeek-R1 showed that RL on checkable problems produces long chains of self-correction: trying, failing, backtracking. SFT can imitate reasoning traces, but RL discovers them. If your custom task has right answers you can verify programmatically, RL post-training beats more SFT.

GRPO in one paragraph

GRPO (Group Relative Policy Optimization) samples a group of responses to the same prompt, scores each with your verifier, and pushes the policy toward the better ones relative to the group average. No value network, no reward model, which makes it far cheaper than PPO. Libraries (TRL, OpenRLHF, veRL) now ship GRPO trainers you can point at your own reward function.

When to use it

  • Your task has objectively checkable outputs: code, math, structured extraction, game play, tool use with defined success.
  • You have already done SFT and hit a plateau. RL squeezes out the next 5–15%.
  • You can afford the compute: RL needs many rollouts per prompt, typically 4–8× the cost of the SFT run that preceded it.

Reward hacking is the tax

The model will find the cheapest way to score, not the way you meant. A code reward based on passing tests produces code that games the tests. Keep verifiers strict, hold out test cases the model never sees during training, and spot-check rollouts by hand. If your reward can be gamed, assume it will be.

7. Knowledge distillation

Distillation trains a small student on the outputs of a large teacher. The teacher generates answers (or reasoning traces, or full token probability distributions), and the student learns to match them. You get much of the teacher's skill in a model you can actually serve.

Three flavors

  • Black-box distillation: prompt a strong model, collect its outputs, SFT your small model on them. This is how most open "reasoning" models were built: teacher writes long chains of thought, student learns the pattern. Check the teacher's terms of service first; some providers prohibit training competing models on their outputs.
  • White-box distillation: train the student to match the teacher's token probabilities (logits), not just its final text. Richer signal per example, needs access to the teacher's weights.
  • Data distillation: the teacher generates your training dataset (synthetic examples, preference pairs, edge cases), and you train normally. This is the most common real-world use.

The honest version of the deal

Distillation transfers behavior, not knowledge the teacher does not reliably have. A student distilled from a teacher's shaky domain knowledge inherits the shakiness. Distill reasoning patterns and style freely; distill facts cautiously.

8. Choosing a base model

The base model sets your ceiling. Pick wrong and no amount of tuning recovers it. Three criteria dominate: capability at your size, license, and context length.

The current open-weight landscape

The 2026 default base is Qwen3: strongest performance per parameter in the open field, Apache 2.0 license, long context, and a large fine-tuning community. Llama 3.3 remains the safe generalist if you want the Meta ecosystem and tooling.

FamilySizesLicenseNotes
Qwen30.5B–72B+Apache 2.0Default fine-tune base; coding, agents, multilingual
Llama 3.31B–70BLlama Community LicenseExcellent ecosystem; license has use restrictions, read before productizing
DeepSeek1.5B–70B+ (distills)Often MIT for distillsReasoning distills are the best cheap starting point for RL work
Gemma1B–27BGoogle terms; check the model cardStrong at small sizes; licensing is inconsistently reported, verify per checkpoint
Mistral7B–24B+Apache 2.0 / MNPLEfficient architectures; check license per release
gpt-oss (OpenAI)20B / 120BApache 2.0OpenAI's open entry; surprisingly strong
Phi3B–14BMITMicrosoft's small models; punchy for their size

How to choose

  • Size: 7–8B is the sweet spot for individuals. Big enough to follow instructions well, small enough to train with QLoRA on one GPU and serve cheaply. Go 1–3B if latency or edge deployment dominates. Go 32–70B only if you have the hardware budget and the task truly needs it.
  • License: Apache 2.0 or MIT means no strings attached commercially. Community licenses (Llama, Gemma) are fine for most businesses but read the terms; some restrict competitive use or trigger obligations at scale.
  • Instruct vs base: start from the instruct version. It already follows instructions; your fine-tune specializes it. Training from a raw base model means re-teaching instruction-following from scratch.
  • Context length: match it to your task. Long-document workflows need 32K+ context. Fine-tuning does not extend context well; pick a model that already has the window you need.

Recency

Model releases move monthly. Before committing, check the current leaderboards (LMArena, Open LLM Leaderboard) at your size class. A model from six months ago is often strictly worse than this month's release at the same size.

9. Building the dataset

Data quality decides the outcome more than any hyperparameter. The repeated finding across the literature: a few thousand excellent examples beat a hundred thousand mediocre ones. The LIMA result is the landmark: 1,000 meticulously curated pairs on a 65B model beat 52,000 lower-quality examples. The refinement since: style saturates at around 100 examples, while new capabilities keep improving with more data.

Format

Most trainers expect instruction pairs or conversations. The modern standard is the chat template: a list of messages with roles.

{"messages": [
  {"role": "user", "content": "Summarize this quarter's churn drivers."},
  {"role": "assistant", "content": "Three drivers explain 80% of churn..."}
]}

Match the base model's own chat template exactly. Every major library applies it automatically; the mistake is hand-rolling a format the model never saw.

How much data

  • Style and format: style itself saturates near ~100 examples; 200–2,000 examples give a robust, generalizing voice shift.
  • A focused task: 1,000–10,000 examples for reliable behavior. Below 200–500, models tend to memorize rather than generalize.
  • Broad domain adaptation: 10,000–100,000+ examples, with replay data mixed in.

Where the data comes from

  • Your own work. For style mimicry this is the gold: your emails, docs, posts, code. Nothing beats it.
  • Synthetic generation. Prompt a strong model to produce examples in your format, then curate hard. Self-Instruct and Evol-Instruct are the classic recipes: seed with a few examples, ask the teacher to invent variations, filter ruthlessly.
  • Distillation. Have the teacher solve your actual task inputs and keep the good outputs (section 7).
  • Human annotation. Slowest, best. Worth it for the final few hundred examples that define quality.

Cleaning checklist

  • Deduplicate aggressively. Near-duplicates teach memorization.
  • Cut anything the model already does well. Training on solved cases wastes capacity.
  • Fix the labels by hand on a sample. If 5% of your examples are wrong, the model learns 5% wrongness.
  • Keep prompt diversity high: vary phrasing, length, and edge cases so the model learns the task, not the template.
  • Check training data against your eval benchmarks for overlap (n-gram contamination checks). Training on the test is the oldest cheat in the book.
  • If one model generates your synthetic data, use a different model family to judge it. A model grading its own output collapses into self-preference.
  • Hold out 5–10% as an evaluation set before training. Never tune on it, never train on it.

10. Mimicking a writing style

This is the most personal use case: a model that writes like you. The good news is that style is one of the cheapest things to transfer. It saturates at around 100 examples, and 2026 work (InMyStyle) showed per-user LoRA adapters working on models as small as 0.5B, with judges rating the outputs over 20% less "AI-sounding." You do not need a giant model to clone a voice.

The pipeline

  1. Collect 200–500 samples of your writing in the target register to start; iterate toward 1–2,000 for polish. Emails if you want email voice; long-form if you want essay voice. One register per adapter; mixing registers blurs the result. Consistency of voice matters more than volume.
  2. Pair each sample with a plausible prompt. A raw essay is not training data. Write the instruction that would have produced it: "Write a memo to the team about the Q3 delay" → your actual memo. For emails, the prompt can be "Reply to this thread: ..." with the thread summarized. A strong model can draft these prompts for you; review them.
  3. Scrub. Remove anything you do not want the model to reproduce: names, private details, company internals, and your bad habits. The model will faithfully learn your typos and your tics. Curate like an editor.
  4. Train LoRA on the instruct checkpoint, not the base. The instruction-tuned conversational priors are worth preserving; full fine-tuning erodes them (researchers observed random topic-switching after full FT on personal corpora). Start at rank 8–16; rank 4 suffices for light tone shifts.
  5. Keep a short style system prompt too. The adapter and the prompt stack: the prompt gets you most of the way for free, the adapter makes it consistent across thousands of calls without eating context.
  6. Evaluate blind. Generate passages from prompts the model never saw. Mix them with your real writing. Ask someone who knows your voice to sort them, or run pairwise A/B judging ("which sounds more like X"). If they cannot tell reliably, you are done. Track a general benchmark alongside to catch capability regression.

Why LoRA preserves style better

Full fine-tuning rewrites the model's general language habits along with everything else, which often washes out the crisp edges of a personal voice. LoRA's small updates steer the existing capability instead of replacing it, so your phrasing habits come through while the model's grammar and reasoning stay intact.

What does not work

  • Too little data, too many epochs. Fifty examples trained for twenty epochs gives you a parrot that repeats your exact sentences, not your voice. More data, fewer epochs.
  • Mixed registers. Training on tweets, whitepapers, and love letters at once produces an average of all three, which is none of them.
  • Expecting facts to transfer. The model will write like you about things you never wrote about, drawing on its own knowledge. It will not know your private opinions on new topics. Style transfers; knowledge does not.

11. Compute and cost

The memory math

Three numbers dominate GPU memory during training: the weights, the gradients, and the optimizer state.

  • Weights: parameters × bytes per parameter. 7B in 16-bit = ~14 GB. In 4-bit = ~3.5 GB.
  • Gradients: only for trainable parameters. LoRA makes this tiny.
  • Optimizer (Adam): ~8 bytes per trainable parameter. Full SFT on 7B needs ~84 GB for optimizer state alone. LoRA on 7B with rank 16 trains ~40M parameters: ~0.3 GB.
  • Activations: grow with batch size and sequence length. Gradient checkpointing trades compute for memory here.
SetupTypical VRAMHardware
QLoRA, 7–8B model10–12 GBSingle RTX 4090 (24 GB)
QLoRA, 13B model~16 GBRTX 4090, or single 48 GB card
QLoRA, 32–33B model~24 GBRTX 5090 (32 GB) / A100 40 GB
QLoRA, 70B model~48 GBSingle A100/H100 80 GB
Full SFT, 7B model~112 GB2× 80 GB or multi-GPU with sharding
Full SFT, 70B model~1.1 TBMulti-node cluster; not an individual project

Renting GPUs

You do not need to own hardware. Spot and on-demand GPU rentals put a card in your hands in minutes. Indicative September 2026 rates on RunPod: RTX 4090 around $0.74/hr ($0.34 community/spot), A100 80GB around $1.59/hr, H100 around $3.49/hr; Vast.ai runs cheaper but host reliability varies. A typical 7B QLoRA run takes about 45 minutes on a 4090-class card: a few dollars per experiment. The expensive part is iterating ten times, so budget for iteration, not for one run. Prices move constantly; check the provider's live pricing before you commit.

Free options

Google Colab (free tier) gives you a modest GPU with time limits, enough for small LoRA experiments on 1–3B models or short 7B runs. Kaggle offers weekly free GPU hours. Both are fine for learning the workflow before you spend money.

Budget for the loop, not the run

Your first training run will not be your last. Plan for 5–15 experiments: data fixes, hyperparameter sweeps, and evaluation iterations. A realistic personal project budget is $50–$300 in compute for a 7B LoRA project, most of it spent on experiments 2 through 10.

12. The tooling stack

ToolWhat it isUse it when
Hugging Face TRLThe reference trainers: SFT, DPO, GRPO, KTOYou want standard, maintained implementations
UnslothHand-optimized kernels; 2–5× faster and ~70% less VRAM than the Hugging Face baselineSingle consumer GPU. The individual-developer default for LoRA/QLoRA
AxolotlYAML-config training for many methods, multi-GPU with FSDP2Reproducible configs, multi-GPU, long context
LLaMA-FactoryWeb UI + CLI for SFT/DPO/RLYou prefer a GUI over scripts
torchtune / LitGPTClean PyTorch-native training stacksYou want readable code you can modify
OpenRLHF / veRLFull RLHF/RL pipelinesYou are doing PPO or GRPO at scale
vLLM / SGLangHigh-throughput inference serversvLLM: production GPU serving with runtime LoRA-adapter swapping (one server, many adapters). SGLang: wins on agent/RAG workloads with repeated prefixes and heavy structured output
Ollama / llama.cppLocal quantized inferenceRunning the model on your own machine
Hugging Face HubModel and dataset hosting, private reposStoring adapters, sharing, versioning

No-code and API options

Several providers offer fine-tuning as a service: OpenAI (about $0.25 per million training tokens), Together AI, Fireworks AI (from about $0.50 per million training tokens), and others let you upload a dataset and get back a tuned model behind an API. You trade control and data privacy for zero infrastructure. A million-token SFT job costs pocket change in training tokens; per-token inference pricing dominates the long-term bill. Read the terms on data retention before uploading anything sensitive, and never train on a provider's outputs against their terms. For style mimicry on non-sensitive text, these are genuinely the fastest path.

# The shortest real training loop (Unsloth + TRL), conceptually:
# 1. pip install unsloth trl datasets
# 2. Load a 4-bit base, attach LoRA adapters (rank 16, all linear layers)
# 3. Train on your chat-formatted JSONL for 1-3 epochs, lr 2e-4
# 4. Save the adapter (a few hundred MB), merge or serve with vLLM

13. Evaluation

Training without evaluation is guessing. Decide before training what "better" means and how you will measure it.

Three layers

  • Task evals you build. 100–500 held-out examples with clear expected outputs. Score them automatically where possible (exact match, regex, unit tests) and by rubric where not. This is the eval that matters most.
  • General benchmarks. MMLU, IFEval, MT-Bench, Arena-Hard. Note that MMLU is saturated above 88% by current models, so treat it as a regression canary (did I break anything?) rather than a differentiator. AlpacaEval's length-controlled win rate corrects for verbosity bias.
  • Human or LLM-as-judge. For style and quality, have a strong model compare outputs pairwise against the base model, blind. Judges reach roughly 80% agreement with humans at about a cent per judgment: noisy but cheap, and reasonable on style tasks.

Checkpoint discipline

  • Save checkpoints every 10–25% of training. The final checkpoint is often not the best.
  • Pick the checkpoint with the best held-out score, not the lowest training loss. Training loss always goes down. That tells you nothing.
  • Watch for the overfitting signature: training loss falling while eval scores stall or drop. Stop early when you see it.

The style eval that works

For voice mimicry, automated metrics fail. The test is blind human sorting: can a reader who knows your writing distinguish your text from the model's? Run it on fresh prompts, not training examples. Fifty passages and one honest reader beats any BLEU score.

14. Serving the finished model

A trained adapter is not a product until it runs somewhere cheap and fast. The standard move is quantization: shrinking the weights so inference needs less memory and runs faster, with minimal quality loss.

FormatUse case
GGUFLocal and edge inference via llama.cpp and Ollama. The universal CPU/laptop format.
AWQ / GPTQGPU inference with 4-bit weights. Pairs with vLLM for serving. AWQ is slightly better at 4-bit; both need calibration data.
bitsandbytes 8-bitQuick local experiments without a conversion step.
  • Merge first, quantize second. Merge LoRA adapters into the base weights before quantizing. Quantizing first and merging after compounds the errors.
  • Merge or not? Merged adapters become one file with zero inference overhead. Kept separate, adapters hot-swap at runtime in vLLM: one server can serve many voices or tasks. Merge for a single-purpose deployment; keep adapters separate for multi-tenant setups.
  • vLLM is the default server for GPU deployment: continuous batching, OpenAI-compatible API, LoRA adapter hot-swapping.
  • Ollama is the default for "it just runs on my machine": one command, GGUF under the hood.
  • Cost reality: a quantized 8B model serves comfortably on a single 24 GB GPU, or on CPU for low traffic. A fine-tuned API from a provider costs per token with no ops burden. Do the per-token math at your expected volume before choosing.

15. Pitfalls and honest caveats

  • Decision order: prompt, then RAG, then fine-tune. Each step costs roughly 10× the engineering of the last. Prompt engineering is free and instant, and it is the mandatory baseline: you cannot claim fine-tuning helped if you never tuned the prompt.
  • RAG vs fine-tuning, again. If the problem is "the model doesn't know my documents," retrieval wins. Fine-tuning teaches behavior. Mixing them up wastes weeks.
  • Fine-tuning does not fix reasoning. A model that cannot do the task at all with a good prompt will not learn it from 500 examples. It will memorize the examples instead. DPO only re-ranks existing behavior; it builds no new capability.
  • Fine-tuning pays off economically at volume. Baked-in behavior replaces long few-shot prompts, cutting per-request token cost 30–70%. The rough break-even is above ten thousand requests.
  • Data licensing is real. Training on text you do not own creates legal exposure, especially if you serve the result commercially. Provider terms from OpenAI, Anthropic, and Google prohibit training competing models on their API outputs; distill from open-weight teachers instead. Your own writing is clean. Everything else needs a license check.
  • Privacy runs one way. Anything in the training data can leak out of the model. Never train on secrets, credentials, or private data about other people. Scrub first, keep style corpora in private repos.
  • Small data lies. Great eval scores on 50 examples mean nothing. Overfit models ace tiny evals and fail in the wild. Make the eval bigger than feels necessary.
  • Base model updates obsolete adapters. A LoRA trained for one base does not transfer to the next release. Budget for retraining when you switch bases.
  • Safety tuning is thin on open models. Your fine-tune can weaken refusal behavior. If you serve the model to others, test adversarial prompts and consider keeping the provider's safety layers.

16. Three playbooks

Playbook A · Write like me (style mimicry)

Goal: a model that drafts in your voice.

1. Collect 200–500 samples of your writing in one register to start; iterate toward 1–2,000 for polish. 2. Write a plausible prompt for each (a strong model can draft these; you review). 3. Scrub names, secrets, and tics you dislike. 4. Format as chat JSONL, hold out 10%. 5. QLoRA on a 7–8B instruct model: rank 8–16, lr 2e-4, 2–3 epochs. 6. Generate on fresh prompts; blind-sort against your real writing. 7. If it is close but loose, add 200–500 DPO pairs (your edit preferred over the model's draft). 8. Quantize to GGUF, run in Ollama.

Budget: one evening of curation, $10–$50 in GPU rental.

Playbook B · A custom workflow (task specialist)

Goal: a model that executes your multi-step workflow reliably, e.g. triaging tickets into your taxonomy with your fields.

1. Define the task as input→output with a strict schema. 2. Build 2,000–10,000 examples: distill from a strong model on your real inputs, then hand-fix the failures. 3. Mix in 10% general instruction data. 4. QLoRA rank 32–64, lr 1e-4, 1–2 epochs. 5. Eval on 300+ held-out cases with programmatic checks (schema valid, fields correct). 6. Mine the failures, add targeted examples, retrain. 7. If outputs are checkable (tests, schemas), add a GRPO pass with your verifier as the reward. 8. Serve merged + quantized behind vLLM.

Budget: a week of iteration, $100–$300 in compute.

Playbook C · Reasoning on your domain

Goal: a small model that reasons carefully about your kind of problem.

1. Start from a reasoning-distilled base (DeepSeek distills or similar). 2. SFT on teacher traces for your problem type so the format is right. 3. Build a verifier: unit tests, a checker script, a rubric with teeth. 4. GRPO with 4–8 rollouts per prompt, a few thousand prompts. 5. Watch for reward hacking from day one; hold out verification cases. 6. Keep the best checkpoint by held-out verifier score.

Budget: the verifier is the real cost (days of engineering); compute $200–$500.

Sources and further reading

Core methods trace to these papers: LoRA (Hu et al., 2021), QLoRA (Dettmers et al., 2023), DPO (Rafailov et al., 2023), RLHF (Ouyang et al., 2022), LIMA (Zhou et al., 2023), DeepSeek-R1 (DeepSeek-AI, 2025), IMPersona (arXiv 2504.04332), InMyStyle (arXiv 2607.29238). Current practice (hyperparameters, VRAM figures, rental pricing, tool comparisons) was gathered from web research in October 2026 at index level, not live-verified. Model releases and prices move fast: re-check leaderboards (LMArena, Open LLM Leaderboard) before choosing a base, and re-check provider price cards before budgeting compute.

  • Hugging Face TRL documentation — SFTTrainer, DPOTrainer, GRPOTrainer references.
  • Unsloth documentation — memory-efficient LoRA guides and VRAM tables.
  • Axolotl and LLaMA-Factory documentation — config-driven training recipes.
  • vLLM and SGLang documentation — serving, quantization support, adapter swapping.
  • mergekit documentation — SLERP, TIES, DARE model merging.

The full research brief behind this report (method-by-method numbers, per-topic source notes) is kept with the author's notes.

Research content is analysis, not investment advice. ASKA does not provide investment advice through this site.