Anasayfa / Software / Fine‑Tune a Pre‑trained Model with Hugging Face: Step‑by‑Step Guide

Fine‑Tune a Pre‑trained Model with Hugging Face: Step‑by‑Step Guide

machine learning fine-tuning

Fine‑tuning a pre‑trained model is the fastest way to get state‑of‑the‑art performance on a niche dataset without starting from scratch. Hugging Face’s 🤗 Transformers library makes this process surprisingly approachable, even if you’re not a PhD‑level researcher. In this guide we’ll walk through everything you need—from setting up the environment to deploying the final model—while highlighting common mistakes and offering pro tips along the way.

What You'll Need

  • Python 3.8 or newer
  • GPU‑enabled machine (CUDA 11.x recommended) or a cloud instance (AWS, GCP, Azure)
  • Basic familiarity with Git and virtual environments
  • A Hugging Face account (free) for model hub access and optional token
  • A small, labeled dataset in CSV or JSONL format

Step 1: Set Up Your Development Environment

First, create an isolated Python environment so your dependencies don’t clash with other projects. We’ll use venv, but conda works just as well.

python -m venv hf‑env
source hf‑env/bin/activate  # Linux/macOS
hf‑envScriptsactivate     # Windows

Next, install the core libraries. The transformers package brings the models, datasets handles data loading, and accelerate optimizes multi‑GPU training.

pip install --upgrade pip
pip install torch torchvision torchaudio   # CPU‑only if you lack a GPU
pip install transformers datasets accelerate sentencepiece tqdm

If you have a CUDA‑capable GPU, verify PyTorch sees it:

python -c "import torch; print('CUDA available:', torch.cuda.is_available())"

Seeing True means you’re ready for accelerated training.

Step 2: Choose and Download a Pre‑trained Model

Hugging Face hosts thousands of models. For this tutorial we’ll fine‑tune distilbert-base-uncased, a lightweight BERT variant ideal for text classification. You can swap it for any model that matches your task (e.g., roberta-base for sentiment, facebook/bart-large for summarisation).

from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "distilbert-base-uncased"
model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=2)
tokenizer = AutoTokenizer.from_pretrained(model_name)

Setting num_labels tells the head how many output classes you have (binary in this case).

Step 3: Prepare Your Dataset

Let’s assume you have a CSV with two columns: text and label. The datasets library can read it directly and split it into train/validation sets.

from datasets import load_dataset
raw_ds = load_dataset("csv", data_files={"train": "train.csv", "validation": "val.csv"})

def tokenize(batch):
    return tokenizer(batch["text"], padding="max_length", truncation=True, max_length=128)

encoded_ds = raw_ds.map(tokenize, batched=True)
encoded_ds.set_format(type="torch", columns=["input_ids", "attention_mask", "label"])

Adjust max_length based on your domain; longer sequences need more GPU memory.

Step 4: Configure the Training Arguments

The Trainer API abstracts most of the boilerplate. Define a TrainingArguments object that controls learning rate, batch size, logging, and checkpointing.

from transformers import TrainingArguments
training_args = TrainingArguments(
    output_dir="./results",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=32,
    learning_rate=5e-5,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    logging_dir="./logs",
    logging_steps=50,
    load_best_model_at_end=True,
    metric_for_best_model="accuracy",
    push_to_hub=False,
    report_to="none",
)

Key points:

  • Batch size: Fit as many samples as your GPU memory allows; you can enable gradient accumulation if you need larger effective batch sizes.
  • Learning rate: 5e‑5 works for most BERT‑style models; lower it for very small datasets to avoid over‑fitting.

Step 5: Define a Metric and Initialise the Trainer

Evaluation metrics guide early stopping and model selection. For classification, accuracy and F1 are common.

import numpy as np
from datasets import load_metric
accuracy = load_metric("accuracy")

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    preds = np.argmax(logits, axis=-1)
    return accuracy.compute(predictions=preds, references=labels)

from transformers import Trainer
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=encoded_ds["train"],
    eval_dataset=encoded_ds["validation"],
    compute_metrics=compute_metrics,
)

Now you’re ready to launch training.

Step 6: Train, Evaluate, and Save the Fine‑tuned Model

Run the training loop with a single command. The accelerate library will automatically detect multiple GPUs if present.

trainer.train()
trainer.evaluate()
trainer.save_model("./fine‑tuned‑distilbert")
tokenizer.save_pretrained("./fine‑tuned‑distilbert")

After training, you’ll have a directory containing pytorch_model.bin, config.json, and the tokenizer files. You can push this to the Hugging Face Hub for easy sharing:

trainer.push_to_hub(commit_message="Add fine‑tuned model for XYZ task")

If you prefer a local inference script, here’s a minimal example:

from transformers import pipeline
classifier = pipeline("text-classification", model="./fine‑tuned‑distilbert")
print(classifier("Your sample sentence goes here."))

Common Mistakes to Avoid

Even experienced practitioners stumble over a few recurring pitfalls:

  • Forgetting to set num_labels: The model head will default to the pre‑trained number of classes, leading to shape mismatches.
  • Using the wrong tokenisation length: Truncating too aggressively discards crucial information; padding too much wastes memory.
  • Hard‑coding the learning rate: A rate that’s too high can cause loss spikes; a rate that’s too low makes training crawl.
  • Neglecting a validation split: Without an unbiased set you can’t tell if you’re over‑fitting.
  • Skipping torch.cuda.empty_cache() after large experiments: GPU memory can become fragmented, causing out‑of‑memory errors later.

Tips and Tricks

Boost your fine‑tuning workflow with these shortcuts:

  • Gradient Accumulation: If your GPU can’t hold a batch of 32, set per_device_train_batch_size=8 and gradient_accumulation_steps=4 to simulate a larger batch.
  • Mixed Precision (FP16): Add fp16=True to TrainingArguments for up to 2× speed‑up on recent GPUs.
  • Early Stopping Callback: Import EarlyStoppingCallback and pass callbacks=[EarlyStoppingCallback(early_stopping_patience=2)] to stop training when validation loss stops improving.
  • Data Augmentation: For text, consider back‑translation or synonym replacement to enlarge tiny datasets.
  • Layer Freezing: Freeze the lower transformer layers (model.base_model.requires_grad_(False)) to reduce training time when your dataset is small.

Frequently Asked Questions

Do I need a GPU?

Training on CPU is possible but extremely slow for anything beyond a few hundred examples. A single NVIDIA RTX 3060 or a free Google Colab GPU will cut training time from hours to minutes.

Can I fine‑tune a model for a multi‑label task?

Yes. Set problem_type="multi_label_classification" in the model config and use BCEWithLogitsLoss. Adjust the metric function to compute f1_score with average="macro".

How do I handle imbalanced classes?

Two common strategies are class weighting (pass weight to the loss function) and oversampling the minority class via datasets ClassBalancedDataset or custom torch.utils.data.WeightedRandomSampler.

Conclusion

Fine‑tuning with Hugging Face blends the power of large pre‑trained models with the flexibility of custom datasets. By following the steps above—setting up a clean environment, picking the right model, preparing data, configuring training, and watching out for typical mistakes—you can produce a production‑ready model in a matter of hours. Remember to experiment with learning rates, batch sizes, and optional tricks like mixed precision; the right combination often yields the best results. Happy training, and may your models always converge!

Photo by Markus Winkler on Unsplash

Etiketlendi: