Projects

MiniGabriel

I fine-tuned an open-source language model on my own Telegram messages to see if it could learn to text like me. It picked up most of my habits, from the lowercase and the Singlish to replying in short bursts.

Style distance to my real replies, lower is closer
Base model, no fine-tuning0.3639
Fine-tuned0.0637
Me vs. me0.0237
Style distance to my real replies, on 1,434 replies from four chats the model never saw. Lower is closer. "Me vs. me" is the floor.

Why I built it

Most language models write like a polite assistant, with capital letters, full stops and one tidy paragraph. I write almost the opposite. In the replies I held out for testing, 94.8% start with a lowercase letter, only 9.3% end with punctuation, and about 60% are split across two or three short messages. There's also a fair bit of Singlish.

I wanted to know whether one person's chats were enough data to move a model toward that style. By style I mean wording, capitalisation, punctuation, slang, emoji, how long replies are, and splitting one reply into several messages. I wasn't trying to make it know things about me, so there's no retrieval or memory in it, only fine-tuning on how I write. It's also something I can check myself, since I know what my own replies look like.

What I built

A pipeline of eight scripts, each producing something I could open and check before moving on.

StageWhat came out
Extract432 chats, 113,053 messages from Telegram
Select58 chats: private chats and groups of 20 or fewer where I'd written at least 100 messages
Build examples17,001 examples: 54 chats for training, 4 held out
Format14,939 training rows in ChatML
TrainA LoRA adapter, after 91 minutes on one A100 on the NUS School of Computing cluster
EvaluateStyle distance on 1,434 held-out replies
ExportA merged, 4-bit GGUF file of 4.79 GB
DeployOllama and a Telegram bot, running on my laptop's CPU

The code that makes decisions (which chats count, where a turn or a conversation ends, how examples are cut, the style metric) is all in plain functions with no Telegram or GPU code in them. That let me test it with 190 tests on made-up data, which run in about two seconds, instead of debugging on my real messages.

The bot splits the model's reply on newlines and sends each line as its own message with a short pause, so it texts in bursts the way I do.

It's written in Python, with PyTorch, Unsloth and Hugging Face's TRL library for training, Slurm for the cluster jobs, and llama.cpp and Ollama for running it.

Decisions

Starting from the base model

I used Qwen3-8B-Base instead of the instruct version. Instruct models are trained toward the polite, capitalised style that my metric scores furthest from mine, so fine-tuning one would mostly be undoing that training.

Dropping Qwen3.5

I first picked the newer Qwen3.5-9B. Reading its config showed it was multimodal (image and text), and the loader I was using is for text-only models, so I switched to Qwen3-8B.

Plain LoRA in bf16

The 8B model needs about 20 GB of GPU memory with activations, which fits on a 40 GB A100. QLoRA is for when a model doesn't fit, and it costs some precision and 20 to 40% of the speed, so I didn't use it. At rank 16 that meant training 43.6 million parameters, 0.53% of the model.

Holding out whole chats

If I'd held out random messages, the model would be tested on conversations it had mostly already read. I set aside four entire chats and never trained on them.

Only training on my replies

Without masking, the model learns to predict everyone's messages, including my friends'. I turned on response-only masking and then decoded the training labels to check it was actually working before starting the run.

The smallest GPU that fits

On the cluster, the wait in the queue was longer than the job. Bigger GPUs use up more of my fair-share allocation (an H200 counts about five times an A100), which would have pushed every later job further back.

Running it on my laptop

For the bot I merged the adapter into the model, quantised it to 4 bits and ran it with Ollama. It went from 16 GB to 4.8 GB, and a reply takes a couple of seconds on CPU. vLLM is built for lots of users at once, and this only ever answers one chat at a time.

How I measured it

Training loss doesn't tell you whether a model sounds like someone. A model that memorised every message would have great loss and be useless.

So I wrote a style metric. It compares ten measurements between the model's replies and mine (length, number of messages, lowercase starts, ending punctuation, emoji, Singlish, questions, and bare replies like "ok") and averages the differences into one number, which I call style distance.

Two halves of my own replies don't score zero against each other. They score 0.0237, because every measurement comes from a finite sample. That's the floor, and it's what the model should be compared to.

Before training anything, I scored the base model with no fine-tuning to see if prompting alone would be enough. It scored 0.3639, about 15 times the floor. It was answering messages like "eh you free later" with around 30 messages and 410 characters.

What went wrong

The cluster's home directory only holds about 120 MB

pip install unsloth failed halfway through one package. quota and df both looked fine because they report different filesystems, and du -sh ~ was the only command that showed the limit. The job scripts now work around it by using the compute node's local disk.

Training took 91 minutes instead of about 25

Padding-free batching was switched off. With an average sequence of 104 tokens, most of the compute went on padding. I've fixed the setting but haven't re-run training yet.

It sends too many one-word replies

13.3% of its replies are bare acknowledgements, compared with 4.2% of mine, and that's the biggest part of the remaining distance. I'd expected something like this: 11% of my own turns are only acknowledgements, so the dataset already kept just a third of them. It wasn't enough.

I tried changing the sampling temperature first, since that doesn't need retraining. Between 0.4 and 0.8 the score only moved from 0.0570 to 0.0637, which is within noise, and the acknowledgement rate didn't change at all, so the fix has to be in the training data. The bot uses 0.5, which I picked by reading the replies.

The metric only measures style

It checks habits like length, lowercase and punctuation, so it can't tell whether a reply fits the conversation. I judge that part by reading the replies, and there's no number for it yet.

One number matched for the wrong reason

The base model's acknowledgement rate (0.041) was almost exactly mine (0.042), but only because it never wrote a short reply at all.

Names inside messages weren't removed

I replaced the names of who sent each message, but not names typed inside the messages. The model can repeat them, so I'm not publishing the weights and the bot isn't public.

Results

Fine-tuning cut the distance by 83%, from about 15 times the floor to 2.7 times, as the chart at the top shows. Most of the habits that make my texts recognisable came through:

HabitMeModel
Starts lowercase94.8%93.3%
Split across several messages59.6%54.0%
Messages per reply1.981.90
Singlish13.1%12.0%
Ends with punctuation9.3%8.1%
Average length (characters)43.932.9
Bare acknowledgements4.2%13.3%
A Telegram chat with the MiniGabriel bot. I ask if it has started the CS2103 assignment. It replies in short lowercase bursts: i finished le, its not that deep, just lots of content, and later, i spent like a whole day on it
A chat with the bot. My messages are the blue ones. It replies in short lowercase bursts the way I do.

There's no blind test yet, so all I can say is that the style numbers are close. I can't say a friend wouldn't be able to tell.

What I learned

Training was the smallest part. It was one 91-minute job, and the decisions that mattered came before it: what counts as a conversation, how to split the data, what to mask, and how to measure the result.

Next time I'd build the evaluation first, including the floor and the no-training baseline. Without those two, the 83% wouldn't mean much.

Checking things by hand paid off. Decoding the labels took a few minutes and would have caught a masking mistake that could have wasted the whole run.

I should have removed names inside messages in the data pipeline. Once they're in the weights, the only option left is not publishing them.

What's next

  • Keep fewer acknowledgements in the training data and retrain.
  • Try different LoRA ranks and epoch counts, now that there's a metric to compare them with.
  • A blind test: mix real and generated replies and see if friends can pick out mine.

I wrote about the project less formally in a note. The repository has the full experiment log and evaluation design.