Distilling omni-moderation into ShieldGemma 2B on one RTX 4090
LoRA on a 24 GB card, an uneven mix of training data, and a moderation model that flagged a person's name as violent
Two weeks before the July 2025 study group, I said I'd fine-tune an 8B or 27B moderation model with plain LoRA, without quantizing the model weights. What I actually trained was ShieldGemma 2B, on one RTX 4090.
One of the most memorable false positives was just two words: Donald Trump. The model flagged the name alone as violent content. Short inputs turned out to be a weak spot.
Here's the story
For the classroom AI platform I worked on at ViewSonic, we prototyped a chat safeguard with two checks:
- OpenAI's
text-moderationendpoint, for its fixed categories. - An LLM prompted with few-shot examples, following the approach in Anthropic's content moderation guide. We wanted it to handle sarcasm, indirect insults, and messages in different languages.
The prototype checked students' messages and warned the teacher. It didn't ship.
I wanted more control over the moderation behavior than the hosted models gave me. In our Azure deployment, the managed serving layer had a safeguard we couldn't disable. When it treated a moderation request as harmful, the call failed instead of returning a classification. I remember those failures returning HTTP 404; that's a recollection of our setup, not a claim about the status code every Azure deployment returns.
Wrapping the request in an LLM chain didn't solve that problem. The content we needed the model to inspect could trigger the surrounding safeguard before we got a usable verdict. That made a local model worth trying.
Here's what I wanted to achieve:
- A small model that agrees with OpenAI's
omni-moderationon its 13 categories closely enough to stand in for it. - It runs locally, for demos and tests, with no cloud API in the loop.
- Something worth presenting at the study group. That one was on the definition-of-done slide.
Why ShieldGemma
Mistral, Llama and plain Gemma were on the list. ShieldGemma matched the prompt shape I wanted: one policy, one piece of content, answer Yes or No. In the demo, I normalize the logits for those two tokens and report the score for Yes.
That shape also means one text becomes 13 training rows, one per category. The model never has to hold the whole taxonomy in a single answer.
The data
I looked at internal chat logs, synthetic examples, and public moderation datasets:
- Our own chat logs. A colleague retrieved archived logs during an infrastructure move. After deduplication, only about 10% of the records remained. The product's chat data contained a lot of repetition, so its raw size overstated how many distinct examples it could give us. The saved training notebook doesn't list these logs among its input files.
- Synthetic data. Some of the training examples were synthetic, with 50 examples per category, each roughly 128โ1024 tokens long. Generating them had its own limits: Gemini would produce roast-style insults, but resisted requests for more harmful material.
- Hugging Face. OpenAI's moderation eval set,
koala_ai_text_moderation_multilingual,ifmain_text_moderation_01,ifmain_text_moderation_02_multilingualandontocord_oig_moderationmade the shortlist. The saved notebook loadsopen-ai-moderation.jsonlandifmain-text-moderation-01.jsonl.
The mix was uneven in both the amount of data and text length. That mattered later: individual names, short phrases, and short sentences were prone to false positives. The input filenames alone don't describe the full provenance or distribution of those examples.
The labeling step used omni-moderation's per-category verdicts as targets. My slide for this step was the Powerpuff Girls intro: system prompt, one category definition, the content, and OpenAI's answer as Chemical X. The intended training format was one row per text per category:
You are a policy expert trying to help determine whether a user
prompt violates a specific safety policy category.
<start_of_turn>
Human Question: {text}
<end_of_turn>
Safety Policy Category: illicit
Description:
<Description>
### Definition
Content that gives advice or instruction on how to commit illicit acts.
A phrase like "how to shoplift" would fit this category.
[...]
</Description>
Does the human question violate the "illicit" policy? Your answer must start [...]
No<eos>
The notebook creates the label in one line:
example["text"] = "Yes" if example["omni_flagged"] else "No"
The dataset recorded in the notebook has 71,678 texts ร 13 categories = 931,814 rows. Expanding each text into 13 rows doesn't add new text examples or fix the imbalance in their lengths.
Training
I used transformers, PEFT and TRL's SFTTrainer. ShieldGemma 2B fit the 24 GB RTX 4090 I had available, so that's the size I worked with. The actual run lasted two epochs.
The saved notebook records the settings below. Its num_train_epochs=60 differs from the two epochs I actually ran, so this snapshot isn't an exact record of the final run. I no longer remember the elapsed training time, the TRL version, or the checkpoint used by the demo.
LORA_CONFIG = LoraConfig(
r=8,
lora_alpha=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
)
TRAINING_ARGS = TrainingArguments(
per_device_train_batch_size=400,
learning_rate=2e-4,
warmup_steps=100,
num_train_epochs=60, # saved setting; actual run: 2 epochs
fp16=True,
optim="paged_adamw_8bit",
# ...
)
These are recorded settings, not a complete recipe for reproducing the run. There is a detail in the saved code that needs checking: text contains only Yes or No, while the complete policy, input, and answer are stored in formatted_input. The trainer doesn't explicitly select that column, and formatting_func is commented out.
In TRL v0.19.0's documentation, the default text column is text. If the actual run used that behavior, the complete classification input may not have reached the trainer. The historical environment and run need checking before attributing any result to that possibility.
The notebook also logs a Transformers recommendation to use attn_implementation="eager" for Gemma 2 training.
The demo page was made in Firebase Studio. Two of four generations produced a working page; the slide next to it compared vibe coding to a slot machine.
What the evaluation showed
I did evaluate the model beyond the example in the demo notebook. Those evaluation results aren't included in that notebook, but the recurring problem was short text: names, individual terms, and short sentences could be classified as harmful even when the input didn't contain harmful content.
The Donald Trump example stood out because the input was only a person's name. It received a violence-related score of 0.68 and was flagged as harmful. A name can appear in texts about violence without being a violent statement on its own.
My interpretation was that the uneven data distribution, including the mix of synthetic examples and text lengths, contributed to those errors. That's an explanation of the pattern I saw, rather than a measured attribution of each error to a particular source of data.
There is currently no base-model-versus-LoRA comparison to report. These observations describe the fine-tuned model's errors; they don't show whether fine-tuning improved or worsened performance relative to the base model.
A saved demo output
The demo notebook runs on a Colab T4 and can switch between the base model and the LoRA adapter. Its saved output shows the adapter scoring a threat across all 13 categories:
User Input: i will find you i will kill you
harassment : 0.8818 (High Risk)
harassment/threatening : 0.8232 (High Risk)
hate : 0.5469 (Medium Risk)
hate/threatening : 0.8667 (High Risk)
illicit : 0.6646 (Medium Risk)
illicit/violent : 0.4976 (Medium Risk)
self-harm : 0.7622 (High Risk)
self-harm/intent : 0.4387 (Medium Risk)
self-harm/instructions : 0.6387 (Medium Risk)
sexual : 0.5254 (Medium Risk)
sexual/minors : 0.6133 (Medium Risk)
violence : 0.2539 (Low Risk)
violence/graphic : 0.4446 (Medium Risk)
The harassment scores make sense for a threat. Other scores are harder to justify: self-harm is 0.76 for a threat aimed at someone else, and sexual/minors is 0.61 for a sentence with nothing sexual in it. Violence is the lowest category at 0.25. Twelve of the 13 fall into the demo's medium or high bands.
Those bands are the demo's display thresholds, not evidence that the scores are calibrated probabilities of harm. This output illustrates a problem with category scores; it doesn't establish overall performance or explain why the errors occurred. The notebook doesn't include the base model's scores for the same input.
The name-only false positive is the example I would keep next to these scores. A moderation model needs to handle the short, ordinary things people actually type, too.