<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://uzunenes.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://uzunenes.com/" rel="alternate" type="text/html" /><updated>2026-10-07T13:22:36+00:00</updated><id>https://uzunenes.com/feed.xml</id><title type="html">Enes Uzun</title><subtitle>Senior Software Architect at Ford Otosan. AI infrastructure, LLM inference and GPU platforms.</subtitle><author><name>Enes Uzun</name></author><entry><title type="html">TRL v1.14.2 fixes prompt-completion training for Gemma 4 12B/26B/31B</title><link href="https://uzunenes.com/2026/10/07/trl-v1-14-2-gemma-4-prompt-completion-fix.html" rel="alternate" type="text/html" title="TRL v1.14.2 fixes prompt-completion training for Gemma 4 12B/26B/31B" /><published>2026-10-07T00:00:00+00:00</published><updated>2026-10-07T00:00:00+00:00</updated><id>https://uzunenes.com/2026/10/07/trl-v1-14-2-gemma-4-prompt-completion-fix</id><content type="html" xml:base="https://uzunenes.com/2026/10/07/trl-v1-14-2-gemma-4-prompt-completion-fix.html"><![CDATA[<p>On 2026-10-06 Hugging Face released TRL v1.14.2. It fixes a bug I reported a week earlier: with the chat template of Gemma 4 12B, 26B-A4B and 31B, prompt-completion training put the loss on the wrong tokens and dropped short answers without an error. It affects SFT, DPO and KTO on TRL 1.14.1 and earlier when thinking is off, the default. The fix is by Quentin Gallouédec, reviewed by Albert Villanova. If you trained one of these models this way and saw the warning “Mismatch between tokenized prompt and the start of tokenized prompt+completion”, upgrade and train again.</p>

<h2 id="issue">Issue</h2>

<p>TRL tokenizes a prompt-completion row twice: the prompt alone with <code class="language-plaintext highlighter-rouge">add_generation_prompt=True</code>, and prompt plus completion as a full conversation. Before the fix it assumed the first is a prefix of the second and set <code class="language-plaintext highlighter-rouge">completion_mask = [0] * len(prompt_ids) + [1] * rest</code>.</p>

<p>Gemma 4 12B-it, 26B-A4B-it and 31B-it share one chat template. When thinking is off, the generation prompt ends with an empty thought block, four tokens: <code class="language-plaintext highlighter-rouge">&lt;|channel&gt;thought\n&lt;channel|&gt;</code>. The full conversation render does not contain it. So the prompt is not a prefix, and the mask started four tokens too late. With <code class="language-plaintext highlighter-rouge">enable_thinking=True</code> the prompt is a prefix of the full render, 22 of 22 tokens, and nothing is lost. Gemma 4 E2B and E4B use a different template and are not affected.</p>

<p>This concerns prompt-completion data only; a dataset with a single <code class="language-plaintext highlighter-rouge">messages</code> column takes a different code path.</p>

<h2 id="impact">Impact</h2>

<p>I ran TRL 1.14.1 (transformers 5.19.0) with the tiny Gemma 4 test tokenizer and the official 12B template on <code class="language-plaintext highlighter-rouge">trl-internal-testing/zen</code>, 17 rows. SFT kept 9 of 17. Completions of four tokens or fewer were fully masked and dropped; the rest trained on the tail of the answer. Three rows before and after (<code class="language-plaintext highlighter-rouge">raw</code> is the dataset row, <code class="language-plaintext highlighter-rouge">proc</code> its index after preprocessing):</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code># TRL 1.14.1
raw[ 0] 'What is better than ugly?': DROPPED (no processed row contains this prompt)
raw[ 7]-&gt;proc[ 1] loss tokens=12: " aren't special enough to break the rules.&lt;turn|&gt;\n"
raw[11]-&gt;proc[ 4] loss tokens= 5: ' to guess.&lt;turn|&gt;\n'

# TRL v1.14.2
raw[ 0]-&gt;proc[ 0] loss tokens= 4: 'Beautiful.&lt;turn|&gt;\n'
raw[ 7]-&gt;proc[ 7] loss tokens=16: "No, special cases aren't special enough to break the rules.&lt;turn|&gt;\n"
raw[11]-&gt;proc[11] loss tokens= 9: 'Refuse the temptation to guess.&lt;turn|&gt;\n'
</code></pre></div></div>

<p>The training sequence itself, <code class="language-plaintext highlighter-rouge">input_ids</code>, was already the full render; only the mask was wrong. DPO failed differently: the chosen and rejected ids of the short rows were empty or a newline, so the pair compared nothing.</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>raw[ 0]-&gt;proc[ 0] chosen_ids='' rejected_ids='\n'
raw[ 1]-&gt;proc[ 1] chosen_ids='' rejected_ids=''
raw[12]-&gt;proc[12] chosen_ids=' only one.&lt;turn|&gt;\n' rejected_ids='.&lt;turn|&gt;\n'
</code></pre></div></div>

<p>KTO lost the first tokens of every completion the same way. No error was raised. The only signal was a generic warning, once per row, 17 times on this dataset: “Mismatch between tokenized prompt and the start of tokenized prompt+completion. This may be due to unexpected tokenizer behavior, whitespace issues, or special token handling. …”</p>

<p>Olmo 3 has the same problem with its <code class="language-plaintext highlighter-rouge">&lt;think&gt;</code> prefill, three tokens: 10 of 17 rows kept, <code class="language-plaintext highlighter-rouge">'.&lt;|endoftext|&gt;'</code> trained where the answer was <code class="language-plaintext highlighter-rouge">'Practicality.&lt;|endoftext|&gt;'</code>. DeepSeek-R1-Distill, Nemotron 3 and Qwen3.5-Think diverge too; Qwen3, Qwen2.5 and Gemma 3 do not in the base case.</p>

<h2 id="fix">Fix</h2>

<p>I opened issue #7449 on 2026-09-29 and PR #7458 on 2026-09-30. My patch kept the inference prompt ids and appended the rest. The same day Quentin opened #7463 with a simpler approach. It keeps the full chat-template render and moves the prompt/completion boundary to where the tokenized prompt and the tokenized prompt+completion diverge; the difference is which side each PR keeps exact. Quentin, in a comment on the PR: “I prefer keeping the sequence a real template render, with less code. A context that differs from inference is a template bug anyway, so there is no ‘right’ way to handle it.” #7458 was closed in favor of #7463 on 2026-10-01. #7463 merged on 2026-10-06 with a new <code class="language-plaintext highlighter-rouge">common_prefix_length</code> in <code class="language-plaintext highlighter-rouge">trl/data_utils.py</code>, covering SFT, DPO, KTO and the experimental TPO trainer. The PR description credits me as co-author.</p>

<p>After the fix all 17 rows are kept and the loss covers the whole assistant turn. DPO and KTO completions are whole turns too. The old warning is replaced by one <code class="language-plaintext highlighter-rouge">warning_once</code> per process. It names the excluded tokens and says “The model is trained on a context that differs from the one it sees at inference.” It also fires for plain-text data when BPE merges tokens across the boundary (<code class="language-plaintext highlighter-rouge">'Hello '</code> + <code class="language-plaintext highlighter-rouge">'world'</code>).</p>

<h2 id="what-remains">What remains</h2>

<p>The fix moves the boundary; it does not put the thought block back. The PR description says so: “Without a patched training template, TRL can’t make the training context match inference for them, but it can stop mis-masking the completion.” With the stock template and thinking off, the training sequence ends with <code class="language-plaintext highlighter-rouge">model\nIt is blue.&lt;turn|&gt;\n</code>, while every inference prompt with thinking off ends with <code class="language-plaintext highlighter-rouge">model\n&lt;|channel&gt;thought\n&lt;channel|&gt;</code>. For DPO, KTO and TPO the prompt ids are also cut to the common prefix. On 1.14.1 their context was already the inference prompt and only the completion was broken; now the sequence is the full render, so for these templates the training context moved away from inference. The PR description documents this trade-off.</p>

<p>To make them match, add two lines to the official template: completed model turns also render the empty block when thinking is off. The <code class="language-plaintext highlighter-rouge">elif</code> branch and the line under it are new; the rest is context:</p>

<div class="language-jinja highlighter-rouge"><div class="highlight"><pre class="highlight"><code>    <span class="cp">{%</span><span class="o">-</span> <span class="k">if</span> <span class="nv">thinking_text</span> <span class="ow">and</span> <span class="nv">thinking_gate</span> <span class="o">-</span><span class="cp">%}</span>
        <span class="cp">{{</span><span class="o">-</span> <span class="s1">'&lt;|channel&gt;thought\n'</span> <span class="o">+</span> <span class="nv">thinking_text</span> <span class="o">+</span> <span class="s1">'\n&lt;channel|&gt;'</span> <span class="o">-</span><span class="cp">}}</span>
    <span class="cp">{%</span><span class="o">-</span> <span class="nv">elif</span> <span class="nv">role</span> <span class="o">==</span> <span class="s1">'model'</span> <span class="ow">and</span> <span class="ow">not</span> <span class="nv">enable_thinking</span> <span class="ow">and</span> <span class="nv">loop.index0</span> <span class="o">&gt;</span> <span class="nv">ns_turn.last_user_idx</span> <span class="o">-</span><span class="cp">%}</span>
        <span class="cp">{{</span><span class="o">-</span> <span class="s1">'&lt;|channel&gt;thought\n&lt;channel|&gt;'</span> <span class="o">-</span><span class="cp">}}</span>
    <span class="cp">{%</span><span class="o">-</span> <span class="k">endif</span> <span class="o">-</span><span class="cp">%}</span>
</code></pre></div></div>

<p>The patched file is <a href="/assets/gemma-4-12B-it.training.jinja">gemma-4-12B-it.training.jinja</a>. Pass it with <code class="language-plaintext highlighter-rouge">SFTConfig(chat_template_path="gemma-4-12B-it.training.jinja")</code>. The prompt is a prefix again, 19 of 19 tokens, there is no warning, and <code class="language-plaintext highlighter-rouge">input_ids</code> end with <code class="language-plaintext highlighter-rouge">\n&lt;channel|&gt;It is blue.&lt;turn|&gt;\n</code>. I measured this with SFTTrainer on single-turn rows only, for completion-only training: <code class="language-plaintext highlighter-rouge">assistant_only_loss=True</code> raises a ValueError with every Gemma 4 template I tried, because TRL has no training template with generation markers for it.</p>

<h2 id="check-it-yourself">Check it yourself</h2>

<p><a href="/assets/verify_gemma4_prompt_prefix.py">verify_gemma4_prompt_prefix.py</a> needs <code class="language-plaintext highlighter-rouge">transformers</code> and <code class="language-plaintext highlighter-rouge">jinja2</code>, no torch. It renders a prompt both ways with the tiny test tokenizer and prints where they diverge:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[official 12B template ]
  generation prompt ends with : 'model\n&lt;|channel&gt;thought\n&lt;channel|&gt;'
  common prefix               : 15/19 prompt tokens
  left out of training        : '&lt;|channel&gt;thought\n&lt;channel|&gt;'
  trained completion          : 'It is blue.&lt;turn|&gt;\n'
[official 12B template {'enable_thinking': True}]
  generation prompt ends with : '?&lt;turn|&gt;\n&lt;|turn&gt;model\n'
  common prefix               : 22/22 prompt tokens
  left out of training        : ''
  trained completion          : 'It is blue.&lt;turn|&gt;\n'
[patched 12B template ]
  generation prompt ends with : 'model\n&lt;|channel&gt;thought\n&lt;channel|&gt;'
  common prefix               : 19/19 prompt tokens
  left out of training        : ''
  trained completion          : 'It is blue.&lt;turn|&gt;\n'
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">--trainer</code> also runs SFTTrainer on both templates; that mode needs <code class="language-plaintext highlighter-rouge">trl</code>, <code class="language-plaintext highlighter-rouge">torch</code> and <code class="language-plaintext highlighter-rouge">datasets</code>.</p>

<h2 id="acknowledgments">Acknowledgments</h2>

<p>I reported the bug and proposed a fix. The merged fix is Quentin Gallouédec’s, reviewed by Albert Villanova.</p>

<h2 id="references">References</h2>

<ul>
  <li><a href="https://github.com/huggingface/trl/issues/7449">Issue #7449</a></li>
  <li><a href="https://github.com/huggingface/trl/pull/7458">PR #7458</a>, mine, closed</li>
  <li><a href="https://github.com/huggingface/trl/pull/7463">PR #7463</a>, merged, 8 files, +172/-53</li>
  <li><a href="https://github.com/huggingface/trl/releases/tag/v1.14.2">TRL v1.14.2</a></li>
  <li><a href="https://huggingface.co/docs/trl/sft_trainer">TRL SFT docs</a></li>
  <li><a href="https://huggingface.co/google/gemma-4-12B-it/blob/main/chat_template.jinja">Gemma 4 12B chat template</a></li>
  <li><a href="/assets/verify_gemma4_prompt_prefix.py">verify_gemma4_prompt_prefix.py</a>, <a href="/assets/gemma-4-12B-it.training.jinja">gemma-4-12B-it.training.jinja</a></li>
</ul>]]></content><author><name>Enes Uzun</name></author><category term="trl" /><category term="gemma" /><category term="fine-tuning" /><category term="chat-templates" /><summary type="html"><![CDATA[With thinking off, the Gemma 4 12B/26B/31B chat template made TRL put the SFT loss on the wrong tokens and drop short answers; short DPO pairs compared nothing. Reported as huggingface/trl#7449, fixed by the maintainers in #7463, released in v1.14.2. What remains, and a two-line template patch.]]></summary></entry></feed>