Tag: Gemma-4-E4b
How Do You Measure Abliteration Damage? I Compared Every Way I Could
How much did that abliteration actually hurt the model? It is the question behind every comparison I run, and for a long time I did not really understand the number we use to answer it. The score is KL divergence, and there is more than one way to calculate it. Different datasets, different token depths, thinking on or off, three tools each with their own method. So I stopped taking the number on faith and compared them. I will explain the maths as plainly as I can too, because it scared me off for years and it really should not.
Gemma4-E4B Abliteration Benchmarked: 23 Variants Under the Microscope
Twenty-three different people tried to remove the safety filters from the same AI model, Google’s Gemma4-E4B. The headline finding is an awkward one. The most popular variant, with 796,000 downloads, is also the most damaged. Meanwhile a surgical edit that touches just 21 of the model’s 719 weight tensors does nearly as well with no measurable harm. This is the biggest abliteration comparison I have run, and the gap between the best and worst is wider than anything I have seen before.