Raphael Fakhri · MSc thesis, AUB

Two papers, one question:
can you trust the highlight?

A model reads a Python file and says "Author 7 wrote this". A highlighter then colours the parts of the code that supposedly made it say so. This page explains, from the ground up, how both of the thesis papers test whether those colours tell the truth. It starts with everything the two papers share, then splits into the first paper and the latest one.

The model on this page is the real trained model from both papers. It runs inside your browser, and its two convolution layers run in WebAssembly. Every number in the paper sections comes from the papers themselves.

The 30-second version

Shared by both papers

1 · The question: plausible or faithful?

Dr. Nassar and Bou Abdo built a tool called "Highlight to explain". A neural network called MalConv reads code as raw bytes and classifies it. They added a cheap built-in highlighter, the XAI point, and showed it in an editor plugin. To check it, they used a Python 2 vs Python 3 task and graded the highlights against patterns a human expects, like print "hello" (Python 2 style).

That grading checks whether the highlight agrees with a human. It doesn't check whether the model used those characters. A model can be right for reasons nobody expected, and some highlighters return roughly the same picture no matter what model you attach them to.

AnalogyAsk a witness why they think the butler did it. "He looks shifty" is a plausible answer. It's only a faithful answer if that's really what convinced them. To find out, you remove the shifty look and ask again. If they still say "the butler", that wasn't the reason.

The thesis moves to a task where no human can supply the answer key: which of 20 programmers wrote this Python file? No one can mark the characters that make a file "belong" to its author. So the only thing you can check is faithfulness, by poking the model.

In one lineBoth papers ask: when a highlighter colours some code, did the model really use that code? And which highlighter should a tool like Nassar's use?
Shared by both papers

2 · The data: 20 Code Jam programmers

Google Code Jam was a yearly programming contest. A public archive holds 378,487 Python solutions to 96 problems from 2017 to 2020. Both papers keep the 20 authors who solved the most distinct problems (with at least 20 training, 4 validation and 8 test files each) and remove exact duplicates.

20authors (classes)
509training files · 57 problems
126validation files · 15 problems
254test files · 24 problems

Why the split is by problem, not by file

If the same problem appears in training and testing, a model can cheat: it recognises the problem ("this is the pancake-flipping one, and Author 3 solved that") instead of the author's style. So whole problems are assigned to one side. Every test file is a problem the model has never seen. A control that splits the same amount of data randomly by file reaches 91.6%, so the problem split isn't hiding a big leak.

In one lineThe model has to recognise how someone writes, not what problem they solved.
Shared by both papers

3 · MalConv, layer by layer (running live)

MalConv was built to spot malware in raw program bytes. Nassar modified it and both papers resize it for 20 authors. It has 28,244 learned numbers (parameters) and reaches 96.1% test accuracy on the seed used below (94.4% ± 1.6 over three seeds). Here is the whole path from a file to a verdict.

bytes 4096→embed 8→2 convolutions ×128→gate→ReLU + L1 = XAI point→max-pool 128→dense 64→output 20 + softmax
Loading model…
1

Bytes: the file becomes numbers 4096 integers, 0–255

A computer stores text as bytes. Each character becomes a number from 0 to 255: p is 112, a space is 32, a newline is 10. MalConv reads the first 4,096 bytes of the file. Shorter files are padded with zeros, longer ones are cut. There is no notion of words, keywords or syntax here. The model sees one long row of numbers.

2

Embedding: each byte becomes 8 numbers 4096 × 8

A byte value like 112 is just an ID. "112 is bigger than 32" means nothing. So the model keeps a lookup table with 256 rows (one per possible byte) and 8 columns. Each byte is replaced by its row: 8 numbers that describe it. That's 256 × 8 = 2,048 learned numbers.

Nobody chose these numbers. Training adjusted them until bytes that play similar roles in an author's style ended up with similar rows. You can check that below: pick a character and see which other characters the model learned to treat as its neighbours.

The whole file as an 8-row picture

Each column is one byte of the file (first 512 shown), each row one of the 8 numbers. Blue is positive, red is negative. Repeated patterns in the code show up as repeated stripes.

3

Convolution: 128 pattern detectors slide along the file two maps, each 4096 × 128

A filter is a small stencil 8 bytes wide. At every position it looks at the 8 bytes around it (3 before, the byte itself, 4 after), multiplies each of their 8 embedding numbers by one of its own 64 weights, adds everything up and adds a bias. A big result means "the pattern I'm tuned for is here". The filter then moves one byte to the right and does it again, 4,096 times.

There are 128 filters, and MalConv has two such layers side by side reading the same input: one says what it found, the other decides how much of it to let through (step 4). That's 2 × 4096 × 128 = over a million results, each a sum of 64 products. This is the heavy part, and it's the part running in WebAssembly in your browser.

4

The gate: a dimmer switch on every detector 4096 × 128

The second convolution goes through a sigmoid, a squashing function that turns any number into something between 0 and 1. That value multiplies the first convolution's result. So at each position, each filter's finding is let through fully (×1), dimmed, or switched off (×0) depending on context. This is called a gated convolution.

gated = conv₁ × sigmoid(conv₂),   sigmoid(x) = 1 / (1 + e−x)

For the filter chosen in step 3, across the whole file. Hover to read values.

5

ReLU + L1: the "XAI point" 4096 × 128, mostly zeros

ReLU sets every negative number to 0 and keeps positive ones. Nassar's idea was to put an L1 penalty right here during training: the model pays a small cost (λ = 0.0001) for every unit of activation. To keep the cost low, the model learns to fire only where it really matters. That turns this map into a built-in highlighter, the XAI point: the positions that light up are "what the model noticed".

With the penalty, only about 2% of these values are non-zero (against 49% without it). Your file right now: … non-zero.

Rows are the 128 filters, columns are positions in the file. Darker means stronger. Nearly everything is blank, which is the point.

The XAI point as a highlight

Each position's strongest filter (max over the 128) is spread over the 8 bytes it read, then each code token takes its strongest byte. Darker orange = higher score.

6

Global max-pool: "did this pattern appear anywhere?" 128 numbers

For each filter, the model keeps only its single highest value over the whole file, and forgets where it was. 4096 × 128 numbers become just 128. This makes the model blind to position: a habit at the top of the file counts the same as at the bottom.

This is also why the XAI point can be misleading: one filter's peak decides everything for that filter, and the other 4,095 positions are thrown away.

The filters that fired hardest on this file, and what they fired on

7

Dense layer: a weighted vote 128 → 64

Each of 64 "hidden" units takes all 128 pooled numbers, multiplies each by its own weight and adds a bias: 64 × 128 + 64 = 8,256 learned numbers. Think of each unit as a judge who weighs the 128 pieces of evidence differently.

As in Nassar's original, there is no activation between this layer and the next. Two linear steps in a row collapse into one, so mathematically the head is a single 20 × 128 table: "how much each filter votes for each author". Paper B checked that adding a ReLU here changes nothing (same 96.1% and the same deletion scores). This collapse is what makes the class-aware XAI point possible (section 4).

8

Output + softmax: one score per author, then probabilities 64 → 20

20 more weighted sums (64 × 20 + 20 = 1,300 numbers) give one raw score, a logit, per author. Softmax turns them into probabilities that add up to 100%: raise e to each score and divide by the total. The biggest wins.

p(author c) = ezc / Σk ezk

How the 28,244 numbers were learned (training)

Start with random numbers. Show the model a batch of 16 training files. For each, compare its probabilities with the true author using cross-entropy (a big penalty when it gives the true author a low probability), add the L1 penalty from step 5, and nudge every number slightly in the direction that lowers the total. The nudging rule is Adam with learning rate 0.001. Repeat over all training files for up to 40 passes (epochs), and stop early if validation accuracy hasn't improved for 8 passes. λ was picked from {0, 10⁻⁵, 10⁻⁴, 10⁻³, 10⁻²} by validation accuracy among penalised models.

LayerShapeLearned numbers
Embedding256 × 82,048
Convolution 1 (what)128 filters × 8 bytes × 8 + 1288,320
Convolution 2 (gate)128 filters × 8 bytes × 8 + 1288,320
Dense128 → 648,256
Output64 → 201,300
Total28,244
Shared by both papers

4 · How a highlight is made: the explainers

An explainer (highlighter) takes the model and a file and gives every code token a score. A token is a word-like run (input, range, T) or a single symbol ((, =). All explainers score the same tokens, so they can be compared fairly. To "remove" a token, its characters are replaced by spaces, so every other byte stays in place and the convolution windows don't shift.

Four you can run right here

Darker orange = higher score. The top 10% of tokens (what the editor plugin would highlight) are underlined.

XAI point (Nassar's)

Score = the strongest filter activation at that spot (step 5). It's free: one pass through the model. But it is class-agnostic: it never looks at which author is being explained. It shows "what lit up", not "what pointed to Author 7".

XAI point, class-aware (thesis)

Because the head is one linear table (step 7), the score for author c is exactly the sum of each filter's peak times that filter's vote for c. So credit each filter's peak × vote back to the spot where it peaked. Same cost, one pass, but now it knows which author it's explaining, and evidence against the author gets a negative score.

Token occlusion

Hide one token at a time and watch the author's probability. The drop is the token's score. Honest and simple, but needs one model run per token.

Random

Random scores. Any explainer worth using must beat this by a lot. It's the floor.

The others in the papers (in plain words)

ExplainerHow it worksCost
Grad-CAMUses the gradient (how much the author's score would change) at the XAI-point layer to weight each filter's map, keeping only positive parts.one forward + one backward pass
Gradient × inputThe gradient with respect to each byte's 8 embedding numbers, times the numbers themselves.one backward pass
Integrated gradientsSlowly morph the file from all-spaces into the real file in 32 steps, add up the gradients along the way. Fixes gradient × input's blind spots.32 passes, ~0.4 s
Sliding windowNassar's other method: blank an 8-byte window, slide it 4 bytes, measure the drop.~1,000 passes
LIMEMake 1,000 copies of the file with random tokens hidden, ask the model about each, then fit a simple straight-line model that predicts the answer from "which tokens were kept". Its weights are the scores.1,000 model calls, ~6 s
LEMNA-likeLIME plus a "fused lasso" rule that pushes neighbouring tokens to get similar scores, so it highlights whole stretches. Same 1,000 samples.1,000 calls, ~10 s
Kernel SHAPGame theory (Shapley values): a token's score is its average contribution over many random coalitions of other tokens. Estimated with 1,000 calls.1,000 calls
Partition SHAPShapley values over a hierarchy (groups of neighbouring tokens split recursively), which spends the 1,000 calls far more efficiently on text.1,000 calls, ~5 s
LLMClaude Sonnet 5 (Paper A adds GPT-5.6 and Gemini Pro) gets the file with line numbers, the predicted author and three other files by that author, and ranks up to 15 lines that "reveal the author". It never sees the classifier, like an assistant asked to justify another tool's verdict.one API call
Model-free referenceRanks tokens by how typical they are of the author in the training data. Ignores the model completely. Shows what "plausible" looks like.free
Greedy deletion(Paper B only) Directly searches for the deletion order that kills the prediction fastest. A reference, not a real explainer.~20 s
Shared by both papers

5 · How a highlight is tested

The main test: delete and watch (ABPC)

Sort the tokens by score. Hide the top 1%, 2%, 5% … 100% and record the probability of the predicted author each time. That's the MoRF curve (most relevant first). It should crash fast. Then do it in reverse, least relevant first: the LeRF curve. It should stay high for a long time. ABPC is the area between the two curves. Big gap = faithful. A random ranking gives about 0.

About 100 model runs, a few seconds.

This is one file, so it's a demo of the method, not a result. The papers average over all 254 test files.

Three quick numbers at 10% (what the plugin would show)

  • Comprehensiveness: how much does the probability drop if you hide the top 10%? (high is good)
  • Sufficiency: how much does it drop if you keep only the top 10%? (low is good)
  • Flip rate: how often does hiding the top 10% change the predicted author? (high is good)

Randomization test: does the highlight even depend on the model?

Scramble the model's weights from the output layer downwards and compare the highlights before and after. A faithful explainer's map should change. Try it on the head:

The class-agnostic XAI point is computed before the head, so it can't change at all (correlation 1.00). That's exactly what Paper B found and why the class-aware version was added.

Retraining check (ROAR)

Deleting tokens creates weird files the model never saw in training, which might fool the deletion test. So: remove each explainer's top tokens from every file, retrain a new model from scratch, and see how much accuracy falls. If the explainer found real evidence, the new model has less to learn from. Explanations are "cross-fitted": no model explains its own training files.

Planted identifiers: the one answer key we can build

Give each author a made-up 8-letter name (say qvxmetrz) and insert a line qvxmetrz = 0 into their files. Now we know a piece of evidence exists. But does the model use it? That's measured directly, never from the explainers: blank the identifier or swap in another author's, and see if the prediction drops. Then a faithful explainer should rank the identifier high exactly in the files where the model relies on it, and not elsewhere.

Statistics, simply

Files by the same author aren't independent, so intervals come from resampling authors, not files (a cluster bootstrap). Pairs of explainers are compared with an exact sign-flip test (one sign per author), corrected for running many comparisons (Holm). Explainers sharing a letter in the results table are not significantly different.

Shared by both papers

6 · First, reproducing Nassar's paper

Before changing anything, both papers rebuild the original setup: the small MalConv (window 7, 64 filters, λ = 0.01, 8 dense units, one yes/no output) on Nassar's released Python 2 vs 3 dataset of single lines, and also load the authors' own released model.

  • It reproduces. The released model gets 99.3% on its training lines; retrained models get 96.7% on lines they never saw. The XAI point "localises" the Python 2 pattern in 90.6% of lines under the released grading rule, inside the published 80–98% range.
  • A catch in the grading. The rule accepts any highlight containing a u string prefix, and 85% of the annotations are that one pattern. On the rarer patterns, with a strict rule, the released model localises 88.7% but retrained models only 58.3%, with big swings between seeds.
  • A first faithfulness test. Blanking the XAI point's highlighted window flips the prediction in 64.0% of held-out lines, against 9.9% for a random window of the same size. Much better than chance, but in about a third of lines the model doesn't need what was highlighted.
In one line"Points at the expected pattern" and "points at what the model used" are different measurements. The rest of both papers measures the second.
Branch 1 · the first paper

Paper B

"Highlight, then verify". 8 pages. One model (MalConv), twelve explainers, one LLM. About 10 rounds of review by several Claude models under Dr. Safa's zero-bad-feedback rule, ending at weak accept. Frozen at git tag planB-weak-accept.

The 28 later "what-if" versions (v4–v28) used made-up numbers to see what reviewers would want. They were never real results.

Branch 2 · the latest paper

Paper A

9 pages, built 2–3 Oct 2026 on branch yield. You approved making the strongest what-if claims real, except C++, a GitHub dataset, Qwen and a 30-file selector. Every number comes from a real run, mostly on pc00. Not yet through the review loop.

Rule of the branch: no hypothetical numbers, and LLM steps use stored Claude outputs or non-Claude models.

Models tested

MalConv only gets the full battery. A character n-gram model (96.9%), a random forest on style features (94.9%) and CodeBERTa (92.5%) appear only as accuracy references.

Three model families get the full battery: MalConv, the character n-gram model and CodeBERTa. The random forest stays an accuracy reference. Their mechanisms:

How the character n-gram model works

1. Cut the file (first 4,096 bytes) into every overlapping chunk of 1, 2, 3 and 4 characters. print gives p r i n t pr ri in nt pri rin int prin rint.

2. Keep the 200,000 most common chunks seen in at least 2 training files. Each file becomes a list of 200,000 numbers, one per chunk, using TF-IDF: how often the chunk appears in this file (log-scaled), times how rare it is across files (common chunks like a space count for little).

3. Logistic regression: for each of the 20 authors, one weight per chunk. Score = sum of chunk value × weight, then softmax. The regularisation strength C is picked from {1, 10, 100} on validation. No layers, no training randomness: same data, same model.

Its weights are a built-in explainer ("LR coefficients"): a token's score comes from the weights of the chunks it contains.

How CodeBERTa works

CodeBERTa-small is a transformer (the same building block as ChatGPT, but small and made for reading code, not writing it), pre-trained by Hugging Face on about 2 million functions from GitHub (CodeSearchNet), then fine-tuned here on the 20 authors.

  1. Tokenizer (BPE): code is split into pieces from a vocabulary of 52,000 common chunks, like def, range, _input. It reads the first 512 pieces of the first 4,096 bytes (35% of test files are cut there).
  2. Embedding: each piece becomes 768 numbers, plus 768 more that encode its position.
  3. 6 transformer layers, each with:
    • Self-attention, 12 heads: every piece looks at every other piece and decides which ones matter to it ("this ) belongs to that ("), then mixes in their information. 12 heads = 12 different ways of looking.
    • Feed-forward: each piece's 768 numbers go through a 3,072-wide layer and back.
    • Shortcuts (residuals) and normalisation keep the numbers stable.
  4. Classification head: the first piece's final 768 numbers (a summary slot) go through a small dense layer to 20 scores, then softmax.

Fine-tuned with AdamW, learning rate 5×10⁻⁵. Unlike MalConv, it knows order and context everywhere, not just inside 8-byte windows.

The faithfulness ranking on MalConv

ABPC on the 254 test files, with 95% intervals (hover for cost and group). The leaders are LEMNA-like, LIME and Partition SHAP. Nassar's XAI point is far above random but significantly below them.

model-based explainerLLMreference / control

Same table, same numbers (Paper A reuses Paper B's MalConv results). Then it adds a check Paper B didn't have:

ROAD: deletion without weird files

Instead of replacing a removed token with spaces, a language model fills in a different plausible token. The files stay realistic. ABPC falls for everyone but the order barely moves (rank correlation 0.97 with the space version):

So the "weird files" worry about space masking doesn't drive the ranking.

Cost

Faithfulness vs seconds per file (log scale, 4 CPU cores). Integrated gradients gets most of the leaders' faithfulness in 0.4 s against 5–10 s. Grad-CAM and the class-aware XAI point cost one pass. Recommendation: integrated gradients or Grad-CAM in an editor; LIME-type explainers for offline audits.

Same cost picture. Paper A's conclusion names integrated gradients as the explainer it would show in the editor for MalConv.

Does the ranking hold elsewhere?

Other seeds and a second pool of 20 authors (criteria fixed in advance): LEMNA-like 0.768, LIME 0.762 and Partition SHAP 0.750 lead again; the XAI point (0.647) trails again; the LLM (0.500) is last again.

Same seeds and second pool, plus 82 authors (870 test files, after removing cross-author duplicates, reused files, files with the author's handle and files under 200 bytes). Accuracy: MalConv 89.9% (5 seeds), n-gram 95.1%, CodeBERTa 90.8%. More authors don't help the deep models. Ranking correlation with 20 authors: 0.92.

And on the other models, the leader changes:

ABPC (spaces)ROAD (infill)

CodeBERTa

n-gram model

Partition SHAP leads on CodeBERTa, token occlusion (and the model's own weights) on the n-gram model. An explainer validated on one model can't be assumed faithful on another. (The n-gram model's scores are lower overall because it spreads its evidence over 200,000 features; one token rarely matters much.)

Randomization and retraining

Randomization: the class-agnostic XAI point doesn't change at all when the head is scrambled (correlation 1.00): it can't depend on which author is explained. Grad-CAM keeps 0.51 because it keeps only positive scores. The test can't separate the rest well on this architecture.

Retraining (ROAR), accuracy after removing each explainer's top tokens and retraining (no removal: 94.3%). * = significantly below random.

Only at 40% removal do the good explainers separate from random. At 10–20% the retrained model copes.

Same MalConv results, now also on the other families:

  • Randomization: every CodeBERTa explainer drops to |ρ| ≤ 0.09; n-gram explainers to |ρ| ≤ 0.03 once the weights are random. They all depend on the model.
  • Retraining, n-gram: removing 40% of the tokens chosen by its own weights, LEMNA-like, LIME or occlusion leaves 66.5–71.7% accuracy, against 96.9% for random removal. Even at 10%: 91.7–93.3% vs 96.9%.
  • Retraining, CodeBERTa: recovers more easily. At 40%, occlusion's tokens leave 87.2%, integrated gradients' 90.9%, random 93.7% (two seeds).
Planted identifiers

120 files, identifier in every file. The model relies on it heavily in only 16. Model-based explainers' ranking of the identifier follows that reliance (ρ 0.82–0.92, Kernel SHAP 0.70). The LLM ranks it first more often (68% of files) because it's the most author-typical string, but doesn't follow the model at all (ρ = 0.09).

When the model barely uses the identifier, the LLM still picks its line in 46% of files if the reference files contain it, and 0% if they don't. It finds it by matching the references, not the classifier.

Rerun on all three families, 254 test files each, with a permutation null. On MalConv, 70 files rely heavily on the identifier.

On MalConv every model-based explainer tracks reliance (ρ 0.59–0.75; null's 95th percentile 0.11) and the model-free baseline doesn't (0.39). On the n-gram model and CodeBERTa the identifier is so author-typical that the model-free baseline tracks reliance too. So this control catches a blind explainer but can't rank good ones there.

The LLM explainers

Claude Sonnet 5 only. Most readable, least faithful (ABPC 0.492). Not just a format problem: LIME forced to rank 15 whole lines still beats it. Asking "which lines did the classifier rely on" changes nothing (0.496); Claude Opus 5.5 is barely better (0.517).

Given the classifier's outputs (the probability when each line is blanked), it jumps to 0.591 and its flip rate from 22% to 51%, but still trails simply sorting those numbers.

Three LLMs, all from different companies, all 254 files:

LLMmodel-based or plain sorting

Repeat runs vary little (0.518 vs 0.512), and a neutral prompt changes little (0.489). What the LLMs lack is access to the model. With it they reach 0.59–0.60, but never beat just sorting the numbers they were handed (0.625). Their only extra is a readable sentence, which wasn't tested with users.

Extra studies

Does the L1 penalty help? Ten seeds per λ. The penalty slightly raises the XAI point's normalised ABPC (0.707 at λ=0 to 0.779 at 10⁻³) while accuracy falls (96.6% to 91.0%). It sparsifies the map but doesn't close the gap to the leaders.

What the explainers highlight: about a fifth of top tokens on input/output handling (stdin, readline, split, format) and a similar share on each author's scaffolding (imports, main guard, test-case loop), against 6% and 8% for random. Nassar's XAI point spends over half its top tokens on punctuation, about as much as random.

The big new idea

What would an analyst learn from the highlights?

Token-by-token scores say "these characters matter". An analyst wants a sentence: "the model recognises this author by how they read input". Paper A builds those sentences mechanically and tests them.

  1. Label every line with one of 16 statement types using Python's own tokenizer: header, comment, import, input reading, output writing, the test-case loop (the loop that prints Case #), function/class definitions, other loops, branches, assignments, augmented assignments, returns, other control flow, calls, other.
  2. Turn highlights into hypotheses. For each correctly attributed file, take the explainer's top 10% tokens. For each type: share of top tokens on that type's lines, minus that type's share of all tokens ("excess share"). Average over files and rank the types. The explainer's hypothesis = its top-3 types.
  3. Test the hypothesis: remove every line of those types from the test files and measure the accuracy drop.
  4. Fair comparison: removing more code always hurts, so each removal is paired with a matched random control that removes the same number of random lines per file. Excess drop = drop(types) − drop(matched random).
  5. Measure it twice: on the original model (what does this model use?) and after retraining on the edited files, 5 seeds (does the task need it?).

The removal rule, the matched control and the pass rule were written down before any explainer's result was computed.

The hypotheses

On MalConv, all eleven model-based explainers say: output writing first, the test-case loop second. Third differs (input reading for Partition SHAP, Grad-CAM and the XAI point; imports for LIME and LEMNA-like). CodeBERTa's explainers say imports, function definitions and input reading or headers. The n-gram's say output writing, input reading, imports.

Result 1: on the original model, the hypotheses are right

remove this statement typeremove as many random lines

Removing Partition SHAP's top 3 types drops MalConv from 96.1% to 64.6%, against 85.0% for the matched control: 20.5 points excess. The other model-based explainers give 13.4–16.1. The random explainer's hypotheses fall 6.3 short of their control, the model-free baseline 15.0 short, and a mutual-information baseline (what's informative in the data) is only 0.8 above. Only model-based explainers find what this model uses.

Result 2: the hypotheses are model-specific

Excess drop (points) when the top-3 types from one model's explainers are removed from another model's test files.

CodeBERTa's own hypotheses: 19.3; MalConv's applied to CodeBERTa: 0.4. The n-gram model barely reacts to any removal: with 200,000 features it doesn't depend on any one statement type.

Result 3: retraining erases it

Retrained on the edited files, every model-based explainer's top-3 types cost no more than random lines (excess −1.8 to −0.1). Even removing all five template types together leaves a retrained MalConv at 84.6% (control 86.1%), while the original model falls to 53.5%. Only explainers that propose input reading pass the pre-registered rule, by 6.3 points [2.9, 9.8].

In one lineHighlights tell an analyst correctly what the model depends on. They don't tell you what the author must be recognised by, because a fresh model learns the author from whatever code is left.
The change of plan (deviation 1), and why it's a finding

The original plan was to rewrite a statement type without changing behaviour (canonical layout + consistent renaming, verified by token-stream equality; 7 of 889 files couldn't be verified and stayed unchanged). Rewriting every line drops the original MalConv from 96.1% to 77.6%, but a model retrained on rewritten files gets 92.8% vs 94.3%. The n-gram model loses 1.2 points, CodeBERTa 1.0. No single type costs more than 1.2.

So the operator couldn't separate any hypothesis from noise, and it was swapped for removal before any explainer's yield was computed, logged as deviation 1 in PLAN.md. The rewrite result stays in the paper as a finding: these models use layout and naming, but the task doesn't need them.

Verdict

Validate the explainer on the model it explains (for instance a quick deletion test against random) before showing highlights as grounds for trust. XAI point: cheap, better than random, less faithful. Class-aware XAI point: same cost, narrows the gap. LLM without model access: plausible, not faithful.

All of Paper B's verdict, holding under ROAD, on more seeds and on 82 authors, plus: the best explainer depends on the model; highlights should be presented as evidence about the model, not about the authors; and giving an LLM the model's numbers helps but never beats the numbers themselves.

8 · Side by side

Paper B (first)Paper A (latest)
QuestionAre the highlights faithful? Which explainer, at what cost?Same, plus: does it depend on the model, and what would an analyst learn?
Models with full testsMalConvMalConv, n-gram LR, CodeBERTa
Authors20 (+ a second pool of 20)20 (+ second pool) + 82
Explainers12 incl. 1 LLM12 incl. 3 LLMs (Claude, GPT-5.6, Gemini)
Deletion testsABPC with spacesABPC with spaces + ROAD (infill)
Retraining / randomization / plantedMalConvAll three families
Statement-type testnot in this paperRemoval + matched control + cross-model matrix + retraining
L1 penalty studyYes, 10 seeds per λdropped for space
Leader on MalConvLEMNA-like, LIME, Partition SHAPSame
Leader elsewherenot testedPartition SHAP on CodeBERTa, occlusion on n-gram
Status~10 review rounds, weak accept, frozenReal numbers, not yet reviewed