Research map · updated 20 July 2026
LLM Distillation & Lineage Observatory
A source-grounded map of methods, repositories, disclosed model dependencies, extraction attacks, defenses, and industry reports for auditing whether large language models were distilled—and from whom.
Core conclusion. The defensible target is not “prove company X copied company Y from output similarity.” It is an evidence-calibrated audit that combines public-artifact reconstruction, black-box behavior, optional internal signals, unknown-teacher support, and explicit abstention.
59indexed papers, tools, disclosures, and reports
4access regimes: outputs, probabilities, representations, parameters
5evidence grades from official disclosure to allegation
Research map
Direct question
Did candidate teacher T influence student S, and through which role?
Adjacent signals
Model family, derivative lineage, tokenizer ancestry, weight similarity, routing habits, and behavioral fingerprints.
Operational evidence
Account clusters, request volume, prompt repetition, payment and infrastructure metadata, and target-capability concentration.
Recommended paper direction
Access-Adaptive Multi-Teacher Distillation Lineage Auditing.
Recover a calibrated teacher set and role assignment under output-only, logprob, hidden-state, and weight access, with hard negatives and an unknown-teacher option.
Navigate by question
| Question | Start here | Typical evidence |
|---|---|---|
| Can we identify a teacher directly? | Direct Detection | Output statistics, likelihood residuals, routing signatures |
| Is the model derived from a known family? | Lineage & Fingerprinting | Weights, gradients, tokenizers, prompts, outputs |
| How are capabilities extracted? | Extraction Attacks | API queries, synthetic data, logprobs |
| Which dependencies are confirmed? | Open/Public Cases | Reports, model cards, repositories |
| What is claimed about closed models? | Commercial/Closed Cases | Official disclosures, telemetry, reporting |