Research map
Research map · updated 20 July 2026

LLM Distillation & Lineage Observatory

A source-grounded map of methods, repositories, disclosed model dependencies, extraction attacks, defenses, and industry reports for auditing whether large language models were distilled—and from whom.

Core conclusion. The defensible target is not “prove company X copied company Y from output similarity.” It is an evidence-calibrated audit that combines public-artifact reconstruction, black-box behavior, optional internal signals, unknown-teacher support, and explicit abstention.
59indexed papers, tools, disclosures, and reports
4access regimes: outputs, probabilities, representations, parameters
5evidence grades from official disclosure to allegation

Research map

Direct question

Did candidate teacher T influence student S, and through which role?

Adjacent signals

Model family, derivative lineage, tokenizer ancestry, weight similarity, routing habits, and behavioral fingerprints.

Operational evidence

Account clusters, request volume, prompt repetition, payment and infrastructure metadata, and target-capability concentration.

Recommended paper direction

Access-Adaptive Multi-Teacher Distillation Lineage Auditing.

Recover a calibrated teacher set and role assignment under output-only, logprob, hidden-state, and weight access, with hard negatives and an unknown-teacher option.

Open the research agenda →

Navigate by question

QuestionStart hereTypical evidence
Can we identify a teacher directly?Direct DetectionOutput statistics, likelihood residuals, routing signatures
Is the model derived from a known family?Lineage & FingerprintingWeights, gradients, tokenizers, prompts, outputs
How are capabilities extracted?Extraction AttacksAPI queries, synthetic data, logprobs
Which dependencies are confirmed?Open/Public CasesReports, model cards, repositories
What is claimed about closed models?Commercial/Closed CasesOfficial disclosures, telemetry, reporting