SAT, SEPTEMBER 05, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude General

Provenance Prompts: 6 for Knowing What You Are Building On

K2 Horizon shipped with its training data published. Kimi, GLM and Qwen did not. A federal court has already held that acquiring training material through piracy is unlawful, and Sony and Warner sued Anthropic on that question last week. These six prompts are for establishing what you can and cannot say about the models in your stack, before someone asks. None of this is legal advice.

⌨️ 6 prompts 🕐 Updated Sep 5, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
Map what is actually in my stack
Most stacks have a model nobody remembers adding. This finds it.
Here is every model my system touches: [LIST, INCLUDING FINE-TUNES AND EMBEDDINGS]

For each, tell me:
- Who published it and under what terms
- What is disclosed about training data
- Whether it is a base model or derived from another
- What I could not establish even if asked

I want the last column filled honestly rather than left blank.
2
Read what the licence actually permits
The last line matters. Silence in a licence is not permission, and a model will fill the gap if you let it.
Here is a model licence: [PASTE]

Tell me, quoting the clause each time:
- Whether commercial use is permitted and at what scale
- Who owns the output
- Whether I may fine-tune and redistribute
- Whether terms could change for future versions
- Anything unusual compared to standard MIT or Apache

If a question is not addressed, say unaddressed rather than inferring.
3
Work out what I could not answer if challenged
Written before a contract, this is diligence. Written after, it is a problem.
A client asks me to warrant that my AI system does not use improperly obtained material.

Based on my stack: [DESCRIBE], tell me:
- What I can establish and how
- What I cannot establish at all
- What I would need from vendors to close the gap
- Whether the warranty is one I could honestly give

Be blunt. I would rather know now.
4
Draft the vendor questions
Indemnity decides who carries the risk. It is the clause most people never raise.
I want to ask an AI vendor about training data provenance before committing.

Draft the questions, and for each tell me what a reassuring answer looks like versus an evasive one.

Include indemnification, and what I should get in writing rather than in a sales call.
5
Design a record I will actually keep
Short records survive. Comprehensive ones get abandoned in month two.
Design a provenance record for every model and dataset entering my pipeline.

Capture enough that in two years someone could establish where each came from and under what terms, without anyone remembering.

Keep it to fields I will realistically fill in. An elaborate schema I abandon is worse than a short one I maintain.
6
Decide whether disclosure changes my choice
The last line is the important one. Most projects do not need this, and pretending otherwise wastes capability.
I am choosing between a stronger model with undisclosed training data and a weaker one with everything published: [DESCRIBE BOTH AND MY USE CASE]

Tell me:
- Whether provenance matters for what I am actually doing
- What specifically I would gain from the disclosed one
- What I would give up in capability
- Whether a hybrid works

If provenance does not matter here, say so plainly rather than defaulting to the cautious answer.