1. Inventory sources and their owners
List the repositories, file types and records needed for the proposed task. Include shared drives, SharePoint libraries, database exports and email attachments only where they are relevant. Record an owner for each source, its purpose and whether it is approved for use.
Our data-readiness service can include a source inventory, quality assessment, preparation pipeline and evaluation material. The scope should specify whether the work covers documents for retrieval, structured records for extraction or examples for model development. These uses need different checks. A document useful for search may be unsuitable as a training example or as an authoritative answer source.
2. Distinguish readable files from usable text
A PDF can contain searchable text, images of pages or a mixture. Optical character recognition, or OCR, converts text in images into machine-readable characters. Its output needs inspection where layout, handwriting, poor scans or unusual characters can change meaning. A visually clear page does not guarantee a clean extraction.
Keep relationships between headings, paragraphs, tables and footnotes. A table converted into a line of disconnected values may lose the connection between a charge and its conditions. Compare extracted text with the original in representative documents. Preserve source locations so an answer or extracted field can be checked against the page that supplied it.
3. Resolve duplicates and conflicting editions
Duplicate material can overweight a source in retrieval and create unnecessary processing. Identify exact copies separately from near-duplicates: a revised policy may look similar but contain a material change. Do not discard it solely because much of the wording matches another document.
Use metadata, meaning information about a record, to track ownership, approval status and applicable dates where those exist. Keep an explicit relationship between superseded and approved editions. If documents disagree, ask the responsible source owner to resolve the conflict. A model should not decide which policy is authoritative by choosing the most confidently worded passage.
4. Minimise personal and confidential information
Data minimisation means limiting personal data to what is necessary for the specified purpose. Remove irrelevant contact details, signatures and sensitive free-text fields before experimentation where possible. Redaction must affect the actual file contents; placing a visual rectangle over text may leave the underlying information recoverable.
Pseudonymisation replaces identifying information but does not necessarily make data anonymous. Keep any re-identification material separate and restrict access. The ICO’s data protection guidance explains the UK requirements. Confirm the lawful basis, supplier arrangements and deletion process with your organisation’s privacy lead before transferring personal records into an AI service.
5. Build representative evaluation material
Collect examples that reflect the task’s actual variation rather than only clean, convenient files. Include incomplete records, unusual layouts, missing fields and inputs outside the intended scope. Specify the expected response, including when the correct result is to flag uncertainty or refuse processing.
Keep evaluation examples separate from material used to tune prompts or models. This helps expose behaviour on unfamiliar inputs. For labelled records, define each label clearly and resolve disagreement between reviewers. If people cannot consistently decide which category applies, the problem may be the category definition rather than the model. Record these uncertainties instead of forcing an artificial certainty.
6. Make preparation repeatable and reversible
A preparation pipeline is the sequence that reads, transforms and stores source material. Document each transformation, keep appropriate source references and record failures. Avoid manual corrections that cannot be reproduced when the source collection is refreshed. A repeatable process makes it easier to investigate why an answer changed.
Plan access changes and deletion across source copies, indexes, evaluation sets and logs. Derived data may still contain confidential information even when it no longer resembles the original document. For an enquiry, provide file types, source locations and known quality problems. A small redacted extract is usually enough to discuss the preparation approach before agreeing any transfer of production material.
Discuss data preparation