AI Readiness Guide · Appendix C

Data Readiness Checklist and Inventory Worksheet.

Identify the information needed for a specific AI use case, understand where it lives, evaluate whether it is usable, and determine what must happen before a responsible pilot can begin.

Artificial intelligence depends on information.

But organizations often overcomplicate the question of data readiness.

They assume they need perfect data, fully integrated systems, a data warehouse, advanced analytics infrastructure, or an enterprise-wide data cleanup initiative.

Usually, that is not where an organization needs to begin.

What data do we need for the specific AI problem we are trying to solve, and is that data usable?

This appendix is designed to help organizations:

  • Identify important data sources.
  • Understand where information lives.
  • Determine who owns it.
  • Evaluate quality.
  • Classify sensitivity.
  • Identify gaps.
  • Assess whether specific data is ready for an AI use case.

The objective is not perfection. The objective is visibility and informed decision-making.

Part 1

Start With the Use Case

Before evaluating data, define the problem.

Example data needs

Sales forecasting may require historical sales, dates, products, customers, quantities, pricing, or market information.

Customer retention may require purchase history, frequency, service interactions, contract status, or behavioral changes.

Predictive maintenance may require maintenance history, equipment type, failure records, operating hours, sensor data, or environmental conditions.

Start with the problem. Then identify the data.

Part 2

Data Source Inventory

Use this worksheet to identify major information sources relevant to the use case.

Data Source System or Location Owner Format History Available Update Frequency Sensitive?

Possible sources may include accounting software, CRM and ERP systems, HR platforms, production systems, spreadsheets, shared drives, documents, surveys, emails, web analytics, equipment logs, customer-service systems, or employee-maintained files.

What do we have?

Part 3

Data Ownership

Every important data source should have someone who understands it.

That person does not necessarily need a formal data title. The owner may be a department manager, finance employee, operations leader, analyst, IT employee, administrative employee, or another knowledgeable person.

If no one can answer these questions, the data may need additional investigation before AI use.

Part 4

Data Accessibility Checklist

Accessibility Question Response
Can we locate the data?
Can we access it without extraordinary effort?
Can it be exported or retrieved in a usable format?
Can relevant sources be combined if necessary?
Do we understand who has permission to access it?
Does access depend on one employee?
Are there technical restrictions?
Are there contractual restrictions?
Are there regulatory restrictions?
Are there privacy concerns that need review?

If several answers indicate uncertainty, accessibility may be the first readiness issue to address.

Part 5

Data Quality Assessment

Rate each relevant dataset from 1 to 5.

1

Poor

The issue significantly limits use.

2

Weak

Major improvement is needed.

3

Usable With Caution

The data may support limited testing.

4

Good

The data is generally reliable for the intended use.

5

Strong

The data is consistent, understood, and suitable for the use case.

Accuracy

Does the data reflect reality?

/5

Completeness

Are important fields populated?

/5

Consistency

Are names, definitions, formats, and categories used consistently?

/5

Timeliness

Is the information current enough for the intended use?

/5

Validity

Do values follow expected rules and formats?

/5

Uniqueness

Are duplicate records appropriately controlled?

/5

Relevance

Is the information useful for the problem being solved?

/5
Data Quality Total

The calculator adds the seven selected quality ratings.

Part 6

Data Quality Review Questions

Use these questions to explore quality issues in more detail.

These issues do not automatically make the data unusable. They need to be understood.

Part 7

Data Definition Worksheet

AI analysis becomes difficult when people do not agree on what important terms mean.

Term 1

Term 2

Term 3

Terms may include active customer, lead, turnover, downtime, defect, completion, retention, revenue, service request, or high-risk account.

Technology cannot solve every disagreement about definitions. Sometimes the organization must agree first.

Part 8

Historical Data Assessment

Many applied AI use cases depend on historical patterns.

Historical Data Question Response
Is the historical data complete across that period?
Did collection practices change?
Did systems change?
Did organizational definitions change?
Are major gaps present?
Does historical information still reflect current conditions?
Is enough history available for the intended use?

Historical data does not need to be perfect, but changes over time should be understood before predictive conclusions are trusted.

Part 9

Structured and Unstructured Data Inventory

Organizations should consider both structured and unstructured information.

Structured

Structured Data

Examples include tables, databases, transaction histories, sales records, inventory, production metrics, and employee schedules.

Unstructured

Unstructured Data

Examples include documents, PDFs, emails, comments, policies, contracts, maintenance descriptions, survey responses, and meeting notes.

Modern AI can often derive value from both.

Part 10

Institutional Knowledge Inventory

Some of the organization's most important information may not exist in a system at all.

Institutional Knowledge Question Response
Could the knowledge be documented?
Could interviews, procedures, videos, or guidance capture it?
Is this knowledge relevant to a future AI use case?

Examples might include customer history, equipment troubleshooting, process exceptions, supplier relationships, informal decision rules, or historical organizational context.

Institutional knowledge should be treated as an organizational asset.

Part 11

Data Sensitivity Classification

Before using data with AI, classify its sensitivity.

1

Public

Information intended for public access, including public website content, published reports, and public marketing materials.

2

Internal

Information intended for organizational use but not highly sensitive, such as internal procedures and routine reports.

3

Confidential

Information that could create harm if disclosed, including customer records, financial data, contracts, or employee information.

4

Highly Sensitive or Regulated

Information requiring heightened protection, including regulated records, credentials, trade secrets, or highly sensitive personal information.

Data Classification Worksheet

Data Source Classification Reason Approved for AI Use? Conditions

Classification should align with existing privacy, security, legal, and compliance requirements.

Part 12

Data Minimization Checklist

Before providing data to an AI system, ask:

Use only what is necessary.

Part 13

Privacy Review Questions

Privacy Question Response
Do we have the right to use this information for the proposed purpose?
Is employee or customer consent required?
Does the information contain personal data?
Does it contain protected or regulated data?
Will a third-party AI provider receive the data?
Do we understand how the vendor uses it?
Will the vendor retain it?
Can the vendor use it for model training?
Can the information be deleted?
Do contractual obligations limit its use?

Questions marked Unsure should be resolved before higher-risk use.

Part 14

Security Review Questions

Security Question Response
Is access role-based?
Is multi-factor authentication used where appropriate?
Is data encrypted in transit where appropriate?
Is data stored securely?
Are shared user accounts being avoided?
Can access be revoked promptly?
Are AI system activities logged where appropriate?
Is there an incident-reporting process?
Do we know what happens after a vendor security incident?

Security requirements should be proportional to the sensitivity and importance of the data.

Part 15

Bias and Representation Review

Historical data may not represent every group, customer, employee, location, or circumstance equally.

Representation Question Response
Are some locations represented more heavily?
Have historical decisions shaped the data?
Could past practices introduce bias?
Could the use case affect people differently?
Are there variables that could create inappropriate proxies?
Does the use case require fairness evaluation?

Higher-risk decisions involving people require greater scrutiny.

Part 16

Missing Data Review

Missing Data Question Response
Is missingness random?
Does one location or department have more missing data?
Did a system or process change create the gap?
Can missing data be recovered?
Would missing information materially affect the use case?

Sometimes missing data reveals a process problem that should be corrected before AI is introduced.

Part 17

Duplicate Data Review

Duplicate Data Question Response
Are duplicate records present?
Can duplicates be identified reliably?
Can they be removed or consolidated?
Would duplicates distort the intended analysis?

Duplicate records can significantly affect counts, trends, customer analysis, and predictive models.

Part 18

Outlier and Anomaly Review

Unusual values may represent errors, rare events, or valuable signals.

Outlier Question Response
Are extreme values present?
Are they likely data-entry errors?
Could they represent meaningful real-world events?
Has someone with operational knowledge reviewed them?

Do not automatically delete unusual data. Sometimes the unusual cases are exactly what an AI system needs to understand.

Part 19

Data Traceability

For important AI-supported decisions, the organization should be able to understand the source.

Traceability Question Response
Can we identify where the data came from?
Can we identify when it was collected?
Can we identify who or what created it?
Can we identify major transformations?
Can we trace an AI result to the underlying information?

Traceability becomes increasingly important as the consequence of decisions increases.

Part 20

Data Preparation Plan

Identify only the improvements necessary for the selected use case.

Issue 1

Issue 2

Issue 3

Improvements may include standardizing customer names, removing duplicate records, defining important terms, correcting date formats, recovering missing data, documenting field definitions, or obtaining privacy approval.

Part 21

Minimum Viable Data

Starting with minimum viable data can reduce complexity, privacy risk, preparation time, and cost.

Part 22

Pilot Data Checklist

Before using the data in a pilot, confirm:

Part 23

Data Readiness Rating

After completing the worksheet, rate the data for this specific use case.

1

Not Ready

Major limitations prevent responsible testing.

2

Significant Preparation Needed

Useful data exists, but substantial work is required.

3

Pilot Ready With Caution

Data may support a controlled experiment when limitations are understood.

4

Ready

Data is sufficiently reliable, accessible, and governed for the intended use.

5

Strong

Data is well understood, managed, and capable of supporting broader AI use.

Part 24

What Prevents a Higher Rating?

This prevents the organization from turning data readiness into an unlimited cleanup project.

Part 25

Data Readiness Decision

After completing the review, choose one.

Part 26

Data Readiness Summary

A Simple Data Readiness Conversation

Organizations that do not need the full worksheet can begin with six questions.

  1. 01

    What information do we need?

  2. 02

    Where does it exist?

  3. 03

    Can we access it?

  4. 04

    Can we trust it enough for this purpose?

  5. 05

    Is it responsible and appropriate to use?

  6. 06

    Who understands what the data actually means?

If the organization can answer those six questions well, it may be more data-ready than it realizes.

What Not to Do

Avoid several common mistakes.

Do Not Clean Everything

Clean the data relevant to the use case.

Do Not Collect Everything

Use only information that serves a legitimate purpose.

Do Not Assume More Data Is Always Better

Relevant data matters more than volume alone.

Do Not Assume a Database Is Accurate

Systems can contain poor-quality information.

Do Not Assume a Spreadsheet Is Bad

Well-maintained spreadsheets can be highly useful.

Do Not Ignore Institutional Knowledge

Employees may understand important context missing from the data.

Do Not Remove Outliers Automatically

They may represent valuable real-world events.

Do Not Ignore Privacy Until Later

Data governance should begin before AI use.

Do Not Treat Historical Data as Objective Truth

It reflects historical behavior, decisions, and context.

Do Not Wait for Perfection

Sufficient data for a controlled experiment may be enough to begin learning.

The Purpose of Data Readiness

The purpose of this appendix is not to make organizations afraid of their data. It is to help them understand it.

Most organizations will discover a mixture of strong data, weak data, unknown data, missing data, useful documents, historical information, and institutional knowledge.

That is normal.

Data readiness means becoming increasingly capable of distinguishing among them.

Artificial intelligence does not need every piece of organizational data. It needs information appropriate to the problem.

The organization also needs enough understanding to know what that information can, and cannot, support.

Is this data ready enough for this use case, under these conditions, with these safeguards?

That question makes data readiness practical.

And practical data readiness makes applied AI possible.

Last reviewed:

This worksheet supports internal organizational review. It does not replace professional legal, privacy, security, regulatory, or technical assessment.