Artificial intelligence depends on information.
But organizations often overcomplicate the question of data readiness.
They assume they need perfect data, fully integrated systems, a data warehouse, advanced analytics infrastructure, or an enterprise-wide data cleanup initiative.
Usually, that is not where an organization needs to begin.
What data do we need for the specific AI problem we are trying to solve, and is that data usable?
This appendix is designed to help organizations:
- Identify important data sources.
- Understand where information lives.
- Determine who owns it.
- Evaluate quality.
- Classify sensitivity.
- Identify gaps.
- Assess whether specific data is ready for an AI use case.
The objective is not perfection. The objective is visibility and informed decision-making.
A Simple Data Readiness Conversation
Organizations that do not need the full worksheet can begin with six questions.
-
01
What information do we need?
-
02
Where does it exist?
-
03
Can we access it?
-
04
Can we trust it enough for this purpose?
-
05
Is it responsible and appropriate to use?
-
06
Who understands what the data actually means?
If the organization can answer those six questions well, it may be more data-ready than it realizes.
What Not to Do
Avoid several common mistakes.
Do Not Clean Everything
Clean the data relevant to the use case.
Do Not Collect Everything
Use only information that serves a legitimate purpose.
Do Not Assume More Data Is Always Better
Relevant data matters more than volume alone.
Do Not Assume a Database Is Accurate
Systems can contain poor-quality information.
Do Not Assume a Spreadsheet Is Bad
Well-maintained spreadsheets can be highly useful.
Do Not Ignore Institutional Knowledge
Employees may understand important context missing from the data.
Do Not Remove Outliers Automatically
They may represent valuable real-world events.
Do Not Ignore Privacy Until Later
Data governance should begin before AI use.
Do Not Treat Historical Data as Objective Truth
It reflects historical behavior, decisions, and context.
Do Not Wait for Perfection
Sufficient data for a controlled experiment may be enough to begin learning.
The Purpose of Data Readiness
The purpose of this appendix is not to make organizations afraid of their data. It is to help them understand it.
Most organizations will discover a mixture of strong data, weak data, unknown data, missing data, useful documents, historical information, and institutional knowledge.
That is normal.
Data readiness means becoming increasingly capable of distinguishing among them.
Artificial intelligence does not need every piece of organizational data. It needs information appropriate to the problem.
The organization also needs enough understanding to know what that information can, and cannot, support.
Is this data ready enough for this use case, under these conditions, with these safeguards?
That question makes data readiness practical.
And practical data readiness makes applied AI possible.