Service 02 · Six to Eight Weeks · ¥40,000
The records your systems will rely on — made consistent, documented, and ready to use
The Data Preparation Engagement works through the records your organisation has accumulated — addressing duplicates, gaps, inconsistent field definitions and cross-system conflicts — and delivers a cleaned dataset, a data dictionary, and validation rules to keep new records consistent going forward.
What this delivers
At the close, your data is in a state you can actually hand to a system — or to another person
Data preparation is often the work that has to happen before anything else can proceed. Automation tools are only as reliable as the records they process. When those records carry duplicates, blank fields where values should be, or inconsistent naming conventions that developed across different departments or reorganisations, the output of any system built on them will reflect those problems.
This engagement addresses the records directly. We work through the dataset with your team, document what we find, apply a structured cleaning process, and deliver the result in an open format along with the rules we used — so that new records added after the engagement don't reintroduce the same issues.
The deliverables are concrete: a cleaned dataset, a documented data dictionary, and a set of validation rules. These are yours to use however the work requires — in an automation project, in reporting, or in planning the next step.
Cleaned dataset
Delivered in an open format. Duplicates resolved, gaps addressed, field definitions made consistent across the dataset.
Data dictionary
A documented record of what each field means, how it was defined, and what values are considered valid — written so anyone in the organisation can refer to it.
Validation rules
A set of rules to apply to new records as they come in, so the cleaned state of the dataset is preserved going forward without manual review of every entry.
Open format delivery
All deliverables are provided in formats that do not require any specific vendor's software to open or use. No lock-in of any kind.
The situation many organisations are in
Records accumulated across years and systems rarely stay consistent on their own
Most organisations have data that comes from more than one source: software systems that were replaced, departments that maintained their own spreadsheets, data migrations that introduced field name changes, and staff turnover that brought different conventions to the same fields. The result, over time, is a dataset that is technically complete but practically difficult to use.
This matters most when a new system — automated or otherwise — is going to rely on that data. An automation tool that processes records with inconsistent customer name formats, or gaps in date fields, or entries that appear three times under slightly different identifiers, will produce output that reflects those problems. Correcting the output is harder than correcting the records before the system runs.
If your organisation is considering automation, or has recently experienced problems with data quality in reports or systems, this engagement is built around that specific problem — and produces outputs you can use regardless of what comes next.
Duplicate records
The same entity appears multiple times under different identifiers — a common result of data migrations and manual entry over time.
Inconsistent field definitions
The same field means different things in different departments, or was renamed during a system change without the underlying data being updated.
Gaps in historical fields
Fields that are required in current records were not always collected, leaving blanks in older entries that affect any analysis or processing that spans the full history.
The approach
Structured work through the dataset, with every decision documented
The engagement follows a defined sequence. Nothing is changed in the dataset without first being recorded and agreed with your team.
Dataset review
We examine the dataset together — identifying duplicate rates, gap frequencies, field inconsistencies, and any definitions that differ between departments or systems.
Definition alignment
Before cleaning begins, we agree on what each field should mean and what values are valid. This forms the basis of both the cleaning work and the data dictionary.
Cleaning and documentation
The dataset is cleaned according to the agreed definitions. Every change is recorded. The original dataset is preserved alongside the cleaned version.
Delivery and handover
Cleaned dataset, data dictionary, and validation rules are delivered in open formats. A handover session walks your team through everything that was done and why.
Working together
What six to eight weeks of data preparation looks like in practice
The engagement begins with access to the dataset — or the subset of it you want addressed — and a conversation with the people who work with it most closely. In most organisations, the people who use the data daily have knowledge about its quirks that isn't written down anywhere. That knowledge is important to the cleaning process, and the engagement is structured to capture it.
The first two weeks are typically review: we catalogue what's there, produce a short summary of what we found, and agree with your team on how each issue should be handled. Nothing changes in the data until this agreement is in place.
Weeks three through six or seven are the cleaning work itself. We work through the dataset according to the agreed decisions, documenting each change as we go. You can review progress at any point — the change log is kept current throughout.
The final week is delivery and handover. You receive the cleaned dataset, the dictionary, and the validation rules. We walk through them together so your team understands what was done and can continue maintaining the dataset going forward without depending on us.
Weeks 1–2
Dataset review and definition alignment. Issues catalogued. Cleaning approach agreed with your team before any changes are made.
Weeks 3–6 or 7
Cleaning work carried out according to agreed decisions. Change log maintained throughout. Original dataset preserved alongside cleaned version.
Week 6–8 (final)
Delivery of cleaned dataset, data dictionary, and validation rules. Handover session with your team. All materials in open formats.
Throughout
Progress visible at any point. Work conducted in English or Japanese. Engagements can proceed partially remotely where the dataset permits.
Investment
¥40,000 for the full engagement
This covers the review, definition alignment, cleaning work, documentation, and handover session. The duration is six to eight weeks depending on the size and condition of the dataset.
What's included
- Initial dataset review with issue catalogue
- Definition alignment sessions with your team
- Full cleaning of the agreed dataset scope
- Documented change log of every decision made
- Cleaned dataset delivered in open format
- Data dictionary covering all fields and valid values
- Validation rules for maintaining consistency in new records
- Handover session in English or Japanese
Cost structure
No recurring costs. Duration varies with dataset size. Payment terms discussed on enquiry.
Who this suits
Organisations whose records have accumulated across several systems, teams, or reorganisations and who need a documented, consistent dataset before any further automation or reporting work can proceed reliably.
Measurement framework
What gets measured before and after, and how you can judge the result
Data quality can be measured in concrete terms. The engagement records the state of the dataset at the start and at the close, so the work done is visible rather than asserted.
| Measured before | Measured after | Comparison point |
|---|---|---|
| Duplicate record count and rate | Same count in cleaned dataset | At project close |
| Gap rate per field (missing values) | Gap rate in cleaned dataset, with documented handling for each field | At project close |
| Number of inconsistent field definitions across departments | Aligned definitions documented in data dictionary | At project close |
| Undocumented field meanings | Fields covered by data dictionary | At project close |
Original preserved
The source dataset is kept intact throughout the engagement. The cleaned version is produced separately. You have both at the close.
Every change logged
The change log records what was found, what was decided, and what was done. It is part of the handover package and can be reviewed at any point during the engagement.
No vendor dependency
All deliverables are in open formats. Kirameki does not sell software that your team would need to continue using the outputs.
How we approach commitment
The scope is agreed before cleaning begins, and nothing changes without your team's sign-off
The engagement is structured so that the cleaning decisions are yours. We bring the analysis and the framework; the choices about how individual issues should be resolved are made together with your team before any changes are applied to the data.
This matters because data cleaning decisions are often judgment calls — whether two records are genuinely duplicates, how a gap in a date field should be treated, what counts as an invalid value for a particular field. Those judgments should be made by people who know the organisation, not applied unilaterally by an outside party.
If you'd like to discuss your dataset before committing to the engagement, a short initial conversation is available without charge. We can tell you whether what you describe sounds like a reasonable scope for this engagement, and what the review phase is likely to find.
Cleaning decisions agreed first
No changes are made to the dataset until the review findings have been discussed and the approach agreed with your team.
Fixed scope, fixed fee
The dataset scope is defined at the start of the engagement. Work does not expand without a separate discussion and agreement.
No-obligation initial conversation
Before any commitment, you can speak with us about the dataset you have in mind. We'll give you an honest view of whether the engagement is a good fit.
How to proceed
The path from enquiry to a running engagement is straightforward
There are no lengthy intake processes before you need to decide anything. An initial conversation is enough to establish whether the engagement makes sense for your situation.
Send a message
Describe what you know about your dataset — which systems it comes from, roughly how large it is, and what problems you've noticed. Any level of detail is fine to start.
Initial conversation
We'll have a short conversation — in English or Japanese — about the dataset. We'll tell you whether the scope sounds workable and what the review phase is likely to find.
Engagement begins
If both sides are satisfied, we agree a start date and access arrangements. The engagement begins with the review phase, and you receive progress updates throughout.
Data Preparation Engagement
If your records have accumulated problems, a conversation is a reasonable place to begin
The initial conversation carries no obligation. If the dataset you describe isn't a good fit for this engagement, we'll say so clearly and suggest what might be more useful.
Six to eight weeks · ¥40,000 · Tokyo consulting, Japan · info@domain.com
Other services
Other engagements from Kirameki
9 weeks · ¥38,000
Workflow Automation Pilot
A contained pilot applying automated handling to one bounded process, with baseline measurement, parallel operation, and a written comparison at the close.
View this service →
5 weeks · ¥31,000
Governance and Policy Setup
Establishing the internal rules under which automated tools are used — approved applications, data limits, review requirements, and a register template.
View this service →