Kirameki
Data preparation — records organised, documented and ready for reliable use

Service 02 · Six to Eight Weeks · ¥40,000

The records your systems will rely on — made consistent, documented, and ready to use

The Data Preparation Engagement works through the records your organisation has accumulated — addressing duplicates, gaps, inconsistent field definitions and cross-system conflicts — and delivers a cleaned dataset, a data dictionary, and validation rules to keep new records consistent going forward.

What this delivers

At the close, your data is in a state you can actually hand to a system — or to another person

Data preparation is often the work that has to happen before anything else can proceed. Automation tools are only as reliable as the records they process. When those records carry duplicates, blank fields where values should be, or inconsistent naming conventions that developed across different departments or reorganisations, the output of any system built on them will reflect those problems.

This engagement addresses the records directly. We work through the dataset with your team, document what we find, apply a structured cleaning process, and deliver the result in an open format along with the rules we used — so that new records added after the engagement don't reintroduce the same issues.

The deliverables are concrete: a cleaned dataset, a documented data dictionary, and a set of validation rules. These are yours to use however the work requires — in an automation project, in reporting, or in planning the next step.

Cleaned dataset

Delivered in an open format. Duplicates resolved, gaps addressed, field definitions made consistent across the dataset.

Data dictionary

A documented record of what each field means, how it was defined, and what values are considered valid — written so anyone in the organisation can refer to it.

Validation rules

A set of rules to apply to new records as they come in, so the cleaned state of the dataset is preserved going forward without manual review of every entry.

Open format delivery

All deliverables are provided in formats that do not require any specific vendor's software to open or use. No lock-in of any kind.

The situation many organisations are in

Records accumulated across years and systems rarely stay consistent on their own

Most organisations have data that comes from more than one source: software systems that were replaced, departments that maintained their own spreadsheets, data migrations that introduced field name changes, and staff turnover that brought different conventions to the same fields. The result, over time, is a dataset that is technically complete but practically difficult to use.

This matters most when a new system — automated or otherwise — is going to rely on that data. An automation tool that processes records with inconsistent customer name formats, or gaps in date fields, or entries that appear three times under slightly different identifiers, will produce output that reflects those problems. Correcting the output is harder than correcting the records before the system runs.

If your organisation is considering automation, or has recently experienced problems with data quality in reports or systems, this engagement is built around that specific problem — and produces outputs you can use regardless of what comes next.

Duplicate records

The same entity appears multiple times under different identifiers — a common result of data migrations and manual entry over time.

Inconsistent field definitions

The same field means different things in different departments, or was renamed during a system change without the underlying data being updated.

Gaps in historical fields

Fields that are required in current records were not always collected, leaving blanks in older entries that affect any analysis or processing that spans the full history.

The approach

Structured work through the dataset, with every decision documented

The engagement follows a defined sequence. Nothing is changed in the dataset without first being recorded and agreed with your team.

01

Dataset review

We examine the dataset together — identifying duplicate rates, gap frequencies, field inconsistencies, and any definitions that differ between departments or systems.

02

Definition alignment

Before cleaning begins, we agree on what each field should mean and what values are valid. This forms the basis of both the cleaning work and the data dictionary.

03

Cleaning and documentation

The dataset is cleaned according to the agreed definitions. Every change is recorded. The original dataset is preserved alongside the cleaned version.

04

Delivery and handover

Cleaned dataset, data dictionary, and validation rules are delivered in open formats. A handover session walks your team through everything that was done and why.

Working together

What six to eight weeks of data preparation looks like in practice

The engagement begins with access to the dataset — or the subset of it you want addressed — and a conversation with the people who work with it most closely. In most organisations, the people who use the data daily have knowledge about its quirks that isn't written down anywhere. That knowledge is important to the cleaning process, and the engagement is structured to capture it.

The first two weeks are typically review: we catalogue what's there, produce a short summary of what we found, and agree with your team on how each issue should be handled. Nothing changes in the data until this agreement is in place.

Weeks three through six or seven are the cleaning work itself. We work through the dataset according to the agreed decisions, documenting each change as we go. You can review progress at any point — the change log is kept current throughout.

The final week is delivery and handover. You receive the cleaned dataset, the dictionary, and the validation rules. We walk through them together so your team understands what was done and can continue maintaining the dataset going forward without depending on us.

Weeks 1–2

Dataset review and definition alignment. Issues catalogued. Cleaning approach agreed with your team before any changes are made.

Weeks 3–6 or 7

Cleaning work carried out according to agreed decisions. Change log maintained throughout. Original dataset preserved alongside cleaned version.

Week 6–8 (final)

Delivery of cleaned dataset, data dictionary, and validation rules. Handover session with your team. All materials in open formats.

Throughout

Progress visible at any point. Work conducted in English or Japanese. Engagements can proceed partially remotely where the dataset permits.

Investment

¥40,000 for the full engagement

This covers the review, definition alignment, cleaning work, documentation, and handover session. The duration is six to eight weeks depending on the size and condition of the dataset.

What's included

  • Initial dataset review with issue catalogue
  • Definition alignment sessions with your team
  • Full cleaning of the agreed dataset scope
  • Documented change log of every decision made
  • Cleaned dataset delivered in open format
  • Data dictionary covering all fields and valid values
  • Validation rules for maintaining consistency in new records
  • Handover session in English or Japanese

Cost structure

Review and definition phase Weeks 1–2
Cleaning and documentation Weeks 3–6/7
Delivery and handover Week 6–8
Total fixed fee ¥40,000

No recurring costs. Duration varies with dataset size. Payment terms discussed on enquiry.

Who this suits

Organisations whose records have accumulated across several systems, teams, or reorganisations and who need a documented, consistent dataset before any further automation or reporting work can proceed reliably.

Measurement framework

What gets measured before and after, and how you can judge the result

Data quality can be measured in concrete terms. The engagement records the state of the dataset at the start and at the close, so the work done is visible rather than asserted.

Measured before Measured after Comparison point
Duplicate record count and rate Same count in cleaned dataset At project close
Gap rate per field (missing values) Gap rate in cleaned dataset, with documented handling for each field At project close
Number of inconsistent field definitions across departments Aligned definitions documented in data dictionary At project close
Undocumented field meanings Fields covered by data dictionary At project close

Original preserved

The source dataset is kept intact throughout the engagement. The cleaned version is produced separately. You have both at the close.

Every change logged

The change log records what was found, what was decided, and what was done. It is part of the handover package and can be reviewed at any point during the engagement.

No vendor dependency

All deliverables are in open formats. Kirameki does not sell software that your team would need to continue using the outputs.

How we approach commitment

The scope is agreed before cleaning begins, and nothing changes without your team's sign-off

The engagement is structured so that the cleaning decisions are yours. We bring the analysis and the framework; the choices about how individual issues should be resolved are made together with your team before any changes are applied to the data.

This matters because data cleaning decisions are often judgment calls — whether two records are genuinely duplicates, how a gap in a date field should be treated, what counts as an invalid value for a particular field. Those judgments should be made by people who know the organisation, not applied unilaterally by an outside party.

If you'd like to discuss your dataset before committing to the engagement, a short initial conversation is available without charge. We can tell you whether what you describe sounds like a reasonable scope for this engagement, and what the review phase is likely to find.

Cleaning decisions agreed first

No changes are made to the dataset until the review findings have been discussed and the approach agreed with your team.

Fixed scope, fixed fee

The dataset scope is defined at the start of the engagement. Work does not expand without a separate discussion and agreement.

No-obligation initial conversation

Before any commitment, you can speak with us about the dataset you have in mind. We'll give you an honest view of whether the engagement is a good fit.

How to proceed

The path from enquiry to a running engagement is straightforward

There are no lengthy intake processes before you need to decide anything. An initial conversation is enough to establish whether the engagement makes sense for your situation.

01

Send a message

Describe what you know about your dataset — which systems it comes from, roughly how large it is, and what problems you've noticed. Any level of detail is fine to start.

02

Initial conversation

We'll have a short conversation — in English or Japanese — about the dataset. We'll tell you whether the scope sounds workable and what the review phase is likely to find.

03

Engagement begins

If both sides are satisfied, we agree a start date and access arrangements. The engagement begins with the review phase, and you receive progress updates throughout.

Data Preparation Engagement

If your records have accumulated problems, a conversation is a reasonable place to begin

The initial conversation carries no obligation. If the dataset you describe isn't a good fit for this engagement, we'll say so clearly and suggest what might be more useful.

Six to eight weeks · ¥40,000 · Tokyo consulting, Japan · info@domain.com

Cookie preferences

Essential cookies

Always active — required for the site to work.

Always on

Analytics cookies

Help us understand how pages are used.

Marketing cookies

Used for advertising and personalisation.

Personalisation cookies

Remember preferences for your visits.