Product catalog enrichment
Standardize brands and categories, fill in descriptions and attributes, and resolve supplier inconsistencies. Review field-level changes before they return to your catalog.
Give every agent an isolated data branch. Review row-level changes, validate the result, then merge approved changes when you’re ready.
Git for your data, built into the database. Powered by MatrixOne.
0 price changes. 0 unreviewed changes in main.
Measured-run replay on synthetic data. Twenty Codex-authored SQL plans; execution limited to three workers at a time. Price violations were deliberately injected to test the approval policy.
Behind the film: read the full experiment →TAKE THE WORKFLOW WITH YOU
Download the open-source MatrixOne skill: isolated branches, complete validation, and separately authorized merges. Use your own database and existing SQL tools.
The workflow
When an agent can rewrite thousands of records, a prompt and a production connection are not a review process. Git4Data gives each task an isolated branch, makes changed rows visible, and lets your team validate the result before merging. Keep your existing system of record; use MatrixOne as a workspace for proposed changes.
Where it pays off
A concrete workflow for teams building agents that fill in, standardize, and repair structured product data. Other use cases are adjacent hypotheses to validate with real workloads.
Standardize brands and categories, fill in descriptions and attributes, and resolve supplier inconsistencies. Review field-level changes before they return to your catalog.
Deduplicate records and normalize company, contact, or account fields while keeping proposed changes reviewable.
Generate or correct structured records where teams can define clear validation rules and an approval step.
Git4Data is a data-mutation workspace alongside your current system of record. Data sync, access controls, export paths, and latency need to be validated for each deployment.
Capabilities
Git4Data brings database-native snapshots, table branches, row-level diffs, and merge into MatrixOne. Agents can work with familiar SQL while your team reviews the proposed changes before merge.
Freeze a table at an instant and give that state a name — the analogue of a commit or a tag. Nothing is copied; the snapshot is the object directory as it stood.
CREATE SNAPSHOT sn1 FOR TABLE mydb T;
Clone a table from a snapshot into a new one that then evolves independently. Inserts, updates and deletes on either side stop affecting the other — exactly the isolation a speculating agent needs.
DATA BRANCH CREATE TABLE TClone FROM T{snapshot='sn1'};
Compare two table versions and inspect the rows that changed. Use the diff to understand an agent run before deciding what to do next.
DATA BRANCH DIFF T{snapshot='sn2'} AGAINST TClone{snapshot='sn3'};
Merge a branch back with an explicit conflict policy. Review and validation around that operation are part of your application workflow.
DATA BRANCH MERGE TClone{snapshot='sn3'} INTO T;
The engine already retains point-in-time history for a recent window, so a past state is queryable by timestamp without anyone having declared it interesting in advance.
SELECT * FROM T{timestamp='2026-08-01 12:34:56'};
Database-native branching avoids materializing a full copy for each candidate state. See the benchmark page for the measured environment, workload, and limits behind published results.
100 GB table branch · 0.20 s · 314 KB — benchmark method →Git4Data provides branching and merge primitives. Add the validation rules, approvals, permissions, audit, and rollback procedures your workload requires; branching alone is not a complete safety system.
How it works
Record a version, branch from it, compare versions, reintegrate the accepted changes. The same four moves you make every day in Git — expressed as SQL your ORM, dbt model or agent can already emit.
Name a past state. It is metadata, not bytes — and the engine already keeps a recent window you can query by timestamp without naming anything.
Clone a table from that snapshot. The clone inherits schema and data, then evolves independently — writes on either side stop touching the other.
Compare two versions and inspect the rows that changed. Use the diff to understand a proposed agent run before deciding what happens next.
Fold the accepted rows back with an explicit conflict policy — or drop the branch and pretend it never happened.
These operations run as SQL in MatrixOne. How data is connected, permissions and validations are configured, changes are approved, and results are returned to an existing system depends on your deployment and application workflow.
Why database-native
Git4Data puts table snapshots, branches, diffs, and merge inside MatrixOne. It is a workspace for proposed data changes and can complement the permissions and governance around your existing systems.
Create a writable table branch from a snapshot so an agent can work against its own candidate state.
Compare versions at row level and make the proposed changes visible to your application or review process.
Merge a branch with an explicit conflict policy, or discard the branch. Your checks and approval gates stay in your control.
Branching and diff do not replace IAM, row-level security, audit systems, backups, or human approval. Confirm the exact semantics and data path for your MatrixOne version and deployment.
Build with Git4Data
Start with the live SQL playground, then run MatrixOne locally and follow the documented branch, diff, and merge workflow.
# 1 · run MatrixOne
docker run -d -p 6001:6001 --name matrixone \
matrixorigin/matrixone:latest
# 2 · connect with any MySQL client
mysql -h 127.0.0.1 -P 6001 -u root -p111
# 3 · your first versioned table
CREATE SNAPSHOT s0 FOR TABLE demo t;