PRODUCTS

KEYWORDS

What is Git4Data?

Earlier this month, the folks at MatrixOrigin published a paper titled “Git4Data: Database-Native Version Control for AI Agents”.

“Git for Data” refers to the concept of using Git-style version control for structured data. To the best of our knowledge, the term was coined by Noms. Here at DoltHub, we’re both spiritually and technically descended from Noms: Dolt began in 2019 as a fork of Noms, and we’ve come a long way since then. So we’ve always been passionate about Git for Data, and we’re excited to see others get passionate about it too.

We’ve seen a surge of interest in “Git for Data” in the last couple years, with multiple other databases claiming to support it, like DuckDB, XetHub, and LakeFS. But all of these different implementations use their own interfaces, and they typically require a custom storage backend.

In the Git4Data paper, the authors propose an extension to the SQL grammar for version control operations, and a method for implementing version-control in existing SQL databases. What they’re proposing sounds tremendously exciting.

But their claim also gave me some pause. See, we’ve put a lot of thought into how to make database storage amenable to version control. And like the other projects in this space, we chose to design our own storage layer from scratch. We didn’t do it for fun; we did it because the science suggested that we had to. Efficient diffing and merging of database branches has specific requirements that traditional solutions don’t meet. Particularly, it needs the following properties:

  • History independence: Two tables with different histories but the same data should have the same underlying representation.
  • Structual sharing - Two tables with only small changes between them should be able to reuse storage for the parts that haven’t changed.

Most databases don’t have these properties. In order to achieve these properties, we make heavy use of a relatively novel data structure called a Prolly Tree, and we’ve written extensively about how it allows us to do efficient operations that wouldn’t be possible on traditional database storage. Every similar project that we’re aware of uses data structures that are similar or analogous.

MatrixOrigin claims that any database built on Log-Structured Merge trees can achieve the same results, but LSM trees don’t have these properties. So if Git4Data is able to achieve similar performance on existing databases, that’s a really big deal. If it’s true, it changes the game. But big claims require big evidence, and I was skeptical.

Then I saw this chart in their paper:

Table 4: BranchBench at Scale Factor 100 runtime (s)
Workflow Git4Data DoltDB Speedup
ColdWarm ColdWarm
software_dev138.9122.11938.81925.615.8×
failure_repro207.6198.91545.51677.38.4×
data_cleaning62.658.61075.21084.218.5×
mcts36.039.8410.4410.210.3×

If MatrixOrigin is claiming that their implementation of Git4Data outperforms Dolt on BranchBench by a factor of 8x, now I’m really skeptical.

I needed to investigate myself.

The Authors#

Like I mentioned in my previous article about evaluating research papers, learning who wrote the paper can often provide a lot of useful context.

MatrixOrigin is a China-based company that makes environments for agentic AI. Their big product is called Astra, and it’s a runtime environment for AI agents designed to make them safe and reliable. That means a good permissions model, robust sandboxing, and solid history tracking. Their goal is to make it possible to not only log all changes made by an agent but to also review, revert, and replay them. To do this, they need the ability to snapshot and version control the entire environment the agent runs in.

Reading about Astra was interesting because they reached the same conclusions as us, but from the opposite direction. While DoltHub is primarily a database company that recognized the value that version-controlled databases provide to agents, MatrixOrigin is an AI company who realized that they needed version-controlled databases.

This is a real company with a real online presence and real products and real code. But what about their results?

Reproducing BranchBench#

Unfortunately, I wasn’t able to corroborate or reject their benchmarks. They implemented Git4Data within their own database, MatrixOne. They have their own fork of BranchBench that supports Git4Dolt, but every time I tried to run their own setup script against MatrixOne, MatrixOne would crash. As a result, I don’t currently have any measurements to compare Dolt against.

That said, I’m willing to believe that the results they measured were genuine. When BranchBench was first released, it concluded that Dolt had better performance for version control operations than the competition but worse general query performance. We identified that this wasn’t an intrinsic property of Dolt but rather a property of our SQL engine failing to optimize specific queries made by BranchBench. We made some improvements and documented our results here, but there’s still more work to do.

Since our last update, we have begun an internal process to identify query patterns disproportionately generated by AI agents and add them to our own benchmarks. Improving query performance in these uncommon cases is one of our top priorities.

But this still leaves the question of how MatrixOrigin was able to achieve comparable performance to other version-controlled databases without needing to make changes to the storage backend. How does Git4Data actually work?

After reviewing the paper, I concluded that Git4Data is real, it works, and it’s a pretty clever design. But it also has some pretty severe tradeoffs that are glossed over by their paper, likely not caught by BranchBench, and would impact performance in many real-world use cases.

Behind the Magic#

In the paper, the authors propose a version-control layer that can be implemented in existing SQL databases and exposed via an extension to the SQL grammar. They outline the minimum set of requirements necessary for a database to support Git4Data efficiently:

  • Immutable data objects: The backing store must consist of immutable objects. Insert or update operations create new data objects without editing existing ones.
  • Deletion Tombstones: Delete operations mark rows as deleted within a table’s metadata, without actually modifying or deleting data objects in the backing store.
  • Multi-version concurrency control: Version control operations must execute as atomic transactions. Concurrent operations must never observe the partial results of a merge operation.

They’ve put a lot of work into it, even publishing a 15-part deep dive on their own website.

(Disclaimer: their articles have some of the signifiers of AI-generated text. But that may be a result of machine translation.)

So how does Git4Data actually work? The core concept is actually pretty straightforward and goes something like this:

  • Row data is stored in immutable data objects. Each data object encodes some number of rows in the table.
  • A table consists of metadata, a list of pointers to data objects, and tombstones marking removed rows.
  • When branching a table, the database copies the above state, without making copies of the data objects. Subsequent insert operations on either branch will result in new data objects that are exclusive to that branch’s list.
  • When comparing two branches: data objects that are shared by both version of the table can be safely assumed to contain the same records. Only the data objects that differ need to be scanned.

Credit where it’s due: I think this is a cool way to implement diff and merge on top of existing data storage. It leverages the same core principle as Dolt: representing snapshots of a database via collections of objects, where common data is stored in objects referenced by both branches. This indeed achieves the structural sharing that makes branches space-efficient, and allows for high-performance diff and merge operations by skipping the objects that are the same and only comparing the objects that are different.

I’m also not surprised that this approach outperforms Dolt in simple toy cases. But it comes with tradeoffs that are given lip-service in the paper but not properly acknowledged.

Compaction#

Compaction is the process by which an LSM-tree combines multiple smaller data objects into one larger data object. Some operations on LSM trees scale with the number of data objects. For example, rows in each data object are typically sorted, but the database doesn’t necessarily know which data object contains a given row and may need to check them all during a lookup.

Compaction also allows the database to delete removed rows from storage instead of relying on tombstone markers.

For both these reasons, compaction is a necessary step for performance. Without it, the number of data objects in a table can grow without bound, which would have a negative impact on performance.

But it’s also clear that compaction would interfere with Git4Data’s ability to perform efficient diffs and merges. Suppose that a user makes a snapshot of a table and then continues to insert rows into the original table until a compaction occurs. After the compaction, the table will now reference a new set of data objects separate from the ones referenced by the snapshot. A subsequent attempt to diff the two table versions will now result in a full table scan of both versions, since there are no longer any common data objects that can be skipped.

Git4Data’s solution to this is to simply avoid compacting any data objects referenced by a named snapshot. This ensures that two snapshots of the same table will share a common set of data objects but at the cost of degraded performance for ordinary queries. And every new snapshot must flush the current set of pending changes—it always creates at least one new data object. As a result, these queries now necessarily scale with the number of snapshots in the table’s history.

How many snapshots might exist in a typical database? Well, one current use case for Dolt is in Beads, a tool for managing context and memory for agents. Beads makes a commit for every write, and commits are snapshots which also reference to their parent commit. Here’s a public Beads repo on DoltHub with 16,000 commits.. And these numbers will only grow with the age of any repo.

BranchBench works by simulating many agents operating in parallel against a fresh database. Performance degredation due to a long commit history or due to a lack of compaction are unlikely to show up in these benchmarks, but they’re important for any real-world use case.

Without compaction, Git4Data can’t eliminate tombstones either. The data objects for a table will necessarily contain every row that was ever captured in a snapshot, even if that row was subsequently edited or deleted. Operations against the live table will have to iterate over these rows during table scans. In the event of a small table whose row values change frequently, the size of the table on disk could easily be many many times the number of actual rows in the table.

History Independence#

We said before that traditional LSM-trees lack history independence. Why does this matter?

Both Git4Data and Dolt achieve efficient diff and merge by filtering out data objects that exist on both branches without scanning them. But this only works if common data is guaranteed to appear in the same data objects. For LSM-trees, this is generally not the case: two trees will only have the same set of data objects if one was explicitly branched from the other. If the same insert is made on both branches after the split, even Git4Data’s optimized diffing method will initially see them as separate changes and will need to iterate over the inserted rows to identify that they are the same.

A similar problem occurs when attempting to diff two tables that don’t have a common history. This is allowed in both Dolt and Git4Data: the “common base revision” of the two table versions is an empty table. But in the case of Git4Data, these two tables will consist of entirely different data objects, requiring a full table scan in order to diff.

Next Steps for Git4Data#

Outside of performance considerations, we think there’s more room to improve the design of Git4Data, based on our own prior experience developing and maintaining a version-controlled database. For example, a standard for version control in a multi-user environment needs to consider user permissions. When using database branches to achieve isolation, you want to make sure that agents aren’t able to make any changes outside of their approved branches. You may want to give an agent mutable access to a particular branch, while preventing it from merging that branch into another one. If an agent spawns sub-agents, you may want to retain version control operations for the dispatching agent or a human.

Dolt has a system of branch permissions that control what version control operations users can take on specific branches. A robust permission system is vital for true sandboxing.

Parting Words#

There was one claim in the Git4Data paper that requires some correction. It writes:

DoltDB […] materializes and compares table contents, so each branch operation scales with table size. Git4Data’s snapshot-and-delta operations scale with the size of the change.

Dolt does not materialize tables in order to diff them, and branch operations scale with the size of the diff, not the table size. In fact, we do one better: merges in Dolt scale with the number of non-overlapping modified ranges, regardless of the size of each of those ranges. This means that if both branches modify distinct ranges of the primary key space, merging those branches is O(logN) on the size of the table, regardless of the size of the change.

Overall, I think Git4Data is still a good idea. There’s value in standardizing the syntax of version-controlled SQL, and there’s value in implementing this functionality on existing databases, even if it comes with tradeoffs. But I think that Git4Data should be more upfront about what those tradeoffs are, and how existing benchmarks may not accurately capture them.

That’s all for now. If you have any additional thoughts about Git4Data, or the concept of “Git For Data” in general, join our Discord and drop us a line. We always love to chat.