
DumboDB is the mash up of MongoDB and Git. Every feature that exists in the product has a direct equivalent in either Mongo or Git. Until now.
We are adding a new feature called “Merge Modes”, which has no equivalent in either product. Today we’ll dive into why we had to invent something new and how it works. Let’s go!
Mongo’s Compare-and-Swap#
One of the common ways that Mongo users perform transactional operations is to use a
compare-and-swap (CAS).
This is done by using an updateOne call to both find a document and update it in one
operation. Like this:
const doc = await db.posts.findOne({ _id: postId });
const result = await db.posts.updateOne(
{
_id: postId,
version: doc.version // Search includes the version you expect to update.
},
{
$set: { title: "Published" },
$inc: { version: 1 } // Update the version, preventing others from getting
// search results with the original version number.
}
);
if (result.matchedCount === 0) {
// someone else moved it first; re-read and retry. It's critical for the application
// to check how many documents it updates, because if 0 documents were updated that
// means you couldn't find the document with the version you expected to update.
}
This is an optimistic lock. Even
though a fair amount of time could pass between the findOne and the updateOne call, it’s okay
because you are effectively saying “update this document if its state is the same as when I read
it.”
Which is great. You read the doc, took some time to process it, and verified that no one else messed with it while you did that work. Simple. Understandable.
The dirty secret is that until the latest version of DumboDB, 0.7.0, this didn’t work at all.
How Dolt Works#
Dolt is the underlying storage system for DumboDB. That means that all writes to a DumboDB database are materialized on disk using a write-ahead journal with transactional updates.
Dolt started out as a data sharing tool, and one of the decisions made long ago was that if you have two branches that update the value of a row to the same value, those two branches can merge without conflicts.
This makes sense, especially when compared to how Git handles source code changes. If two edits to the code have the same end results, that’s great. Everyone is in agreement about the correct state of the file.
In Dolt, this applies to short-lived branches as well, specifically when two edits happen at the same time by two different sessions. I wrote about this a long time ago, and it still applies today.
The mechanism that performs the 3-way merge is deep in the Dolt database. It’s low enough level that the journal update happens in unison with the implicit 3-way merge. It’s an optimistic lock, so as multiple sessions are attempting to update the data, the storage system will actually perform retries when the lock acquisition fails.
Given that MongoDB uses the Dolt storage engine to perform its transactional update of the
journal, this 3-way merge mechanism has been in DumboDB since the first release. To put a finer
point on it, if two sessions attempted to use the compare and swap mechanism described above, then
both sessions would bump the version number from n to n+1, and the Dolt storage system will
happily allow both session writes to succeed because they incremented to the same value.
And that’s the bug. No MongoDB application that uses this pattern is safe.
Introducing Merge Modes#
One option was to make DumboDB behave like MongoDB and call it a day. But there is actually more nuance to this problem than you might expect. First, many applications don’t use this pattern, and the current behavior makes lots of sense for the data sharing use case. Second, there are actually holes in the Mongo model exposed by the fact that DumboDB adds branching. For example, an application that is built on top of DumboDB’s branching system may want to consider the entire document as an atomic unit that should never be altered by two branches at the same time.
The key piece is to be able to dynamically look at the document being merged, and determine if we can automatically merge or if a conflict requires resolution before we can proceed.
Instead of forcing one model, we chose to support four. We are calling this collection setting
mergeMode, and it’s configured just like validators and collations.
A mode’s name is two choices joined together. The first half is the unit the
merge compares — the whole document, or a single field. The second half is
the trigger — whether a side merely Touched it (wrote it at all) or the
two sides are Divergent (wrote it to different values). The four combinations
are the four modes:
Touched — a side wrote it at all | Divergent — the sides wrote it differently | |
|---|---|---|
document — the whole document | documentTouched | documentDivergent |
field — a single field | fieldTouched | fieldDivergent |
And here’s what each one actually does:
| mode | conflicts when |
|---|---|
fieldTouched | New default behavior. Both sides wrote the same field, even to the same value. |
fieldDivergent | Old Dolt default behavior. Both sides wrote the same field, and the values differ. |
documentTouched | Both sides wrote anything in the same document. |
documentDivergent | Both sides wrote the same document, and the results differ. |
Upon close inspection, each level is unique, and they aren’t even strictly ordered.
documentTouched strictest: any concurrent write
/ \
fieldTouched documentDivergent incomparable with each other
\ /
fieldDivergent loosest: only differing values
The two middle modes, fieldTouched and documentDivergent, are the surprising part.
Neither is stricter than the other; they measure different things. fieldTouched asks
whether both sides wrote the same field, while documentDivergent asks whether the two
sides ended up with different documents. It takes two examples to show that neither one
implies the other.
First, a case where fieldTouched conflicts but documentDivergent merges. Both branches
write the same field to the same value:
base { a: 0, b: 0 }
ours { a: 0, b: 1 } both wrote b, so fieldTouched conflicts
theirs { a: 0, b: 1 } but the documents are identical, so documentDivergent merges
Now the reverse: documentDivergent conflicts but fieldTouched merges. Each branch writes
a different field:
base { a: 0, b: 0 }
ours { a: 0, b: 1 } no field in common, so fieldTouched merges
theirs { a: 1, b: 0 } but the documents differ, so documentDivergent conflicts
Making a document divergent only requires touching a field, but fieldTouched cares
whether both sides touched the same field. Those aren’t the same question, and that’s why
these two modes sit side by side on the lattice instead of one above the other.
There may be additional merge modes that make sense. We think this covers the primary ways people build applications, but if you need something more, let us know!
Example#
Let’s walk through documentTouched, the strictest mode, since it makes conflicts easy to
trigger and shows the whole resolution flow end to end. This mode treats the entire document
as a single unit: if two branches both write to a document, it conflicts, even when they
touch completely different fields or the documents are identical.
Run DumboDB#
Grab the v0.7.0 release. Any
earlier version will not support this demo! We publish pre-built binaries for
Linux, MacOS, and Windows on the
releases page; drop the one for your
platform somewhere on your PATH. There’s a Docker image and build-from-source
instructions in the README if you
prefer either of those.
In a terminal, run the following command to start the server:
$ dumbodb --data-dir /tmp/dumbodb_data
Let that terminal run in the background. In another terminal, connect to the server with
mongosh:
$ mongosh mongodb://localhost:27017
Set Up a Collection#
Create a collection with the documentTouched merge mode and commit a baseline
document on main. Note that the merge mode is set when the collection is created.
db = db.getSiblingDB("blog@main")
// The mergeMode option selects a non-default mode.
db.createCollection("posts", { mergeMode: "documentTouched" })
db.posts.insertOne({ _id: 1, title: "Draft", author: "neil", status: "wip" })
db.runCommand({ dumboCommit: 1, message: "baseline post" })
Make Conflicting Edits on Two Branches#
Now branch off, and have each branch touch a different field of the same document. main
changes the title, and the editor branch changes the author.
// An editor branches off to work in parallel.
db.runCommand({ dumboBranch: 1, action: "add", branch: "editor" })
// main touches the `title` field.
db.posts.updateOne({ _id: 1 }, { $set: { title: "Launch Day" } })
db.runCommand({ dumboCommit: 1, message: "set title on main" })
// Editor branch touches the `author` field.
var editor = db.getSiblingDB("blog@editor")
editor.posts.updateOne({ _id: 1 }, { $set: { author: "grace" } })
editor.runCommand({ dumboCommit: 1, message: "set author on editor" })
These two edits don’t overlap at all. Under fieldTouched they would merge cleanly
into a document that has both changes. But we asked for documentTouched,
which treats the whole document as the unit, so the merge conflicts.
Merge and Hit the Conflict#
db.runCommand({ dumboMerge: 1, mergeIn: "editor", message: "merge editor" })
{
conflicts: [ { collection: 'posts', count: 1 } ],
ok: 0,
code: 96,
errmsg: 'dumboMerge: unresolved conflicts in 1 collection(s)'
}
The merge stopped and left us a conflict to resolve. dumboConflicts shows us exactly what’s
in dispute:
db.runCommand({ dumboConflicts: 1 })
{
conflicts: [
{
conflictId: '7lM6Ptl5Re/mF/24rnxqNA',
type: 'document',
collection: 'posts',
reason: {
code: 'bothModified',
message: "branch 'main' (ours) and branch 'editor' (theirs) both modified document 1"
},
base: { _id: 1, doc: { _id: 1, author: 'neil', status: 'wip', title: 'Draft' } },
ours: { _id: 1, doc: { _id: 1, author: 'neil', status: 'wip', title: 'Launch Day' }, diffType: 'modified' },
theirs: { _id: 1, doc: { _id: 1, author: 'grace', status: 'wip', title: 'Draft' }, diffType: 'modified' }
}
],
ok: 1
}
You can see the whole story in the base, ours, and theirs documents. ours (the main
branch) changed the title, and theirs (the editor branch) changed the author. No single
field is in dispute, but documentTouched flagged it anyway because both branches wrote the
document.
Resolve the Conflict#
Since we actually want both edits, we resolve the conflict with a custom resolution: we hand
DumboDB the exact document we want, keeping the new title and the new author.
db.runCommand({
dumboResolveConflict: 1,
collection: "posts",
conflictId: "7lM6Ptl5Re/mF/24rnxqNA",
resolution: "custom",
value: { _id: 1, title: "Launch Day", author: "grace", status: "wip" }
})
{ ok: 1 }
With every conflict resolved, we tell the merge to finish up with the continue flag:
db.runCommand({ dumboMerge: 1, continue: 1 })
{
commitId: '1jvfavh54dhu0m1hvhqrkp9ajctrgs0q',
message: "Merge branch 'editor' into 'main'",
...
ok: 1
}
And the final document on main has both changes, exactly as we resolved it:
db.posts.findOne({ _id: 1 })
{ _id: 1, author: 'grace', status: 'wip', title: 'Launch Day' }
That’s the full loop: two branches wrote the same document, documentTouched refused to guess,
and we made the call ourselves. Swap in a different merge mode at collection creation and the
exact same edits would have been handled differently, all the way from a silent merge under
fieldDivergent to this hands-on resolution under documentTouched.
Update the Merge Mode#
The merge mode isn’t locked in at creation time. Just like a validator or any other collection
option, you can change it on an existing collection with collMod. Say we decide posts should
use the default fieldTouched behavior after all:
db.runCommand({ collMod: "posts", mergeMode: "fieldTouched" })
{ ok: 1 }
You can confirm the change with getCollectionInfos, which reports the mode in force:
db.getCollectionInfos({ name: "posts" })[0].options
{ mergeMode: 'fieldTouched' }
The new mode applies to every merge from that point forward. History already committed under the old mode is untouched.
Conclusion#
There you have it. More flexibility for application builders to leverage branching and merging
in their code. Different policies may make sense for each collection, so you can configure
each one individually. And like all collection modifications, collMod is used to update the setting
on an existing collection.
What’s next? We’re grinding on TLS support!
If you want to nerd out about version control databases, join the DoltHub Discord!