Last week, I wrote about the TEST Lab at National University of Singapore and an overview of the work they’ve done with Dolt over the years. Today, I will be reviewing “Scaling Automated Database System Testing” by Suyang Zhong of the TEST Lab. This paper was published earlier this year as part of ACM’s International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ‘26) and was actually cited by “Automated Discovery of Test Oracles for Database Management Systems Using LLMs”, which described Argus.
SQLancer++#
The paper primarily outlines SQLancer++, a testing platform that aims to scale SQLancer. SQLancer is a database testing framework that focuses on finding logic bugs via test oracles and was developed by the TEST Lab in the past. One of SQLancer’s approaches is to generate two logically equivalent queries and making sure a database returns the same result for both queries. Logic bugs, or correctness bugs as we call them at DoltHub, are dangerous in database systems, because rather than erroring out, the database silently returns incorrect results which can be hard to detect and can propagate incorrect data throughout a system. One of the drawbacks of the original SQLancer is that it requires a custom implementation for each database to integrate with it, since each database has different supported and unsupported features. This can be tedious for database developers who may want to use SQLancer and does not easily scale. SQLancer++ tries to tackle this scaling problem by creating an automated system that uses an adaptive statement generator.
SQLancer++’s adaptive statement generator generates random SQL queries and runs them against a database for an empirically-determined number of iterations. It then uses Bayesian inference to identify which queries failed because the features tested were unsupported and which failed because due to an actual bug. Queries involving unsupported features are subsequently suppressed, allowing only supported features to be tested.
So does it use AI?#
That’s the big question on everyone’s minds in 2026. In the colloquial 2026 sense of the term, no, SQLancer++ does not use “AI”, in that it does not use a large language model (LLM). However, SQLancer++ does use Bayesian inference, a statistical method used in machine learning, so it does fall under the umbrella of “old-school” artificial intelligence. And even though SQLancer++ generates SQL queries, these queries seem to be created using hard-coded rules and randomness as opposed to using a generative neural network. So again, it’s arguably generative AI in the old-fashioned sense, but it’s not what we would typically consider generative AI in our current AI era.
Impact on Dolt#
SQLancer++ tested 18 databases total. Many of these databases were ones that had already been tested with the original SQLancer, but Dolt was added because of its many stars on GitHub. At the time the paper was written, Dolt had 16.9k stars – we now have 24.5k, and you can help make that an even bigger number.
While the paper says SQLancer++ found 28 bugs in Dolt, I was actually able to find 29 GitHub issues filed by Suyang. 9 of those bugs involved a panic, making them particularly serious issues for us. These bugs were all filed between November 2023 and February 2024, before I started working at DoltHub. And while the paper says all but one of the bugs were fixed, every bug has now been fixed, and the final bug was fixed a year ago by… me (wow). I clicked through several of these issues, and they’re mostly related to type coercion, which is a very well-known issue to us at this point.
Not only did SQLancer++ find a lot of important bugs in Dolt, it also laid the foundation for later TEST Lab projects with Dolt, like Zhaokun Xiang’s work on testing joins and Yibo Dong’s WINdow Equivalence (WINE), as well as Argus. In the upcoming weeks, I will be blogging about those projects as well.
We here at DoltHub are honored that Dolt is considered important enough to be included in research studies and super grateful for all the bugs the TEST Lab has found for us. Are you an academic researcher working on a project involving Dolt or another DoltHub database? Join our Discord community – we’d love to hear from you.