|
Authors
Volodymyr Sabadosh, Vladyslav KotsovskyAbstract
Context: LLM-based Text-to-SQL has advanced quickly, but benchmark and training datasets may contain defects that distort evaluation and fine-tuning. Prior audits remain fragmented, addressing dimensions in isolation.
Objective: This paper proposes text2sql-dataset-analyzer, an open-source framework that audits Text-to-SQL datasets across five complementary quality dimensions within a single reproducible pipeline.
Method: The framework covers five dimensions: database schema integrity, SQL syntactic structure and complexity, execution testing, antipattern detection, and semantic correspondence. Semantic correspondence is evaluated via an LLM-as-a-judge committee with majority voting. An analytical database stores the resulting metrics for direct querying and Markdown report generation, while structured JSONL output records per-item annotations.
Results: Auditing all 11,840 Spider 1.0 examples reveals quality issues despite 99.97% of queries passing execution checks. Schema and data checks identify 51 structural foreign-key errors and 41,927 row-level referential-integrity violations, 41,913 of them in just three of 206 databases. At the item level, the audit flags unanimously Incorrect (8-10%), disputed (23-27%), and Unanswerable NL-SQL pairs, together with SQL antipatterns associated with correctness, robustness, and portability concerns. Manual review of 367 flagged items confirms genuine defects in 263, including 26 consensus Unanswerable questions and all 17 Cartesian-product join bugs.
Conclusions: Multi-dimensional validation exposes dataset defects missed by executability-based checks. We release the open-source framework and a prioritized remediation roadmap for a popular Text-to-SQL benchmark. The downstream impact of these defects on model training and benchmark scores remains future work.
Keywords
Text-to-SQL system, dataset quality, LLM, SQL validation, semantic evaluation, empirical software engineering, benchmark auditing, LLM-as-a-judgeReferences
1. J. Fürst, C. Kosten, F. Nooralahzadeh, Y. Zhang, and K. Stockinger, “Evaluating the data model robustness of Text-to-SQL systems based on real user queries,” in Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025) , 2025, pp. 158–170. [Online]. https://doi.org/10.48786/edbt.2025.13
2. L. Shi, Z. Tang, N. Zhang, X. Zhang, and Z. Yang, “A survey on employing large language models for Text-to-SQL tasks,” ACM Computing Surveys , Vol. 58, No. 2, 2025, pp. 1–37. [Online]. https://doi.org/10.1145/3737873
3. D. Gao, H. Wang, Y. Li, X. Sun, Y. Qian et al., “Text-to-SQL empowered by large language models: A benchmark evaluation,” Proceedings of the VLDB Endowment , Vol. 17, No. 5, 2024, pp. 1132–1145. [Online]. https://doi.org/10.14778/3641204.3641221
4. Yale LILY Group, “Spider 1.0 – leaderboard,” 2025, accessed: 2025-11-08. [Online]. https://yale-lily.github.io/spider
5. V. Shkapenyuk, D. Srivastava, T. Johnson, and P. Ghane, “Automatic metadata extraction for Text-to-SQL,” CoRR , Vol. abs/2505.19988, 2025. [Online]. https://doi.org/10.48550/arXiv.2505.19988
6. BIRD Benchmark Team, “BIRD – leaderboard,” 2025, accessed: 2025-11-08. [Online]. https://bird-bench.github.io/
7. T. Yu, R. Zhang, K. Yang, M. Yasunaga, D. Wang et al., “Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and Text-to-SQL task,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Brussels, Belgium: Association for Computational Linguistics, 2018, pp. 3911–3921. [Online]. https://doi.org/10.18653/v1/D18-1425
8. Snowflake Inc., “Using Snowflake Copilot inline,” Snowflake Documentation, 2025, accessed: 2025-11-08. [Online]. https://docs.snowflake.com/en/user-guide/snowflake-copilot-inline
9. Microsoft, “What is an AI/BI Genie space,” Microsoft Learn (Azure Databricks Documentation), 2025, accessed: 2025-11-08. [Online]. https://learn.microsoft.com/en-us/azure/databricks/genie/
10. Microsoft, “Copilot in Fabric in the SQL database workload,” Microsoft Learn, 2025, accessed: 2025-11-08. [Online]. https://learn.microsoft.com/en-us/fabric/database/sql/copilot-sql-database
11. V. Zhong, C. Xiong, and R. Socher, “Seq2SQL: Generating structured queries from natural language using reinforcement learning,” CoRR , Vol. abs/1709.00103, 2017. [Online]. https://doi.org/10.48550/arXiv.1709.00103
12. J. Li, B. Hui, G. Qu, J. Yang, B. Li et al., “Can LLM already serve as a database interface? a BIg bench for large-scale database grounded Text-to-SQLs,” in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track , 2023, pp. 42 330–42 357. [Online]. https://proceedings.neurips.cc/paper _files/paper/2023/hash/83fc8fab1710363050bbd1d4b8cc0021-Abstract-Datasets _and _Benchmarks.html
13. Gretel.ai, “Synthetic Text-to-SQL dataset (Gretel-Synth),” Hugging Face Datasets, 2024, accessed: 2025-11-08. [Online]. https://huggingface.co/datasets/gretelai/synthetic _text _to _sql
14. H. Li, S. Wu, X. Zhang, X. Huang, J. Zhang et al., “OmniSQL: Synthesizing high-quality Text-to-SQL data at scale,” Proceedings of the VLDB Endowment , Vol. 18, No. 11, 2025, pp. 4695–4709. [Online]. https://doi.org/10.14778/3749646.3749723
15. S. Chang, J. Wang, M. Dong, L. Pan, H. Zhu et al., “Dr.Spider: A diagnostic evaluation benchmark towards Text-to-SQL robustness,” in Proceedings of the 11th International Conference on Learning Representations (ICLR) , 2023. [Online]. https://openreview.net/forum?id=Wc5bmZZU9cy
16. Z. Yao, G. Sun, Ł. Borchmann, Z. Shen, M. Deng et al., “Arctic-Text2SQL-R1: Simple rewards, strong reasoning in Text-to-SQL,” in Findings of the Association for Computational Linguistics: ACL 2026 . San Diego, California, United States: Association for Computational Linguistics, 2026, pp. 26 966–26 995. [Online]. https://doi.org/10.18653/v1/2026.findings-acl.1345
17. T. Jin, Y. Choi, Y. Zhu, and D. Kang, “Pervasive annotation errors break Text-to-SQL benchmarks and leaderboards,” Proceedings of the VLDB Endowment , Vol. 19, No. 5, 2026, pp. 931–944. [Online]. https://doi.org/10.14778/3796195.3796206
18. A. Mitsopoulou and G. Koutrika, “Analysis of Text-to-SQL benchmarks: Limitations, challenges and opportunities,” in Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025) . OpenProceedings.org, 2025, pp. 199–212. [Online]. https://doi.org/10.48786/edbt.2025.16
19. P. Pandey, D. Patel, S. Mandvikar, and N. Kota, “Ensuring data accuracy in Text-to-SQL systems: A comprehensive validation framework,” International Journal of Computer Trends and Technology , Vol. 72, No. 12, 2024, pp. 17–24. [Online]. https://doi.org/10.14445/22312803/IJCTT-V72I12P103
20. X. Liu, S. Shen, B. Li, N. Tang, and Y. Luo, “NL2SQL-BUGs: A benchmark for detecting semantic errors in NL2SQL translation,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’25) , 2025, pp. 5662–5673. [Online]. https://doi.org/10.1145/3711896.3737427
21. L. Zheng, W.L. Chiang, Y. Sheng, S. Zhuang, Z. Wu et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track , 2023.
22. C.M. Chan, W. Chen, Y. Su, J. Yu, W. Xue et al., “ChatEval: Towards better LLM-based evaluators through multi-agent debate,” in Proceedings of the 12th International Conference on Learning Representations (ICLR) , 2024.
23. B. Karwin, SQL Antipatterns: Avoiding the Pitfalls of Database Programming . Pragmatic Bookshelf, 2010.
24. P. Dintyala, A. Narechania, and J. Arulraj, “SQLCheck: Automated detection and diagnosis of SQL anti-patterns,” in Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data , 2020, pp. 2331–2345. [Online]. https://doi.org/10.1145/3318464.3389754
25. Z. Abedjan, L. Golab, and F. Naumann, “Profiling relational data: A survey,” The VLDB Journal , Vol. 24, No. 4, 2015, pp. 557–581.
26. M. Memari, S. Link, and G. Dobbie, “SQL data profiling of foreign keys,” in Proceedings of the 34th International Conference on Conceptual Modeling (ER 2015) , 2015, pp. 229–243.
27. S. Fatehi, “SchemaCrawler: Free database schema discovery and comprehension tool,” 2026, accessed: 2026-06-06. [Online]. https://www.schemacrawler.com/
28. Python Software Foundation, “Python 3.11 documentation,” 2025, accessed: 2025-11-08. [Online]. https://docs.python.org/3.11/
29. T. Mao, “SQLGlot: Python SQL parser and transpiler,” 2025, accessed: 2025-11-08. [Online]. https://github.com/tobymao/sqlglot
30. M. Bayer et al., “SQLAlchemy 2.0 documentation,” 2025, accessed: 2025-11-08. [Online]. https://docs.sqlalchemy.org/en/20/
31. DuckDB Foundation, “DuckDB documentation,” 2025, accessed: 2025-11-08. [Online]. https://duckdb.org/docs/stable/
32. R. Mogylatov, “Dependency Injector: Dependency injection framework for Python,” 2025, accessed: 2025-11-08. [Online]. https://python-dependency-injector.ets-labs.org/
33. PyYAML Developers, “PyYAML – YAML parser and emitter for Python,” 2025, accessed: 2025-11-08. [Online]. https://pyyaml.org/
34. C.J. Date, SQL and Relational Theory: How to Write Accurate SQL Code , 2nd ed. O’Reilly Media, 2012.
35. R. Ramakrishnan and J. Gehrke, Database Management Systems , 3rd ed. McGraw-Hill, 2003.
36. Information technology – Database languages – SQL – Part 2: Foundation (SQL/Foundation) , ISO/IEC Std. 9075-2:2016, 2016. [Online]. https://www.iso.org/standard/63556.html
37. SQLite, “SELECT – section 2.5: Bare columns in an aggregate query,” 2025, accessed: 2025-11-08. [Online]. https://sqlite.org/lang _select.html#bare _columns _in _an _aggregate _query
38. PostgreSQL Global Development Group, “The GROUP BY and HAVING clauses,” 2024, accessed: 2025-11-08. [Online]. https://www.postgresql.org/docs/current/queries-table-expressions.html#QUERIES-GROUP
39. A. Molinaro and R. de Graaf, SQL Cookbook: Query Solutions and Techniques for All SQL Users , 2nd ed. O’Reilly Media, 2020.
40. H. Chen and S. Goldfarb-Tarrant, “Safer or luckier? LLMs as safety evaluators are not robust to artifacts,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2025, pp. 19 750–19 766. [Online]. https://aclanthology.org/2025.acl-long.970/
41. SQLite, “Foreign key support,” 2025, accessed: 2025-11-08. [Online]. https://sqlite.org/foreignkeys.html
42. OpenAI, “Models documentation (API reference),” 2025, accessed: 2025-11-08. [Online]. https://platform.openai.com/docs/models
43. Google, “Gemini API documentation,” 2025, accessed: 2025-11-08. [Online]. https://ai.google.dev/gemini-api/docs
44. Oracle Corporation, “MySQL handling of GROUP BY,” 2024, accessed: 2025-11-08. [Online]. https://dev.mysql.com/doc/refman/8.0/en/group-by-handling.html
45. V.M. Sabadosh and V.M. Kotsovsky, “Automated detection and comparative analysis of SQL antipatterns in Text-to-SQL datasets,” Ukrainian Journal of Information Technology , Vol. 8, No. 1, 2026, pp. 135–143. [Online]. https://doi.org/10.23939/ujit2026.01.135
46. J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement , Vol. 20, No. 1, 1960, pp. 37–46.
47. J.R. Landis and G.G. Koch, “The measurement of observer agreement for categorical data,” Biometrics , Vol. 33, No. 1, 1977, pp. 159–174.
48. L.D. Brown, T.T. Cai, and A. DasGupta, “Interval estimation for a binomial proportion,” Statistical Science , Vol. 16, No. 2, 2001, pp. 101–133. [Online]. https://doi.org/10.1214/ss/1009213286
49. V.M. Kotsovsky and A. Batyuk, “Towards the design of bithreshold ANN regressor,” in Proceedings of the 19th IEEE International Scientific and Technical Conference on Computer Sciences and Information Technologies (CSIT 2024) . Lviv, Ukraine: IEEE, 2024, pp. 1–4.







