e-Informatica Software Engineering Journal Text-to-SQL Dataset Quality Assessment: A Multi-Dimensional Validation Framework

Text-to-SQL Dataset Quality Assessment: A Multi-Dimensional Validation Framework

  1. Volodymyr Sabadosh, Vladyslav Kotsovsky, Text-to-SQL Dataset Quality Assessment: A Multi-Dimensional Validation Framework, In: e-Informatica Software Engineering Journal, vol. 21, no. 1, pp. 270102, 2027. DOI: 10.37190/e-inf270102.
Download article (PDF) Downolad BibTeX file

Authors

Volodymyr Sabadosh, Vladyslav Kotsovsky

Abstract

Context: LLM-based Text-to-SQL has advanced quickly, but benchmark and training datasets may contain defects that distort evaluation and fine-tuning. Prior audits remain fragmented, addressing dimensions in isolation.

Objective: This paper proposes text2sql-dataset-analyzer, an open-source framework that audits Text-to-SQL datasets across five complementary quality dimensions within a single reproducible pipeline.

Method: The framework covers five dimensions: database schema integrity, SQL syntactic structure and complexity, execution testing, antipattern detection, and semantic correspondence. Semantic correspondence is evaluated via an LLM-as-a-judge committee with majority voting. An analytical database stores the resulting metrics for direct querying and Markdown report generation, while structured JSONL output records per-item annotations.

Results: Auditing all 11,840 Spider 1.0 examples reveals quality issues despite 99.97% of queries passing execution checks. Schema and data checks identify 51 structural foreign-key errors and 41,927 row-level referential-integrity violations, 41,913 of them in just three of 206 databases. At the item level, the audit flags unanimously Incorrect (8-10%), disputed (23-27%), and Unanswerable NL-SQL pairs, together with SQL antipatterns associated with correctness, robustness, and portability concerns. Manual review of 367 flagged items confirms genuine defects in 263, including 26 consensus Unanswerable questions and all 17 Cartesian-product join bugs.

Conclusions: Multi-dimensional validation exposes dataset defects missed by executability-based checks. We release the open-source framework and a prioritized remediation roadmap for a popular Text-to-SQL benchmark. The downstream impact of these defects on model training and benchmark scores remains future work.

Keywords

Text-to-SQL system, dataset quality, LLM, SQL validation, semantic evaluation, empirical software engineering, benchmark auditing, LLM-as-a-judge

References

1. J. Fürst, C. Kosten, F. Nooralahzadeh, Y. Zhang, and K. Stockinger, “Evaluating the data model robustness of Text-to-SQL systems based on real user queries,” in Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025) , 2025, pp. 158–170. [Online]. https://doi.org/10.48786/edbt.2025.13

2. L. Shi, Z. Tang, N. Zhang, X. Zhang, and Z. Yang, “A survey on employing large language models for Text-to-SQL tasks,” ACM Computing Surveys , Vol. 58, No. 2, 2025, pp. 1–37. [Online]. https://doi.org/10.1145/3737873

3. D. Gao, H. Wang, Y. Li, X. Sun, Y. Qian et al., “Text-to-SQL empowered by large language models: A benchmark evaluation,” Proceedings of the VLDB Endowment , Vol. 17, No. 5, 2024, pp. 1132–1145. [Online]. https://doi.org/10.14778/3641204.3641221

4. Yale LILY Group, “Spider 1.0 – leaderboard,” 2025, accessed: 2025-11-08. [Online]. https://yale-lily.github.io/spider

5. V. Shkapenyuk, D. Srivastava, T. Johnson, and P. Ghane, “Automatic metadata extraction for Text-to-SQL,” CoRR , Vol. abs/2505.19988, 2025. [Online]. https://doi.org/10.48550/arXiv.2505.19988

6. BIRD Benchmark Team, “BIRD – leaderboard,” 2025, accessed: 2025-11-08. [Online]. https://bird-bench.github.io/

7. T. Yu, R. Zhang, K. Yang, M. Yasunaga, D. Wang et al., “Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and Text-to-SQL task,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Brussels, Belgium: Association for Computational Linguistics, 2018, pp. 3911–3921. [Online]. https://doi.org/10.18653/v1/D18-1425

8. Snowflake Inc., “Using Snowflake Copilot inline,” Snowflake Documentation, 2025, accessed: 2025-11-08. [Online]. https://docs.snowflake.com/en/user-guide/snowflake-copilot-inline

9. Microsoft, “What is an AI/BI Genie space,” Microsoft Learn (Azure Databricks Documentation), 2025, accessed: 2025-11-08. [Online]. https://learn.microsoft.com/en-us/azure/databricks/genie/

10. Microsoft, “Copilot in Fabric in the SQL database workload,” Microsoft Learn, 2025, accessed: 2025-11-08. [Online]. https://learn.microsoft.com/en-us/fabric/database/sql/copilot-sql-database

11. V. Zhong, C. Xiong, and R. Socher, “Seq2SQL: Generating structured queries from natural language using reinforcement learning,” CoRR , Vol. abs/1709.00103, 2017. [Online]. https://doi.org/10.48550/arXiv.1709.00103

12. J. Li, B. Hui, G. Qu, J. Yang, B. Li et al., “Can LLM already serve as a database interface? a BIg bench for large-scale database grounded Text-to-SQLs,” in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track , 2023, pp. 42 330–42 357. [Online]. https://proceedings.neurips.cc/paper _files/paper/2023/hash/83fc8fab1710363050bbd1d4b8cc0021-Abstract-Datasets _and _Benchmarks.html

13. Gretel.ai, “Synthetic Text-to-SQL dataset (Gretel-Synth),” Hugging Face Datasets, 2024, accessed: 2025-11-08. [Online]. https://huggingface.co/datasets/gretelai/synthetic _text _to _sql

14. H. Li, S. Wu, X. Zhang, X. Huang, J. Zhang et al., “OmniSQL: Synthesizing high-quality Text-to-SQL data at scale,” Proceedings of the VLDB Endowment , Vol. 18, No. 11, 2025, pp. 4695–4709. [Online]. https://doi.org/10.14778/3749646.3749723

15. S. Chang, J. Wang, M. Dong, L. Pan, H. Zhu et al., “Dr.Spider: A diagnostic evaluation benchmark towards Text-to-SQL robustness,” in Proceedings of the 11th International Conference on Learning Representations (ICLR) , 2023. [Online]. https://openreview.net/forum?id=Wc5bmZZU9cy

16. Z. Yao, G. Sun, Ł. Borchmann, Z. Shen, M. Deng et al., “Arctic-Text2SQL-R1: Simple rewards, strong reasoning in Text-to-SQL,” in Findings of the Association for Computational Linguistics: ACL 2026 . San Diego, California, United States: Association for Computational Linguistics, 2026, pp. 26 966–26 995. [Online]. https://doi.org/10.18653/v1/2026.findings-acl.1345

17. T. Jin, Y. Choi, Y. Zhu, and D. Kang, “Pervasive annotation errors break Text-to-SQL benchmarks and leaderboards,” Proceedings of the VLDB Endowment , Vol. 19, No. 5, 2026, pp. 931–944. [Online]. https://doi.org/10.14778/3796195.3796206

18. A. Mitsopoulou and G. Koutrika, “Analysis of Text-to-SQL benchmarks: Limitations, challenges and opportunities,” in Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025) . OpenProceedings.org, 2025, pp. 199–212. [Online]. https://doi.org/10.48786/edbt.2025.16

19. P. Pandey, D. Patel, S. Mandvikar, and N. Kota, “Ensuring data accuracy in Text-to-SQL systems: A comprehensive validation framework,” International Journal of Computer Trends and Technology , Vol. 72, No. 12, 2024, pp. 17–24. [Online]. https://doi.org/10.14445/22312803/IJCTT-V72I12P103

20. X. Liu, S. Shen, B. Li, N. Tang, and Y. Luo, “NL2SQL-BUGs: A benchmark for detecting semantic errors in NL2SQL translation,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’25) , 2025, pp. 5662–5673. [Online]. https://doi.org/10.1145/3711896.3737427

21. L. Zheng, W.L. Chiang, Y. Sheng, S. Zhuang, Z. Wu et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track , 2023.

22. C.M. Chan, W. Chen, Y. Su, J. Yu, W. Xue et al., “ChatEval: Towards better LLM-based evaluators through multi-agent debate,” in Proceedings of the 12th International Conference on Learning Representations (ICLR) , 2024.

23. B. Karwin, SQL Antipatterns: Avoiding the Pitfalls of Database Programming . Pragmatic Bookshelf, 2010.

24. P. Dintyala, A. Narechania, and J. Arulraj, “SQLCheck: Automated detection and diagnosis of SQL anti-patterns,” in Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data , 2020, pp. 2331–2345. [Online]. https://doi.org/10.1145/3318464.3389754

25. Z. Abedjan, L. Golab, and F. Naumann, “Profiling relational data: A survey,” The VLDB Journal , Vol. 24, No. 4, 2015, pp. 557–581.

26. M. Memari, S. Link, and G. Dobbie, “SQL data profiling of foreign keys,” in Proceedings of the 34th International Conference on Conceptual Modeling (ER 2015) , 2015, pp. 229–243.

27. S. Fatehi, “SchemaCrawler: Free database schema discovery and comprehension tool,” 2026, accessed: 2026-06-06. [Online]. https://www.schemacrawler.com/

28. Python Software Foundation, “Python 3.11 documentation,” 2025, accessed: 2025-11-08. [Online]. https://docs.python.org/3.11/

29. T. Mao, “SQLGlot: Python SQL parser and transpiler,” 2025, accessed: 2025-11-08. [Online]. https://github.com/tobymao/sqlglot

30. M. Bayer et al., “SQLAlchemy 2.0 documentation,” 2025, accessed: 2025-11-08. [Online]. https://docs.sqlalchemy.org/en/20/

31. DuckDB Foundation, “DuckDB documentation,” 2025, accessed: 2025-11-08. [Online]. https://duckdb.org/docs/stable/

32. R. Mogylatov, “Dependency Injector: Dependency injection framework for Python,” 2025, accessed: 2025-11-08. [Online]. https://python-dependency-injector.ets-labs.org/

33. PyYAML Developers, “PyYAML – YAML parser and emitter for Python,” 2025, accessed: 2025-11-08. [Online]. https://pyyaml.org/

34. C.J. Date, SQL and Relational Theory: How to Write Accurate SQL Code , 2nd ed. O’Reilly Media, 2012.

35. R. Ramakrishnan and J. Gehrke, Database Management Systems , 3rd ed. McGraw-Hill, 2003.

36. Information technology – Database languages – SQL – Part 2: Foundation (SQL/Foundation) , ISO/IEC Std. 9075-2:2016, 2016. [Online]. https://www.iso.org/standard/63556.html

37. SQLite, “SELECT – section 2.5: Bare columns in an aggregate query,” 2025, accessed: 2025-11-08. [Online]. https://sqlite.org/lang _select.html#bare _columns _in _an _aggregate _query

38. PostgreSQL Global Development Group, “The GROUP BY and HAVING clauses,” 2024, accessed: 2025-11-08. [Online]. https://www.postgresql.org/docs/current/queries-table-expressions.html#QUERIES-GROUP

39. A. Molinaro and R. de Graaf, SQL Cookbook: Query Solutions and Techniques for All SQL Users , 2nd ed. O’Reilly Media, 2020.

40. H. Chen and S. Goldfarb-Tarrant, “Safer or luckier? LLMs as safety evaluators are not robust to artifacts,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2025, pp. 19 750–19 766. [Online]. https://aclanthology.org/2025.acl-long.970/

41. SQLite, “Foreign key support,” 2025, accessed: 2025-11-08. [Online]. https://sqlite.org/foreignkeys.html

42. OpenAI, “Models documentation (API reference),” 2025, accessed: 2025-11-08. [Online]. https://platform.openai.com/docs/models

43. Google, “Gemini API documentation,” 2025, accessed: 2025-11-08. [Online]. https://ai.google.dev/gemini-api/docs

44. Oracle Corporation, “MySQL handling of GROUP BY,” 2024, accessed: 2025-11-08. [Online]. https://dev.mysql.com/doc/refman/8.0/en/group-by-handling.html

45. V.M. Sabadosh and V.M. Kotsovsky, “Automated detection and comparative analysis of SQL antipatterns in Text-to-SQL datasets,” Ukrainian Journal of Information Technology , Vol. 8, No. 1, 2026, pp. 135–143. [Online]. https://doi.org/10.23939/ujit2026.01.135

46. J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement , Vol. 20, No. 1, 1960, pp. 37–46.

47. J.R. Landis and G.G. Koch, “The measurement of observer agreement for categorical data,” Biometrics , Vol. 33, No. 1, 1977, pp. 159–174.

48. L.D. Brown, T.T. Cai, and A. DasGupta, “Interval estimation for a binomial proportion,” Statistical Science , Vol. 16, No. 2, 2001, pp. 101–133. [Online]. https://doi.org/10.1214/ss/1009213286

49. V.M. Kotsovsky and A. Batyuk, “Towards the design of bithreshold ANN regressor,” in Proceedings of the 19th IEEE International Scientific and Technical Conference on Computer Sciences and Information Technologies (CSIT 2024) . Lviv, Ukraine: IEEE, 2024, pp. 1–4.

  • 2026-08-28

Design © 2015-2026 by e-Informatyka.pl

Built on WordPress Theme: Mediaphase Lite by ThemeFurnace.