← Writing

Large Language Models and Synthetic Data for Monitoring Dataset Mentions in Research Papers

Presented at the ICLR 2025 Workshop. This work explores how large language models combined with synthetic training data can automatically detect dataset mentions in research papers and classify how those datasets are used.

arXiv:2502.10263 · ICLR 2025 Workshop