CEFR-Cymraeg: A Dataset and Baseline Models for Welsh Language Proficiency Assessment
April 1, 2025
·
1 min read
Welsh Government-funded project building the first CEFR-annotated dataset and baseline models for automatically assessing the proficiency level of Welsh texts.
Funder: Welsh Government
Period: April 2025 – March 2026
Role: Co-Investigator
Research theme: Welsh Language Technology & Multilingual NLP
Automated tools for assessing the proficiency level of Welsh learner texts have been lacking, limiting support for Welsh-language education and assessment. This project addressed that gap with two objectives:
- Collect and annotate a comprehensive dataset of Welsh texts classified by CEFR level
- Develop baseline machine learning models for predicting the CEFR level of Welsh texts
The result, CEFR-Cymraeg, is the first CEFR-annotated language proficiency dataset for Welsh (A1–B2), together with baseline classification models, enabling automated language proficiency assessment for Welsh learners.
