CEFR-Cymraeg: A Dataset and Baseline Models for Welsh Language Proficiency Assessment

April 1, 2025 · 1 min read
projects

Welsh Government-funded project building the first CEFR-annotated dataset and baseline models for automatically assessing the proficiency level of Welsh texts.

Funder: Welsh Government
Period: April 2025 – March 2026
Role: Co-Investigator
Research theme: Welsh Language Technology & Multilingual NLP


Automated tools for assessing the proficiency level of Welsh learner texts have been lacking, limiting support for Welsh-language education and assessment. This project addressed that gap with two objectives:

  1. Collect and annotate a comprehensive dataset of Welsh texts classified by CEFR level
  2. Develop baseline machine learning models for predicting the CEFR level of Welsh texts

The result, CEFR-Cymraeg, is the first CEFR-annotated language proficiency dataset for Welsh (A1–B2), together with baseline classification models, enabling automated language proficiency assessment for Welsh learners.


Selected Publications

Fernando Alva-Manchego
Authors
Researcher in Natural Language Processing
My research interests include text simplification, readability assessment, multilingual NLP, Welsh language technology, and NLP for education and social care.