This project addresses the urgent need for accurate, inclusive, and accessible scientific communication by enabling multilingual access to scientific knowledge. In a global research landscape dominated by English, language barriers prevent many researchers and public audiences from fully engaging with scientific developments. By empowering users to access and share scientific information in their native languages, the project promotes linguistic equity, reduces information silos, and fosters a more inclusive and globally connected research environment.
Our core innovation is a terminology-aware machine translation (MT) framework tailored to the specific challenges of scientific texts, where terminological accuracy and contextual consistency are critical. Scientific language is dense and specialized, often resisting general-purpose translation tools. To overcome this, we will leverage Large Reasoning Models (LRMs)—an advanced category of Large Language Models (LLMs) that treat translation as a reasoning task. LRMs employ chain-of-thought prompting, self-correction mechanisms, and document-level understanding to ensure terminological precision and coherence across extended texts, maintaining the integrity and nuance of original scientific content.
To further enhance translation quality, we will build an integrated pipeline combining Quality Estimation (QE) and Automatic Post-Editing (APE). QE models will detect terminology-related errors and provide targeted feedback. This feedback will guide APE modules, which will apply reinforcement learning techniques—such as direct preference optimization—to refine translations. Crucially, insights from QE will not only improve post-editing but also feed back into the training of LRMs, creating a continuous loop of improvement. The pipeline will be trained on carefully curated corpora annotated with terminology errors to maximize relevance and effectiveness.
Beyond ensuring translation accuracy, the project will address the accessibility of scientific content for diverse audiences. Recognizing that scientific materials need adaptation for non-experts, we will implement post-translation text augmentation strategies, including simplification and explanatory techniques. These adaptations will help tailor scientific outputs for education, public engagement, and citizen science, broadening the impact and reusability of translated materials.
The project focuses on five languages: English, Spanish, Catalan, Estonian, and Irish. These represent a balance of well-resourced and under-resourced languages, allowing us to tackle translation challenges across different linguistic contexts. The Life Sciences domain has been chosen for its societal relevance and demand for clear, accurate communication. Real-world validation will be ensured through pilot collaborations with the Centre de Recerca Genòmica (CRG) in Barcelona, the Institute of Family Medicine and Public Health in Tartu, and Conradh na Gaeilge, providing direct user feedback and practical testing scenarios.
Bridging advances in MT, NLP, and scientific expertise, the project will apply a thorough evaluation strategy combining adapted automatic metrics with human assessments involving domain experts and general users. By the project’s conclusion, we aim to deliver a system validated at Technology Readiness Level (TRL) 5–6, capable of substantially improving the accuracy, inclusivity, and trustworthiness of scientific translation. In doing so, the project will contribute to a more equitable and accessible scientific ecosystem, empowering broader participation in global research and facilitating the free flow of knowledge across languages.
Start date: (36 months)
Funding support: 967 291 €