Skip to main content

This project aims to develop a cross-lingual information retrieval system tailored to scientific literature in under-represented languages. It addresses critical challenges in language technologies, including the absence of annotated scientific datasets, the lack of language-specific models for less-resourced languages, and the dominance of English in scientific communication. The project will focus on constructing comparable corpora in targeted domains (linguistics, medicine, mathematics, geography, and jurisprudence), training language models adapted to scientific discourse, and aligning them into a shared multilingual embedding space. These models will power a retrieval system that enables structured information access across languages, supported by retrieval-augmented generation (graph-RAG) and multi-agent architectures.
Scientifically, the project introduces a novel approach to multilingual scientific computing by bridging language gaps at both the formal and conceptual levels. It investigates the structure of scientific discourse across typologically and sociolinguistically diverse languages, contributing to both NLP and intellectual history. The system will support linguistic self-representation in scientific contexts, enrich terminological resources, and provide tools for analyzing the diffusion and transformation of knowledge across cultures and time. Methodologies will ensure reproducibility and rigorous evaluation, laying the groundwork for inclusive, linguistically diverse knowledge systems.
Practically, the project will result in: (1) curated datasets in the targeted scientific domains for seven languages – Belarusian, Estonian, Punjabi, Slovak, Taiwanese (Tâigí), Ukrainian, Yiddish; (2) encoder and decoder language models fine-tuned for scientific texts in each selected language; (3) shared multilingual vector spaces for scientific information alignment; (4) an interactive AI-based platform for scientific information retrieval and generation; and (5) benchmarks and evaluation protocols suitable for multilingual scientific discourse.
The system developed will not only facilitate cross-lingual research but also support terminology development, translation practices, and multilingual education. It will provide open access tools to educators, researchers, and policy makers in communities historically excluded from global scientific discourse. Through documentation and open-source sharing, the project promotes replicability and adoption across other low-resource contexts.
The project directly addresses the objectives of the “Science in Your Own Language” (SOL) call by developing technologies that support scientific work in less-resourced languages. It aligns with the SOL call’s emphasis on democratizing access to knowledge by enabling researchers to write, read, and retrieve scientific content in their own languages. The project builds upon cross-disciplinary collaboration among experts in AI, NLP, linguistics, and intellectual history, ensuring that both the technical and humanistic dimensions of scientific communication are taken into account.
By integrating AI methods with humanities-driven insight into the dynamics of scientific knowledge, this project contributes to the development of equitable and inclusive scientific infrastructures. It aims to empower speakers of under-represented languages to participate in global knowledge production and to ensure that scientific advancement is truly multilingual and culturally grounded.
We hope that this initiative will serve as a significant step toward realizing the vision of “Science in Your Own Language,” contributing to both the revitalization of minoritized languages and the diversification of AI. The project demonstrates how under-resourced languages can enrich global scientific ecosystems and improve the adaptability and fairness of AI technologies.

Call Topic: Science in your Own Language (SOL), Call 2025
Start date: (36 months)
Funding support: 661 893 €