Skip to main content

Despite English being adopted as lingua franca in almost every scientific discipline, there is a remarkable amount of scientific production in other languages (e.g., in biomedicine in French or in ecology in Spanish). The use of other languages allows for the inclusion of diverse cultural and epistemological perspectives, which enriches the process of knowledge generation. For instance, local languages ​​can offer unique concepts and ways of understanding natural phenomena that are not present in dominant languages ​​such as English. However, the predominance of English limits the visibility of research conducted in other languages, which in turn affects international collaboration and the dissemination of results. The lack of resources to access scientific results in multiple languages hampers access to global knowledge and de-incentivize authors to publish in their own language.

The primary objective of this project is to facilitate smooth and interoperable access to mono/multilingual scientific and technological data hubs and repositories for stakeholders to engage with them in their native language, to make research more accessible and inclusive. These tools will be showcased through their implementation in the climatology use case (climate extreme events). 

In order to achieve such a goal, neuro-symbolic AI approaches will be developed and applied, in combination with a number of linguistic services, for: (a) The extraction and semantic annotation of scientific data and documents in several languages, to build a knowledge graph of interlinked multilingual scientific knowledge for the climatology domain, (b) Cross-lingual access to such knowledge graph and their underlying annotated data independently of the natural language used, and (c) translation of the retrieved data and documents into the user language.

Call Topic: Science in your Own Language (SOL), Call 2025
Start date: (36 months)
Funding support: 698 688 €