Close

AI semantics for biomedical data integration

James McLaughlin, Aleix Puig-Barbe, Arwa Ibrahim, Diego Pava, Zoe May Pendlington, Nicolas Matentzoglu, Elliot Sollis, Amy Foreman, Robert Wilson, Federico Lopez Gomez, Laura Harris, Yeyejide Adeleye, Satwant Kaur, Birgit Meldal, Damian Smedley, Helen Parkinson

Posted on: 5 October 2026

Preprint posted on 3 September 2026

Crouching AI, hidden connections; The preprint authors assemble and deploy an artificial intelligence (AI)-based semantics workflow to facilitate data integration across heterogeneous databases.

Selected by Theodora Stougiannou

Crouching AI, hidden connections; The preprint authorsassemble and deploy an artificial intelligence (AI)-based semantics workflow to facilitate data integration across heterogeneous databases.

 

Setting the scene: What are the authors looking at?

 

The authors of this preprint investigated how to enable biomedical data integration across multimodal datasets, diverse model organisms and different biological scales; to evaluate a research question, in practice, knowledge across different databases must be integrated. [1].

 

Databases, including those applied for the biomedical sciences, can have heterogenous schemas and annotation practices; furthermore, they may often have incompatible application programming interfaces (API). These characteristics render interoperability between different biomedical databases difficult, especially in an era when biomedical research requires the integration of these heterogenous databases, for example, interconnection between human Genome Wide Association Study (GWAS) study catalogs, i.e., the GWAS Catalog (NHGRI-EBI GWAS Catalog), and associated mouse research databases, i.e. the International Mouse phenotyping Consortium (IMPC) [1].

 

Bridging the gap: bring a dreaming LLM back to scientific reality

 

Large language models (LLM) can be quite useful in bridging the gap between these varied and fragmented biomedical databases, with their syntactic and semantic heterogeneity, though hallucinations can occur. In this context, the LLM may incorrectly apply ontology identifiers, or alternatively, completely fabricate them; application of appropriate constraints is therefore required [1]. To address this, the flexibility of LLMs can be combined with the structured nature of biomedical ontologies, enabling database integration through semantic alignment, while at the same time, keeping the resulting inferences grounded in scientific reality.

 

The unknowns & the experimental questions

 

The authors studied whether this combination of LLM flexibility with ontology rigidity can effectively automate curation of biomedical databases, resolve semantic proliferation and ground the combination into scientific truth.  Through this study, the authors thus aimed to assemble and deploy an artificial intelligence (AI)–based semantics workflow to facilitate data integration across heterogeneous databases. They used this workflow to:

  1. Curate biomedical data across heterogenous databases, either through utilization or de novo creation, if the term does not exist, of ontology terms, via a specialised workflow termed as the Ontology Lookup Service (OLS)
  2. Expand LLM embedding practices to include all ontologies, currently present in the OLS

 

The experimental tools

 

To this end, the authors devised a computational workflow, employing multiple large language model (LLM) agents (curator agent, importer agent, ontologist agent, and all coordinated via the coordinating agent), which were in turn grounded in scientifically truthful knowledge via Model Context Protocol (MCP) servers (e.g., OLS MCP server for access to validated ontology terms and identifiers, art1 MCP server for the retrieval and parsing of Europe PMC content).

To generate vector embeddings for all OLS terms, a staggering 11 mil. Terms, for semantic alignment between databases, they employed four generalist models (llama-embed-nemotron-8b, harrier-oss-v1-27b, text-embedding-3-large, text-embedding-3-small); in this manner they calculated cosine similarity across various phenotypes and traits. They also implemented GrEBI (Graphs@EBI, Neo4j KG and MCP server), which makes use of parameterized, Cypher query templates that allows for bridging between human GWAS data and IMPC-derived, mouse ortholog phenotypes [6].

 

How the workflow works: ontologies, embeddings and semantic alignment

Ontologies for Database annotation

 

In the field of bioinformatics and computational biology, an ontology can be defined as a formal, structured framework with standardized vocabulary that characterizes and describes concepts in biological and medical sciences; it establishes specific terms, as well as the relationship between them [2] [3].

Ontologies have also been defined mathematically; the Gene ontology (GO), for example, has been defined as “directed acyclic graphs (DAG)”. In this framework, individual biological concepts/terms are the “nodes”, while the hierarchical relationships between them are the “directional edges” [4].

  1. The vocabulary is standardized across different locations, in order to be identifiable by any computer program employing deterministic methods (thus referred to as “machine-readable”); the symbols employed may be symbolic or text-based but always adhere to a rigid set of rules.
  2. Should the requirement be that such standardized terms be identifiable by machine-learning (ML) algorithms instead, then this text–based information must be transformed into numerical representations. These representations can include continuous vectors, matrices, or multi–dimensional tensors, which can be then further processed via mathematical or statistical operations. The transformation step can either be tokenization,normalization or embedding.

 

Semantic alignment

 

Through the process of semantic alignment, a machine-readable ontology term (e.g., a string label in the Experimental Factor Ontology [EFO], an ontology developed by the EMBL–EBI) will be processed into a large language model (LLM) embedding (feature vectors), i.e., a dense numerical vector, which can then be readable by machine–learning (ML) algorithms. Using this representation, mathematical operations can be then carried out:

  • In this manner, scientists can evaluate how closely related two concepts are; for example, through calculation of the cosine similarity between two embedding vectors, they can align and connect two non-identical terms, across different databases, based on their conceptual proximity [1] [5] [6].

 

Table summarising the key findings and applications of the study

 

Building and grounding the framework

Evaluating performance

Biological applications

Infrastructure enabling data integration 

 

Why this work is interesting

 

The work by McLaughlin et al., 2026, presents a valuable, and useful toolkit, employing multi-agent AI curation, semantic embeddings, and a unified knowledge graph (KG) implementation, through which incompatible biological and medical databases can be coordinated to allow for a unified response to a specific research question [1]. Nowadays, many research projects, and research questions involve the combined use of more than one database, oftentimes spanning different species; a unified layer of ‘translation’ between those databases is sure to make this a lot easier.

 

Questions to the authors:

 

  • On training data contamination: In order to prevent the LLM from training on ontology terms generated by the AI-based workflow, you mention that you plan to tag these terms with the editor property ‘AI agent‘. However, given the fact that LLM development regularly involves data scraping (i.e., the automated extraction of information from publicly available data):
    • Is application of an internal annotation property enough, or do you think additional barriers that affect automated data extraction should be implemented, especially given the fact that often, automated processes/scripts that scrape metadata tags may be involved in the processes applied in neural network training?

 

References

 

[1]. McLaughlin, J., Puig-Barbe, A., Ibrahim, A., Pava, D., Pendlington, Z. M., Matentzoglu, N., Sollis, E., Foreman, A., Wilson, R., Gomez, F. L., Harris, L., Adeleye, Y., Kaur, S., Meldal, B., Smedley, D., Parkinson, H. (2026) AI semantics for biomedical data integration [Preprint], BioRχiv, doi:  https://doi.org/10.64898/2026.08.03.742514

[2]. Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA, Hill DP, Issel-Tarver L, Kasarskis A, Lewis S, Matese JC, Richardson JE, Ringwald M, Rubin GM, Sherlock G. (2000) Gene ontology: tool for the unification of biology. Nature Genetics. 25(1):25-9. doi: 10.1038/75556.

[3]. Alghamdi, S.M, Schofield, P.N., Hoehndorf, R. (2022) Contribution of model organism phenotypes to the computational identification of human disease genes. Disease Models & Mechanisms. 15 (7): dmm049441. doi:  https://doi.org/10.1242/dmm.049441

[4]. Yang, X., Li, J., Lee, Y., Lussier, Y.A. (2011) GO-Module: functional synthesis and improved interpretation of Gene Ontology patterns, Bioinformatics, 27 (10): 1444–1446, doi: https://doi.org/10.1093/bioinformatics/btr142

[5]. Kulmanov, M., Smaili, F.Z., Gao, X., Hoehndorf, R. (2021) Semantic similarity and machine learning with ontologies, Briefings in Bioinformatics, 22 (4): bbaa199, doi: https://doi.org/10.1093/bib/bbaa199

[6]. Smaili, F. Z., Gao, X., Hoehndorf, R. (2019) OPA2Vec: combining formal and informal content of biomedical ontologies to improve similarity-based prediction, Bioinformatics, 35 (12): 2133–2140, doi: https://doi.org/10.1093/bioinformatics/bty933

 

 

Read preprint (No Ratings Yet)

Have your say

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Sign up to customise the site to your preferences and to receive alerts

Register here

Also in the bioinformatics category:

Clonal embeddings allow exploratory analysis of lineage-resolved single-cell data

Sergey Isaev, Alek G Erickson, Igor Adameyko, et al.

Selected by 23 September 2026

Marine Secchi

Bioinformatics

Additive baselines furnish no evidence for epistasis learning by MULTI-evolve

Gian Marco Visani, Aayush Verma, William S. DeWitt

Selected by 08 September 2026

Leonie Brüne

Bioinformatics

Temporal degradation of PRC2 uncovers specific developmental dependencies

Ming-Kang Lee, Sebastian D. Mackowiak, Daniel Felismino, et al.

Selected by 19 May 2026

María Mariner-Faulí

Developmental Biology

Also in the scientific communication and education category:

AI semantics for biomedical data integration

James McLaughlin, Aleix Puig-Barbe, Arwa Ibrahim, et al.

Selected by 05 October 2026

Theodora Stougiannou

Bioinformatics

Expanded implementation of Fast & Fair paid peer review reduces time to first decision without reducing review quality in a biology journal

Daniel A. Gorelick, Alejandra Clark

AND

Fast & Fair peer review: a pilot study demonstrating feasibility of rapid, high-quality peer review in a biology journal

Daniel A. Gorelick, Alejandra Clark

Selected by 21 September 2026

Jonathan Townson

Scientific Communication and Education

Who chooses open peer review and is it an indicator of article quality? An observational study of PLOS journals

Adrian Barnett, Matt Spick

Selected by 31 August 2026

Chee Kiang Ewe

Scientific Communication and Education

preLists in the bioinformatics category:

Keystone Symposium – Metabolic and Nutritional Control of Development and Cell Fate

This preList contains preprints discussed during the Metabolic and Nutritional Control of Development and Cell Fate Keystone Symposia. This conference was organized by Lydia Finley and Ralph J. DeBerardinis and held in the Wylie Center and Tupper Manor at Endicott College, Beverly, MA, United States from May 7th to 9th 2025. This meeting marked the first in-person gathering of leading researchers exploring how metabolism influences development, including processes like cell fate, tissue patterning, and organ function, through nutrient availability and metabolic regulation. By integrating modern metabolic tools with genetic and epidemiological insights across model organisms, this event highlighted key mechanisms and identified open questions to advance the emerging field of developmental metabolism.

 



List by Virginia Savy, Martin Estermann

‘In preprints’ from Development 2022-2023

A list of the preprints featured in Development's 'In preprints' articles between 2022-2023

 



List by Alex Eve, Katherine Brown

9th International Symposium on the Biology of Vertebrate Sex Determination

This preList contains preprints discussed during the 9th International Symposium on the Biology of Vertebrate Sex Determination. This conference was held in Kona, Hawaii from April 17th to 21st 2023.

 



List by Martin Estermann

Alumni picks – preLights 5th Birthday

This preList contains preprints that were picked and highlighted by preLights Alumni - an initiative that was set up to mark preLights 5th birthday. More entries will follow throughout February and March 2023.

 



List by Sergio Menchero et al.

Fibroblasts

The advances in fibroblast biology preList explores the recent discoveries and preprints of the fibroblast world. Get ready to immerse yourself with this list created for fibroblasts aficionados and lovers, and beyond. Here, my goal is to include preprints of fibroblast biology, heterogeneity, fate, extracellular matrix, behavior, topography, single-cell atlases, spatial transcriptomics, and their matrix!

 



List by Osvaldo Contreras

Single Cell Biology 2020

A list of preprints mentioned at the Wellcome Genome Campus Single Cell Biology 2020 meeting.

 



List by Alex Eve

Antimicrobials: Discovery, clinical use, and development of resistance

Preprints that describe the discovery of new antimicrobials and any improvements made regarding their clinical use. Includes preprints that detail the factors affecting antimicrobial selection and the development of antimicrobial resistance.

 



List by Zhang-He Goh