AI semantics for biomedical data integration
Posted on: 5 October 2026
Preprint posted on 3 September 2026
Crouching AI, hidden connections; The preprint authors assemble and deploy an artificial intelligence (AI)-based semantics workflow to facilitate data integration across heterogeneous databases.
Selected by Theodora StougiannouCategories: bioinformatics, scientific communication and education
Crouching AI, hidden connections; The preprint authorsassemble and deploy an artificial intelligence (AI)-based semantics workflow to facilitate data integration across heterogeneous databases.
Setting the scene: What are the authors looking at?
The authors of this preprint investigated how to enable biomedical data integration across multimodal datasets, diverse model organisms and different biological scales; to evaluate a research question, in practice, knowledge across different databases must be integrated. [1].
Databases, including those applied for the biomedical sciences, can have heterogenous schemas and annotation practices; furthermore, they may often have incompatible application programming interfaces (API). These characteristics render interoperability between different biomedical databases difficult, especially in an era when biomedical research requires the integration of these heterogenous databases, for example, interconnection between human Genome Wide Association Study (GWAS) study catalogs, i.e., the GWAS Catalog (NHGRI-EBI GWAS Catalog), and associated mouse research databases, i.e. the International Mouse phenotyping Consortium (IMPC) [1].
Bridging the gap: bring a dreaming LLM back to scientific reality
Large language models (LLM) can be quite useful in bridging the gap between these varied and fragmented biomedical databases, with their syntactic and semantic heterogeneity, though hallucinations can occur. In this context, the LLM may incorrectly apply ontology identifiers, or alternatively, completely fabricate them; application of appropriate constraints is therefore required [1]. To address this, the flexibility of LLMs can be combined with the structured nature of biomedical ontologies, enabling database integration through semantic alignment, while at the same time, keeping the resulting inferences grounded in scientific reality.
The unknowns & the experimental questions
The authors studied whether this combination of LLM flexibility with ontology rigidity can effectively automate curation of biomedical databases, resolve semantic proliferation and ground the combination into scientific truth. Through this study, the authors thus aimed to assemble and deploy an artificial intelligence (AI)–based semantics workflow to facilitate data integration across heterogeneous databases. They used this workflow to:
- Curate biomedical data across heterogenous databases, either through utilization or de novo creation, if the term does not exist, of ontology terms, via a specialised workflow termed as the Ontology Lookup Service (OLS)
- Expand LLM embedding practices to include all ontologies, currently present in the OLS
The experimental tools
To this end, the authors devised a computational workflow, employing multiple large language model (LLM) agents (curator agent, importer agent, ontologist agent, and all coordinated via the coordinating agent), which were in turn grounded in scientifically truthful knowledge via Model Context Protocol (MCP) servers (e.g., OLS MCP server for access to validated ontology terms and identifiers, art1 MCP server for the retrieval and parsing of Europe PMC content).
To generate vector embeddings for all OLS terms, a staggering 11 mil. Terms, for semantic alignment between databases, they employed four generalist models (llama-embed-nemotron-8b, harrier-oss-v1-27b, text-embedding-3-large, text-embedding-3-small); in this manner they calculated cosine similarity across various phenotypes and traits. They also implemented GrEBI (Graphs@EBI, Neo4j KG and MCP server), which makes use of parameterized, Cypher query templates that allows for bridging between human GWAS data and IMPC-derived, mouse ortholog phenotypes [6].
How the workflow works: ontologies, embeddings and semantic alignment
Ontologies for Database annotation
In the field of bioinformatics and computational biology, an ontology can be defined as a formal, structured framework with standardized vocabulary that characterizes and describes concepts in biological and medical sciences; it establishes specific terms, as well as the relationship between them [2] [3].
Ontologies have also been defined mathematically; the Gene ontology (GO), for example, has been defined as “directed acyclic graphs (DAG)”. In this framework, individual biological concepts/terms are the “nodes”, while the hierarchical relationships between them are the “directional edges” [4].
- The vocabulary is standardized across different locations, in order to be identifiable by any computer program employing deterministic methods (thus referred to as “machine-readable”); the symbols employed may be symbolic or text-based but always adhere to a rigid set of rules.
- Should the requirement be that such standardized terms be identifiable by machine-learning (ML) algorithms instead, then this text–based information must be transformed into numerical representations. These representations can include continuous vectors, matrices, or multi–dimensional tensors, which can be then further processed via mathematical or statistical operations. The transformation step can either be tokenization,normalization or embedding.
Semantic alignment
Through the process of semantic alignment, a machine-readable ontology term (e.g., a string label in the Experimental Factor Ontology [EFO], an ontology developed by the EMBL–EBI) will be processed into a large language model (LLM) embedding (feature vectors), i.e., a dense numerical vector, which can then be readable by machine–learning (ML) algorithms. Using this representation, mathematical operations can be then carried out:
- In this manner, scientists can evaluate how closely related two concepts are; for example, through calculation of the cosine similarity between two embedding vectors, they can align and connect two non-identical terms, across different databases, based on their conceptual proximity [1] [5] [6].
Table summarising the key findings and applications of the study
Building and grounding the framework

Evaluating performance

Biological applications

Infrastructure enabling data integration

Why this work is interesting
The work by McLaughlin et al., 2026, presents a valuable, and useful toolkit, employing multi-agent AI curation, semantic embeddings, and a unified knowledge graph (KG) implementation, through which incompatible biological and medical databases can be coordinated to allow for a unified response to a specific research question [1]. Nowadays, many research projects, and research questions involve the combined use of more than one database, oftentimes spanning different species; a unified layer of ‘translation’ between those databases is sure to make this a lot easier.
Questions to the authors:
- On training data contamination: In order to prevent the LLM from training on ontology terms generated by the AI-based workflow, you mention that you plan to tag these terms with the editor property ‘AI agent‘. However, given the fact that LLM development regularly involves data scraping (i.e., the automated extraction of information from publicly available data):
- Is application of an internal annotation property enough, or do you think additional barriers that affect automated data extraction should be implemented, especially given the fact that often, automated processes/scripts that scrape metadata tags may be involved in the processes applied in neural network training?
References
[1]. McLaughlin, J., Puig-Barbe, A., Ibrahim, A., Pava, D., Pendlington, Z. M., Matentzoglu, N., Sollis, E., Foreman, A., Wilson, R., Gomez, F. L., Harris, L., Adeleye, Y., Kaur, S., Meldal, B., Smedley, D., Parkinson, H. (2026) AI semantics for biomedical data integration [Preprint], BioRχiv, doi: https://doi.org/10.64898/2026.08.03.742514
[2]. Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA, Hill DP, Issel-Tarver L, Kasarskis A, Lewis S, Matese JC, Richardson JE, Ringwald M, Rubin GM, Sherlock G. (2000) Gene ontology: tool for the unification of biology. Nature Genetics. 25(1):25-9. doi: 10.1038/75556.
[3]. Alghamdi, S.M, Schofield, P.N., Hoehndorf, R. (2022) Contribution of model organism phenotypes to the computational identification of human disease genes. Disease Models & Mechanisms. 15 (7): dmm049441. doi: https://doi.org/10.1242/dmm.049441
[4]. Yang, X., Li, J., Lee, Y., Lussier, Y.A. (2011) GO-Module: functional synthesis and improved interpretation of Gene Ontology patterns, Bioinformatics, 27 (10): 1444–1446, doi: https://doi.org/10.1093/bioinformatics/btr142
[5]. Kulmanov, M., Smaili, F.Z., Gao, X., Hoehndorf, R. (2021) Semantic similarity and machine learning with ontologies, Briefings in Bioinformatics, 22 (4): bbaa199, doi: https://doi.org/10.1093/bib/bbaa199
[6]. Smaili, F. Z., Gao, X., Hoehndorf, R. (2019) OPA2Vec: combining formal and informal content of biomedical ontologies to improve similarity-based prediction, Bioinformatics, 35 (12): 2133–2140, doi: https://doi.org/10.1093/bioinformatics/bty933
Read preprint
Sign up to customise the site to your preferences and to receive alerts
Register hereAlso in the bioinformatics category:
Clonal embeddings allow exploratory analysis of lineage-resolved single-cell data
Marine Secchi
Additive baselines furnish no evidence for epistasis learning by MULTI-evolve
Leonie Brüne
Temporal degradation of PRC2 uncovers specific developmental dependencies
María Mariner-Faulí
Also in the scientific communication and education category:
AI semantics for biomedical data integration
Theodora Stougiannou
Expanded implementation of Fast & Fair paid peer review reduces time to first decision without reducing review quality in a biology journal
AND
Fast & Fair peer review: a pilot study demonstrating feasibility of rapid, high-quality peer review in a biology journal
Jonathan Townson
Who chooses open peer review and is it an indicator of article quality? An observational study of PLOS journals
Chee Kiang Ewe
preLists in the bioinformatics category:
Keystone Symposium – Metabolic and Nutritional Control of Development and Cell Fate
This preList contains preprints discussed during the Metabolic and Nutritional Control of Development and Cell Fate Keystone Symposia. This conference was organized by Lydia Finley and Ralph J. DeBerardinis and held in the Wylie Center and Tupper Manor at Endicott College, Beverly, MA, United States from May 7th to 9th 2025. This meeting marked the first in-person gathering of leading researchers exploring how metabolism influences development, including processes like cell fate, tissue patterning, and organ function, through nutrient availability and metabolic regulation. By integrating modern metabolic tools with genetic and epidemiological insights across model organisms, this event highlighted key mechanisms and identified open questions to advance the emerging field of developmental metabolism.
| List by | Virginia Savy, Martin Estermann |
‘In preprints’ from Development 2022-2023
A list of the preprints featured in Development's 'In preprints' articles between 2022-2023
| List by | Alex Eve, Katherine Brown |
9th International Symposium on the Biology of Vertebrate Sex Determination
This preList contains preprints discussed during the 9th International Symposium on the Biology of Vertebrate Sex Determination. This conference was held in Kona, Hawaii from April 17th to 21st 2023.
| List by | Martin Estermann |
Alumni picks – preLights 5th Birthday
This preList contains preprints that were picked and highlighted by preLights Alumni - an initiative that was set up to mark preLights 5th birthday. More entries will follow throughout February and March 2023.
| List by | Sergio Menchero et al. |
Fibroblasts
The advances in fibroblast biology preList explores the recent discoveries and preprints of the fibroblast world. Get ready to immerse yourself with this list created for fibroblasts aficionados and lovers, and beyond. Here, my goal is to include preprints of fibroblast biology, heterogeneity, fate, extracellular matrix, behavior, topography, single-cell atlases, spatial transcriptomics, and their matrix!
| List by | Osvaldo Contreras |
Single Cell Biology 2020
A list of preprints mentioned at the Wellcome Genome Campus Single Cell Biology 2020 meeting.
| List by | Alex Eve |
Antimicrobials: Discovery, clinical use, and development of resistance
Preprints that describe the discovery of new antimicrobials and any improvements made regarding their clinical use. Includes preprints that detail the factors affecting antimicrobial selection and the development of antimicrobial resistance.
| List by | Zhang-He Goh |






