What are you looking for?
Birds of a Feather (BoF) Session April 16, 2025

Towards reproducible and transparent computational modelling in health and biomedical research

Plenary: RDA 25th Plenary Meeting [part of International Data Week 2025]

Submitted by

Meeting objectives

Collaborative notes https://docs.google.com/document/d/1FhdNWCx0

 

The formation of the Computational Modelling of Health Data working group has been funded through the RDA TIGER Project. This BoF session will explore the formation of this Working Group focused on improving the reproducibility, transparency, and accessibility of computational models developed for health research. The session will engage the community in shaping a curated resource of best practices, datasets, workflows, and standards for modelling approaches ranging from simulation and statistical models to AI and language models developed for health research. It will also explore opportunities for collaboration across existing RDA groups and broader communities working at the intersection of health, data science, and computational modelling. The specific objectives of this session are given below –

  1. Introduction: Present the rationale, goals, and initial scope of a proposed RDA Working Group focused on improving transparency, reproducibility, and accessibility in computational modelling for health research.
  2. Understand community interest: Assess the level of interest within the RDA and broader research communities, and identify individuals or organisations keen to contribute to or collaborate with the proposed working group.
  3. Explore case studies of reproducibility and alignment with existing RDA and external groups
  4. Identify opportunities for collaboration and complementarity with relevant RDA Interest and Working Groups, such as the Health Data IG, FAIR4ML IG, and Sensitive Data IG.
  5. Outline next steps towards group formation: Share the formal process for establishing a new Working Group within RDA, invite expressions of interest, and define immediate follow-up actions following the BoF session
  6. Feedback on scope:  Facilitate discussion to refine the thematic focus, explore gaps or emerging needs, and ensure the group’s direction reflects shared priorities and adds value to the community.
  7. Discuss potential sub-groups and outputs: Gather input on proposed sub-working groups (e.g. standards and metadata, tools and datasets, workflows, ethics and governance) and the types of outputs that would be most impactful.

Meeting presenters

Priyanka Nair-Turkich, Christian Page, Francis Crawley, Matthias Konig

Meeting agenda

  1. Computational modelling for health data – Priyanka Nair-Turkich (15 minutes)
    • Acknowledgment of country
    • Welcome and introduction to the session goals
    • Overview of the motivation for the proposed Working Group
    • Types of computational modelling in health data – machine learning, AI, agent-based models, and gaps
  2. Discussion (10 minutes)
  3. Reproducibility challenges and experiences – Christian Pagé (10 minutes)
    • Introduction – Main concept
    • Background – Reproducibility in Climate Modelling
    • Challenges – Interdisciplinarity and Sensitive Data
    • New Frontiers – AI in Data Analytics
  4. Discussion (10 minutes)
  5. Reproducibility by Design: Data and Workflows for Pharmacokinetic Digital Twins – Matthias König (Recorded presentation (10 minutes))
    • Introduction – Digital Twins  in pharmacokinetic
    • Data integration – Data curation and standardization
    • Application – Glimepiride example
    • Reproducibility – COMBINE standards and reproducible workflows
  6. Governing the digital health future: hardwiring ethics and crisis resilience into computational models for the SDGs – Francis P. Crawley (10 minutes)
    • Introduction – An implementation vehicle for the Digital SDGs Programme
    • The integrated framework – A layered architecture for trust
    • From principles to practice – Hardwiring governance into WG deliverables
    • A Strategic contribution to the Digital SDGs Programme (DSP)
  7. Discussion and closing remarks (20 minutes)

Have you presented a session on the same topic at any previous plenaries?

No

Additional links to informative material

Draft case statement –

  1. Executive summary

This Computational Modelling of Health Data Working Group aims to establish a curated, community-driven resource to support the development and dissemination of computational models in health research. It will provide detailed guidance on how to develop models using appropriate inputs related to health research domain —such as data features, synthetic and simulated datasets, open-source software, and parameters—ensuring reproducibility and accessibility. By consolidating best practices, standard protocols, and examples across a range of model types including simulation, statistical, machine learning, and large language models, the group will support users in creating transparent and reusable workflows in health research.

The resource will be particularly valuable in data-scarce environments, where high-quality synthetic datasets and standardised tools can accelerate research while reducing development time. It addresses challenges in computational  health research including lack of reproducible workflows, inconsistent data preparation practices, and the perception of computational models as opaque “black boxes”.

Target audiences include clinicians, biomedical researchers, data scientists, policy-makers, and software developers. Engagement with existing RDA groups and other relevant initiatives will ensure alignment with ongoing global efforts. Through workshops, collaborative writing sessions, and stakeholder outreach, the group will foster adoption and real-world application.

2. WG Charter:

The working group will develop a curated community resource on how to develop and share a computational model developed for health research, detailing which inputs such as data features, synthetic and simulated data, software, parameters and settings, are used to produce results. The curated community resource will also include different types of computational models used in health research and the associated data features, synthetic and simulated health datasets for research, open-source software, standardised protocols, and reproducible workflows. By creating and maintaining these standards, the group will ensure that computational modelling of health data remains accessible, well-documented, and easily discoverable.

  1. Value Proposition:

Computational methodologists, researchers, practitioners, and healthcare professionals will gain access to a well-structured, high-quality resource that simplifies the process of developing computational models and workflows, particularly in data-deficient environments. By providing resources on data standards and synthetic and simulated datasets, open-source tools, and best-practice workflows, users can accelerate research, enhance reproducibility, and reduce the time and effort required to develop their own computational models.

  1. Challenges addressed 
  • Many researchers do not have the curated list of resources on best practices, synthetic datasets, and data standards to improve interoperability or reproducibility of their models.
  • In the absence of clear guidelines, inconsistencies in data preparation, model building, and validation can lead to unreliable results.
  • Researchers often struggle to replicate models or adapt workflows across different contexts.
  • The absence of readily available resources means significant time is spent on data preprocessing and workflow setup., as well as parameterisation.
  • Computational models can be often viewed as experimental black boxes and having resources that help design, outputs, parameters and working of computational models accessible  and reproducible improves the understanding of people who do not actively develop models but would like to learn how they can apply these models for their research.
  1. Methodology themes
  • Simulation models
  • Statistical/Mathematical models
  • Machine learning/Deep learning models
  • Language models
  • Large multimodal models

 

  1. Domains
  • Clinical trials and workflows
  • Population and global health
  • Biomedical engineering and research
  • Social, environmental, and cultural health indicators of health
  • Computer science
  1. Target audience
  • Clinicians wanting to learn more about how statistical, AI and machine learning models can help improve clinical workflows
  • Biomedical researchers
  • Healthcare data analytics organisations
  • Policy-makers using data for decision-making
  • Research software developers interested in the best practice for health
  1. Engagement with existing work in the area: The proposed membership will include some members from existing RDA working and interest groups, as well as experts from outside of RDA. We will also engage with existing groups within the RDA.
  • Health data interest group
  • COVID-19 working group
  • Sensitive data interest group
  • Raising fairness in health data and health research
  • FAIR for Machine Learning (FAIR4ML) IG
  • https://aiaudit.org/
  1. UN Sustainable Development Goals (SDGs): Goal 3 – Ensure healthy lives and promote well-being for all at all ages

 

  1. Work plan 

 

Community Engagement (Months 1–12, Ongoing)

Objectives: Facilitate discussions, knowledge sharing, and collaborative writing to refine outputs and ensure broad engagement.

  • Regular workshops and meetings (Months 1–12, Ongoing)
    • “Shut up and write” sessions for collaborative document development.
    • Thematic discussions on the working group’s goals and objectives.
  • Stakeholder engagement (Months 3, 6, 9, 12)
    • Presentations at RDA plenaries and meetings.
    • Engagement with organisations such as EOSC and CODATA.
    • Regular engagement to researchers, policymakers, and practitioners in health and computer science to encourage adoption.
  • Feedback (Months 3–6, Ongoing)
    • Engage the community and RDA stakeholders for continuous refinement of outputs.
    • Develop interactive versions of outputs with periodic updates incorporating feedback.

Engaging with wider audience

Objectives: Demonstrate how the working group’s outputs can be applied in real-world settings.

  • Facilitate proof-of-concept projects (Months 9–12) – Needs to be organic and if people decide to work together then it could be a potential outcome. Does not need to go in the case statement explicitly but can be one of the possible outcomes as “encourage people to pursue collaborative projects”
    • Identify and initiate research projects demonstrating the use of synthetic and simulated datasets, standardised workflows, and metadata curation.
    • Share findings and best practices with the broader community.

 

Standards and best practices (Months 2–10)

Objectives: Develop and document guidelines to improve the accessibility, interoperability, and reproducibility of computational health research.

  • Assessment of existing models (Months 2–6)
    • Evaluate existing health data modelling frameworks for biases and limitations.
    • Develop guidelines for validating and reproducing health data modelling studies.
  • Metadata and ontologies (Months 4–8)
    • Define best practices for metadata standardisation.
    • Curate a list of ontologies to support reproducibility in biological and health research.
  • Frameworks for reproducibility and interoperability (Months 6–10)
    • Compile best practices for using frameworks such as Fast Healthcare Interoperability Resources (FHIR).
    • Protocols like the ODD (Overview, Design concepts, Details) protocol is a set of documentation guidelines designed to describe agent-based models.
    • Provide recommendations for improving data exchange across multiple systems.

 

  1. Sub-working groups and deliverables

 

Four sub-working groups focussed on – 

  • Standards, metadata and best practices
  • Tools and datasets
  • Workflows
  • Ethics and governance

Estimate of the required venue room capacity

Up to 30

Applicable Pathways

Training, Stewardship, and Data Management Planning
Semantics, Ontology, Standardisation
Data Lifecycles - Versioning, Provenance, Citation, and Reward

Please indicate at least (3) three breakout slots that would suit your meeting.

Breakout 3. Wednesday, 15 October 2025, 01:30-03:00 UTC
Breakout 4. Wednesday, 15 October 2025, 23:00-00:30 UTC
Breakout 5. Thursday, 16 October 2025, 03:30-05:00 UTC