Senior Data Scientist · Genomics Researcher · Educator
Working at the intersection of data science, clinical AI, and personalised genomics. Research Fellow at UCL Great Ormond Street Institute of Child Health. Founder of CESGEN.
I am a Senior Data Scientist (Grade 7) and Research Fellow at the UCL Great Ormond Street Institute of Child Health, Population, Policy and Practice Department. My career spans more than 30 years at the interface of population genetics, molecular diagnostics, clinical decision support, and education — always with the aim of making data useful for health.
My primary research focus is the Neotree project — a digital health platform for neonatal clinical decision support deployed in Zimbabwe and Malawi. I lead the statistical modelling, machine learning pipelines, interrupted time series analysis, and data governance components.
Alongside research, I am deeply interested in nutrigenomics and personalised medicine — the science of how individual genetic variation shapes responses to diet, environment, and lifestyle. I develop and teach courses on this topic, and build genomic analysis tools through CESGEN and open-source contributions to the ClawBio bioinformatics library.
I have taught mathematics, statistics, and data science at university level for over two decades — in English, Spanish, and German — across institutions in the UK, Spain, and internationally.
My research sits at the intersection of machine learning, statistics, and global health — asking how AI and rigorous methods can improve clinical decisions in high-stakes, resource-limited settings.
A digital health platform deployed in neonatal units in Zimbabwe (Sally Mugabe Central Hospital, Harare) and Malawi (Kamuzu Central Hospital). It captures structured clinical data at point of care and supports clinical decisions, outcome monitoring, and quality improvement.
My role spans interrupted time series analysis evaluating platform impact on neonatal outcomes, machine learning models for early prediction of hypothermia and sepsis, master data pipelines (local and NHS TRE versions), and data governance.
Visit NeotreeMore than two decades of work on the genomic basis of individual variation in health, metabolism, and disease risk — from population genetics and molecular evolution through to SNP-based panels for metabolic nutrition profiling and clinical interpretation pipelines.
Current work focuses on genomic risk scoring, gene-diet interactions, and building scalable open-source analysis pipelines through CESGEN and the ClawBio bioinformatics library.
Visit CESGENBuilt over 30 years across biology, statistics, and computing — applied to research, teaching, and industry.
Precise, accessible, grounded in real examples. Two decades of teaching in English, Spanish, and German, across the UK, Spain, and internationally.
An applied course covering the use of large language models, machine learning pipelines, and AI tools in genomic data analysis. Aimed at researchers, clinicians, and bioinformatics practitioners who want to work effectively with AI in a genomics setting.
View course at CESGENA comprehensive course on the science of personalised nutrition — how individual genetic variation shapes responses to diet, lifestyle, and environment. Covers SNP interpretation, metabolic pathways, gene-diet interactions, and the evidence base behind personalised health recommendations.
Join the mailing listI write primarily in Python and R, with a focus on genomics analysis, clinical data pipelines, and machine learning. All public work is on GitHub (51 repositories).
Pipeline for interpreting SNP-based genotype data in the context of metabolic nutrition. The analytical core of the CESGEN personalised genomics platform — linking individual genotype profiles to evidence-based nutritional recommendations.
View on GitHubOpen-source library of skills for bioinformatics AI agents, built on the OpenClaw runtime. Local-first, reproducible, designed for researchers who want to integrate AI into genomics workflows without losing methodological transparency.
View on GitHubData pipelines and statistical analyses for the Neotree neonatal health project — data ingestion, cleaning, interrupted time series analysis, and machine learning models for clinical prediction. Both local and NHS Trusted Research Environment versions.
View repositoriesScience communication, genomics, data science, and everything in between. Find me on the platforms below.
I am always happy to hear from researchers, collaborators, students, and anyone with a genuine interest in data science, genomics, or global health.