Glucose and everyday records
Sensor readings, meter readings, recorded insulin, meals and activity. Identify how each measurement was collected and which parts of the day are missing.
Data library
Our ambition is a world-leading diabetes dataset and evidence repository, connecting eligible research sources, laboratory context and molecular knowledge with clear permissions.

Two files can use the same column name for different measurements. We are building tools to check units, dates, missing information and study populations, while retaining the original source and its usage restrictions.
A global diabetes evidence library
Our ambition is to build the world’s largest diabetes-focused dataset and evidence repository: connecting eligible records, published studies, laboratory knowledge and molecular references in one research foundation.
That is a long-term target. Today, we are building the tools to qualify sources, preserve permissions and make each connection useful. Coverage, quality and independent usefulness will define the scale we can substantiate.
Explore the data foundation
Explore by evidence type
Collection specification
Compare patterns across an eligible study without inventing unobserved readings.
Sensor readings, meter readings, recorded insulin, meals and activity. Identify how each measurement was collected and which parts of the day are missing.
Laboratory reports, published studies and trial records. Keep the original wording, study population and limits alongside each extracted finding.
Research on islets, proteins, genetics, metabolites and food composition. Keep the tissue, measurement method and source attached so different kinds of evidence are not treated as equivalent.
Longitudinal and population studies can support questions about changing risk. Researchers still need to check whom the data represents and whether follow-up is adequate.
A measurement, a document and a molecular finding each contribute a different kind of evidence. Our goal is to connect them without losing what makes them meaningful.
Time-stamped readings, device context and recorded meals or activity.
Which patterns deserve a closer look?
Timing · Units · Missing observations
Reported test values, dates and documented clinical context.
How does an observation change over time?
Assay · Reference context · Record date
Scientific papers, study descriptions and authorized reports.
What was studied, in whom, and with what result?
Source · Population · Study design
Ingredient composition, biological pathways and measured molecular findings.
Which connections justify further investigation?
Identity · Preparation · Tissue or assay
Illustrative data categories, not an inventory of acquired records. Access, permissions and fitness for a specific study must be established separately.
Compute & security roadmap
We plan to scale from controlled local research to high-memory systems, GPU-accelerated training and distributed experiments as the data, funding and evaluation justify them.
Planned controlled research environment
Infrastructure as code · Separate research and production environments · Tested recovery
Our planned workloads span machine learning, deep neural networks, computational models and later large language models. We will select hardware against memory needs, reproducibility, training time and cost, with capacity and benchmarks documented as infrastructure is commissioned.
AWS GovCloud (US) is a deployment option we intend to evaluate for eligible workloads. Account eligibility and regional service availability must be confirmed. This is an infrastructure plan, not a claim of an existing GovCloud deployment or certification.
Our planned controls include least-privilege access, encryption in transit and at rest, isolated workloads, audit trails, restriction propagation and tested backups. Cloud hosting does not replace application security or governance.
Record where data came from
Check permission for the study
Describe gaps and measurement limits
Keep source versions available for review
These are the types of sources we plan to work with, not an inventory of acquired datasets. This website does not accept health records.
Discuss a data partnership