Research and software for a better future with diabetes.

Data library

A global foundation for diabetes data.

Our ambition is a world-leading diabetes dataset and evidence repository, connecting eligible research sources, laboratory context and molecular knowledge with clear permissions.

Data scientists reviewing example diabetes research charts and a world map

Know what a record means before combining it.

Two files can use the same column name for different measurements. We are building tools to check units, dates, missing information and study populations, while retaining the original source and its usage restrictions.

A global diabetes evidence library

Bring the world’s diabetes knowledge into reach.

Our ambition is to build the world’s largest diabetes-focused dataset and evidence repository: connecting eligible records, published studies, laboratory knowledge and molecular references in one research foundation.

That is a long-term target. Today, we are building the tools to qualify sources, preserve permissions and make each connection useful. Coverage, quality and independent usefulness will define the scale we can substantiate.

Explore the data foundation
Data scientists reviewing example diabetes research charts and a world map
Glucose & clinical context · Literature · Molecular evidence · Population research
Foundry/ Evidence library

Interface concept · Example content

Explore by evidence type

Collection specification

Glucose & daily context

Compare patterns across an eligible study without inventing unobserved readings.

Context to preserve
Measurements · units · collection time · missing intervals
Access & evaluation
Study-specific permission and participant-aware evaluation.
Availability
Planned collection type · No dataset offered in this preview

Glucose and everyday records

Sensor readings, meter readings, recorded insulin, meals and activity. Identify how each measurement was collected and which parts of the day are missing.

Reports and research papers

Laboratory reports, published studies and trial records. Keep the original wording, study population and limits alongside each extracted finding.

Cells, genes and ingredients

Research on islets, proteins, genetics, metabolites and food composition. Keep the tissue, measurement method and source attached so different kinds of evidence are not treated as equivalent.

Health over time

Longitudinal and population studies can support questions about changing risk. Researchers still need to check whom the data represents and whether follow-up is adequate.

Different data. Different questions.

A measurement, a document and a molecular finding each contribute a different kind of evidence. Our goal is to connect them without losing what makes them meaningful.

01 / Evidence family

Glucose & daily context

Time-stamped readings, device context and recorded meals or activity.

The question

Which patterns deserve a closer look?

Timing · Units · Missing observations

02 / Evidence family

Laboratory & clinical records

Reported test values, dates and documented clinical context.

The question

How does an observation change over time?

Assay · Reference context · Record date

03 / Evidence family

Language & literature

Scientific papers, study descriptions and authorized reports.

The question

What was studied, in whom, and with what result?

Source · Population · Study design

04 / Evidence family

Food & molecular evidence

Ingredient composition, biological pathways and measured molecular findings.

The question

Which connections justify further investigation?

Identity · Preparation · Tissue or assay

Illustrative data categories, not an inventory of acquired records. Access, permissions and fitness for a specific study must be established separately.

Compute & security roadmap

Serious computing. Defined boundaries.

We plan to scale from controlled local research to high-memory systems, GPU-accelerated training and distributed experiments as the data, funding and evaluation justify them.

Planned controlled research environment

  1. AccessIdentity · study purpose · permitted use
  2. DataEncrypted stores · versioning · traceable transformations
  3. ComputeCPU preparation · GPU training · isolated model evaluation
  4. ReleaseReviewer approval · scoped APIs · audit records

Infrastructure as code · Separate research and production environments · Tested recovery

High-performance computing with a purpose.

Our planned workloads span machine learning, deep neural networks, computational models and later large language models. We will select hardware against memory needs, reproducibility, training time and cost, with capacity and benchmarks documented as infrastructure is commissioned.

AWS GovCloud (US), where appropriate.

AWS GovCloud (US) is a deployment option we intend to evaluate for eligible workloads. Account eligibility and regional service availability must be confirmed. This is an infrastructure plan, not a claim of an existing GovCloud deployment or certification.

Our planned controls include least-privilege access, encryption in transit and at rest, isolated workloads, audit trails, restriction propagation and tested backups. Cloud hosting does not replace application security or governance.

What every tool needs to make clear.

Record where data came from

Check permission for the study

Describe gaps and measurement limits

Keep source versions available for review

These are the types of sources we plan to work with, not an inventory of acquired datasets. This website does not accept health records.

Discuss a data partnership

Tell us what your team needs.

Contact us