1. About This Guide 2. AI/ML Foundations 3. Medical Foundations 4. AI in Medicine Today 5. Meet the Speakers 6. The Five Fault Lines 7. Key Tensions 8. Glossary 9. Further Reading ← Back to Roundtable

Section 1About This Guide & the Helix Format

Verification Notice

Facts, links, and quantitative claims last verified: April 11, 2026. Regulatory counts, citation totals, and organizational metrics may change over time.

This guide is designed for college-educated attendees (MS, PhD, MD, or equivalent) who may be deeply familiar with either AI/machine learning or medicine/clinical practice — but not both. It provides concise background on the concepts, vocabulary, and research questions you’ll need to follow the roundtable discussion productively.

The Helix Center Format

Helix Center roundtables are moderated conversations, not lecture panels. Five experts with different (often opposing) perspectives sit in a semicircle and engage in real-time dialogue for 90 minutes. No PowerPoints, no prepared talks. The audience observes the conversation and joins during a Q&A period. The format is designed to surface productive disagreement — and this guide will help you anticipate where those disagreements lie.

How to Use This Guide

Download PDF 537 KB · printable version

Section 2AI/ML Foundations for Medical Professionals

If you’re a clinician, researcher, or medical professional, this section provides the AI/ML vocabulary and concepts you’ll need to follow the roundtable’s technical discussions.

What “AI” Means in This Context

When we say “AI in medicine,” we almost always mean machine learning (ML): algorithms that learn patterns from data rather than following hand-coded rules. The key categories relevant to this roundtable:

How Clinical AI Models Work (Simplified)

The Basic Pipeline

1. Data collection → EHR records, imaging, genomic data, wearable sensors
2. Preprocessing → Cleaning, de-identification, formatting
3. Training → The model learns patterns from thousands/millions of examples
4. Validation → Testing on held-out data the model hasn’t seen
5. Deployment → Integration into clinical workflow (the hardest part — ask Aphinyanaphongs)
6. Monitoring → Ongoing surveillance for model drift and bias

Key Concepts You’ll Hear

Section 3Medical & Clinical Foundations for AI Researchers

If you’re an AI researcher, engineer, or technologist, this section provides the medical context you’ll need to understand the clinical stakes of the roundtable discussion.

How Clinical Decisions Are Made

Clinical medicine is fundamentally about decision-making under uncertainty. A clinician integrates:

AI enters this process as an additional input, not a replacement for it. The question is: where in this decision chain does AI add value, and where does it introduce risk?

The Electronic Health Record (EHR)

The EHR is the central nervous system of modern hospitals. It stores clinical notes, lab results, medications, imaging, billing codes, and more. For AI researchers, two things to know:

Key Medical Domains in This Roundtable

Five Domains, Five Speakers

Digital mental health (Choudhury): Using smartphones and wearables to detect stress, depression, and behavioral changes in real time.
Genomics & protein science (Brandes): Predicting how genetic mutations affect protein function and disease risk.
Neuropathology (Crary): Diagnosing brain diseases by examining tissue under a microscope — now with AI.
Health equity (Chunara): Ensuring AI doesn’t worsen existing disparities in who gets what care.
Clinical operations (Aphinyanaphongs): Actually deploying AI tools in a working hospital system.

Why “Deploying AI” in Medicine Is Hard

Section 4The Intersection: AI in Medicine Today

This section maps the current landscape — what’s working, what’s not, and where the roundtable speakers sit within it.

Where AI Is Already in Clinical Use

Where AI Is Struggling

Clinical AI Evidence & Reporting Standards

Several reporting frameworks now set the benchmark for evaluating clinical AI evidence:

  • CONSORT-AI / SPIRIT-AI — for clinical trials involving AI interventions
  • DECIDE-AI — for early-stage clinical deployment evaluation
  • STARD-AI — for diagnostic accuracy studies
  • TRIPOD-AI / PROBAST-AI — for prediction models and risk-of-bias assessment

These standards are curated by the EQUATOR Network and referenced in FDA guidance on AI/ML-based software as a medical device.

Privacy, Consent & Data Governance

Governance Essentials

Clinical AI development raises critical governance questions that go beyond model performance:

  • HIPAA & GDPR: U.S. and EU frameworks impose strict requirements on how patient data is collected, stored, and shared for AI training and inference.
  • De-identification: Methods such as Safe Harbor and expert determination aim to protect patient privacy, but re-identification risks grow as datasets are combined.
  • IRB oversight: Research involving patient data typically requires Institutional Review Board review or documented waivers, even for retrospective studies.
  • Secure enclaves: Controlled-access computing environments allow AI training on sensitive data without exposing raw records to researchers.
  • Foundation model training: Training large models on clinical data raises distinct governance and consent questions, particularly when patients did not originally consent to AI use of their records.

The Regulatory Landscape

Key Regulatory Facts

As of the FDA’s March 2026 update (covering authorizations through the end of 2025), the FDA had authorized about 1,451 AI/ML-enabled medical devices, with radiology accounting for roughly 1,104 (~76%). The agency’s 2021 “Action Plan” for AI/ML-based SaMD proposed a framework for continuously learning systems, but implementation is still evolving. The EU AI Act entered into force in 2024, with obligations phasing in over subsequent years; many medical-AI systems are likely to fall into high-risk categories, creating additional compliance obligations alongside existing medical-device rules. No single global standard exists.

Post-Market Surveillance & Drift Monitoring

Clinical AI requires ongoing real-world performance monitoring, drift detection, and update governance. For adaptive systems, regulators have emphasized predetermined change control plans and post-market surveillance rather than one-time validation alone. The FDA’s AI/ML SaMD Action Plan outlines the evolving expectations for how deployed models should be monitored and updated over their lifecycle.

Section 5Meet the Five Speakers

Each speaker brings a distinct lens to the question of AI in medicine. Understanding their research and perspective will help you follow the roundtable’s dynamics.

The Sensor Scientist — Digital Mental Health & Ubiquitous Computing

Tanzeem Choudhury, PhD

Professor of Information Science, Jacobs Technion-Cornell Institute, Cornell Tech & Cornell University; Co-founder, HealthRhythms

Choudhury is a pioneer in using smartphone and wearable sensor data to monitor mental health. Her lab developed systems that detect stress, depression, and behavioral changes through passive sensing (accelerometer, microphone, GPS, typing patterns) without requiring active user input. She co-founded HealthRhythms to translate this research into clinical tools, and her work on the StudentLife dataset is one of the most cited studies in mobile health.

Questions she’ll raise: Can continuous sensing detect mental health crises before they happen? Who has access to this data, and who should? What happens when AI knows you’re depressed before you do?

Preparation Reading

The Molecular Decoder — Genomic AI & Biological Data Governance

Nadav Brandes, PhD

Assistant Professor, Center for Human Genetics & Genomics, NYU Grossman School of Medicine

Brandes develops protein language models to predict how genetic mutations affect protein function. His 2023 Nature Genetics paper predicted effects of all ~450 million possible missense variants in the human genome. He created ProteinBERT (Bioinformatics, 2022) and PWAS (Genome Biology, 2020), and his collaboration with Jennifer Doudna and Carl June on mitigation of CRISPR-Cas9-induced chromosome loss in clinical T cells was published in Cell. He is also an active voice on AI safety and biological data governance, contributing to LessWrong and the AI Alignment Forum and co-authoring a 2026 Science article on data governance.

Questions he’ll raise: Who controls biological data when AI can decode it? What are the long-term safety implications of increasingly capable AI? How should we govern genomic data in an age of foundation models?

Preparation Reading

The Diagnostic Architect — AI Neuropathology & Brain Aging

John F. Crary, MD, PhD

Professor of Pathology, Molecular and Cell-Based Medicine, Neuroscience, and AI & Human Health; Vice Chair for AI, Dept. of Pathology, Icahn School of Medicine at Mount Sinai

Crary is a board-certified neuropathologist who created HistoAge (the first AI to predict biological brain age from histopathology slides) and co-developed an AI system detecting alpha-synuclein pathology in peripheral biopsy tissue with reported sensitivity and specificity near 99%. His 2014 definition of PART (Primary Age-Related Tauopathy) is widely cited (over 1,000 citations, Google Scholar, accessed April 2026) and changed how the field classifies brain aging. He directs a major Alzheimer’s brain bank.

Questions he’ll raise: When AI can read a slide with near-perfect accuracy, what is the role of the human pathologist? How do we validate AI diagnostic tools for tissues we’ve never seen before? What does “trust” mean when the AI sees something the human can’t?

Preparation Reading

The Equity Architect — Algorithmic Fairness & Population Health

Rumi Chunara, PhD

Associate Professor of Biostatistics and Computer Science and Engineering, NYU; Director, Center for Health Data Science

Chunara pioneered social media-based disease surveillance, developed influential frameworks for equitable AI in health (Nature Machine Intelligence, 2021), and has directly measured algorithmic bias at NYC Health + Hospitals. Her 2024 PNAS paper demonstrated that utilizing big data without domain knowledge impacts public health decision-making. She is an ACM Distinguished Member, NSF CAREER awardee, and MIT TR35 honoree and is widely cited (Google Scholar, accessed April 2026).

Questions she’ll raise: Whose data trains the algorithm, and who is missing? Can fairness be “added on” after a model is built, or must it be designed in? When AI scales a biased system, who bears the harm?

Preparation Reading

The Systems Builder — Clinical AI Deployment at Scale

Yindalon (Yin) Aphinyanaphongs, MD, PhD

Research Professor, Dept. of Population Health & Medicine; Director, Division of Applied AI Technologies (DAAIT), NYU Langone Health

Aphinyanaphongs co-led the development of NYUTron, a BERT-like clinical language model trained on 7.25 million clinical notes, published in Nature (2023). His mortality prediction system, deployed as interruptive EHR alerts, has been associated with a substantial increase in goals-of-care conversations and a reduction in the Vizient Mortality Index. He has deployed more than 30 operational AI models at NYU Langone, led the rollout of GenAI Studio (a HIPAA-compliant generative AI platform with broad internal adoption), and oversees the Ultraviolet HPC cluster, one of the largest academic GPU clusters in the U.S.

Questions he’ll raise: What does it take to move an AI model from publication to production? When is the evidence “enough” to deploy? What organizational infrastructure sustains clinical AI at scale? Can AI improve not just efficiency but the quality of care conversations?

Preparation Reading

Section 6The Five Fault Lines

The roundtable is organized around five interlocking tensions. For each, we provide the core question, which speakers are most directly engaged, and what to think about beforehand.

1. Evidence & Deployment

Core question: What counts as sufficient evidence to deploy AI in clinical care?

Aphinyanaphongs has deployed more than 30 models; Chunara has shown that similar models can encode bias in safety-net hospitals. The threshold for “enough” evidence looks very different depending on whether you prioritize speed of impact or equity of outcomes.

Think about: How many patients need to benefit before deployment is justified? Who decides, and how do you account for patients who might be harmed?

2. The Body as Data

Core question: What is gained and what is lost when the patient becomes a data substrate?

From Choudhury’s smartphone stress sensors to Crary’s brain tissue slides to Brandes’ protein sequences, AI increasingly “knows” the patient through computational representations rather than clinical encounters.

Think about: Does an AI model that reads your genome “know” you? Does it know you differently than a model that reads your clinical notes? What about one that listens to your voice for signs of stress?

3. Equity & Scale

Core question: How do we prevent AI from scaling existing health disparities?

Chunara’s research shows bias is not hypothetical — it is measurable in real systems. Choudhury studies whether digital mental health tools work equitably across populations. Aphinyanaphongs deploys at scale across a major urban system.

Think about: If a model improves outcomes for 80% of patients but worsens outcomes for a vulnerable 20%, should it be deployed? Who makes that call?

Practical transparency tools: Model Cards, Datasheets for Datasets, and fairness-audit toolkits such as AIF360 and Fairlearn offer concrete approaches to bias documentation and mitigation.

4. Trust Across Scales

Core question: Can a single framework for “trust in medical AI” span from molecules to health systems?

Trust means something different when a clinician trusts a protein variant prediction (Brandes), an AI pathology reading (Crary), a smartphone behavioral assessment (Choudhury), an EHR mortality alert (Aphinyanaphongs), or a fairness audit (Chunara).

Think about: Is “trust” even the right word? When you “trust” a molecular prediction, is that the same kind of trust as when you trust a doctor?

5. Governance from the Inside

Core question: What happens when the people building AI are also the ones governing it?

Without a dedicated regulator on the panel, governance questions fall to the builders. Brandes writes on data governance, Chunara on responsible AI, Aphinyanaphongs on institutional AI policy, Choudhury on navigating FDA/IRB processes as an entrepreneur.

Think about: Can builders effectively regulate themselves? What does “governance” look like when it comes from practitioners rather than external regulators? Is this better or worse?

Section 7Key Tensions to Watch

The most generative conversations at Helix roundtables emerge from productive disagreements between speakers. Here are the pairings to watch:

Aphinyanaphongs
The deployer: “We have 31 models live. The evidence supports action.”
vs.
Chunara
The auditor: “Similar models encode bias in safety-net hospitals. We must audit first.”
Choudhury
Patient-facing AI: wearables, apps, direct-to-patient sensing
vs.
Aphinyanaphongs
Clinician-facing AI: EHR alerts, predictive models, workflow tools
Brandes
Foundational science: protein language models, pre-clinical tools
vs.
Aphinyanaphongs
Operational deployment: the “last mile” from model to bedside
Brandes
Molecular scale: AI reads protein sequences
vs.
Crary
Tissue scale: AI reads brain slides
Choudhury
Sensing as empowerment: early detection, patient agency
vs.
Chunara
Sensing as surveillance: who is monitored, who benefits?
Aphinyanaphongs
Near-term: maximize clinical impact now
vs.
Brandes
Long-term: existential AI safety risks demand caution
The Central Debate

The Aphinyanaphongs–Chunara tension is the panel’s defining axis. Both are at NYU. Both care about patient outcomes. They approach deployment from opposite directions — scale versus scrutiny, action versus audit. The question of whether to deploy first and audit later, or audit first and deploy later, is the roundtable’s most consequential operational disagreement. Listen for it early and throughout.

Section 8Glossary of Key Terms

AI & Machine Learning

Large Language Model (LLM)
A neural network trained on massive text data to predict and generate language. Clinical language models like NYUTron are trained on medical records for predictive tasks; NYUTron specifically is a BERT-like encoder model, distinct from generative chatbot-style LLMs.
Protein Language Model
A neural network trained on amino acid sequences (instead of words) to learn the “grammar” of proteins, enabling prediction of how mutations affect function.
Foundation Model
A large AI model trained on broad data that can be adapted to many downstream tasks. ESM (proteins) and GPT (text) are examples.
Computer Vision
AI systems that analyze images. In medicine: reading radiology scans, pathology slides, retinal photos, or dermatology images.
Multiple Instance Learning (MIL)
An ML approach for analyzing very large images (like pathology slides) by dividing them into patches and learning patterns across the collection. Used by Crary’s HistoAge.
Algorithmic Bias
Systematic errors in AI predictions that disadvantage certain groups, typically arising from biased training data or biased problem framing.
Model Drift
Degradation in model performance over time as the real-world data distribution shifts away from the training data.
AI Alignment
The research field focused on ensuring advanced AI systems act in accordance with human values and intentions. Brandes contributes to this field alongside his genomics work.
Passive Sensing
Continuous data collection from devices (smartphones, wearables) without requiring active user input. Core to Choudhury’s work.
Digital Biomarker
A measurable health indicator derived from digital data (e.g., voice patterns indicating stress, typing speed indicating cognitive state).

Medical & Clinical

EHR (Electronic Health Record)
Digital system storing patient data: clinical notes, lab results, medications, imaging, billing. Primary data source for clinical AI.
Clinical Decision Support (CDS)
Any system that provides clinicians with knowledge or patient-specific information to enhance decision-making. AI-powered CDS is what Aphinyanaphongs deploys.
Alert Fatigue
Clinician desensitization from excessive electronic alerts, causing important warnings to be dismissed. A major barrier to clinical AI adoption.
Missense Variant
A single-letter change in DNA that alters one amino acid in the resulting protein. Can be benign or disease-causing. Brandes’ models predict which.
GWAS (Genome-Wide Association Study)
A study scanning the entire genome of many individuals to find genetic variants associated with a disease or trait.
Tauopathy / PART
Diseases involving abnormal tau protein in the brain. PART (Primary Age-Related Tauopathy), defined by Crary, is a common age-related form distinct from Alzheimer’s.
Alpha-Synuclein
A protein whose abnormal aggregation is a hallmark of Parkinson’s disease and Lewy body dementia. Crary’s AI detects it in peripheral biopsies.
Histopathology
The study of tissue under a microscope to diagnose disease. “Digital pathology” converts this process to AI-analyzable images.
Social Determinants of Health (SDOH)
Non-medical factors (housing, education, income, neighborhood) that strongly influence health outcomes. Central to Chunara’s equity work.
Safety-Net Hospital
A hospital that provides care regardless of ability to pay, serving predominantly low-income and uninsured populations. NYC Health + Hospitals is the largest in the U.S.
Advance Care Planning
Conversations between patients, families, and clinicians about goals of care, especially near end of life. Aphinyanaphongs’ AI has been associated with a substantial increase in these conversations.
Vizient Mortality Index
A standardized metric comparing a hospital’s actual mortality to expected mortality, adjusting for patient severity. Used to benchmark care quality.

Section 9Further Reading & Courses

Books

Free Courses

Institutional Resources

Clinical Evidence & Governance

Bias, Fairness & Real-World Implementation

← Back to Medicine and AI