Time: 8:30 AM – 10:00 AM
Title: Statistical Challenges and Opportunities in the AI Agentic Era
AI has advanced so rapidly over the past decade that it is reshaping nearly every part of the world around us. It is no longer limited to simple prediction tasks. Today, in the pharmaceutical industry, AI - especially agentic AI - can support knowledge retrieval, automation, scientific hypothesis generation, workflow optimization, diagnosis, and much more. As these capabilities continue to expand, their impact on the workforce, including statistical roles, is likely to be profound.
For statisticians, this shift may feel disruptive, even unsettling. When AI agents can perform complex inference, automate analytical work, and generate high-level insights in minutes, or even seconds, it is natural to ask what role remains for human statisticians. Yet this new era also underscores why statistical thinking matters more than ever. Statistical principles are essential for understanding, validating, and governing agentic AI systems, and for ensuring that they are not only powerful, but also reliable, transparent, and scientifically grounded. This keynote will explore both the challenges and the opportunities that AI presents for pharmaceutical statisticians in this rapidly changing landscape.
Title: Solving the Diagnostic Odyssey: Bridging Macro-Level EHR Subphenotypes and Micro-Level Missense Variants
The journey toward true precision medicine has long been fragmented by a vast analytical divide: we observe macro-level clinical presentation in our health systems, yet we isolate micro-level genetic variations in our laboratories, frequently failing to translate the two into a timely diagnosis. This disconnected paradigm directly fuels the global "diagnostic odyssey" for hundreds of millions of rare disease patients, where sparse documentation, clinical pleiotropy, and millions of genetic variants of uncertain significance leave up to half of all suspected monogenic conditions entirely unresolved. To bridge this chasm, we introduce a unified, open-source translational framework that connects longitudinal health system phenomics with molecular deep learning by simultaneously aligning the phenotypic and genomic scales. At the clinical macro-level, our approach utilizes a transformer architecture and an iterative, self-refining weak-supervision loop to decode messy, real-world electronic health records—moving beyond traditional binary tracking to map continuous disease evolution and uncover highly distinct, prognostically critical patient subphenotypes. At the molecular micro-level, the framework maps these rich clinical phenotypes directly to specific single amino acid alterations by leveraging multi-modal contrastive learning to align protein sequence features and medical knowledge graphs within a shared metric space. Validated on real-world diagnostic dilemmas within major hospital networks and the Undiagnosed Diseases Network, this integrated pipeline routinely surfaces the definitive clinical diagnosis out of thousands of possibilities and pinpoints the true causal missense mutation as a top candidate. By closing the loop from raw clinical trajectories to residue-level biophysical alterations, this paradigm shifts AI-driven medicine past abstract pathogenicity scores and into the realm of precise, actionable, and scalable diagnostic discovery.
Title: Control of Unconditional Type I Error with Dynamic Bayesian Design: A Hybrid Frequentist-Bayesian Approach
Bayesian designs aim to improve trial efficiency by incorporating informative priors, which is mathematically equivalent to borrowing external data. Exchangeability and Type I error control have been major challenges in Bayesian designs. Conditional Type I error has been the de facto metric in the literature for Type I error in external-data borrowing. Research has shown that controlling the conditional Type I error at the alpha level will disallow external-data borrowing. Several authors have suggested that the main reason for the challenge is that conditional Type I error may not be the proper metric. Gao et al. (2025) proposed a hybrid approach combining Bayesian and frequentist ideas (Efron, 2005) and metrics for unconditional Type I error in external-control borrowing. These metrics can be applied in parallel to Bayesian designs. We propose an adaptive-design perspective in which a Bayesian design can be considered a frequentist adaptive design with an interim analysis at which the choice of prior is made. With this perspective, we introduce metrics for unconditional frequentist-style Type I error in Bayesian design and propose a dynamic Bayesian design. The unconditional Type I error can be controlled even when exchangeability may not strictly hold. The design allows external-data borrowing, provides more power than a frequentist design, and reduces sample size.
Time: 10:30 AM – 12:10 PM
Session Title: Innovative Analysis Methods in Rare Disease Research
Title: A Bayesian Dynamic Borrowing Approach to Support Regulatory Approval of Marstacimab in Pediatric Hemophilia
Pediatric drug development in rare diseases is often limited by small patient populations, ethical constraints, and the need to avoid delaying access to potentially beneficial therapies. These challenges are especially important in pediatric hemophilia, where early and effective prophylaxis is critical to reduce bleeding burden and long-term joint morbidity. When disease biology, pharmacology, and treatment response are sufficiently similar across age groups, pediatric extrapolation can provide a scientifically justified pathway to leverage evidence from older populations while preserving rigorous evaluation of pediatric data.
This presentation describes a prospectively planned Bayesian extrapolation framework developed to support pediatric evaluation of marstacimab, a once-weekly subcutaneous non-factor therapy for hemophilia prophylaxis. The approach used a robust mixture prior to dynamically borrow information from a completed adult/adolescent study while explicitly accounting for uncertainty in the transportability of evidence to pediatric patients. Rather than assuming full exchangeability, the framework allowed the degree of borrowing to adapt based on consistency between pediatric and external evidence. Borrowing was quantified using the expected local information ratio-based effective sample size, enabling transparent interpretation of external information on the pediatric sample-size scale.
A key component of the framework was simulation-based evaluation of operating characteristics before database release. Simulations examined the long-run probability of meeting Bayesian success criteria across clinically relevant treatment-effect scenarios, including null and alternative settings, and were used to calibrate prior weight and decision thresholds with attention to Type I error control. Sensitivity and tipping-point analyses further assessed robustness to assumptions about borrowing strength and decision criteria.
This case study illustrates how Bayesian extrapolation can serve as a transparent regulatory-science framework for rare disease development. The approach aligns with principles emphasized in FDA guidance on Bayesian methodology, including pre-specification, assessment of external-data relevance, operating-characteristic evaluation, and sensitivity analyses. More broadly, it highlights how simulation-guided Bayesian borrowing may accelerate pediatric rare disease evidence generation while maintaining safeguards against inappropriate reliance on external data.
Title: Comparison of Methods for Handling Death and Missing Data in Survivors in the Analysis of Functional Outcomes in Amyotrophic Lateral Sclerosis (ALS)
There are different analysis strategies to handle death events in the analysis of primary functional outcomes in ALS clinical trials. We compare the performance of several methods such as joint ranking, joint modeling, and ordinal modeling, in terms of bias, type I error rate, and power with a focus on the impact of missing data. We illustrate that the handling of missing data in the original formulation of the joint rank test of function and survival, also called the Combined Analysis of Function and Survival (CAFS) in ALS, is not adequate, only controlling the type I error under the strong null hypothesis. We identify the overall best performing method based on simulated data consistent with ALS natural history data.
Title: Bayesian Mixed Models for Repeated Measures with Informative Priors
The mixed model for repeated measures (MMRM) is a standard
approach for analyzing continuous longitudinal endpoints in
clinical trials. Despite growing interest in Bayesian borrowing
in clinical trials, particularly in rare disease settings,
direct application to the MMRM has remained limited, relying,
when pursued, on custom implementations. Specifying informative
priors on clinically meaningful quantities in a model with
treatment-by-visit interactions and additional covariates as in
the MMRM is inherently difficult, and in practice borrowing is
often reduced to synthesizing evidence at a single time point in
a simpler Bayesian model, leaving the longitudinal structure
largely unexploited. This talk presents a workflow that
addresses these challenges, covering planning, fitting,
assessing, and reporting of Bayesian MMRMs with informative
priors. Central to the workflow is the concept of
informative prior archetypes: standardized model
reparameterizations in which fixed-effect parameters correspond
directly to clinically interpretable quantities, enabling
deliberate and transparent prior assignment across visits and
treatment arms. This correspondence reduces the gap between the
statistical model and domain knowledge, and facilitates prior
elicitation and communication of assumptions to clinical and
regulatory stakeholders. The R package brms.mmrm
provides a flexible implementation of the proposed workflow,
lowering the barrier to applying Bayesian MMRMs with informative
priors.
Title: Bayesian Hierarchical Dose-Response Model Averaging in Small Clinical Trials for Decision-Making
In cell and gene therapy (CGT) and rare diseases, early-phase (Phase 1/2) clinical trials face unique statistical challenges when making Go/No-Go decisions for pivotal (Phase 3) development. These trials typically involve a limited number of investigational dose levels and very small sample sizes at each dose level. Furthermore, the Phase 2 proof of concept part does not include a control group for comparison. As a result, inferences relying solely on observed data from each dose level are not only inefficient but also can introduce significant bias, leading to an increased risk of incorrect decisions on progressing to Phase 3 (the pivotal phase) of development. Bayesian techniques offer more effective and informative statistical methodologies by integrating prior knowledge from real-world data or natural history study data and enhancing flexibility in decision-making. We propose a Bayesian hierarchical dose-response model averaging (BHDRMA) method for small clinical trials with a continuous primary efficacy endpoint. The proposed approach integrates data across all dose levels from both parts of the Phase 1/2 study, incorporates historical control information into the prior, and employs posterior weighted dose-response estimates through Bayesian model averaging and posterior probability-based Bayesian decision criteria to aid in Go/No-Go decision-making. In a simulation study, the proposed BHDRMA method consistently improves the precision in dose-response estimation compared with the conventional parametric bootstrap model averaging (BootsMA) approach, achieving substantial reductions in both average absolute prediction error and root-mean-square error.
Time: 1:10 PM – 3:00 PM
Session Title: AI and Machine Learning in Rare Disease Research
Title: Sample Size Reduction by Applying ML-Based Causal Inference Methods
We conducted a comprehensive comparative analysis of causal machine learning (ML) methods to assess their utility in improving the efficiency of clinical trial designs with or without historical data. Specifically, we compared standard ANCOVA analysis used in a randomized controlled trial (RCT) with several causal ML methods, including PROCOVA, TMLE, DML, and GRF. PROCOVA is gaining popularity in RCT design and requires historical data for prognostic scores, but other methods can be applied with or without such data. Our primary focus was on a strict RCT setting without borrowing historical control data, though we also explored the impact of borrowing data. The historical data consisted of placebo data from two Phase 3 ophthalmology studies with a continuous primary endpoint.
We employed a generative AI approach, specifically generative adversarial networks (GANs), to simulate RCT data from the historical data under various scenarios, varying treatment effects with and without treatment-effect heterogeneity, RCT sizes, and bias. Results showed that causal ML methods can increase power even without borrowing historical data. For example, TMLE increased effective sample size by 21% in one scenario. In scenarios with borrowing of controls, PROCOVA increased power while controlling Type I error, showing robustness to model misspecification.
Title: Leveraging AI to Enhance Study Design for a Pivotal Phase 3 Rare Disease Program
Background: Phase 3 rare disease trials often face small patient populations, heterogeneous treatment effects, and operational challenges. Adaptive enrichment designs can refine the target population using interim data but may be complex to evaluate, optimize, and communicate. We developed a human-in-the-loop, AI-enabled workflow to evaluate and optimize a two-stage adaptive enrichment design and support selection of a preferred design.
Methods: The workflow integrated a validated simulation engine with AI-assisted coding, scenario management, quality control, and results communication. Statisticians defined the design assumptions, optimization objectives, and population-selection rules based on Bayesian posterior predictive probability, while final inference preserved frequentist Type I error control through combination testing and multiplicity-control procedures. Statisticians reviewed AI-assisted code and verified the statistical reasoning and methodologies presented in the AI-generated report. Candidate designs were compared using power, population-selection performance, estimation properties, expected sample size, and false-positive control.
Results: AI-assisted automation accelerated the evaluation and optimization of a broad range of design options across plausible assumptions, reduced programming and reporting effort, and improved consistency across scenarios. Independent statistical review confirmed the design parameters, code, interim decision rules, and statistical validity for inference. Comprehensive operating-characteristic assessments, presented in standardized visual summaries, enabled development teams to effectively compare design trade-offs and select a preferred design.
Conclusions: A human-in-the-loop, AI-enabled workflow enhanced the efficiency and transparency of adaptive enrichment design evaluation and optimization. AI assisted with coding, visualization, and reporting, while statisticians retained responsibility for validation, interpretation, and final design decisions. This workflow can also facilitate rigorous evaluation and optimization of other innovative and complex designs, enable effective comparison of design trade-offs, and strengthen the evidentiary foundation for informed decision-making in late-stage clinical development.
Title: From Modeling and Simulation to Decision: Architecting Human-AI Collaboration for Trial Design
Clinical trial design is a complex decision-making process that extends far beyond numerical optimization. Although it often requires computationally intensive evaluation of multiple scenarios, the real challenge lies in integrating scientific understanding, regulatory acceptance, operational feasibility, and strategic program goals. As AI becomes increasingly capable of writing code, executing simulations, and summarizing results, an important question emerges: how should AI workflows be designed so they support, rather than oversimplify, biostatistical decision-making?
Using group sequential design as a practical example, this talk examines what effective human-AI collaboration could look like in clinical trial design. We will discuss how AI can support coding, simulation execution, synthesis of results, and structured exploration of scenario space. However, the aim is not to let AI evaluate endless combinations of possibilities. Instead, AI can be used as a thought partner to help narrow the problem to a manageable set of decision-relevant scenarios based on human-provided scientific context, strategic priorities, and hard constraints, and to help visualize and interpret the resulting trade-offs.
Humans remain central in defining the question, assessing the credibility and relevance of assumptions, and judging trade-offs across risk, cost, timeline, and evidentiary strength in the context of broader development goals. Through this example, we aim to show that the value of AI in trial design lies not simply in automation, but in creating workflows that make complex options easier to compare, discuss, and decide upon.
This work is not purely conceptual. Through a joint effort between the R Consortium and BBSW, we are actively translating these ideas into practice: from developing an open skill for group sequential design, to creating benchmark test cases, to designing human-in-the-loop workflows for realistic evaluation and use. This emerging effort is intended not only to demonstrate what AI can do, but to test how AI should be used in biostatistical workflows. We welcome community contributions to help expand the benchmark and shape this open-source effort.
Title: Agents Statisticians Can Trust in Regulated Domains: Rare Disease Trial Design as a Stress Test
AI assistants are increasingly handed work where a confident wrong answer is worse than no answer at all. This talk uses a live demonstration - an AI agent that helps design experiments - to show what it actually takes to trust one. The idea is simple: the agent is not allowed to make up numbers. It reasons about your problem and asks the right questions, but every calculation is handed off to tested, reproducible tools, and each result is checked before it is ever shown to you. Just as important, the agent knows the limits of what it can do and says so plainly instead of bluffing. We will watch it work, see how it catches its own mistakes, and step back to a broader lesson that applies far beyond this one tool: trustworthy AI is less about how smart the model sounds and more about how it is built - separating judgment from calculation, verifying every answer, and being honest about what it does not know.
Time: 3:30 PM – 5:10 PM
Session Title: Emerging Topics and Case Studies in Rare Disease Development
Title: TBD
Abstract: TBD
Title: Opportunities for Innovation Are Hiding in Plain Sight, Let's Find Them
The 1962 Kefauver-Harris Amendment was a miracle of politics, passed in the wake of the thalidomide tragedy. It catalyzed a revolution in public health by requiring "substantial evidence" from "adequate and well-controlled investigations" (21 CFR 314.126) for new therapies. The scaffolding for the modern clinical trial followed: a control group, randomization, blinding, and pre-specification of the analysis plan. While the fruits of neural nets and deep learning solve some problems previously addressed by statisticians, their utility in the controlled clinical trial seems limited, at least for now. Rare diseases benefit tremendously from careful thought on design and analysis, since each patient provides a proportionally large amount of information. ICH E9(R1) in 2019 defined the estimand framework, but why did it take 57 years for a basic idea like the estimand to appear in regulatory guidance?
While many fundamental ideas in mathematical and applied statistics have long existed, their applications, modifications, and amalgamations with other ideas result in an infinite set of future possibilities. A motivating example: is it possible to reliably test a hypothesis that has changed during a study? Many rare diseases do not have established, well-understood endpoints. Can we adapt to evolving information in an ongoing, blinded trial and still control Type I error? Another example: what should be done when a disease is so rare that traditional sample-size and Type I error requirements become infeasible? In a regulatory setting, the customary application of equipoise dictates that only data generated in the trial be used to test the primary hypothesis. When can external data reliably inform the primary estimand?
We will explore these questions and consider inspiration for new ideas with two case studies: (1) Hu, Yung, and Mackey (2025, Statistics in Biopharmaceutical Research) formalize an idea in support of a clinical trial on a rare, genetically driven form of early-onset Alzheimer's disease conducted in South America; and (2) a pediatric stroke trial shaped by the FDA's 2026 draft guidance on the use of Bayesian methods in clinical trials.
Title: Designing a Pivotal Trial Where a Placebo Is Impossible: External Controls and a Surrogate-Endpoint Accelerated-Approval Strategy for Gene Therapy in Sanfilippo Syndrome Type A
MPS IIIA (Sanfilippo syndrome type A) is an ultra-rare, fatal lysosomal storage disease of early childhood. Deficient sulfamidase lets heparan sulfate (HS) accumulate in the brain, driving relentless neuronal loss: affected children develop briefly, then regress, losing cognition, speech, and motor function, with death in the second decade. No therapy is approved.
That reality drives the design. A placebo arm is neither ethical nor feasible when decline is certain, and families will not accept sham gene therapy. The trial is therefore an open-label, single-arm study of a one-time intravenous AAV9 vector delivering a functional SGSH gene, with an external control drawn from two prospective natural history cohorts. Because treated infants were younger than the natural history patients, cognitive trajectories were compared using propensity-score (stabilized IPTW) weighting on baseline age and a non-linear growth-curve model, aligning the comparison on developmental slope rather than raw means.
The regulatory strategy runs on two tracks. Reduction of CSF HS, the primary disease-causing biomarker, is treated as a surrogate reasonably likely to predict clinical benefit - the basis for accelerated approval - and serves as the primary endpoint in one region, with the neurodevelopmental (BSITD-III Cognitive) endpoint primary in the other and the two transposed per health-authority feedback. Endpoints, analysis sets, the biomarker exposure metric, and the gated testing hierarchy were pre-specified in a statistical analysis plan aligned with regulatory agreements and finalized before database lock, with confirmatory testing on pooled data across the treatment and long-term follow-up phases.
Title: TBD
Abstract: TBD
Time: 8:30 AM – 10:00 AM
Title: Interim Monitoring in snSMART Designs with Multiple Endpoints
The small-n, Sequential, Multiple Assignment, Randomized Trial (snSMART) design improves efficiency in rare disease trials by allowing participant re-randomization across treatment stages. This talk extends the snSMART framework to address interim monitoring when multiple correlated endpoints jointly define trial success. We propose a Bayesian predictive probability of study success (PrSS) framework that supports early stopping for futility or efficacy and removal of ineffective treatment arms, while accounting for endpoint correlation through joint modeling. Simulation studies demonstrate that joint modeling outperforms independent endpoint analyses in Type I error control and adaptive arm selection.
Title: Advancing Rare Disease Drug Development: DRDMG Perspective
Clinical trials for rare diseases face fundamental challenges in generating reliable evidence of effectiveness when patient populations are small and highly heterogeneous, the natural history of the disease is poorly characterized, established endpoints are few or lacking, and disease progression is variable. This presentation surveys recent FDA guidance and case examples on generating substantial evidence of effectiveness, as well as the use of innovative trial designs and analysis methods, including master protocols, adaptive designs, and Bayesian methods, that collectively reflect flexible and evidence-driven approaches to addressing challenges in drug development for small populations.
Time: 10:30 AM – 12:10 PM
Session Title: Innovative Clinical Trial Design in Rare Diseases
Title: Reframing BASIS as a Basket Trial Master Protocol: Lessons for Rare Disease Drug Development
The BASIS trial (NCT03938792) was a Phase 3 study evaluating marstacimab in hemophilia A and B. Although not originally described as a basket trial, BASIS incorporated many of the defining features of a basket trial master protocol, evaluating three clinically distinct patient populations under a single protocol while leveraging shared infrastructure, coordinated site operations, and a common regulatory strategy. Patients served as their own controls through a six-month observational lead-in period, and statistical analyses were conducted independently within each cohort without cross-cohort borrowing, reflecting important clinical and biological differences between patient populations.
Using BASIS as a case study, this presentation reframes the trial through a basket-trial lens and explores how master-protocol principles can be successfully applied in rare disease drug development. The discussion will highlight the scientific and operational efficiencies gained through a common protocol structure, including streamlined study execution, harmonized evidence generation, and coordinated regulatory interactions. The experience from BASIS demonstrates that basket-trial approaches can provide substantial value even when patient populations are analyzed independently and information sharing across cohorts is not appropriate.
More broadly, the presentation will examine opportunities and challenges for expanding the use of basket trials and master protocols in rare diseases and other settings involving small patient populations. Particular attention will be given to considerations for defining cohorts, balancing operational efficiency with population heterogeneity, and identifying circumstances in which a common development framework can accelerate evidence generation while maintaining scientifically rigorous evaluation of individual patient populations.
Title: A Quantitative Framework for the Design and Optimization of N-of-1 and N-of-Few Trials in Rare Neurological Diseases
The design of clinical trials for rare neurological diseases is challenging due to small and heterogeneous patient populations, limited natural history data, and the unique demands of disease-modifying therapies (DMTs). This presentation introduces a quantitative framework for the systematic design and optimization of N-of-1 and N-of-few trials - study designs that are increasingly relevant for ultra-rare diseases and precision DMTs such as antisense oligonucleotides and gene therapies, where evaluating treatment effects at the individual level is both scientifically compelling and often the only feasible approach. We will illustrate the framework's utility by in-silico comparison of several design and analysis strategies for a hypothetical DMT trial in Autosomal-Recessive Spastic Ataxia of Charlevoix-Saguenay (ARSACS).
Title: From Dose Finding to Decision Making: Comparing 3+3, BOIN, and BF-BOIN Designs and Integrating Go/No-Go Criteria in Early-Phase Development
Choosing the appropriate study design is critical in early-phase drug development as it affects how efficiently and reliably a safe and effective dose can be identified. The conventional 3+3 design is simple to implement but it is often not statistically efficient; it may select an inappropriate dose and does not fully leverage accumulating data to guide decision-making. A more robust design can identify the optimal dose more efficiently and with fewer patients, shortening development timelines and supporting regulatory submission. Beyond the study design, a clear decision rule is needed to determine whether a program should continue. An innovative Bayesian Go/No-Go framework can be used to establish pre-specified rules for declaring Go, No-Go, or Indeterminate (insufficient evidence to decide). These rules should be defined before data review and based on clinically and commercially meaningful thresholds, enabling faster and more transparent decisions.
We present a simulation-based comparison of 3+3, the Bayesian Optimal Interval (BOIN) design, and its backfilling extension (BF-BOIN) across different scenarios. In our simulations, BOIN and BF-BOIN picked the correct dose much more often than 3+3 and stopped early less often. Sample sizes were similar across all three designs, so the better accuracy did not require a bigger trial. BF-BOIN also collects more safety data through backfilling, though it is more complex to perform. We also apply the Bayesian Go/No-Go framework to the expansion phase using different thresholds and prior distributions to assess assurance and determine whether the program should stop or continue while maintaining transparency and statistical rigor.
In conclusion, model-assisted designs and Bayesian decision rules provide a robust and operationally efficient framework for accelerating early-phase development. This innovative approach can help sponsors avoid continued investment in drugs unlikely to succeed, provide regulators with a clearer rationale, and support patients to access effective treatments sooner.
Title: Doing More with Fewer: What Hierarchical Composite Endpoints and Win Statistics Offer Rare Disease Trials, an ALS Illustration
Rare and neurodegenerative disease trials face a persistent tension: patients are few, yet several outcomes all matter, including survival, clinical function, and biological markers of disease. Hierarchical composite endpoints (HCEs) analyzed with win statistics, though increasingly established in cardiovascular research, remain relatively unfamiliar in the rare disease community, where they may offer a promising solution. Using amyotrophic lateral sclerosis (ALS) as the primary illustration, we introduce this framework and show how it speaks to several challenges of small, heterogeneous populations.
Because the framework is modular, sponsors can fold additional endpoints of interest - functional scales, validated or emerging biomarkers, imaging, or digital measures - into a single clinically ordered hierarchy. This can increase statistical power or reduce the required sample size while enabling earlier detection of a treatment signal, an advantage that matters most when patients are scarce. FDA guidance for ALS already encourages combining survival and function into a single overall measure, citing the joint rank (CAFS) test as one such approach. Building on that foundation, we use simulation studies to show how win statistics fulfill the same objective while offering greater interpretability and comparability across trials of different sizes. The ordering of the hierarchy invites its own discussion: because the sequence encodes clinical priority, arranging the same outcomes differently - for instance, prioritizing survival over daily function - may answer distinct scientific questions and reflect different definitions of patient benefit. Looking ahead, we also outline emerging work integrating digital-twin methods to improve trial efficiency further. Throughout, our aim is to bring an interpretable, practical framework to the rare disease community and invite discussion on its design.