Federated training for secure database management and screening

Federated training with diplotype models addresses data fragmentation in newborn genetic screening, reducing false positives and enhancing accuracy and scalability while maintaining privacy, enabling efficient and cost-effective screening for genetic disorders.

WO2025231188A1PCT designated stage Publication Date: 2025-11-06RADY CHILDRENS HOSPITAL RES CENT
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/027182
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-04-30
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Current newborn genetic screening methods face challenges due to fragmented data storage and sharing, leading to high false positive rates, high costs, and under-representation of certain ancestries, making it difficult to identify genetic disorders accurately and provide timely therapeutic interventions.

Method used

A federated training approach where remote compute devices analyze genetic data independently, generating and transmitting processed information to a central system for dataset refinement, using diplotype models to reduce false positives and enhance screening accuracy while maintaining privacy.

Benefits of technology

Achieves a 97% reduction in false positives and maintains 99% sensitivity, enabling scalable, affordable screening for millions of babies within two weeks, with improved accuracy and reduced costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025027182_06112025_PF_FP_ABST
    Figure US2025027182_06112025_PF_FP_ABST
Patent Text Reader

Abstract

A dataset comprising a plurality of genetic variants is received. A first set of data generated at a first remote compute device that has access to a first set of genetic data and not a second set of genetic data and based on the first set of genetic data and not the second set of genetic data is received. A second set of data generated at a second remote compute device that has access to a second set of genetic data and not the first set of genetic data and based on the first set of genetic data and not the second set of genetic data is received. At least one misclassified genetic variant from the plurality of genetic variants is identified. The dataset is updated based on the at least one misclassified genetic variant.
Need to check novelty before this filing date? Find Prior Art

Description

ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 FEDERATED TRAINING FOR SECURE DATABASE MANAGEMENT AND SCREENING CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Patent Application No.63 / 640,749, filed April 30, 2024 and titled “NEWBORN GENOME BASED SCREENING,” and U.S. Patent Application No. 63 / 669,629, filed July 10, 2024 and titled “NEWBORN GENOME BASED SCREENING,” the contents of each of which are incorporated herein in their entirety. GOVERNMENT LICENSE RIGHTS

[0002] This invention was made with government support under grant numbers UM1TR004407 and R01HD101540 awarded by the National Institutes of Health (NIH). The government has certain rights in the invention. FIELD

[0003] Some implementations relate generally to federated training for secure database management and screening. Some implementations relate generally to Newborn Genetic Screening (NBS) with rapid Whole Genome Sequencing (rWGS) and more specifically to methods for evaluation at birth for early childhood onset of a severe disease for which effective treatment may be known. BACKGROUND

[0004] Newborn screening (NBS) is performed worldwide in ~140 million newborns annually to identify severe congenital disorders and initiate treatments at or before onset of symptoms. While NBS can greatly improve health outcomes, the number of genetic disorders screened has not kept pace with genomic or therapeutic innovation. Between 2006 and 2022, the number of core disorders that were recommended for NBS of dried blood spots (DBS) in the United States by the Recommended Uniform Screening Panel (RUSP) increased from 27 to 35, and the number of affected infants identified increased from 6,439 to 6,466. However, there are ~7,200 known genetic diseases and hundreds of new targeted treatments that have been approved or are in clinical trials. Over the past decade, rapid WGS (rWGS®) has developed into an effective diagnostic test (Dx-ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 rWGS®) for many heritable diseases and is gaining acceptance as a first-tier test for critically ill newborns with suspected genetic diseases.

[0005] Further, accurate screening of individuals (e.g., newborns) can be more difficult than accurate diagnostic of individuals. Screening can include testing individuals to identify those at risk for specific conditions, and attempts to detect subtle indicators without causing unnecessary stress or harm. Diagnosis, on the other hand, focuses on confirming a condition in individuals who already exhibit symptoms or possess known risk factors, making the process more targeted and precise.

[0006] Further, some known screening methods face challenges due to data being siloed across various systems that make it difficult to effectively share information because of, for example, privacy or legal restrictions. This fragmentation results in a lack of comprehensive data, making it harder to identify patterns or gain meaningful insights. Without unified access to information, the effectiveness of screening processes can be compromised.

[0007] More advanced methods are needed for automated screening of newborns for rare genetic diseases that either have an effective treatment or that are amenable to development of a genetic therapy in order to implement improved, etiology-informed management at or before onset of symptoms. SUMMARY

[0008] In an embodiment, a dataset comprising a plurality of genetic variants is received. A first set of data generated (1) at a first remote compute device that has access to a first set of genetic data and not a second set of genetic data and (2) based on the first set of genetic data and not the second set of genetic data is received without receiving the first set of genetic data. A second set of data generated (1) at a second remote compute device that has access to a second set of genetic data and not the first set of genetic data and (2) based on the first set of genetic data and not the second set of genetic data is received without receiving the second set of genetic data. At least one misclassified genetic variant from the plurality of genetic variants is identified using a model and based on the first set of data and the second set of data. The dataset is updated based on the at least one misclassified genetic variant to generate an updated dataset that achieves an improved result for a genetic screening.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 BRIEF DESCRIPTION OF THE FIGURES

[0009] Fig.1: Genome-based newborn screening (gNBS) has substantial potential to improve outcomes in hundreds of severe, childhood, single locus genetic disorders (SCGD). However, a major impediment to gNBS is imprecision due to variants classified as pathogenic (P) or likely pathogenic (LP) that are not SCGD causal. gNBS with 53,855 P and LP variants, 342 genes, 412 SCGD, and 1,603 efficacious therapies was positive in 73% of UK Biobank (UKB470K) adults, suggesting 97% false positives. The phenomenon of purifying hyperselection was used, which acts to decrease the frequency of SCGD causal diplotypes, to improve precision of gNBS. Training of gene-disease-inheritance pattern-diplotype tetrads in 618,290 subjects identified 280 variants contributing more positive diplotypes than consistent with purifying hyperselection and with little or no evidence of SCGS causality. Upon their removal, 1.9% of UKB470K adults were positive. In contrast, trained gNBS was positive in 7.2% of 3,118 critically ill children with suspected SCGD and in 7.9% of 705 infant deaths. When compared with diagnostic genome sequencing (DGS), gNBS had 98.4% recall, and, in 8 children with positive results by DGS and gNBS, was projected to decrease time-to-diagnosis by a median of 121 days, avoid life-threatening disease presentations in 4 children, organ damage in 6, save ~$1.25 million in healthcare cost, and ten (1.4%) infant deaths. Federated training in large cohorts based on purifying hyperselection provides a general framework for achievement of acceptable precision in gNBS and provides a privacy-preserving mechanism for analyses across clinical trials which can be used to qualify gNBS across diverse genetic ancestries.

[0010] Fig.2A: Technical approach to structured, adaptive development of a newborn genome sequencing (NGS) SCGD screening, diagnosis, and treatment platform. a. Development of a structured SCGD molecular and treatment knowledgebase and screening model that is trained in multicentric, large diplotype models. Federated training identifies variants inNGS genes contributing to diplotypes with frequencies … inconsistent with purifying hyperselection, such that…, are greater than P, the population prevalence of the corresponding genetic disease(s) after correction for penetrance (p), expressivity (e), diplotype heterogeneity (d), and locus heterogeneity (l). b. Highly automated platform for scalable population screening,ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 diagnosis, and treatment that is empowered by the knowledgebase and trained model. GS, genome sequence; SME, subject matter expert; Rx, treatment.

[0011] Fig. 2B: Technical approach to structured, adaptive development of the NGS genetic disease screening, diagnosis, and treatment platform. a. Development of a structured rare disease molecular and treatment knowledgebase and screening model that is informed by large diplotype models. For simplicity, a monocentric workflow with centralized learning is depicted. In some implementations, however, a multicentric workflow with federal or split learning can be employed. b. Highly automated platform for standardized, scalable population screening, diagnosis, and treatment that is empowered by the large diplotype model MD, physician; SME, subject matter expert; Rx, treatment.

[0012] Fig. 3: Training of the NGS genetic disease screening model (e.g., algorithm) in multicentric, large diplotype models. a. Federated training in large GS cohorts flags P or LP variants for evaluation as non-severe disease causing in childhood (NSDCC) based on absence of purifying hyperselection evidenced by contributing diplotype frequencies (f) that are greater than those expected based on the sum of the corresponding disease prevalences (P) following correction for penetrance (p), expressivity (e), diplotype heterogeneity (d), and locus heterogeneity (l). b. Manhattan plot of counts of 2,785 diplotypes that were gNBS positive in UKB470K. 122 diplotypes with count >49 in UKB470K (frequency >1 in 9,592) are indicated in a first color 302 if disease causal (n=20), and in a second color 304 if determined to be NSDCC (n=102) using the method of a. The top 109 cystic fibrosis-CFTR diplotypes (with counts >3, 1 in 118,000) are also indicated as the first color 302 if disease causal (n=4) and red if not (n=105). The top 18 glucose 6-phosphate dehydrogenase deficiency-G6PD diplotypes (with counts >39, 1 in 12,051) are also indicated as the first color 302 if disease causal (n=7) and the second color 304 if not (n=11). The UKB470K cutoff count corresponding to the cutoff corrected UK adult diplotype frequency (fcorrected) for cystic fibrosis-CFTR (3.9 / 100,00, count <18, Table 1) and glucose 6-phosphate dehydrogenase deficiency-G6PD (53.3 / 100,00, count <251, Table 1) are shown in dotted lines having a third color 306. c. Rank ordering of 2,785 diplotype counts in UKB470K from largest (left) to smallest (right). The top 10 (third color 306) and 100 diplotypes (fourth color) accounted for 91% and 97%, respectively,ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 of the total diplotype count. The 122 diplotypes with frequencies >1 in 9,592 (counts >49) are indicated in the first color if disease causal (n=20), and in the second color 304 (n=102) if determined to be NSDCC using the method of a indicating the power to reduce false positives.

[0013] Fig. 4. illustrates a system block diagram to update a dataset based on variants generated using separate compute devices, according to an embodiment.

[0014] Fig. 5A illustrates an example flow comparing an observed prevalence to a corrected expected prevalence, according to an embodiment.

[0015] Fig.5B illustrates an example flow comparing a corrected observed prevalence to a corrected expected prevalence, according to an embodiment.

[0016] Fig. 6 illustrates an example of a flow iteratively analyzing data derived from genetic data, according to an embodiment.

[0017] Fig. 7 illustrates a flowchart of a method to identify a misclassified gene, according to an embodiment.

[0018] Fig. 8 illustrates a flowchart of a method to generate a more accurate dataset, according to an embodiment.

[0019] Fig.9 illustrates a flowchart of a method to determine whether a variant is a risk factor for a genetic disease, according to an embodiment. DETAILED DESCRIPTION

[0020] Implementations disclosed herein provide a platform with scalability to allow, for example, millions of babies to be screened for genetic disorders within, for example, two weeks from birth at affordable costs. In some embodiments, a 97% reduction in false positives was achieved for severe childhood diseases (positive rate decreased from 74% to 2%) while maintaining a 99% sensitivity / recall rate compared to “the gold standard” of rapid diagnostic genome sequencing (RDGS). Implementations disclosed herein achieve these results using a combination of three aspects: 1) increasing genetic information content by analyzing diplotypes instead of alleles or genotypes, 2) increasing diversity and volume of genetic data while maintaining patient privacy byATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 using query federation instead of data aggregation, and 3) significantly decreasing false positive rates while maintaining sensitivity using purifying hyperselection to model an iterative process.

[0021] Compared to screening, diagnosis can be more targeted because the focus is on confirming a condition in individuals who already show symptoms or have clear risk factors. Screening, on the other hand, can be more challenging as the process involves evaluating a broader population to identify potential risks. Thus, screening for genetic diseases can be difficult to pinpoint and more susceptible to inaccuracies, such as false positives. Further, known screening methods face significant limitations due to the fragmented nature of data storage and sharing. Genetic data is often siloed across multiple remote systems, with privacy laws and ethical considerations preventing data exchange. These restrictions hinder the ability to compile comprehensive datasets that could otherwise enhance the accuracy and efficiency of screening processes. Without centralized or unified access, some known screening methods struggle to identify patterns or make informed decisions, resulting in missed opportunities for advancements in diagnostics, research, and preventive measures. The lack of collaboration between isolated systems ultimately leads to gaps in identifying pathogenic and non-pathogenic genes, limiting progress in critical areas such as healthcare and genetic research. To address this challenge, some implementations relate to a framework where remote computers analyze their genetic data independently, ensuring privacy and compliance with legal restrictions. Instead of sharing raw genetic data, these computers generate and transmit processed information, such as statistical data or machine learning model weights, to a central computer. This central system maintains a “master” dataset of pathogenic and non-pathogenic genes and uses an iterative process to refine its accuracy based on the incoming data. By aggregating insights from multiple remote systems without compromising privacy, the central computer can improve the reliability of its dataset, eliminating false positives and enhancing the overall effectiveness of screening methods.

[0022] Current newborn genetic screening methods suffer from high cost, high false positive rates, lack of power for ultra-rare disease, and under-representation of non- European ancestries. Some implementations address high cost by achieving over 95% automation, including pre-validation of every gene, disease, inheritance pattern, andATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 variant in the screening process. Further, some implementations significantly reduce false positive rates using query federation with diplotype models tuned for penetrance, expressivity, locus heterogeneity, diplotype heterogeneity, disorder heterogeneity, and age / affected status of cohort. Additionally, some implementations increase power of detection for rare diseases by oversampling populations enriched for specific disorders. Some implementations address under-representation of certain ancestries by using query federation with diverse population genetic architectures and genetic ancestries. In some implementations, query federation refers to a method of analyzing genomic data remotely without data being moved or shared.

[0023] Disclosed herein are methods (e.g., performable by a processor, such as processor 402 in Fig. 4) for newborn screening comprising: a) performing genome sequencing on a sample from a newborn individual; b) querying the genomic sequence with a set of variants associated with severe childhood genetic disorders, wherein the set of variants has been prequalified by training in large cohorts to achieve high recall and positive predictive value; c) identifying positive findings based on the query results and predefined inheritance patterns; d) generating a report of the positive findings; and e) providing an electronic management guidance system to facilitate the translation of positive screening results into effective therapeutic interventions.

[0024] In some embodiments, prequalifying the set of variants comprises: (a) calculating an expected prevalence for each disorder based on population data; (b) observing a genetic prevalence for each disorder in one or more large genomic datasets; (c) comparing the observed genetic prevalence to the expected prevalence; and (d) removing or adjusting variants associated with disorders where the observed genetic prevalence significantly exceeds the expected prevalence.

[0025] In some embodiments, calculating the expected prevalence comprises adjusting for penetrance, expressivity, and / or locus heterogeneity of the disorder.

[0026] In some embodiments, observing the genetic prevalence comprises summing frequencies of diplotypes containing variants associated with each disorder.

[0027] In some embodiments, comparing the observed genetic prevalence to the expected prevalence uses a statistical threshold to identify significant differences.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0028] In some embodiments, the large cohorts comprise genomic data from multiple sites and federated learning is used for prequalification.

[0029] Some embodiments further comprise supplementing the variant query with identification of novel loss-of-function variants.

[0030] Some embodiments further comprise using an artificial intelligence tool to assist in interpreting query results.

[0031] In some embodiments, the electronic management guidance comprises information on confirmatory testing, therapeutic interventions, and / or their efficacy.

[0032] In some embodiments, the severe childhood genetic disorders comprise at least 300 disorders associated with at least 300 genes.

[0033] Also disclosed herein are systems for newborn screening of genetic disorders, comprising: a) a sequencing device configured to obtain genomic sequences from newborns; b) a database containing a prequalified set of variants associated with severe childhood genetic disorders; c) a processor configured to query the genomic sequences with the prequalified set of variants and identify positive findings based on predefined inheritance patterns; d) a report generator configured to produce a report of the positive findings; and e) an electronic management guidance module configured to provide information on therapeutic interventions for positive findings.

[0034] Some embodiments further comprise a federated learning module configured to prequalify the set of variants using genomic data from multiple sites.

[0035] Some embodiments further comprise an artificial intelligence module configured to assist in interpreting query results.

[0036] In some embodiments, the database is regularly updated based on new genomic data and clinical knowledge.

[0037] Also disclosed herein are methods (e.g., performable by processor, such as processor 402) for prequalifying variants for use in newborn genomic screening, comprising: a) obtaining genomic data from multiple large cohorts; b) calculating an expected prevalence for each disorder of interest based on population data; c) observingATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 a genetic prevalence for each disorder in the genomic data; d) comparing the observed genetic prevalence to the expected prevalence; e) identifying variants associated with disorders where the observed genetic prevalence significantly exceeds the expected prevalence; and (f) removing or adjusting the identified variants in a screening database.

[0038] Also disclosed herein are methods (e.g., performable by processor, such as processor 402) for achieving optimal outcomes in individuals with genetic diseases comprising presymptomatic population screening by genomic sequencing comprising automated interpretation and electronic management guidance to translate positive screens into effective therapeutic interventions.

[0039] In some embodiments, the automated interpretation is accomplished by querying variants in the genomic sequence of individuals with genetic diseases with a prequalified set of variants, diplotypes, haplotypes, genes, disorders, inheritance patterns or any combination thereof.

[0040] In some embodiments, the prequalified set of variants, diplotypes, haplotypes, genes, disorders, inheritance patterns or any combination thereof, is subjected to a prequalification by training in genomic sequences of large cohorts of individuals to achieve sufficiently high recall and positive predictive value to be acceptable for presymptomatic population screening.

[0041] In some embodiments, the prequalified set of variants, diplotypes, haplotypes, genes, disorders, inheritance patterns or any combination thereof is subjected to a prequalification by training in genomic sequences of large cohorts of individuals enables identification of variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns, or any combination thereof, that do not cause severe disease with sufficient penetrance and expressivity to be acceptable for presymptomatic population screening.

[0042] In some embodiments, the prequalification by training in genomic sequences of large cohorts of individuals leads to removal of variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns, or any combination thereof, that do not cause severe disease in individuals with genetic diseases.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0043] In some embodiments, variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns, or any combination thereof, that do not cause severe disease in screened individuals are identified based on having higher observed genetic prevalence in genomic sequences of large cohorts of individuals than an expected prevalence of one or more disorders to be screened in individuals with genetic diseases.

[0044] In some embodiments, the expected prevalence of one or more disorders to be screened and associated with a gene is determined by the sum of the frequency of the prevalence of each of those disorders in the individuals with genetic diseases.

[0045] In some embodiments, the expected prevalence of one or more disorders to be screened and associated with a gene is corrected for the penetrance, expressivity, and locus heterogeneity of those disorders in the individuals with genetic diseases.

[0046] In some embodiments, the observed genetic prevalence of one or more disorders associated with a gene to be screened is determined by the sum of the observed frequencies of diplotypes containing many putatively pathogenic variants in combinations that fit a pattern or patterns of inheritance of those disorders in genomic sequences of large cohorts of individuals.

[0047] Some embodiments further comprise, prequalification of a pattern or patterns of inheritance of one or more disorders associated with a gene to be screened for which the observed genetic prevalence in genomic sequences of large cohorts of individuals exceeds their expected prevalence in a population to be screened is achieved by changing from dominant to recessive.

[0048] In some embodiments, the expected prevalence of each putatively pathogenic variant associated with one or more disorders associated with a gene to be screened is the product of diplotype heterogeneity in the population and the expected prevalence calculated according to any of the preceding embodiments.

[0049] In some embodiments, the expected prevalence in the population to be screened is adjusted under a suitable distribution and one-tailed confidence interval.

[0050] In some embodiments, the variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns or any combination thereof, that do not cause severe disease inATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 screened individuals are identified or refined by the additional step of identification of other evidence of lack of disease causality.

[0051] In some embodiments, a plurality of genomic sequences of large cohorts of individuals is used.

[0052] In some embodiments, the plurality of genomic sequences of large cohorts of individuals are aggregated at a single site.

[0053] In some embodiments, the plurality of genomic sequences of large cohorts of individuals are at multiple sites and federated or distributed learning is used for prequalification.

[0054] Also disclosed herein are methods (e.g., performable by processor, such as processor 402) for achievement of optimal and / or improved outcomes in individuals with genetic diseases through presymptomatic population screening by genomic sequencing with automated interpretation and electronic management guidance to translate positive screens into effective therapeutic interventions.

[0055] In some embodiments, automated interpretation is accomplished by querying the variants in an individual’s genomic sequence with a prequalified set of variants, genes, disorders, and inheritance patterns. In some embodiments, variants, diplotypes, haplotypes, genes, disorders, and / or inheritance patterns are prequalified by training in genomic sequences of large cohorts of individuals to achieve sufficiently high recall and positive predictive value to be acceptable for presymptomatic population screening.

[0056] In some embodiments, training in genomic sequences of large cohorts of individuals enables identification of variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns that do not cause severe disease in screened individuals with sufficient penetrance and expressivity to be acceptable for presymptomatic population screening.

[0057] In some embodiments, prequalification by training in genomic sequences of large cohorts of individuals leads to removal of variants, diplotypes, haplotypes, genes, disorders, and / or inheritance patterns that do not cause severe disease in screened individuals.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0058] In some embodiments, variants, diplotypes, haplotypes, genes, disorders, and / or inheritance patterns that do not cause severe disease in screened individuals are identified based on having higher observed prevalence in genomic sequences of large cohorts of individuals than their expected prevalence in the population to be screened.

[0059] In some embodiments, the expected prevalence of one or more disorders to be screened and associated with a gene is determined by the sum of the frequency of the prevalence of each of those disorders in the population.

[0060] In some embodiments, the expected prevalence of one or more disorders to be screened and associated with a gene is corrected for the penetrance, expressivity, and / or locus heterogeneity of those disorders in the population. In some embodiments, the observed prevalence of one or more disorders associated with a gene to be screened is determined by the sum of the observed frequencies of diplotypes containing many putatively pathogenic variants in combinations that fit the pattern or patterns of inheritance of those disorders in genomic sequences of large cohorts of individuals.

[0061] In some embodiments, prequalification of the pattern of inheritance of one or more disorders associated with a gene to be screened for which the observed prevalence in genomic sequences of large cohorts of individuals exceeds their expected prevalence in the population to be screened is achieved by changing from dominant to recessive.

[0062] In some embodiments, the expected prevalence of each putatively pathogenic variant associated with one or more disorders associated with a gene to be screened is the product of diplotype heterogeneity in the population and the expected prevalence calculated according to the embodiments described herein.

[0063] In some embodiments, the expected prevalence in the population to be screened is adjusted under a suitable distribution and one-tailed confidence interval.

[0064] In some embodiments, variants, diplotypes, haplotypes, genes, disorders, and / or inheritance patterns that do not cause severe disease in screened individuals are identified or refined by the additional step of identification of other evidence of lack of disease causality.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0065] In some embodiments, a plurality of genomic sequences of large cohorts of individuals is used.

[0066] In some embodiments, the plurality of genomic sequences of large cohorts of individuals are aggregated at a single site.

[0067] In some embodiments, the plurality of genomic sequences of large cohorts of individuals are at multiple sites and federated or distributed learning is used for prequalification.

[0068] An initial, general solution to the problem of rapidly progressive childhood genetic diseases is rapid diagnostic WGS. Rapid WGS mitigated the problem of unknown etiology, wherein it was difficult to make a molecular diagnosis for most genetic diseases during hospitalization. Since then, rapid WGS has increased in speed, diagnostic performance, and scalability. Rapid WGS now allows concomitant evaluation of almost all differential diagnoses which may number over 1,000 genetic disorders in a single patient. Rapid WGS has started to be implemented nationally for inpatient diagnosis of genetic disease in England, Australia, and Wales and in several US states.

[0069] Rapid WGS mitigated the rare disease diagnostic odyssey bottleneck, but exposed another downstream bottleneck in the therapeutic odyssey that results in missed opportunity for clinicians. Clinical trials of rapid WGS have repeatedly shown gaps between expected and observed clinical utility. Several factors contribute to missed clinical utility. Firstly, exponential advances in genomics have outpaced medical education. Many healthcare providers lack adequate genomic literacy to practice genomic medicine unaided. Neonatologists, intensivists, and hospitalists are often dependent upon other subspecialists, particularly medical geneticists, for translation of rapid WGS results into treatment recommendations. In quaternary hospitals, this leads to treatment delays. In front-line settings, however, these delays can greatly limit the clinical utility of rapid WGS. Secondly, many genetic diseases were either discovered only recently, or are ultra-rare, and therefore evidence-based treatment guidelines have not yet been developed. Moreover, effective management strategies are often interspersed across the literature in the form of case reports, case series or small cohort studies. Information resources pertaining to management of rareATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 genetic diseases are incomplete, lack interoperability, and are typically not targeted toward acute ICU treatment or front-line physicians. Upon receipt of a rapid WGS- based diagnosis, these factors put an unsupportable burden on front-line physicians to search and synthesize the available evidence for rare genetic diseases, many of which they may have never encountered previously. Therapeutic unfamiliarity may continue to increase as new diseases and new effective, n-of-few, genetic therapies proliferate. Thirdly, failure to order rapid WGS as a first-tier test frequently leads to return of results around time of hospital discharge, when management plans have been solidified or, for rapidly progressive diseases, too late to have full clinical utility.

[0070] While historically it took an average of twelve years for drug development and approval, effective genetic therapies for genetic diseases can be developed and receive expanded-access investigational clinical protocol authorization by the Food and Drug Administration in as little as one year (such as milasen for Batten disease); Thus, newborn screening by WGS should consider not only those conditions for which current treatments exist, but also conditions for which novel genetic therapies can be developed in a timeframe that is pertinent to disease progression. 2. Genetic therapies can delay progression and death in patients with fatal genetic diseases (such as onasemnogene abeparvovec for the treatment of symptomatic patients with spinal muscular atrophy and eteplirsen for Duchenne Muscular Dystrophy. However, these genetic therapies do not reverse damage to organs. Instead, they prevent disease progression. Thus, it can be desirable to identify these conditions at birth, rather than at onset of symptoms in order to have improved (e.g., maximal) efficacy. 3. While the cause of more than 6,100 genetic diseases are known, for most genomic diagnoses, frontline physicians will never have encountered a patient with that disorder, indicating the desire to provide management guidance as part of a system of newborn screening by WGS. 4. While effective treatments exist for ~600 genetic diseases, for the vast majority the evidence of effectiveness is limited to case reports and case series, preventing ready access to such knowledge by frontline physicians, further indicating the desire to provide management guidance as part of a system of newborn screening by WGS.5. Some known population newborn screening is currently limited to 35 core conditions on the Advisory Committee on Heritable Disorders in Newborns and Children’s Recommended Uniform Screening Panel. As a result, there remain hundreds of genetic diseases with effective treatments that are not currently screened for.6. TheATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 cost of research-quality human genome sequencing has dropped to $689 and is projected to be $100 in a few years. 7. California Newborn screening currently costs $129 per subject, and therefore genome sequencing is approaching the cost effectiveness needed for newborn screening.8. Genome sequencing is now possible in 13.5 hours, a turnaround time that is sufficient for newborn screening. 9. Genome sequence analysis can be automated and therefore scaled to populations, as needed for newborn screening of approximately 4 million US births per year. 10. Genome sequencing, analysis and virtual treatment guidance can be automated; otherwise, genome sequencing is not practically performable as populations scale.

[0071] Some known newborn screening is focused on screening for a small number of individual diseases for which strong evidence supports the effectiveness of currently available treatments. Previous attempts to convert newborn screening to whole exome sequencing have also focused on screening for the same types of diseases. They have failed in part because of insufficient sensitivity relative to traditional newborn screening. For example, newborn screening of 48 disorders by whole exome sequencing had a sensitivity of 88% compared to 99.0% for traditional newborn screening. In another example, newborn screening by whole exome sequencing had a sensitivity of 88% in children with metabolic disorders and 18% in children with hearing loss. The system of newborn screening by WGS disclosed herein instead focuses on screening ~600 genetic diseases with effective treatments and genetic diseases for which novel genetic therapies can be developed in a timeframe that is pertinent to disease progression. It should be noted that novel genetic therapies are often designed based not on the disorder pathology but rather on the class of genetic variant that causes the condition. Thus, patients with any disorder that is caused by variants that create premature stop codons may potentially be effectively treated with antisense allele specific oligonucleotide therapies that alter exon skipping. Newborn screening by WGS if focused on tens of thousands of variant diplotypes that are known to be pathogenic or known and likely to be pathogenic (defined by a subset of the American College of Medical Genetics criteria) that map to the ~600 genetic diseases with effective treatments and the genetic diseases for which novel genetic therapies can be developed in a timeframe that is pertinent to disease progression. Thus, newborn screening by WGS achieves cost effectiveness and clinical utility in aggregate across tens of thousands of variants and hundreds screened conditions, rather than on a condition-by-ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 condition basis. Insensitivity for any single screened variant or condition (“missing” true positives) is acceptable provided the aggregate clinical utility and cost effectiveness across all conditions and variants is acceptable. This is because the incremental cost of adding a new condition or variant to newborn screening by WGS is negligible, whereas in known newborn screening the cost was substantial.

[0072] Genome sequencing results in identification of 6 million variants per subject, most of which do not cause disease. Previous attempts to convert newborn screening to whole exome sequencing have utilized conventional interpretation methods that were frustrated by the need for many hours of interpretation and many false positives (low precision). For example, newborn screening of 48 disorders by whole exome sequencing had specificity of 98.4%, compared to 99.8% for some known newborn screening. For population newborn screening to be effective, having an extremely low rate of false positives (high precision) is desirable. By focusing on tens of thousands of variant diplotypes that are known to be pathogenic or known and likely to be pathogenic (defined by a subset of the American College of Medical Genetics criteria), and that have population frequencies that are less than that of the condition being tested for, newborn screening by WGS system achieves the requisite precision for population implementation.

[0073] Some known newborn screening is designed for a rather static set of disorders, with at best annual changes in screened disorders, decided upon by a federal committee. Likewise, previous attempts to convert newborn screening to whole exome sequencing and panel tests have been static. In addition, neither such known screening nor previous attempts to convert newborn screening to whole exome sequencing and panel tests include disorders for which novel genetic therapies could be developed in a meaningful timeframe. The methods of newborn screening by WGS disclosed herein are highly dynamic conditions and variants can be added or subtracted frequently (e.g., hourly, daily, monthly, etc.).

[0074] Some known previous attempts to convert newborn screening to whole exome sequencing or panel tests used static variant pathogenicity assertions and were not designed to be self-learning. With few exceptions, sensitivity and specificity were not improved based on learning from tested subjects. The methods of newborn screening by WGS disclosed herein were designed to be self-learning: Each individual patient’sATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 set of variants are uploaded into a master database of variants (e.g., updated dataset 408 of Fig. 4) and the allele frequency of each variant in the entire tested population is recalculated repeatedly (e.g., hourly, daily, monthly, etc.). Thus, likely false positive variants that occur more frequently in the population than the incidence of the corresponding genetic disease can be identified and blocked from consideration in subsequent patients tested. Likewise, upon confirmatory testing, true positive and false positive results are uploaded into the master database and the pathogenicity assertion is updated for subsequent patients tested. Self-learning cannot be dynamically retrofitted to a conventional, dense array database in which each patient adds 6 billion null values and six million non-null values. Instead, a non-obvious, sparse array database solution can be used that features exceptionally fast read / write capability and that is designed to support self-learning with regard to variant frequency and confirmatory test results. The database solution disclosed herein features sparse array representation of only the six million non-null WGS variants and of the ~30,000 variants that are screened for that is optimized and / or improved for exceptionally fast read / write capability and designed to support self-learning with regard to variant frequency and confirmatory test results on a per subject basis. The attributes of data storage managers sufficient for screening of millions of newborns per year for hundreds of genetic diseases by WGS.

[0075] Some known previous attempts to convert newborn screening to whole exome sequencing or panel tests were predicated upon selection of disorders that were “actionable,” meaning likely to result in a change in clinical management of the subject. As noted above, effective management strategies for individual genetic diseases are often interspersed across the literature in the form of case reports, case series or small cohort studies. It is thus non-obvious which disorders should be included in newborn screening by WGS. In the methods disclosed herein, inclusion of disorders and variants were predicated on expert curation of the subset of those variants for which an effective genetic therapy can be developed and the efficacy and / or quality of evidence of efficacy of available treatments for the set of disease-causing genes.

[0076] Described herein is a healthcare delivery platform for childhood genetic diseases that starts to meet these requirements, sometimes referred to as “NGS.” NGS development is adaptive and versioned in response to rapid advances in understanding of rare genetic disease prevalence, penetrance, expressivity, causal variants, efficaciousATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 therapies, and technical evolution of GS, informatics, and artificial intelligence (AI). In some implementations, the technical approach to adaptive NGS development has three elements:

[0077] 1. A structured rare disease molecular and treatment knowledgebase that grows by ~50 new disorders, ~300 new treatments, and ~30,000 additional variants per version 16-19. The knowledgebase informs a GS screening model that is trained by distributed or federated learning from multisite, large diplotype models to achieve analytic performance suitable for population risk screening (Fig. 2A). Large diplotype models are co-occurring, gene-specific, genotype pairs in hundreds of thousands of individual GS from subjects of known age, sex, and / or genetic ancestry, and, in some implementations, of known affected status for the diseases to be screened. Akin to training with large language models to identify rare disease phenotypes, large diplotype models can identify variants curated as pathogenic but which do not cause severe, childhood genetic diseases.

[0078] 2. A highly automated platform for standardized, scalable population screening, diagnosis, and treatment that is empowered by the knowledgebase, screening model, and an imputation model (e.g., a transformer) (Fig.2B). The former informs electronic clinical decision support (eCDS, called Genome-to-Treatment, GTRx) that disseminates results and management guidance in a manner understandable to frontline providers nationwide to upskill them to translate positive results into optimal and / or improved outcomes. The latter two provide automated, supervised GS interpretation, which can be used for scalability.

[0079] 3. Adaptive clinical trials that employ elements 1 and 2 to fill knowledgebase gaps, further train the screening model, and evaluate clinical utility and cost effectiveness across diverse demographics and geographies, together with acceptability to parents and providers (see, for example, two clinical trials having ClinicalTrials.gov ID numbers NCT06276348 and NCT06306521, the contents of each of which are incorporated herein in their entirety).

[0080] In addition to methods validation, herein results of evaluation of NGS regarding clinical requirements for population screening such as maintainability, reliability, usability, testability, recall, specificity, and clinical utility are reported.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0081] Disclosed herein are methods (e.g., performable by a processor, such as processor 402 of Fig. 4) for newborn population screening comprising obtaining a sample from a newborn individual; performing genome sequencing on the sample from the newborn individual; interpreting the genome sequencing data using large diplotype models to identify genetic variants associated with a set of genetic diseases; automatically classifying the identified genetic variants based on their potential to cause severe disease in the newborn individual; and providing an electronic management guidance system to facilitate the translation of positive screening results into effective therapeutic interventions.

[0082] In some embodiments, the large diplotype models are used to determine the recall, positive predictive value, clinical utility, and cost-effectiveness of the genome sequencing-based newborn population screening.

[0083] Some embodiments further comprise removing variants, diplotypes, haplotypes, inheritance patterns, and disease-gene dyads from the screening that are determined not to cause severe disease with sufficient penetrance and expressivity.

[0084] In some embodiments, non-severe-disease-causing variants are identified by comparing their genetic prevalence in the large diplotype models to the corresponding prevalence in the population to be screened.

[0085] In some embodiments, the genetic prevalence of disease-gene dyads is determined by summing the frequency of the positive diplotypes of the potentially pathogenic variants for that locus in the large diplotype models.

[0086] Some embodiments, further comprise correcting the genetic prevalence for the occurrence of more than one positive diplotype per gene.

[0087] In some embodiments, the disorder population prevalence is established as: the upper bound of prevalence after adjustment for genetic disease heterogeneity, genetic penetrance, and expressivity.

[0088] In some embodiments, genetic heterogeneity is defined as the proportion of disease-affected subjects associated with a specific gene-disease dyad.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0089] In some embodiments, penetrance is defined as the proportion of subjects who screen positive for a disease-gene dyad and are affected by that disease.

[0090] In some embodiments, expressivity is defined as the proportion of subjects who screen positive and are affected in whom the disease is of sufficient severity to warrant therapeutic intervention.

[0091] Some embodiments further comprise changing the screened pattern of inheritance of disease-gene dyads with higher genetic prevalence than adjusted population prevalence from dominant to recessive.

[0092] Some embodiments further comprise identifying and removing individual non- severe-disease-causing variants in disease-gene dyads from the screen.

[0093] In some embodiments, non-severe-disease-causing variants are identified by comparing their diplotype frequency in a large diplotype model to the maximum credible diplotype frequency in a matched population.

[0094] In some embodiments, the maximum credible population diplotype frequency is calculated by adjusting the disorder population frequency for diplotype heterogeneity.

[0095] In some embodiments, diplotype heterogeneity is defined as the upper bound of the proportion of genetic prevalence attributable to diplotypes containing a specific variant.

[0096] In some embodiments, maximum credible population diplotype frequency is adjusted using a suitable distribution and one-tailed confidence interval.

[0097] Some embodiments further comprise identifying non-severe-disease-causing variants in disease-gene dyads by additional evidence of lack of disease causality.

[0098] In some embodiments, a plurality of large diplotype models is used.

[0099] In some embodiments, the large diplotype models are aggregated at a single site.

[0100] In some embodiments, federated or distributed learning is used to query large diplotype models at multiple sites.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0101] Fig. 4 illustrates a system block diagram to update a dataset based on variants generated using separate compute devices, according to an embodiment. Fig.4 includes compute devices 400, 420, and 440, each operatively coupled to one another via network 180.

[0102] Network 480 can be any suitable communications network for transferring data, for example operating over public and / or private communications networks. For example, network 480 can include a private network, a Virtual Private Network (VPN), a Multiprotocol Label Switching (MPLS) circuit, the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a worldwide interoperability for microwave access network (WiMAX®), an optical fiber (or fiber optic)-based network, a Bluetooth® network, a virtual network, and / or any combination thereof. In some instances, network 480 can be a wireless network such as, for example, a Wi-Fi® or wireless local area network (“WLAN”), a wireless wide area network (“WWAN”), and / or a cellular network. In other instances, the network 480 can be a wired network such as, for example, an Ethernet network, a digital subscription line (“DSL”) network, a broadband network, and / or a fiber-optic network. In some instances, network 480 can use Application Programming Interfaces (APIs) and / or data interchange formats, (e.g., Representational State Transfer (REST), JavaScript Object Notation (JSON), Extensible Markup Language (XML), Simple Object Access Protocol (SOAP), and / or Java Message Service (JMS)). The communications sent via network 480 can be encrypted or unencrypted. In some instances, network 480 can include multiple networks or subnetworks operatively coupled to one another by, for example, network bridges, routers, switches, gateways and / or the like.

[0103] Each of the compute devices 400, 420, and 440 can be any type of compute device, such as a server, desktop, laptop, tablet, phone, and / or the like. Compute device 400 includes processor 402 operatively coupled to memory 404 (e.g., via a system bus), compute device 420 includes processor 422 operatively coupled to memory 424 (e.g., via a system bus), and compute device 440 includes processor 442 operatively coupled to memory 444 (e.g., via a system bus).

[0104] Each of the processors 402, 422, and / or 442 can be, for example, a hardware- based integrated circuit (IC) or any other suitable processing device configured to runATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 and / or execute a set of instructions or code. For example, processors 402, 122, and / or 442 can each be a general-purpose processor, a central processing unit (CPU), an accelerated processing unit (APU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic array (PLA), a complex programmable logic device (CPLD), a programmable logic controller (PLC) and / or the like. In some implementations, processors 402, 422, and / or 442 can be configured to run any of the methods and / or portions of methods discussed herein.

[0105] Each of the memories 404, 424, and / or 444 can be, for example, a random- access memory (RAM), a memory buffer, a hard drive, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), and / or the like. Memories 404, 424, and / or 444 can be configured to store any data used by processors 402, 422, and / or 442 (respectively) to perform the techniques (methods, processes, etc.) discussed herein. In some instances, memories 404, 424, and / or 444 can store, for example, one or more software programs and / or code that can include instructions to cause processors 402, 422, and / or 442 (respectively) to perform one or more processes, functions, and / or the like. In some implementations, memories 404, 424, and / or 444 can include extendible storage units that can be added and used incrementally. In some implementations, memories 404, 424, and / or 444 can be a portable memory (for example, a flash drive, a portable hard disk, a SD card, and / or the like) that can be operatively coupled to processors 402, 422, and / or 442 (respectively). In some instances, memories 404, 424, and / or 444 can be remotely operatively coupled with a compute device (not shown in Fig. 1). In some instances, memories 404, 424, and / or 444 is a virtual storage drive (e.g., RAMDisk), which can improve I / O speed and in turn, accelerate image reading and writing. In some implementations, the term “memory” can be used interchangeably with “storage device.”

[0106] Memory 424 of compute device 420 includes (e.g., stores) genetic data 428. Genetic data 428 includes information derived from one or more organisms’ nucleic acid sequences, which contains the instructions that govern biological development and function. This data includes sequences of nucleotides (A, T, C, G) that can be analyzed to identify patterns, mutations, or variations.

[0107] Memory 444 of compute device 440 also includes (e.g., stores) genetic data 448. Genetic data 448 includes information derived from one or more organisms’ nucleicATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 acid sequences, which contains the instructions that govern biological development and function. This data includes sequences of nucleotides (A, T, C, G) that can be analyzed to identify patterns, mutations, or variations.

[0108] Genetic data 448 is different than genetic data 428. In some implementations, genetic data 428 represents genetic data of a first population, and genetic data 448 represents genetic data of a second population mutually exclusive from the first population. In some implementations, genetic data 428 includes information derived from a group of organisms having a first trait, and genetic data 448 is derived from a group of organisms having a second trait different than the first trait. In some implementations, genetic data 428 includes information derived from a group of organisms having a first trait but not a second trait, and genetic data 448 is derived from a group of organisms having the second trait but not the first trait. As an example, genetic data 428 can represent genetic data collected from East Asians, while genetic data 448 can represent genetic data collected from Northern Europeans. Other examples of traits include geographic region (e.g., continent, country, region, etc.), health and medical traits (e.g., presence of genetic mutations, markers for inherited diseases, genetic predispositions), physical attributes (e.g., height, hair texture, skin pigmentation, etc.), and / or the like.

[0109] Compute device 420 does not have access to genetic data 448, and compute device 440 does not have access to genetic data 428. This separation can be due to, for example, data security, privacy policies, and / or organizational boundaries. In some implementations, compute device 420 is associated with (e.g., owned by, used by, accessible by) a first entity and not a second entity, while compute device 440 is associated with (e.g., owned by, used by, accessible by) the second entity and not the first entity.

[0110] Memory 424 of compute device 420 includes (e.g., stores) data 426. Data 426 can be, for example, statistical data and / or machine learning model weights. Data 426 can be generated at, for example, compute device 420 and not compute device 400 and / or 440. Data 426 can be generated based on genetic data 428 and not genetic data 448. In some implementations, a machine learning model is trained based on genetic data 428 and not genetic data 448; for example, genes from genetic data 428 can be used as input training data for the machine learning model and indications of whetherATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 those genes are pathogenic can be used as target training data such that the machine learning model, once trained, is configured to receive a gene and output an indication of whether the gene is pathogenic. Weights from the trained machine learning model can then be represented as data 426. Additionally or alternatively, data 426 can represent statistical data derived from genetic data 428, such as diplotypes of genetic data 428, zygosities of genetic data 428, genes of genetic data 428, patterns of inheritance (POIs) of genetic data 428, positive counts of genetic data 428, metadata of genetic data 428, and / or the like. In some implementations, genetic data 428 includes information not included in data 426; for example, genetic data 428 can include sensitive information (e.g., personal identifiers, medical conditions, biological relationships, etc.) that is not included in data 426.

[0111] Memory 444 of compute device 440 includes (e.g., stores) data 446. Data 446 can be, for example, statistical data and / or machine learning model weights. Data 446 can be generated at, for example, compute device 440 and not compute device 400 and / or 420. Data 446 can be generated based on genetic data 448 and not genetic data 428. In some implementations, a machine learning model is trained based on genetic data 448 and not genetic data 428; for example, genes from genetic data 448 can be used as input training data for the machine learning model and indications of whether those genes are pathogenic can be used as target training data such that the machine learning model, once trained, is configured to receive a gene and output an indication of whether the gene is pathogenic. Weights from the trained machine learning model can then be represented as data 446. Additionally or alternatively, data 446 can represent statistical data derived from genetic data 448, such as diplotypes of genetic data 448, zygosities of genetic data 448, genes of genetic data 448, patterns of inheritance (POIs) of genetic data 448, positive counts of genetic data 448, metadata of genetic data 448, and / or the like. In some implementations, genetic data 448 includes information not included in data 446; for example, genetic data 448 can include sensitive information (e.g., personal identifiers, medical conditions, biological relationships, etc.) that is not included in data 446.

[0112] Compute device 400 can receive data 426 from compute device 420 and receive data 446 from compute device 440 (e.g., without receiving genetic data 428 or genetic data 448). Compute device 400 can then update dataset 406 based on data 426 and dataATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 446 to generate updated dataset 408. Dataset 406 can represent a list of genes and an indication and / or likelihood as to whether those genes are predicted to be pathogenic or not; in some implementations, at least one gene from the list of genes is misclassified, for example, as a false positive (e.g., indicating that the gene is pathogenic when the gene is not actually pathogenic). Model 410 can be applied to data 426 and 446 to identify any misclassified genes in dataset 406. The misclassified genes in dataset 406 can be correctly classified by model 410 to generate updated dataset 408. Said differently, updated dataset 408 is a more accurate version of dataset 406 (e.g., with one or more false positives removed).

[0113] In some implementations, model 410 is a machine learning model and data 426 and 446 represent machine learning model weights. Thus, through federated training, model 410 can be generated and updated based on data 426 and 446. Thereafter, model 410 can receive dataset 406 and correct and / or identify (e.g., without human intervention) any misclassified genes in dataset 406 to generate updated dataset 408.

[0114] In some implementations, model 410 is configured to perform an iterative process to identify misclassified genes. In some implementations, data 426 and / or 446 (1) comprises genes associated with genetic diseases and (2) is from a cohort. The iterative process can include, for each gene, calculating an expected prevalence( ) for each genetic disease and an observed prevalence ( ) for each geneticdisease. Genes where the observed prevalence exceeds the expected prevalence beyond a predetermined acceptable threshold are identified and a diplotype is selected. Anexpected maximum frequency ( ) and an observed frequency ( ) ofthe diplotype is calculated, and diplotypes where the observed frequency exceeds the expected maximum frequency are detected. Variants that contribute to the detected diplotypes are then categorized into allowed variants (e.g., a whitelist) and blocked variants (e.g., a blocked list). Updated values for the observed prevalence and / or expected prevalence can then be calculated, and the above process can repeated until the observed prevalence is within the predetermined acceptable range of the expected prevalence. During each iteration, the observed prevalence and / or expected prevalence at the preceding iteration can be updated during that iteration; thus, in some implementations, each iteration updates the observed prevalence and / or expectedATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 prevalence until the observed prevalence is within the predetermined acceptable range of the expected prevalence.

[0115] In some implementations, updated dataset 408 includes a list of genes and / or variants, and an indication of whether each gene and / or variant is associated with one or more genetic diseases. Additionally or alternatively, updated dataset 408 indicates a risk and / or likelihood associated with each gene and / or variant for a genetic disease(s). In some implementations, updated dataset 408 includes an allow list (e.g., white list of genes and / or variants that are not pathogenic) and / or blocked list (e.g., of genes and / or variants that are pathogenic). Thus, in response to receiving a sample (e.g., genetic sample from an individual, such as a newborn or infant), the sample can be analyzed based on updated dataset 408 (e.g., and not dataset 406) to determine a risk factor for one or more genetic diseases. Said differently, the sample can be analyzed to determine an individual’s risk for a genetic disease(s). For example, compute device 400 can receive an electronic signal (e.g., from compute device 420, compute device 400, or a compute device not shown in Fig. 4) representing the sample, and the sample can be looked-up in updated dataset 408 to determine if the sample is blocklisted or allow listed. In some implementations, updated dataset 408 includes multiple blocked lists and / or multiple allow lists, such as a blocked list and / or allow list for each category from a plurality of categories; for example, updated dataset 408 can include a block list for northern Europeans and a separate block list for east Asians. Additionally or alternatively, updated dataset 408 can include a single blocked list and / or single allow list associated with a plurality a categories (e.g., northern Europeans, east Asians, etc.). In some implementations, updated dataset 408 is updated. For example, additional data can be received at compute device 400 (e.g., from a compute device not shown in Fig. 1) and that data can be used to update updated dataset 408. In some implementations, the additional data represents genes identified as pathogenic or not (or a likelihood of being pathogenic or not) based on genetic data different from genetic data 428 and 448. In some implementations, model 410 is updated. For example, if model 410 is a machine learning model trained via federated learning based on weights provided by compute device 420 and 440, compute device 400 can receive additional weights from compute device 420, 440, and / or a compute device not shown in Fig. 4 and updateATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 model 410 based on those additional weights. As another example, model 410 can be updated based on updated dataset 408 and not dataset 406. Because updated dataset 408 more accurately indicates risks associated with genes than dataset 406, model 410 can be more accurate if trained using updated dataset 408.

[0118] In some implementations, the iterative process can occur at compute device 420 and / or 440 instead of and / or in addition to compute device 400. For example, compute device 420 can store model 410 and model 410 can be configured to perform the iterative process based on data 426 to identify one or more pathogenic variants; compute device 400 can receive an indication of the identified one or more pathogenic variants from compute device 420 and update dataset 406 such that the identified one or more pathogenic variants are indicated as pathogenic. Similarly, compute device 440 can store model 410 and model 410 can be configured to perform the iterative process based on data 446 to identify one or more pathogenic variants; compute device 400 can receive an indication of the identified one or more pathogenic variants from compute device 440 and update dataset 406 such that the identified one or more pathogenic variants are indicated as pathogenic. This can result in, for example, a distributed system with multiple local models that are trained based on information from other models but that doesn’t receive and / or store data from the other models and / or that were used to train the other models. For example, compute device 420 can store a first model and compute device 440 can store a second model, where (1) the first model can be (a) trained based on genetic data 428 and (b) trained and / or updated based on data 446 and / or weights from the second model but not genetic data 448 and (2) the second model can be (a) trained based on genetic data 448 and (b) trained and / or updated based on data 426 and / or weights from the first model but not genetic data 428.

[0119] Although Fig. 4 illustrates three compute devices, in other implementations, more or less compute devices can be used. For example, the functionalities of compute device 400 can be split across multiple compute devices (e.g., a first compute device updates dataset 406 to generate updated dataset 408 and a second compute device stores model 410).

[0120] Fig. 5A illustrates an example flow comparing an observed prevalence to a corrected expected prevalence, according to an embodiment. For a given gene—CFTRin this example—the observed prevalence ( ) and corrected expected prevalenceATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001( ) can be determined based on data (e.g., data 426 and / or 446) derived fromgenetic data (e.g., data 428 and / or 448). If the ratio of the observed prevalence and corrected prevalence is outside a predetermined acceptable range (e.g., 60:1), the iterative process can occur until a stopping criterion is met (e.g., observed prevalence and expected prevalence being within the predetermined acceptable range). Thecorrected prevalence ( ) can be determined using , where represents anactual prevalence, represents an estimated locus heterogeneity, represents an estimated penetrance, and represents an estimated expressivity.

[0121] Fig.5B illustrates an example flow comparing a corrected observed prevalence to a corrected expected prevalence, according to an embodiment. For a given gene— CFTR in this example—genetic variants that contribute to diplotypes whose expected population frequency (fcorrected) was lower than the observed diplotype frequency beyond a predetermined acceptable threshold can be identified. Such genetic variants can be block-listed, and the iterative process can occur until a stopping criterion is met (e.g., corrected observed prevalence and corrected expected prevalence being within the predetermined acceptable range). A given diplotype’s expected population frequency (fcorrected) can include corrections for diplotype heterogeneity (d), penetrance (p), estimated expressivity (e), and estimated locus heterogeneity (l).

[0122] Fig. 6 illustrates an example of a flow iteratively analyzing data derived from genetic data, according to an embodiment. A query can be made (e.g., by compute device 400) to one or more remote compute devices (e.g., compute device 420 and / or 440) requesting data (e.g., data 426 and / or 446). For example, a query can be made to a United Kingdom-based bio bank database (e.g., UKB470K database) storing genetic data of adults from the United Kingdom, and a query can be made to a RCIGM database storing genetic data of children. Each remote compute device can analyze the genetic data (e.g., genetic data 428 and / or 448) stored at and / or is accessible by that compute device, and return (e.g., to compute device 400) the requested data (e.g., data 426 and / or 446) as the “result set.” As illustrated at Fig. 6, the result set can include information about diplotypes, zygosities, genes, POIs, counts, metadata (e.g., sex, genetic ancestry, affected status), and / or the like. In response to receiving the result set, the compute device (e.g., compute device 400) can perform the iterative analysis until a stoppingATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 criteria is met (e.g., ), resulting in, for example, a block list (e.g., for updated dataset 408). EXAMPLES EXAMPLE 1 - Clinical utility of newborn screening for 412 genetic diseases by genome sequencing with federated learning from large diplotype models

[0123] Building on biochemical newborn screening (NBS) success, genome sequencing (GS)-based NBS has the potential to improve outcomes for hundreds of severe single locus disorders with early childhood onset and effective treatments by identification and early intervention. A prototype GS-based NBS platform for 388 genetic diseases (NGS) has been described. NGS features distributed learning in diverse, large diplotype models (610,948 exomes from the UK Biobank (UKB470K) or Mexico City Prospective Study, and 7,342 children and parents who received rapid diagnostic GS (RDGS) in our clinical laboratory). The learning model is derived from a knowledgebase of 412 genetic diseases that includes estimated prevalence, penetrance, expressivity, locus heterogeneit, 54,168 curated pathogenic variants, inheritance modes, and 1,603 effeicacious therapeutic interventions. Twenty of 342 NGS genes had higher summated diplotype frequencies in UKB470K, corrected for locus and diplotype heterogeneity, penetrance, and expressivity, were examined for evidence of disease causality and blocked from NGS. Training decreased the NGS positive rate in UKB470K from 73% to 2.8%. In contrast, NGS was positive in 7.2% of 3,118 critically ill children with suspected genetic disease and had 98.4% recall when compared with RDGS. Counterfactual analysis of 8 critically ill children with positive results by NGS and RDGS suggested that NGS would have shortened time to diagnosis by a median of 121 days, avoided life-threatening disease presentations in 4 children, organ damage in 6, and save $1.25 million in healthcare cost. Using archived NBS dried blood spots, NGS was positive in 7.9% of 705 infant deaths in San Diego. Counterfactual analysis suggested that NGS would have avoided ten deaths. Distributed prevalence assessments in large diplotype models appear to provide a general framework for qualification of gene-disorder dyads, variants, and inheritance modes for NBS. Distributed learning provides a privacy-preserving mechanism for stratified analyses of clinical utility across numerous clinical trials of GS-based NBS.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0124] Genome-based newborn screening (gNBS) has substantial potential to improve outcomes in hundreds of severe, childhood, single locus genetic disorders (SCGD). However, a major impediment to gNBS is imprecision due to variants classified as pathogenic (P) or likely pathogenic (LP) that are not SCGD causal. gNBS with 53,855 P and LP variants, 342 genes, 412 SCGD, and 1,603 efficacious therapies was positive in 73% of UK Biobank (UKB470K) adults, suggesting 97% false positives. The phenomenon of purifying hyperselection, which acts to decrease the frequency of SCGD causal diplotypes, was used to improve precision of gNBS. Training of gene- disease-inheritance pattern-diplotype tetrads in 618,290 subjects identified 280 variants contributing more positive diplotypes than consistent with purifying hyperselection and with little or no evidence of SCGS causality. Upon their removal, 1.9% of UKB470K adults were positive. In contrast, trained gNBS was positive in 7.2% of 3,118 critically ill children with suspected SCGD and in 7.9% of 705 infant deaths. When compared with diagnostic genome sequencing (DGS), gNBS had 98.4% recall, and, in 8 children with positive results by DGS and gNBS, was projected to decrease time-to-diagnosis by a median of 121 days, avoid life-threatening disease presentations in 4 children, organ damage in 6, save ~$1.25 million in healthcare cost, and ten (1.4%) infant deaths. Federated training in large cohorts based on purifying hyperselection provides a general framework for achievement of acceptable precision in gNBS and provides a privacy- preserving mechanism for analyses across clinical trials which is critical to qualify gNBS across diverse genetic ancestries.

[0125] Newborn screening (NBS) is risk evaluation at birth for early childhood onset of a severe disease (e.g., severe, childhood, single locus genetic disorders (SCGD) with early onset) for which effective treatment / therapeutic interventions are available to reduce (e.g., minimize) morbidity and mortality by early intervention. Since inception in 1963, NBS has expanded worldwide to ~140 million newborns / year and up to 80 SCGD. In the US, NBS identifies ~12,500 affected infants per year. However, expansion of NBS to keep abreast of new therapeutic interventions for SCGD is impeded by the lengthy administrative process and / or evidence generation to merit inclusion on the Recommended Uniform Screening Panel (RUSP), the need for new custom assay development for most added disorders, and state-by-state implementation. A longstanding goal (e.g., in genomics and genetics) has been to supplement RUSP NBS ~10-fold or ~20-fold by inclusion of SCGD with effectiveATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 treatments that are detectable by genome sequencing (GS). Phenotype agnostic GS- based NBS of populations can be different from established phenotype-driven diagnostic GS of individuals. Early efforts at genome-based NBS (gNBS) demonstrated safety and feasibility, but were confounded and / or impeded by GS cost, immaturity of population bioinformatics resources, and lack of variant pathogenicity knowledge. Recent advances in these three areas enabled development of prototypic methods for parentally consented, gNBS of 388 SCGD, 342 genes, 29,771 P and LP variants, confirmatory diagnosis, and ~1,500 treatments by virtual, management guidance that utilized recent advances in these three areas. Recently, numerous groups worldwide have started to undertake independent clinical trials to evaluate the clinical utility of gNBS for SCGD. They employ manual methods of variant interpretation, return of results, guidance regarding confirmatory testing, and therapeutic interventions. For population implementation, however, gNBS uses a much more scalable framework that can include, for example: 1. Very low-cost, clinical-grade GS from dried blood spots (DBS) at a scale of millions with high precision, specificity, sensitivity, and recall; 2. Automated GS clinical analysis, interpretation, and reporting for hundreds of SCGD; 3. Translation of positive screens into confirmatory tests by tens of thousands of geographically dispersed primary care pediatricians who lack genomic literacy; and 4. Precision medicine intervention implementation without delay by the same workforce. Some embodiments described herein are related to a healthcare delivery platform for SCGD that starts to meet these requirements. Another feature is that NGS development is adaptive and versioned in response to rapid advances in understanding of SCGD prevalence, penetrance, expressivity, causal variants, new efficacious therapies, and rapid evolution of GS, informatics, and artificial intelligence (AI). In some implementations, for example, the technical approach to adaptive NGS development has three elements:

[0126] 1. A structured rare disease molecular and treatment knowledgebase that grows by ~50 new disorders, ~300 new treatments, and ~30,000 additional variants per version. This knowledgebase informs a gNBS model that is trained by distributed or federated learning from multisite, large diplotype models to achieve analytic performance suitable for population risk.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0127] 2. A highly automated platform for standardized, scalable population screening, diagnosis, and treatment that is empowered by the knowledgebase, screening algorithm, and an imputation model, such as a transformer(Fig.2B). The former informs electronic clinical decision support (eCDS, called Genome-to-Treatment, GTRx) that disseminates results and management guidance in a manner understandable to frontline providers nationwide to upskill them to translate positive results into optimal outcomes. The latter two provide automated, supervised GS interpretation, which can be used for scalability.

[0128] 3. Adaptive clinical trials that employ elements 1 and 2 to fill knowledgebase gaps, further train the screening model, and evaluate clinical utility and cost effectiveness for each SCGD across diverse demographics and geographies, together with acceptability to parents and providers (ClinicalTrials.gov ID NCT06276348, NCT06306521).

[0129] Herein a model for differentiation of SCGD causal and non-causal diplotypes in large cohorts and results of evaluation of NGS regarding clinical requirements for population screening such as maintainability, reliability, usability, testability, recall, specificity, and / or clinical utility is described. Subjects and Methods

[0130] Research Participants: Research subjects were 469,902 UKB470K participants, 141,046 Mexico City Prospective Study (MCPS) participants, and 7,342 children, parents, or siblings who received rapid diagnostic GS (RDGS) for a suspected genetic disease at Rady Children’s Institute for Genomic Medicine (RCIGM). Deidentified UKB470K exomes and phenotypes were queried through the UK Biobank Research Analysis Platform under application 82213 as previously described. At enrollment, UKB470K participants were aged 4069 years, and 86%, 10.1%, 1.3%, and 1.1% were of white British, other European (EUR), African (AFR), and East Asian (EAS) ancestry, respectively. At enrollment, MCPS participants were aged >35 years, of which 66%, 31%, and 1.1% were of indigenous Mexican, EUR, and AFR, respectively. Deidentified MCPS ancestry-specific allele frequencies were received (downloaded) from Regeneron Genetics Center. Retrospective analysis of genomes, phenotypes, and electronic medical records of critically ill newborns and children, their parents, andATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 siblings at RCIGM was approved by the Institutional Review Board (IRB) of Rady Children’s Hospital / University of California San Diego (UCSD). Analysis of genetic contributors to infant death was performed in 1,000 consecutive infant deaths in San Diego County from 20102018 in the San Diego Study of Outcomes in Mothers and Infants (SOMI) with approval from the IRBs of UCSD and the California Department of Public Health. The genetic ancestries of SOMI infants were 10.4% AFR, 27.8% American ancestry, 7.7% EAS, and 53.9% EUR. 32% of SOMI infants had 30% ancestry admixture.

[0131] Selection of Disease-Gene Dyads, Therapeutic Interventions, and Inheritance Modes for NGS queries: Disorder and intervention curation for the Genome-to- Treatment (GTRx) management guidance system have been described in detail for NGS (Fig.2A). Briefly, the efficacy of therapeutic interventions available for 645 childhood- onset, single locus genetic disorders that met the following criteria were examined: acute presentations that were likely to lead to neonatal, pediatric or cardiovascular ICU admission; having somewhat effective treatments; high likelihood of rapid progression without treatment; and, diagnosable by GS. Publications relating to ~10,000 interventions associated with these disorders were extracted with custom scripts (Rancho Biosciences) and curated manually for relevance. The interventions were adjudicated by six pediatric clinical and biochemical geneticists using a modified Delphi technique and electronic data capture (RedCap). Consensus was used for inclusion of interventions and disorders regarding: 1. Age groups in which the intervention was indicated; 2. Optimal time of intervention initiation; 3. Contraindications; 4. Efficacy category (curative, effective, ameliorative); and 5. Level of evidence supporting efficacy. For NGS, the therapeutic interventions available for the 388 NGS disorders and 77 new gene-disorder dyads were re-evaluated. The suitability of these gene-disease dyads for NBS as described with the same expert panel, electronic data capture, and Delphi methods were also evaluated. The panel comprised five pediatric clinical and biochemical geneticists representing hospitals in four states. To reach consensus regarding inclusion of gene-disease dyad in NGS, the panel considered six questions and clarifying sub-questions 1. Is the natural history of this genetic disease well-understood? a. Is there at least one well-established gene-phenotype association?ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 b. Is there significant variation in expressivity? Is expressivity sufficient in children to be characterized as a severe disease? c. Is there reduced penetrance (see below)? d. Is inheritance (autosomal dominant (AD), autosomal recessive (AD), X-linked, mitochondrial) well understood (see below)? e. Is pathogenicity of at least a subset of DNA variants well understood (gain vs loss of function)? f. Is genotype - phenotype correlation sufficient for those variants to predict disease course? g. Can variability in outcome or disease severity be clarified by additional investigation (such as an analyte, enzyme, biomarker, or functional test)? 2. Is this genetic disease a significant risk for morbidity and mortality in infants or young children? a. Is penetrance high enough such that identification of clinically insignificant disease is minimal or causes minimal harm? 3. Is a treatment or intervention available that is effective and accepted? a. Is a treatment available that can affect outcome? b. Is a treatment effective for all affected individuals? c. Is response to treatment consistent for a given recognized pathogenic variant(s)? d. Is treatment effective for all symptoms of a disorder? e. If no specific treatment available, would making a diagnosis change management some other way? f. Is a treatment widely available and are there sufficient providers, facilities, and resources to accommodate all identified individuals? g. Is a treatment acceptable to the majority population? Considerations include cost, morbidity of the treatment, and religious or political beliefs. For example, does this intervention require use of fetal derived tissue? 4. Does early treatment improve outcome? a. Is there a latent phase during which initiation of treatment leads to improved outcome or prevent complications? b. Does delayed diagnosis lead to poorer outcome or serious complications? c. Does early diagnosis and treatment lead to improved outcome over reactive care following symptom-onset? 5. Do the benefits of early intervention clearly outweigh the risks? a. Are False Positives problematic with this gene?ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 b. Might NB S adoption of this condition have a negative net benefit? Considerationsinclude the proband, family, and the general population. c. Do concerns exist regarding identification of carriers?6. For genes with more than one associated disorder, do their treatments differ, and canthey be distinguished by rWGS or additional testing? Some implementation of NGS are designed to meet criteria for clinical testing. An expert panel of laboratory directors, genetic counselors, and doctoral level genome analysts met weekly for six months to address five additional questions related to the GS testing performance and test reporting of each disease-gene dyad:1. Did the gene have several phenotype associations that represent a single phenotypespectrum? The expert consensus was to amalgamate such phenotypes if the mode of inheritance and therapy were identical (see Supplementary Results described herein).2. Was the gene-disorder association of uncertain significance (GUS)? The expertconsensus was that inclusion required each gene-disorder dyad to have been adequately described in >5 independent families (see Supplementary Results described herein).3. Can a majority of affected individuals be identified by short-read GS? Sensitivitymay be low in ‘dead zone’ exons (such as adjacent to LINE elements), genes with highly homologous pseudogenes, or difficult variants (such as inversions). The expert consensus, based on comparisons with other screening tests, was that an overall NGS sensitivity of >90% was desired. This implied that only a few NGS gene-disorder dyads should have lower sensitivity. In such disorders, GS identification of variants associated with >50% of affected individuals was required for retention (see Supplementary Results described herein).4. Is childhood penetrance sufficient for population screening? The expert consensuswas that an average >40% positive predictive value during early childhood was desired for NGS (see Supplementary Results described herein).5. Several NGS-eligible phenotypes are restricted to a very small subset of observedvariants, such as rare gain of function or dominant negative variants. While these genes were not included in NGS, causal variants were incorporated in an “Allow List” that permits screening for these disorders and inheritance patterns delimited to these variants.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 A web resource integrated the NGS and GTRx information resources and the adjudicated interventions of 412 retained disorders associated with 342 genes and 1,603 interventions. The patterns of inheritance used in NGS queries of gene-disorder dyads were also re- evaluated as part of the transition from a first implementation of NGS e.g., research- grade prototype) to second implementation of NGS (e.g., clinical test for NBS) by three questions.1. Would the burden of follow-up of asymptomatic positives overwhelm primary carepediatricians? For mild dominant gene-disease dyads disorders that have severe recessive forms, should NGS queries be limited to the severe recessive form?2. For autosomal disorders in which Mendelian Inheritance in Man (MIM) listsinheritance as both AD and AR, which inheritance mode should be employed in NGS queries? In some implementation of NGS, these were queried in a dominant manner.3. For X-linked recessive (XR) disorders, does Lyonization lead to sufficiently severedisease in female carriers to warrant population screening? BeginNGS.1 screened for these disorders under a recessive model.

[0132] Table 1 shows the 412 retained NGS disorders, 342 genes, and inheritance patterns. Periodic review of these allocations can be desirable considering new knowledgeATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 DIITS,AD.AoAIISGE&.AECITFENFA EE oC P D C --LRANfeUTIBCIY ILYDIC -ELdL. onC AIM AIT XY C CI EFu.L A DE1R UJ1E 0C 2A E ASN EG SANCEmnDmysC2L 3NI AIDI bc LPE NAIAACIA K i NA, streLYATS Ldr AC IU MC,A MP IY ANHN T LE ,.DnoiceN AUS EAIOTHEGC- EGECT 2 G A tafeoRANTI NIC YCIR C U RC- O GO AC EP EOmdsiEFODT E LL RLNNLAA MR NR OIN GY -TmcRNINEEPGOITIaigRE UD ODTE AT LLE lfolAF TM YOP LASDLCIDYLY D ON HH YHCEU DniotaEEH.Y Y NUEZHMC O NEREAR L RCoBDItum I N MLO M DED Y O ,C a eAYM HFV HM. S ,si hLA A MFH C RTO OO. FOTEEGHEtilER EEHTIF.NuMEF M- .R-LcsNPM DE ELaED HEV G C TATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrriIIrEP II-ALYSDE T D HPVOISOPEOLTE 8HREC51OI 2. EB Y OE ARAIRP IYSCA.NTIDSILIN P A.HTNAD N YION ET YR STU OESY YSIS WHE ER N E YLIAFS DTKCAPYSI DIA M DCIAI ENE TNLSCH C NII,C N AIMOT DSC O OAreOAEN MNIdrRR GEGS EOTIN,AEBE SMMY AC KDICIoHUAE ENT.Y D A NNI EI IS SGU ODELOUERNsiTPH R RTIOSNSFEDEEHSAAA R RIOYPNI FALAIC C D RPESTS T TAGHISNINIM C H I C C D CITUP TSA YIHOATMLDO-TA A H -P.DENE C Y HPN NIG O R WRTCUAASON MEU OM PS S LA1H MSO AE .CINIU.O U7.T Y NM G G MH TR G OH LAF.OL IO O N RLR C OE YEXF HP SM G R R O.NPO REN HOPNIG HYL CP SYODIROOP EOSDO C O RECPYMO C A C R T.G HN YH Y HIOT NTU A RE PH E D U UEM M A A M HMIATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun3GTIET N N SNIIWSO EASR A2SSTIRD.IOI.F4TFAT EOE SIOE P I EATATAID AIDIHSLB O CDIN N M3M E CLOGeAIRA E ME E E TAD NmD PN APOS EM M.NNI RRIU CIARe orCAsdn IASNITSEA N KNIIL ELEFEILE1L DU NrL AEaey NS IKUP P JOeLdrULREUNISHsi s EUF ERNIEB MD MPESB OPosR B RSAL LdnrGD OI O.UiTIASOOPU OU ALM IU OOUAseoRPIND HALGC,OC,ODIG O TS LB Rkhl SI C NR A C DC nH Y CT - EAAIRAINTI LNUAETUEeatiP NpE ISNDG MMG GIMTMNSSMiN CIAEOE EOMESNAE T LNIccHR M N NIAML RU ARD NE CD A CLON OOEG NY GA AB GEAI IALAHERSK TTI LN ENIO A R RHN NL PBER O O X M SIW ASL- HPD A C C O N C D / . TXEN WSINA ADF FATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun.F.F.F.N EE EY T YFTEOERS ciDD D D. .D B 56 7FRUEI tpOESTNT T EFE9TAZC.ITEI eM. TA Lli S 1 F HEN N DD S8N FOE8.Sp,AeE I E ERD AIDUT 9 8NENENTNT EN NTD.EA N N MMel YSRi 0t5yE IVO RMRDN EO AY1.1.rO O OEOIOn hEIS-F Fed P P PNENPUV Y S a tDSNNfna SYCTLC RAE EroM M MsiO O OOC C CPOMEP OS TY CAIAQ HOB ipo LHT A A UTD D TRO O EA ylr laAA CF NI EN N DT T TM O MTMR G OT CI ahTEpAROPCI TB SEU U N ENENE C O CNED NY N O MISOG;ecNAPYSN YNIM M MT T EOHL TD OL GnEM M N N MMNLDeOOER NE HIPDRC T ONMOIMILELE E E EAARC- Y A MIP PLPM M LPM Y D HNIOHM M MELEA M R RUE A HTA O O O CPLMPO C CEPA N CTSO M C YMIY C ORC C HPATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun,6. . .N - - -W- -WY WT. .T. .36. .S OSW O CFWSCFAIAIFEFECI LSAFLO SSEOLA O SD S F SESDM M D D N, ,B,AR, ,AR EB D. EA1A,OA B3,ONEFIN LIR RE SLO IH1T. .2N.C2T3P. .CTN3 Pr U UTOCTDSNIOSN R A YdS LY N SLYS L.NEN CEY Y.SL S LNECeB B Cr O O AFAFD YERBIM CINCENEYECEYENICINSRCININSRosi LGLGT TE F.NANE N N ANACI ENN ANE N ACI ED A AN NPCIG NEHH HHEHNNIEH HHNNIM ME EO RTSOTCTSC H T CE LH HOTCTE LSC HO M MMEMEPYCSA ASATSHSA ATH A ALCC CSCG GLA APLPXIMTY Y Y ALY Y ALMPM M.MYYM M.YY R R O O A.G.M T.GM TA A C C N G N G. EG Y N C N. ES O N G N O GC O O A A C C C N O C N E R O C O P C CATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun.- -WY WT.O DIOSC.F 1 7 6 8 4 5M.LA OE 9 .F. . . . .AFEESSF S D 1 5. EF F F F FAD A ,, SR.NE E E E9EARIRA BA4.4. ,CO4 2ND D D D DTSY Y0 0 0 0 0 E DIP0R AEEFN4 P ISSISS S1Q1Q1Q1Q1Q Y1CQ YS SN Y. EO OC CIT H OTA T N r YL LNCI E E E E E . E PedSESEAr CEYER R N NM M N RER RT T E EM M M Y MOUH ToCINsNINS EPEPH HYZYZYZYZYZS YZPiN DT LEANE N ACI ENO O TSTSN N N N NCNOAN Y Y D HH HHNTCTCEIL EOTET A AESY Y OEOEOEOEIOT ERO OP ISO OM E TIS SH A ATHSOO M M C C C C CRC CE TMYSCR. .Y Y Y Y Y H Y.N OALY D G R RP R AM MAL.YY A A G R R R R N N A A A A AEAEH M O N HPS P.GM TOMIMIM M M MMOEG N N. EC C GCR RI I I IAH NIO OP PRPRPRPRPRPTC C N AEIOPN C Y R H A CATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunES SESHOP PO OA AEA TIXTE .FY T YT T71. ESISI ESIWESUEE TE. UFEU A ND DSESDED siIY D D S D Y CYEA AD.soNSSE U USU NX,A P46 IS SADIDA IISEA SISFEtaO2K CO O EI 4T TOTCA LLX X A AD mYO O L A L Eore Rd Ts.iFE .sEBN A A AIFETL PY M MF EARXLO YL P LY RYPX RShtAnaro SOMU LSER H.H.E ELxoUD Eni O AS OL LOL S PRPDT FET FEPO RPYssiNtsN R GMIU U UNIMEY V HYEDEY Y D HD HXuoDE yR C U - D H N N N OHM M Y OniEM A ALCE L - E E LHLRdV MN W R A AR R A- N A- ADnU RGRN N1O NNANYetE IL G GEOE1OESSC CICRLRD R RRHPERH-orbRRIEN NIDA DI E EDLD1erO NA I AOTST A A A2 eAMO I R R O. TR GR.RGEOSO. -H7.C ATCIC G1G H C C H C N OPNIN N OS T TO RLC C R R O C C A X R A O C O CATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunSSY N HAL ETI SE Y A SHA TNYHOI TEHP T IM O 03RONE2 WUG.ATIAEDI . TAIN GPO C O RE 1.1.W, EWNID.AYS LLSICFEP EO RTFN N D G RTSNIY YE O W,HTS AL EDAE.ODO NN ER DSYS ST TD E7N O Y H YDTNCAA IC L R EO D NPSMTO O YSE re NIYI S TACEDRE I I ETDTdrRSMALN EE SID.RAL R2.NENED,KA L&ML EosU NA- Yi PN LX GHS EFEAL R U UN H H YDID,AAIAB NRAROTDID UCSC Y TSTSHTN,DM -ES ID UECOI A CE SA AAY N D RIBSY MIA MSUR A R H RA YTSOTAR O - DRA U ALEOMEY YPMMOIAT ED MHPI ESEMC NIKS PI MRE gNI- .G.Y GH APOTH CPE N L OEPR DEA ATYLD OEKNR MY O O REEN NOILYA AE O- HPO O DLMR YMRLDCP X YEC Y C C ROOEW HLPI LR H B UHOIKSTI D R A C W D,RI LAE D A R AA U C H MATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunNESESESN 1 OIA Y OIAI TA A ,N N N H TITI TNAMF E E E TA N N.AETNUPU G G GPU UFTAI P E SO O O O R R ROB BE INE N OORLU UD EEPR MTD D D AS S NI PMTEN GY Y Y H A BA BYEr OU LP ENHedREMQEHEHEP ,I,IB MAIAIT,LPrTNPMOI D.D.D. ECIIIIA OL LIAMosiUE D OUGI TAF A F A FNX XRIHHIOPU AC,OP AT o E o E o EER RHP PMEC,O DN.AR ANC- D C- D C- DCIO OT TO O O NAR CIGIG M L NME LYLYL TM M IIG Y N C C RE E SMC OEREM O A A PC N DEC C CL F F .H H O REN Y CEARIOLE NRPA A AAF FGY AME E E M O ONT IV OE LXOPLIPLIPI L . .O YFN E C CTEFE C O LTLTL H D D C S N U U UTN AEFM M MAFATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunN N N N N N N NEOIOIOIO O O O O V.IFIIITI I I I IT EATTs ATNu TATTAT T T TTATATATAT AD EREPY ElaNENENENENENENEEF SAT,l.atiM hpIM M M M M M ML TFnEeLc E E E E E E EO A HEegrPorLPLP 2LPLPedMBFLPLP ILPRPP D nS NocrPdyMC P MD MEGLPMPMPMPMPO.OO,aoOhiCUhOU OPU OU OU OU OU OUHPNHISis,O tiC,OC,C,OC,OC,OC,OC,OYPSEmeDAIR w GAIR GAO I RAIR GAIR GAIR GAIR GAIR M GYS IB Hn- DeMLG E RM M M M M M ML6 A goNEE E E E E E E E ,ATN N N N N N N N1-E niET rbI CAIAIAIAIAIAIAU I MSYifN A N N NOC A V N N N N O OTO O O O O O OMIC C C C C C C COC K N N N N N N N NTU U A URFAFAFAFAFAFAFAFAFELATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun..N oitNTIG GIES ps e A,.N N V AAn nae WgAG O OI ,E Y,N C; CLTFE rt,sCLI HNaO Ca.;.Het eiEa.TT PO&.T al cn PNOIT ,I.YSI;ai iGaaimmemNi iGam mNiDN YofleiOcLYAR FT A HD neenO eCe,neenOmeCe,nYSarifAeHEb e.Hd GPNIS L,Y UTAYEPHG reednegog a g g a gTDerOD GTA nonii ononi on A IT . ec .ensN ORAILERO NARrogoir rmeirirmeir POFEofeomrCTSNIAS I POsinirbifbifnebi bifnebiOEt dmr o,.YEN YROTSfs gfgfYLDe o fFPDCR Dbi Asyyd. os AosMu hEDOTOE No Gniyrd.GniyCdrdEUnyrDLIHPU O NDTNED.pyNbOiofpyNbiopyGNIoitatiOA O MH UNE EG O G H CsyH OfCsA y H RDarutN N NDEe iUYLTIM MYL C Y O.D.OTNnepdMLWIOLC G GSINge enMILRLXPG N N DE diL EO O OIDor bL C- C CPIL Auemo ETN C C- D T AATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun4MA I 2TH IE ITN.NTELai ICINA ESE r WTE Y UyPCAAS h,EIOI. E II V.OTaHAi pr DEOIII II eIRTF SIII I IsIFETI LP S me opKPOe eIIeITYL EDA AI is AIDWREI naciNIRpypypeprUEODE I oM MdiME / OTId te Lt -H etytytedr N MDPGE E raENIH T / Lno io XTAI sa eseseo . E 6AsiGH GRS ShSN CO OTOT cc OT ID aTAEWNA mp,AMaRior ,Y htARMeNEEs aieaNd si e sdsiaedsiD OI OTC CsCLD- ySr r dCT TS A Ayl AR CX AIO Nn re EPY AeherhereEYENRCUE L L opLAL ,NaEBfD kla OcuchuchcEO D G A A GocG AAI PAcatiT TYUa auauaVREEO GuRMOl nC MB eG G G G S HE ERTBgCO n OH PYN LRoBTE AUCIRSG CEM N W A N O O R NHTATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunISH ETI IV 1A B T2WC EAINAEC 1SCI PYCGOIBI .UMITP 2 T I;ITSM AIMNTI SIM E 8 MTEN DE3Y1 2NSLDI E6SNLIYIS E4AIY D NI LN1I AIC LEOAIAII AI PHO NI AIre Md E EEIAr DIVI CIAL T S LMI eAU UAS EN sEae MRX XLM ETDNTNC Hsi SYEEHTLEPL UPS EE ANCLIP I LMRUOS EANCosiCSFNE LYOIEL IRYdLTS yr TEO KEKIERYLTNLHIRYL DADL ECIOP-4NNE PG O AbBPEG A CEG Ya AY R E R EPO AFH CPORSEEHHTM YPMF IH. P PYPNIPEASYPATRBPNEH.Y H.D L G Y YHY H.HYCY H.Y H U -LAR N M G H PEPAANAN MLNLM TOAREO AG O Y MFO A CFAP FD H R C NEO EPOC E U N MATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunH.AL L3 N AIIAMIFEAo y EoB CEM DCc H-A AAIO DE-nNELei TLN N O O NL ICA. SYci FY.I I EGA AIAFE A RfeOCFET T POISCI D T A d SADC C OM a CIMEAT E TAEYEN U.N F U.RFT Eai iNNILYLESHTUL.1 TE3XSF F U Hme O PA NFFBEreOI E I EdmdRAen L E SOGNPr N R D R DENna. ss AATSYM NIYSL EYD H AIYoDsE TNI TNI .iYA SI alll MYDEK OEHE otDT,HGLAELAEGIDahetCLCCIHXSTSeu TE G D -3O RI TROI TN RO OMI CE-Sat eYO lkHTMPCITEAHL EAdSN YM-YNU ND Y DRPDRP S e cTO O Y3Lm O A N NEBiERASM HO T COT X O -siOYni - Y Y HHCEO OE LDPO - D H H C CVAEE HPNE REB XlY AS EU R OusTIH D A RnirR G O O NT T S T-T PC DeUOI IR ASOYpTLM M A HXPLNLHyA L B O O - H M A C N H3ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunG . Ny E-2 2WUc4OnH 3 ET. SESES.YeiTcF .A N1 2 A AD Nif O DEAIAI ESESN YF e IIAGII .M MI IAS O d SRSIIOIAEAE4T 1 ESPARsissiN s.YENED DH I NL L TIN OIY TF ELDo o S . LIE EWTTreL EN BEPY TPYdidrirN YEINLU U W W3AINdUBH AIY,R EHE . ahah SEYSB B O O6.RUroBAIostDT,APD FciUTDeTGME YcccR H CEODE LO B B FEEFM Dasas E S IE GLG Y Y DI MD OT uNE d ESN N HI ylylLR - R RLR H A AO OOOIO ESm NUILO OSARopopUELCSM MT TNRP TRNsIOin O- Y O RNE o oH A A U U ET c cuR M M MO A N- i YluY TIYRS uU A A M MHY M M H G G M MMI PO CTI snTirR D UAXA AA A M R RL LNRe .OF FYA U pFT y TA GR A A NDNLINIA H M OY R R M C H A AATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunR A..SGO.SEHI E .1NF S S.P- FEEETT P-E ETEN .FELDLDE FYLL D OETYSE O BOT,EN BEBCINOCNAN KUIS.AICUADEFT0 EMKIAEDIME2 SUNML E176 DENM D ,ML LNI AI LE1.2.rA4BE.D .I 2 3 L IN NeM,. MIdr MI S T SFNI SA A MNYAU AI PYEOICODET SUT SUUS E- YS SoODsENTIDTDS AP- A N NTIATI NCEGT TiDP- N OL T ,OEN HCI L IOLNL IRYL NQQ DENLLIBEL ESG N U D RLE B MEL OEL EG A GTI ENECA E CNENE PO LR B MTM NU MLOTM M YPN O M O BOCN- O YMIET A B CTV ,N N HD . Y NOLHO ,. EY Y O.GE E EH ASC LGENTIERC E A R OSI ERNENIM V AS LNA LX N- V MUT K - LEM AF ELER LSEA U EPE L SR ERV R ML ECRa P TECJT TATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun51 ,E3.4.AIESEASE EAP PS BF FRL LY Y CEOEDED.U Y YT TI RLF D7T U PLCITENE E IX X A B NDC1A. .II1E 7 ZIA.O O1b1bLET PE41O OE AIYCIR N N YdA A nFR Rc,c,PHSNLM MS IY OIY R R A NUS SIaIED A C AAIAIrET LARPEHT O OPIedE L ONICTIHsCR R Se DIA AU UrLAP T FLELAPH HLD ASNE ApytOo . o .DIDIoI OAsN AI OY Y R RIMY HGI,sCIC- FEC- FEC Ci TNL TAON NLA A C ALC OTSH is T LR YDLYDA ADAER F HN U AA HTITI LY - HM A KodiO N1N2CICINPIE . T FNPEU T U T ATE O Y HAI soCO ON N OT TO OYC M CILNEAFO NYCIPIPMMLN O MD.DnDDSEN GEnHaC O O U R C RL LC A A R,RE E EOPA N C ML L LM MAN GYSA NINSY O G Y YL LE INPEBIBYLTC H H Y Y EEH H BL F T TI M ML E ETPO OETEEC C A C M- M 3 -3M MATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunE H .nO.F bICI PEYCI PTIFiEarD DT E, N AEPNEED.e TF T TY ApyS ST W..FE4.D bE htSiB AIYAPU TAIYET DS EADPt U, ANAN C n2LO o.AILODN T AIOSRA YSSAwN .p.p ;.R RATRO itF O OT SOIH o olN UC D1UD1IRUOTR Ibc G al EyDLATLATATR P UE T lR NahahYS,DI bcU ,D MCNsA oDIG NA EG N 9E. EDIU YS pepe EreC CdA ArIA AIICAoAFOIcylOEMEMNCYECLIE cncnG Aos CRCRAC- Oi ICTgfCI M- ELM- ELSRAATEC FAE.eokK NUIDONINU NCI LAYTAMTorRI PAI PCININ W M Gue AELT OILTNN UNEedO RM R UO UM ON IELOO RAOlTR- / R ASASOO N rC NCNC HOLRP +BMY YL L EMosOI I TH A AUL T aNL CMLC AA M DE iCTGT EBL dlUS1Slb SCLVE MG ESme EYO YOLMPaLYbcYcAYLNd GHM HMLYMtiG C,C,Y MEYe ETOTO YYLOnOA AIM TENT OM E HEH H TH O Ceg IO.C OEY JIM METEMnoM OMMME G A A CLN M M CEON NA- R HN H O A A C NAEATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunESESEMEA EASA4TSSE EA HG UINNIIS SE TNIO O H TADIDID H WZS R OIH. 1 2RI . oNTIFEC CENFEC- USUSU R OO O AR IG O,IW / D EDSE TSS 4 3R OLALO D DI .AA ECFEIT T TD OWIA A DS SOSA R A A AEYLTr.N4.OE ECSeM M MVISI LR RTEAL ODU E Ndr OIASO HN I FEIAE ETD D O V VUY X OSOosL 1OL 2OL 3TPPEiU U U RESTADR K KP E EDM -3A RTECIDIO CCICIY HR R AIA:BAEKD N N N A A AOSO&D A OO P-P- OXEXEMRoFCS LR R R B RN GG G A COIR CILN N DS SEA - N Y A N N U NCLNCIC CL IMT TRRE A AEX Y SX,X,OSYA REH NINIA NM.A D O M M MNNCN E E P 6464AI T PC O O O.WR OI I ID M ARC R R R GDHIA C M N N AIRTUCH C H NTCOCEU E EP SSR R R C RLGPY A A AYT HATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunycn S SSL -6ycyc eic 2 3 TA N7 LEBDneneif.F. F E EC NIAI iciMif ci efdeEFDEOK GADI1M RCIODEeA EAIAe e sMd d at0101 ERRMERpTDNCDyt ITN ME- aQ3 IY NI6 F yVNINE-D DlE1 hpQ EERO OI 2.MHL3.O.cn,T TrEeI I IEs TF EEUF F e E NdC AL C CesesohM MSIAL EDTD.AE E ici S ErIoF A A A ananpe YZYZD.YO AE F OD OD feNDsi E LCICIegeD D- Yg sN NGSN HPT EDLG NE d O-NE4NN NoO OIrd oradneE ENO g O O OC USA M ORE A USAnMeYgLPE HEIy yo C C C Y HMNI oR D BHP P PO Ohehe rRdY YTI LGMI PC Y MMIK niA EE RPPRPdetda eyR ROtheA AEP LA G GEm Tsa ,Y Yva dHuvru eyr tMIMPY IYTHOAAlH R VP SPEPy aPvRuPRPDLPSA U rLyX O R YIPPHP P EATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun12E. .E.T LFTFSFA AEUSEAEHTD IDE 2OEDR DIEA CPSNE2OERN.CI2H STI T EI NFSES2 A ASI I2HPGESO NTI-. IM Y YSWLO NAIAM MM ATSE EEPEF.D WRS. OERSCOOCIOTR OM RRA R OTHENTNNI TNIYN NNTrEe S N HGYHOPS EE E,.dMIDN YI ENE AC HO GIOTNNI .OC AILC ALNroRsPESS DG Yi-5A O R HH IIB Y R MF H AIF AIF AY R&D SI ATP TSOSIWAA AEDPCIE L E LIDEDIE SE OTIAMH62 CITI E E TA D- Y D- Y NLNLIXX ILRU YEO.FG UNINI 4N4LREHEHE EM OIAA- AETMTIMSP.H.HGP EODLTIRER BHOP ESCPBHRPC RSIXYODM N M ORDSO NE ER OENEAFYNEO HAP PG DILT ANIO CLUUENIHP PPSY Y R N B M N BSO H H YPA MMOIM O OH HPC CPATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunLE AA C. .-&RMWT .FOO.FN72 8 1 65 6 9 4CEMLYA 1 1 1 1IA A A A AIAIAIAI / NIUU OD SMI I I IMM M M. VHNSASREAOT EM M M MNE E E E ENENE EWMY DR1 TPY N N N N N N CTN.1ECA A A A A A A A AAI IGN.O N N N N N N N NreENAH YDIDINC HdNANAroE RU RT S C CYERP FAFAFAFAF FAFAFAFsP EiOVMAIDHEMLIPS U WNS-S,OLS N1-2SEM K K K K K C K K K K E G GCINI YLAC C C C C C C C AA A ATLTMARARNL E LALALALAL L L L LMNCEO RB B B B B B B B BYL OIUE EO HH- - - - - - - - -L SAC FETA D SC BN D D D D D D D D N N N N N N N L NDEAED ALINEYY TIO N O O O O O O O O M M M M M CP I-TXB NME MU. EECPYAM M M M IAIA A A AIAIAIAIM G ATD DIDIDID D D D D / O CMIN O CATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunH TIH TL SBIS SI1 301WSI31 8,I I1I3 A YSAA A5 SA A.C W,CIRLYLMMI1OAI I INIMT M MM YTPYCSS I IA GTPA R A EEAI S ER 1 S 6PU1VA R N NEOENE ,E ELE E 1HPA AAN AMESN O DLYA A H T IY PHLI PLPLIIPYTIC PN N ND NTN ET PSEE HICIredAFAN FAFALN A RY OA ERMTWD D FAFEPTIMELPDUEHELAP,L4 OIO 2roK KsK NAiC CAIK K YLIAITO E ZI .II ITOACR R EPDALACFC LALKAC C BDAHI- NLAZI EM NLCEAO PEYCFAL L T TPN AAFHLASAF AOFPF HF C CTB-B-B-A OLB-B- N AEMN PIERELI,EN P. IEMID D DLU D D NCH CE IMEM N NBNSCY N NRNIYC ALEO O N O - DAIN G O OI UM ML SALRE EA GEALN FLRE ,A A A M ND M M A W IA AONA A H AE2FR GIA YK K ESPRODIDID MAI IMSEAPMEEP PIM D D YLIY Y DTPH E HATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunEI I I2.NPO5 E E EN YSPYPYPY1.F AIY VET T. ENS61F1KCET T.OY CIR NDOCT UM M MYISE C T NN Y AHTOLPE31ZI SISNISNINE S I SAMYEH STIACP3.BLIYESO INE G OT PH E RO R O V RITTE7L NDA-TSR N RET L E E EA AILIM MSIA redNrEG Y NU MYSAEE A ITTSTSTSR E MUY SENoHsOBOTHiT CNALIPEI TON O O OFI OSI NICO M.1DQ. INLAFDLDLDLL -RYLSN G2DSAAYI LR G M A NUA NA AMF FH N PIA A A OE.O O O RSEP APGIN O K O C .OEC O NI EMP P PO R YPR GTDL EO VI YC A Y Y Y H H HHCHY ACP U.HPIT NEYT S LNEFO O OSMEPO ASR N D D D M.ALIA C M ALI ERAEGIU U UYLGFNTN R D GN E E EALN Y P OE SPSPSPO X R B CAFSEPR R RN RA A AI PATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun,OE .NP cI ViN TI FE YSY T D EIebOICSN DAIC A. E ZIp I TNOERININyNL t, ESAUFPNI -UEM- -YIAnALSSETI2-DIHEYN AIAIS TI .F RE1 2 oit ESSYR D-NLCTE R& AS ILMEM AIN REDN. .al IO E N NyD CredMrSNI )A C -2-CALIY UNI ER A2NU RC E G 1Y YsoMR HONc EYIL SCIY S.S.ylGL nGIIoIsLMEPCID A.3 TII TIMLYA CiOAIY HSP F F gAFEM DTG2CINMLCE E fR AOT ED D o O OPD B H ATT E ENUL N OTRAU R- LT / TSAPLI1 1reTR YTENIS IB GC ESO RR ETEINNOPETUTUdr SED T Y Y S MXCIN OEP PCITIAIRDI L L osiNE R MOOOT -YYO NTG GdG O EIY N CRP LH HM R OlIB(R D AAOaOSIAT. tiCDM2A . Y NTA H CCne YL.AINMIH YSNSUgnG G N H TY RSP EROS oO PEC N CATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrun. .N NetI eNTO Y YPNrIS SRO- opTAtY8paT Ds y TH 1 ipP EREORSBESOynOht ac NP .ea rtneEO NsR OEAEAALAE Hpo ni3iAciMERY TS oteycArI SBL LASID LrumatIfeAIL A A A CI udneieRdrUAN N L A AMDI enioEEHvyXcEDtRP PRkci kci N,ciCritlnL roU MPAP-P-Eevifes Ni IAV-V- SDotueiDTSMO OONEcPpO sNICUL n nHs dVomifKEn T,O U CnanaTSno esYE TTTT TOI mltnedRarSAIRSm m Aps atCTAE EC AL- SNaredEP TY G CMUeieiYer cuLLOALALOI Oehp neeYni EM N NM -.apderFI.VIAZ Pi R -V G / A SErpHteedaN eALA GodE N - NEFRP - rCININ,aW WS -muN O OPCinH O N O O CIR RTid SooCtsB BULOI SN AyFD G BATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPrunA 8IkiIASIeDyAcNLpEP 4 58.yNneP R CFt, 2 ASiciOE IC PT SI ISI II .AE nosis 1 2S S TfeYS TYSE FEMEDIita7oni ISISCE .NdRT Y U H COT COT. PYDNI . lLO O NySIcsO OFYEnGY GY NT H ILNIYsoS ufR E R EEDS. ireEdA AC AC Yr N S,ATC U MScyO RopL L LFmato .N HOI HOIIAAAAlgsT ilCSCE iSADviGERP TSP TS HTMDOLLA Wfo Edi NIOdD N O O DIO HIH RENEG B Ar POorSUSUTNetaCA ME OME O AIS TA O GeEdrE ecO OSE UloE .R GH.H H HB OALMCS S os TS la RERE TMsiENP . PR M O M M MY SNidOnorB BNIMIhtiV O EC AS DFYLAM FYT IA A L GRlati ueUTUT O R w aIAT neNTSixDOgnAatAPIoG A L CATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001IIrrsiIPIPyreuR R R R R R R D Q A A A A X X X XSG NeTDn23 1P5eE14GBC MSSPASASAPA U NAW W WIX UV VN OITAAT, 5I2.3 52 AIN E.N N C EIS .NP1NYSMTIE YSNEOAOYSP IY ORNSEL CTrPOYCICIRTUEPH V CITedMTGrPACN TP TUENOITR A RoOUHOIEAE .THNNG YDLEsiC,OPR OSITSY.CFG NA-IDAIGMMEH ASO O O YEN C BTLTO EH RN.HPMPOEM ORAM.CRO KPO IAM GE ERS HYNRE VHIT PNF LO VESW M O C CES LYN XLAFATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0133] Variant Selection: The input for NGS was 53,855 germline variants mapping to 412 disease-gene dyads (Table 3). They were 43,064 ClinVar P, LP, and variants with conflicting pathogenicity assertions that included at least one P or LP assertion (May 4, 2023 XML release). 15,624 P or LP published variants were identified in 119 NGS genes using a genomic intelligence platform (e.g., Mastermind by Genomenon®), with evidence curation and variant interpretation according to standard ACMG clinical guidelines.4,833 variants were common to ClinVar and Genomenon sets. Variants of uncertain significance (VUS) were excluded, except for variants common to ClinVar and Genomenon that were annotated as P or LP in one.

[0134] Rapid diagnostic whole genome sequencing: Clinical GS methods from EDTA blood samples and DBS were as described. Briefly, genomic DNA was isolated from blood with the EZ1 DSP DNA Blood Kit (from Qiagen®) or from five 3mm archived, deidentified, NBS DBS punches with the DNA Flex Lysis Reagent Kit (from Illumina®). Sequencing libraries were prepared with KAPA HyperPlus PCR-free library kits (from Roche®). Libraries with concentration >3nM and large fragment size were sequenced (2x101 nucleotide) on NovaSeq 6000 instruments (from Illumina®). Quality controls wereQ30 80%, error rate 3%, and >120Gb sequence generated per sample. GS were alignedto human genome GRCh37 and variants identified and genotyped with the DRAGEN platform (v3.9, from Illumina®). Structural variants were filtered to retain those affecting coding regions associated with SCGD and with allele frequencies <2% in the RCIGM database. GS variant quality controls included: 1) identity tracking by CODIS short tandem repeats by capillary electrophoresis (from ThermoFisher®) and in silico from GS; 2) <15% duplicates, 3) >98% aligned reads; 4) Ti / Tv ratio 2.0-2.2); 5) Hom / Het variant ratio 0.40- 0.61); 6) >90% of OMIM genes with >10-fold coverage of all coding nucleotides; 7) sex match; 8) Coverage uniformity by GC bias, standard deviation of coverage normalized to average coverage, and the total length of the reference genome with read coverage. Variants were interpreted according to standard guidelines by clinical molecular geneticists with GEM and Enterprise software (from Fabric Genomics®) using the variant call file (vcf), list of observed human phenotype ontology terms, and individual metadata. Variant diplotypes were ranked according to phenotypic match with the associated genetic disease, pathogenicity classification, and rarity in population databases. Variants were confirmedATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 by Sanger sequencing, multiplex-ligation-dependent probe amplification, or chromosomal microarray, as appropriate.

[0135] Re-Pipelining of WGS, TileDB Development and Queries: CSI and TBL files for 7,342 children, their relatives, and SOMI infants who received GS at RCIGM to the GRCh38 reference genome using DRAGEN (v3.9) were realigned on Illumina Connected Analytics (ICA) as described. Array-based data models were developed for genomic variants and metadata extracted from Fabric Enterprise, Ensembl, Gnomad, Clinvar, and variant effect prediction (VEP). The resultant VCFs were ingested into a TileDB array (v2.8) on AWS S3 using TileDB-VCF (v0.15). Metadata fields from prior RDGS at were extracted from Fabric Enterprise and interpretation reports were de-identified, lifted to GRCh38 coordinates, and ingested into TileDB-Cloud (v0.7.41), together with Ensembl (v104), Gnomad (v3.1.1), Clinvar (downloaded 2022-5-20) and VEP (v105) metadata for each variant.342 NGS genes were parsed and the VCFs queried with the variants selected above. Multi-allelic variant rows were flattened. High quality variants were retained and the query results annotated with gene information, project-specific subject codes, gender, and disorder pattern of inheritance. Custom scripts were used to calculate variant zygosity and to determine whether genotypes represented NGS positives based on diplotypes and disorder pattern of inheritance. Completeness of query results was assessed by comparison with results of prior diagnostic interpretation. Among individuals who had been diagnosed with an NGS disorder, additional NGS positive individuals were sought by analysis of VCFs using an automated interpretation tool (e.g., transformer in Fabric Enterprise). In SOMI infants, GEM was performed with a Bayes Factor-based cutoff of >0.1 and the phenotype death in infancy, HP:0000152).

[0136] UK Biobank Queries: NGS gene regions were extracted from UKBB pVCFs, after which multiallelic rows were split, indels normalized, and low-quality variants filed out as described. ClinVar variants with clinical significance (CLNSIG) of “Likely pathogenic” or “Pathogenic” that mapped to the NGS gene regions were retrieved. The two variant sets were intersected and positive individuals identified based on pattern of inheritance and individual zygosity (Heterozygous for dominant disorders, and Compound Heterozygous, Hemizygous, or Homozygous for recessive disorders). Where Mendelian Inheritance inATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 Man indicated the pattern of inheritance to be mixed dominant and recessive, only individuals exhibiting recessive patterns of inheritance were retained. The aggregated International Statistical Classification of Diseases and Related Health Problems (ICD)- 9 / 10 codes, Read v2 medication codes, Death Register codes, and self-reported medical condition data provided for UK Biobank subjects were used to identify those affected by specific conditions, including Hemophilia A.

[0137] Federated learning of variants that are unlikely to cause severe, childhood disease. Supervised, federated training and / or learning in the UKB470K, MCPS, and RCIGM large diplotype models was performed to identify variants that had been curated as pathogenic (P) or likely pathogenic (LP) but that were unlikely to be causal of severe childhood disease (non-severe disease causal in childhood, NSDCC). Likely NSDCC variants were identified and removed comprehensively in a generalizable manner to increase the precision of NGS gNBS queries (Fig. 2A). In some implementations, diseases whose prevalence in a population considerably lower than their calculated genetic prevalence are identified in one or more of the large diplotype model cohorts drawn from that population; thus, for n pathogenic variants in a NGS gene that contributed one or more positive diplotypes in a cohort receiving GS-based NBS, seeking those in which:where the population disease prevalence was corrected to account for x single locus disorders being associated with that gene in the population form which the cohort was drawn (disease plurality). Locus heterogeneity was the proportion of the disease-affected subjects associated with that specific disease-gene dyad. Locus heterogeneity and disease plurality values were obtained from Mendelian Inheritance in Man. Prevalence, penetrance, and expressivity values were obtained from the literature. Prevalence, penetrance, and expressivity of severe childhood single locus diseases vary with age, leading to matching the population and cohort values. The second step was to examineATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 NGS genes meeting step 1 criterion to identify potential NSCDCC variants on the basis of their contributing to positive diplotypes in the cohort for which:where the population disease prevalence was further corrected for diplotype heterogeneity, which was the greatest proportion of subjects affected by a disease-gene dyad that were associated with a specific diplotype. Variants meeting the step 2 criterion were further evaluated potential NSDCC variants based on pathogenicity assertions in ClinVar, Mastermind, and American College of Medical Genetics classification (Artificial Intelligence Classification Engine, Fabric Genomics) and functional consequences in the literature and disease databases including, for CFTR (MIM:602421)- associated cystic fibrosis (CF, MIM:219700) the Clinical and Functional TRanslation of CFTR (CFTR2) database, and similar databases for F2 (MIM:176930)-associated hypoprothrombinemia / dysprothrombinemia (MIM:613679), F8 (MIM:300841)- associated hemophilia A (MIM:306700), F9 (MIM:300746)-associated hemophilia B (MIM:306900), FGA (MIM:134820), FGB (MIM:134830), and FGG (MIM:134850)- associated dysfibrinogenemia / hypodysfibrinogenemia (MIM:616004) and hypofibrinogenemia / afibrinogenemia (MIM:202400), G6PD (MIM:305900)-associated glucose-6-phosphate dehydrogenase deficiency (MIM:300908), GLA (MIM:300644)- associated Fabry disease (MIM: 301500), OTC (MIM:300461)-associated ornithine transcarbamylase deficiency (MIM:311250), RYR1-associated malignant hyperthermia risk, and SCN5A-associated heart rhythm disorders49-60. Variants classified as NSDCC by this evaluation were removed (block listed) from the NGS set, and step 1 was repeated iteratively the prevalence of all NGS disorders in populations was in equilibrium with their calculated genetic prevalence in matched large diplotype model cohorts.280ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 variants were classified as NSDCC in NGS. Further automation of these methods will be required to undertake this comprehensively.

[0138] In some implementations, an iterative process is performed. For each NGS gene, the first step is to calculate P, the sum of the prevalence of the corresponding x genetic disease(s) associated with that gene (range 1-9) in a population (Fig. 2A, Table 4), where:x was obtained from MIM. Disease prevalence values were taken from the literature. Where possible, they were specific to the genetic ancestry or ancestries being evaluated. Where prevalence estimates differed the most authoritative was selected. For each NGS gene, the second step is to calculate Pcorrected, the prevalence of the corresponding x genetic disease(s) in the population corrected for disease penetrance (p, range 0.25 0.75), disease expressivity (e, range 0.25 0.75), and locus heterogeneity (l, range 0.01 1). Thus:where values for p and e were obtained from the literature for that gene and population. Locus heterogeneity was proportion of the disease-affected subjects attributable to that specific gene. Values for l were derived from the number of genes associated each disease in MIM, together with the relative proportion of disease attributable to that gene from GeneReviews and the literature. For each NGS gene, the third step is to calculate O, the Observed genetic prevalence, calculated by the sum of the frequencies of n unique diplotypes containing P and LP variants with appropriate pattern of inheritance for the disorder mapped to that gene in theGS of individuals in a cohort derived from the population (Fig. 3, Table 4):Step 3 then identified NGS genes for which O was considerably greater than :ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 For such genes, the fourth step is to calculate , the expected maximum diplotype frequency in the population, based on adjustment of for diplotype heterogeneity (d):d (range 0.10.5) was derived from the number of unique diplotypes containing NGS P and LP variants in that gene, together with the relative proportion of disease attributable to that diplotype from MIM, GeneReviews, the literature, and locus specific databases. It should be noted that the prevalence, penetrance, and expressivity of severe childhood single locus diseases may vary with age and genetic ancestry, necessitating matching the population and cohorts for these. The fourth step then examines…, the frequency of each of the n diplotypes mapped to that gene to identify those whose observed frequency in the cohort exceeded :or:In the fifth step, variants contributing to diplotypes where…exceed are further evaluated as potential NSDCC variants based on pathogenicity evidence in ClinVar, Mastermind, and American College of Medical Genetics classification (Artificial Intelligence Classification Engine, Fabric Genomics) and functional consequences in the literature and disease databases including, for CFTR (MIM:602421)-associated cystic fibrosis (CF, MIM:219700) the Clinical and Functional TRanslation of CFTR (CFTR2) database, and similar databases for F2 (MIM:176930)-associated hypoprothrombinemia / dysprothrombinemia (MIM:613679), F8 (MIM:300841)-associated hemophilia A (MIM:306700), F9 (MIM:300746)-associated hemophilia B (MIM:306900), FGA (MIM:134820), FGB (MIM:134830), and FGG (MIM:134850)-associated dysfibrinogenemia / hypodysfibrinogenemia (MIM:616004) and hypofibrinogenemia / afibrinogenemia (MIM:202400), G6PD (MIM:305900)-associated glucose-6-phosphate dehydrogenase deficiency (MIM:300908), GLA (MIM:300644)- associated Fabry disease (MIM: 301500), OTC (MIM:300461)-associated ornithineATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 transcarbamylase deficiency (MIM:311250), RYRl-associated malignant hyperthermia risk, and SCN5A-associated heart rhythm disorders. Variants classified as NSDCC by this evaluation were removed (block listed) from the NGS set. The fourth and fifth steps were repeated iteratively until O, the sum of the observed diplotype frequencies in that gene in the cohort was in equilibrium with , the corrected, expected prevalence in the population. As a result of these steps, 280 variants were classified as NSDCC in NGS.

[0139] NGS queries: In some implementations, supervised, automated GS analysis and interpretation by NGS has, for example, two steps. The first is an automated query that returns the positive diplotypes in an individual GS that contain NGS variants (after removal of block-listed variants) in the NGS genes that obey NGS inheritance rules. The second is a two-step process designed to recover disease causing diplotypes composed of one or more non-NGS variants. This second step consists of GEM and a transformer (e.g., from Fabric Genomics®), run in tandem. First GEM is run in Newborn Sequencing mode, without phenotype data and with variant analyses restricted to the NGS genes and inheritance patterns (Table 1). The GEM output is input to the transformer. The transformer can be, for example, an AI Intelligent Agent that models a human reviewer’s actions when interpreting GEM results. The transformer is trained on a corpus of GEM results and their corresponding clinical diagnoses. The transformer uses a Probabilistic Graphical Model (PGM) to discover and model the data features used by clinical geneticists to interpret GEM results for diagnosis. The transformer can model a single institution’s interpretation SOP, or when trained on a multi-institutional corpus, can create a consensus interpretation machinery. For the analyses reported here, The transformer was trained using prior clinical diagnoses for a benchmark dataset of WGS for 119 probands. One of the Ttransformer’s features is its ability to solve simple cases on its own, and to request human assistance for difficult cases. When used in tandem, GEM and the transformer provided fast, scalable, and highly effective means for automated review of NBS results. The mean number of candidate pathogenic genotypes per proband identified by the transformer was 0.05 (median 0). The transformer automatically and correctly interpreted >98% of positive cases and requested human interpretation help for <5% of cases. In some implementations, the transformer is stored at, trained at, and / or otherwise used by compute device 420, 440, and / or 400.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0140] Retrospective Clinical Utility Assessment: The potential added clinical utility of NGS was evaluated retrospectively by comparing the actual age at diagnosis with counterfactual return of NGS results on day of life (DOL) 10 in 152 of the 3,118 critically ill children who had previously received RDGS with disease detection by both modalities. By review of the electronic medical record (EMR), their ages at symptom onset, first hospitalization, and diagnosis were calculated by RDGS. For children in whom the age at diagnosis was considerably later than DOL 10, further EMR review was undertaken to ascertain the severity and acuity of disease presentation, whether there were definitive changes in treatment upon diagnosis, the total number of hospitalizations and days hospitalized before diagnosis, and whether the child suffered disease-related organ damage. The observed presentations and organ damage were compared with the clinical features of the specific genetic disease in the literature to determine which were attributable to that molecular diagnosis. Based on the assessed efficacy of each indicated intervention for that disorder in GTRx, the impact on the observed clinical features of disease of starting those interventions at the actual age of diagnosis by RDGS was compared with that at the counterfactual age at NGS return of result (DOL 10). The degree of confidence for these assertions were assessed on a scale of low, medium, high, and very high. In SOMI infants for whom a genetic disease was identified by NGS a similar analysis was performed to determine whether infant death was likely to be attributable to that disorder. Based on the efficacy of each indicated intervention for that disorder in GTRx a determination was made whether infant mortality might have been avoided by starting those interventions in the neonatal period. Results

[0141] Disorder Selection for Version 2 of NGS: Some implementations of NGS are configured to meet criteria for clinical gNBS. Of 77 new gene-SCGD dyads evaluated by expert review of published evidence according to Wilson and Jungner principles, 44 (57%) were added to NGS (Table 1). Critical re-appraisal of the 388 NGS disorders led to removal of 34 (9%) due to insufficient penetrance, expressivity, or severity to warrant population NB S, consolidation of disorders comprising a phenotypic spectrum, evidence as genes of uncertain significance, or insufficient identification of causative variants by short-read GSATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 (see Supplementary Results described herein). In toto, NGS included 412 (73%) of 568 gene-SCGD dyads evaluated (Table 1). It should be noted that the consensus to retain a disorder did not imply that the evidence was sufficient for inclusion in a public health gNBS. Rather, it indicated that the benefit-to-harm ratio was sufficient for inclusion in gNBS research studies.

[0142] The patterns of inheritance used in NGS queries were also re-appraised for transition from research to clinical test (see Supplementary Results described herein). Inheritance was re-evaluated in 29 disorders for which MIM lists inheritance as both AD and AR, 6 disorders for which Lyonization leads to sufficiently severe disease in female carriers to warrant population screening, and disorders for which severe childhood disease is limited to homozygotes and compound heterozygotes (see Supplementary Results described herein). For example, G6PDD queries in NGS were changed from X-linked dominant to recessive and limited to variants associated with more severe manifestations (WHO classifications II IV). As a result, the pattern of inheritance used in NGS queries was changed in 36 (9%) of 412 gene-disorder dyads to limit screening to SCGD (Table 1). Periodic review of gene-disorder dyads and inheritance mode used for gNBS can be desirable considering new knowledge.

[0143] Therapeutic interventions for NGS disorders: One reason for disorder selection for gNBS is the likelihood of improved outcomes by early therapeutic interventions in identified, affected individuals. NGS results are returned together with the GTRx eCDS that contains structured evaluations of the efficacy, evidence of efficacy, indications, contraindications, and urgency of initiation of the therapeutic interventions corresponding to the screened disorders. Since the pace of therapeutic development and approval for childhood genetic disorders is accelerating, therapeutic literature was reviewed for new NGS disorders and prior NGS disorders, respectively, as previously described. Fifty interventions were added for 31 NGS disorders, and 140 interventions removed from 46 NGS disorders, representing 12% change in an 18-month interval. For 44 new NGS disorders, 119 interventions were added. In total, NGS features 1,603 beneficial interventions for 412 SCGD (Table 4). Of note, only 16.1% of 9,965 interventions reviewed to date were adjudged to have sufficient evidence of efficacy for inclusion. ThisATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 reflects the paucity of clinical trials for existing SCGD therapies and the extensive use of off-label treatments and therapies based on case report and case series evidence. TABLE 2ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001TABLE 3

[0144] Variant Selection: To enable gNBS to scale to whole populations at very low cost, GS interpretation can be performed by supervised AI (e.g., model 410) rather than manually, the current clinical standard. Some implementations of NGS featured an automated query with ~30,000 ClinVar Pathogenic (P) or Likely Pathogenic (LP) variants that were manually trained by root cause analysis of UK Biobank exomes, identifying 94 variants with evidence against disease causality (Not Severe Disease Causal in Childhood, NSDCC). Upon removal, the estimated specificity of BeginNGS.1 was >99.7% in the UK Biobank. In some implementations, the input for NGS was almost twice as large 53,855 ClinVar and Mastermind variants and featured variants that were curated to be P, LP, and with conflicting pathogenicity assertions (that included at least one P or LP assertion). Since for many variants, the associated condition was not delineated, P and LP assertions were not necessarily for the 412 NGS SCGD. Additionally, P and LP assertions are generally developed in affected children receiving diagnostic GS. Thus, they generally ignore penetrance and expressivity, which are critically important considerations in risk assessment for SCGD in healthy newborns. Likewise, for disorders with variable inheritance, penetrance and expressivity are not considered in affected children receiving diagnostic GS, whereas they are pertinent in risk assessment for SCGD in healthy newborns. These factors were anticipated to exacerbate imprecision due to contaminatingATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 NSDCC variants. Indeed, query of the UKB470k exomes with 53,855 variants identified 73.5% (345,443) subjects as positive for 2,785 diplotypes (containing 2,173 (4%) variants) that matched the inheritance patterns of the 412 disease-gene dyads (after correction for positive subjects with >1 diplotype per disease-gene dyad, Table 4). This exceeded the combined prevalence of the 412 SCGD in UK adults aged >40 years 37-fold, suggesting 97% to be false positives.

[0145] To solve the problem of imprecision associated with NSDCC variants comprehensively in a reproducible, structured manner that met clinical requirements, manual root cause analysis was replaced with supervised federated learning and / or training in several large diplotype models. Federated learning and / or training (queries of datasets at different sites, with or without machine learning) identified high likelihood NSDCC variants comprehensively in a generalizable privacy-protecting manner and on the basis of population frequences that violated the effects of purifying hyperselection on SCGD. Diplotypes, rather than genotypes, were queried to provide direct counts and distinguish compound zygosity states from haplotypes with more than one variant. Following NSDCC variant removal, automated, supervised interpretation of gNBS in individual subjects was performed by both by variant diplotype query and AI-based pathogenicity prediction models (GEM and a transformer; Fig. 2A, see Methods). The trained, adjudicated variant set was used as a true positive training set for refinement of AI-based pathogenicity prediction models to supplement direct GS-based NBS queries with novel variant identification (e.g., transformer), which supplemented direct gNBS queries by identification of novel variants with high likelihood of disease causality to achieve high recall.

[0146] Federated learning and / or training in large diplotype models occurred in several (e.g., two) stages. NGS genes were identified for which the expected prevalence of the associated disease in the UK adult population was considerably lower than the calculated and / or observed genetic prevalence in UKB470K adults using 53,855 variants and the patterns of inheritance discussed above (Fig.3, Table 4). Where more than one disease was associated with a gene, their combined prevalence was used. The observed genetic prevalence was the sum of the observed NGS diplotype frequencies in UKB470K. TheATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 expected genetic prevalence was derived from the observed clinical prevalence by correcting for estimated penetrance, expressivity, disease plurality, and locus heterogeneity. Where more than one gene was associated with a disorder, locus heterogeneity was the proportion of disease-affected subjects associated with that gene (e.g., disease gene dyad; disease plurality can be accounted for by the sum of the prevalence of the individual disorders associated with that gene, if more than one). An exception was RYR1 (MIM:180901)-associated malignant hyperthermia susceptibility (MIM:145600), where P and LP variants and prevalence excluded those associated with RYR1-associated congenital myopathy (MIM:255320) and King-Denborough syndrome (MIM:619542). Prevalence, penetrance, and expressivity values were estimates from the literature unless results of prior population screening were available. The prevalence, penetrance, and expressivity of SCGD often vary with age. For example, the prevalence of Duchenne Muscular Dystrophy (DMD, MIM:310200) in males aged 15-19 years is 2-fold higher than aged <5 years (due to increased expressivity) and 7-fold higher than aged >25 years (due to early death). Among the 25 NGS genes with highest observed, corrected, genetic prevalence in UKB470K, 24 were higher than the corresponding expected UK adult disease prevalence (Table 4). For example, the observed genetic prevalence of AR CFTR-CF (the sum of the frequency of 499 diplotypes comprising 160 of 1,261 NGS CFTR variants) in UKB470K was 2,349 / 100,000. NBS indicates the prevalence of cystic fibrosis (CF, MIM:219700) in UK newborns to be 40 / 100,000, and median survival is 47 years, suggesting an adult population prevalence of 20 / 100,000. CFTR (MIM:602421) is one of four loci associated with bronchiectasis and elevated sweat chloride (BESC). The others are SCNN1B (MIM:600760)-associated BESC1 (MIM:211400), SCNN1A (MIM:600228)-associated BESC2 (MIM:613021), and SCNN1G (MIM:600761)- associated BESC3 (MIM:613071). Since BESC1-3 are much less common than CF, the locus heterogeneity of CFTR-CF was estimated at 0.95. The adult penetrance and expressivity of CFTR-CF are high and were set to 0.7. These values gave an expected genetic prevalence of CFTR-CF m UKB470K of 39 / 100,000, which was 60-fold lower than the observed genetic prevalence in UKB470K, suggesting some of the 160 variants to be NSDCC (Table 5).ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-20014ELBATATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001..4 5 4 5 1 2 1 1 2 7 1 2 1 4 4 1 1 3 1 47 7 7 7 6 5 7 7 75 5. . . . . . . . .7.5.2.7.5.7.7 7 7 7 70 0 0 0 0 0 0 0 0 0 0 0 0 0 0.0.0.0.0.05 5 5 5 6 5 5 5 5 55 2. . .5 5 5 5 5 5 7 7 5 50 0 0.0.0.0.0.0.0.0.0.0.0.0.0.0.0.0.0.0942198373 6 6 2 1 43 3 2 242815131210119097868388614 9 0619 9 1 1 056..4 0 12. 034 1 5 5 3 0 1 2 1 1 2 0 5 7 9 9 9.7 115D D D D D R D R D D R D D D D D R R R D A A A A X A A X A A A A A A A A X X A A C 15A D51A81 A 1 1A1Q C NA I 42X *NXAN NA 1 RH2IR AS M6CC PS SC C D KBAL EGRPLG8F C N COGCSLS P RH5CYR CRNE S AC HDIDLS ACATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0147] The second stage was to examine the disease gene dyads with higher genetic prevalence in UKB470K than the corrected UK prevalence further and identify potential NSDCC variants (the latter of those whose diplotypes frequency in UKB470K was higher than the corrected maximum UK diplotype frequency) and / or to identify NGS variants that contributed to diplotypes whose expected UK adult population frequency was much lower than the observed, diplotype frequency in UKB470K, and evaluate these as potential NSDCC variants (Fig. 3, Table 4). The expected maximum UK diplotype frequency was corrected for diplotype heterogeneity, penetrance, expressivity, disease plurality, and locus heterogeneity. Diplotype heterogeneity was the largest proportion of subjects affected by a disease-gene dyad associated with a specific diplotype. For example, the UK diplotype heterogeneity for CF-CFTR was 0.6, since diplotypes containing the most frequent P variant, CFTR F508del, are identified in 30-80% of patients, dependent upon genetic ancestry. For CF-CFTR, this yielded a maximum expected UK diplotype frequency of 23 / 100,000, a value exceeded by 19 of 499 CFTR diplotypes observed in UKB470K (Fig. 3). Potential NSDCC variants were also evaluated based on pathogenicity assertions in ClinVar, Mastermind, and American College of Medical Genetics classification (Artificial Intelligence Classification Engine, Fabric Genomics) and functional consequences in the literature and disease databases. For example, the compound heterozygous diplotype [chr7- 117590400-G-C];[chr7-117592169-C-T] in CF-CFTR had a frequency of 1,301 / 100,000. The corresponding variants, p.Gly576Ala (ClinVar:7165) and p.Arg668Cys (ClinVar:35835) were part of a known haplotype and not a CF causal diplotype according to the Clinical and Functional TRanslation of CFTR (CFTR2) database. After evaluation, 21 CF-CFTR variants were classified as NSDCC and removed. 93 unique CF-CFTR diplotypes remained, with a total UKB470K genetic prevalence of 46 / 100,000. The most frequent remaining positive CF diplotype was within the corrected, expected UKB470K diplotype frequency. It was compound heterozygosity for CFTR p.Arg117His and p.Phe508del, which are two, relatively common, practice guideline CF-causative alleles [Chr7-17530975-G-A];[Chr7-117559590-ATCT-A] (Fig. 3). Additional examples are discussed in the Supplementary Results described herein.

[0148] The results of federated training in large diplotype models in two other disease- gene dyads with well understood prevalence, locus heterogeneity, allelic heterogeneity,ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 and comprehensive diplotype-phenotype assessment databases were checked (Table 5). G6PDD has a prevalence of 300 / 100,000 in the UK. The raw UKB470 genetic prevalence of G6PDD was 852 / 100,000. Following removal of 4 of 239 NGS variants, the adjusted genetic prevalence was 311 / 100,000. Following removal, the maximum UKB470K G6PDD positive diplotype frequency was within the corrected UK diplotype frequency. It was hemizygous G6PD p. Va198Met (ChrX-154536002-C-T), which is associated with WHO Class III G6PDD. The prevalence of dystrophinopathies in the UK population >40 years of age is 7.3 / 100,000, which includes Becker muscular dystrophy (MIM:300376) and dilated cardiomyopathy 3B (MIM:302045) with a penetrance of ~72% in adult European populations. Following removal of one of 1,682 variants, the genetic prevalence in the UKB470K cohort was 7.0 / 100,000 subjects. Of 50 variants contributing to diplotypes observed at least 195 times in UKB470K (>1 in 2,400), only 6 were childhood disease- causal, all of which were associated with G6PDD (Fig.3). This congruity demonstrates the power of training based on purifying hyperselection of SCGD diplotypes.

[0149] Federated learning and / or training was extended to 96,811 MCPS exomes with different genetic ancestry to UKB470K. For example, IDS (MIM:300823)-associated mucopolysaccharidosis II (MPS2, MIM:309900) is an X-linked recessive disorder without expressivity in carrier females. While the birth prevalence of iduronate-2-sulfatase deficiency is 10.4 / 100,000 by biochemical screening, this includes severe MPS2 (MPS2A, 0.7 per 100,000), attenuated MPS2 (MPS2B, 0.7 / 100,000), and pseudodeficiency (9.0 / 100,000). Since enzyme replacement therapy with idursulfase (EC 3.1.6.13) has been available only since 2006, individuals with MPS2A would not have survived to be enrolled in UKB470K or MCPS. Thus, the prevalence of iduronate-2-sulfatase deficiency in the UKB470K and MCPS cohorts should be ~9 per 100,000. The raw genetic prevalence of IDS-associated disorders in UKB470K and MCPS were 41 / 100,000 and 51 / 100,000, respectively (Table 4). Assuming 75% penetrance and expressivity in adults, and a maximum contribution per diplotype of 30% of IDS positives, the corrected adult diplotype frequency limit was 5 / 100,000. Of 170 UKB470k hemizygotes for IDS c.641C>T (ClinVar:92622, 1 P assertion, 2 LB assertions, 6 B assertions), 163 were of African ancestry (AFR diplotype frequency 2100 / 100,000) and only 6 of European ancestry (EUR diplotype frequency 1.4 / 100,000). MCPS had 49 IDS c.641C>T hemizygotes (MCPSATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 frequency 50 / 100,000). Thus, federated training and / or learning with diverse genetic ancestries flagged this variant as NSDCC. After blocking this variant, the genetic prevalence of IDS-associated disorders in UKB470K was 1.7 / 100,000. Thus, differences in ancestry between the large diplotype models allowed assessment of the specificity of NGS in populations reflective of the diverse US population and detection of ancestry- specific NSDCC.

[0150] Of 370 variants evaluated in this manner, 90 were retained because of functional or case-based evidence that they were disease causal, while 280 were classified as NSDCC and block listed. They included 103 variants determined to lack a ClinVar P or LP assertion, 108 variants not associated with a NGS gene, disorder, or inheritance mode, 52 variants which functional data demonstrated to be NSDCC, and 9 variants that were in haplotypes (linkage disequilibrium) with other LP or P variants. After blocking 280 NSDCC variants, a NGS query of UKB470K identified 1,919 of the 53,575 (53,855-280) remaining variants in 1,903 diplotypes and 13,207 (2.8%) UKB470 subjects, a 26-fold reduction in positive rate. After blocking, 12 of the 25 most frequently NGS positive genes still had a higher genetic prevalence in UKB470K than the corrected UK adult prevalence, indicating that variant evaluation is incomplete (Table 4). AD disorders accounted for the largest number of gNBS positive UKB470K subjects after NSDCC blocking (69%, Table 5), suggesting that this group warranted additional evaluation. The pattern of inheritance with the largest proportionate decrease in gNBS positive UKB470K subjects was AR (29% reduction in variants, 99% reduction in positive subjects, Table 5). Variants categorized as having conflicting pathogenicity assertions in ClinVar were those with the largest proportionate adjustment in positive subjects (98%) but remained the group with the largest number of UKB470K positive subjects (Table 6), suggesting that additional evaluation was needed in this group. Removal of the variants with higher diplotype frequencies in UKB470K than the corrected maximum UK diplotype frequency in the 25 most prevalent genes reduced the UKB470K positive rate to 1.9% (9,062 subjects).ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 V A D ar d Poa jnu sts i ecrst teidveeaPV seosariinitivane tsAD1159 2% (61%)AR520 29% (27%)XR170 14% (9%)XD54 5% (35) Total or 1903 12% Proportio 3.6% or variant APssernoC soi c i r i c si rr rtVeioa rtitertiv or tiv e jectivor tive iannr Poan me ri e cceted ts erVec ected tstn Sh du tob eSdub U Cotrare V CdarUorgPeajntec U jeKr iaUiaKrt KctBecnKnBec ihsBs4te tsBts4tecog4 770K d4 770K dity eni 0c K 0KityConflicting None 342,688 5,963 98% 542 454 16% Pathogenic None 12,937 4,929 62% 1063 939 12% Pathogenic Pathogenic 1,258 950 24% 268 228 15% Uncertain Pathogenic 995 668 33% 70 66 6%ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0151] NGS performance in critically ill newborns with suspected SCGD. Critically ill newborns with suspected SCGD in intensive care units (ICU) increasingly receive RDGS. The precision and recall of NGS in 3,118 such children were evaluated. Phenotype- informed, singleton or parent-child trio RDGS reported diagnostic findings in ~30% of these children. Of these, 187 variants (152 diplotypes) were associated with the 412 NGS gene-disorder dyads (Table 7). Query of 3,118 singleton GS (without phenotype information) by the retained 53,375 NGS variants identified 147 (78.6%) of the RDGS- reported variants and (124 (81.6%) of RDGS-reported diplotypes (Table 7). The 40 variants detected by RDGS but not NGS were either absent from ClinVar or Mastermind or present without P or LP assertions. The query was supplemented by identification of novel loss-of-function variants (in disorders with known loss-of-function genetic mechanism) and by utilizing a transformer automated interpretation tool (e.g., Fabric Genomics). The transformer identified 37 additional RDGS-reported variants. Thus, the NGS recall with these additions was 184 (98.4%) of 187 RDGS-reported variants (150 (98.7%) of 152 RDGS-reported diplotypes; Table 7). Two remaining unidentified variants by the transformer were VUS in recessive disorders that had been reported by RDGS because they were in compound heterozygous diplotypes with pathogenic variants in gene- disorder dyads that were a good fit for the proband phenotypes. In addition, the query with 53,375 retained variants identified 86 variants (82 diplotypes) that had not been reported by RDGS in the 3,118 probands (following removal of nine variants—ClinVar, 189817, 420109, 431442, 575744, 918368, 29895, 464187, 199877, 504421—that were rejected since not P or Lop for a NGS condition and inheritance pattern). Of these, 5 variants (6 diplotypes) were query false positives (not associated with a NGS pattern of inheritance or disorder), giving a NGS query PPV of 97% (200 of 206 diplotypes, Table 7). The 76 newATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 NGS diplotypes were disorders in which early detection allows monitoring for problems, earlier treatment, and better outcomes. Of 228 total diplotypes identified by RDGS or NGS in the 412 gene-disorder dyads in 3,118 GS, the recall of NGS and RDGS were 99.1% and 66.7%, respectively (Table 8). Thus, in a cohort of critically ill infants with suspected, underlying SCGD, the prevalence of positive NGS screens was 7.2% (226 of 3,118), which was significantly greater than in the UKB470K (2.8%, p<.00001; Table 8).13 (6%) of the 228 findings were NBS RUSP core conditions. TABLE 7TABLE 8ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0152] Akin to RUSP NBS, a potential benefit of NGS (e.g., in newborns) is the possibility of effective treatment at or before symptom onset. In contrast, treatment is delayed in RDGS until severe illness has developed, admission to an ICU occurs, evaluation suggests a broad genetic differential diagnosis, RDGS is ordered, and results are returned. The potential added clinical utility of NGS was evaluated by comparing the counterfactual return of NGS results on day of life (DOL) 10 with the actual time of diagnosis by RDGS in children with SCGD detected by both modalities. In 8 children, RDGS was performed outside the neonatal period and there was sufficient knowledge of the natural history of the disorder with and without early treatment to make quantitative assessment possible (Table 9). In those children, symptom onset was at a median DOL 70 (average 105 days, range 0 313, Table 11). A median of 2 (average 2, range 1-7) hospitalizations occurred before RDGS diagnosis, and diagnosis occurred after a median of 8 hospital days (average 23, range 2-97). NGS would potentially have prevented 19 hospitalizations and 181 hospital days in the 8 children (Table 11). NGS would also potentially have shortened the time to diagnosis by a median of 121 days (average 474 days, range 3 - 2,203 days, total 3,791 days, Table 10). Five of the 8 children had life threatening disease presentations, of which 4 would potentially have been avoided by NGS (Table 11). All the children had organ damage, which would potentially have been avoided in six by NGS. In 3 families, parents were evaluated by social services for potential child abuse due to failure to thrive that would potentially have been avoided by NGS (Table 11).

[0153] NGS performance in infant deaths and first-degree relatives of critically ill children. The NGS.253,576 variant query was also evaluated in singleton GS of 3,519 parents and siblings of the 3,118 children who had received RDGS. Following removal of 3 false positive variants (not pathogenic in the heterozygous state) there were 126 positive diplotypes (3.6%, 98% precision, Table 8). Twenty-seven of these were G6PDD. Thus, the proportion of positive NGS screens in young adult first degree relatives of children with suspected SCGD was greater than in the UKB470K (2.8%, p=.01).

[0154] The NGS 53,576 variant query was also evaluated among 705 infant deaths in San Diego County (SOMI) for which archived NBS DBS were available (Table 8). They were a random subset of San Diego infant deaths between 2005 and 2018. NGS identified 61ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 positive diplotypes, of which 5 were query false positives (not pathogenic in the heterozygous state), giving a precision of 92% (56 of 61, Table 8). The 56 remaining positive individuals yielded a similar positive rate (7.9%) to that in children with suspected SCGD who received RDGS (7.2%, p=.52) but significantly greater than in parents and siblings of those who had received RDGS (3.6%, p<.00001, Table 8). In 18 of the 56 infant deaths, NGS identified G6PDD, which was unlikely to have contributed to mortality. In 10 infant deaths, however, NGS identified SCGD that are known to cause infant death and that have effective therapies (Table 10). Clinicopathologic correlation was not possible in these individuals since the archived DBS were de-identified. However, one of them was almost certainly an affected older sister of a 5-week-old boy diagnosed in 16-hours by RDGS with SLC79A3-associated biotin / thiamine-responsive basal ganglia disease (BTRBGD, MIM: 607483, homozygous c.597dup) who received immediate, effective therapy, and had a good outcome. His older sister had died in infancy of progressive encephalopathy during the SOMI study period without a molecular diagnosis. Had these 10 infants received NGS and the indicated therapeutic interventions, their deaths could potentially have been avoided (Table 12). The remaining 28 infant deaths with non-G6PDD NGS findings were positive for SCGD in which early detection allows monitoring for problems, earlier treatment, and better outcomes. Only 8 (14%) of the 56 infant deaths with positive NGS findings would have been positive by RUSP NBS. Thus, NGS might potentially decrease infant mortality by 1.4-5.3%.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 tnaiss S Sbg G A ,hrz son isze rv ira rPujei l, , alyal yal mhCt,Sl aya lit ucciEtiecvtia ut lr yiBd ciltar oortnso irehst,a t esenrdhe,,ee sah ruegndreaAe e s titk l rtisc Bys e aacai eylg tvdal siMliDe De AtibaLooP nenL inr,t uj noa g,ag ,e ziE nyidybc ep aa fhmhoaproa,adnrtMesgag nuauetxsmrel rnnpae I, oc ss ,e ruarha ncgn arco otc,ec feisi aanhhE ’tDwogngn isr eplymcuse oenrig git poo iel arsedeotmDorrDci utlpsmocseeOUICaLaLePHoC NyEhr eRaFO mefniehG matsydnipsybc yhc yhdxeittpta it taeppe pDal liopla lioplaSuti E h E hGp S . pe. peDacrGvReNe cDnEB v6e cDnE41xeSM M DI 1 2ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001,noitirtunlameirolaRhx ,aFct p s c o l t te mu c n-nTG,iJImxu iret riacaaF P h tosuor l,,ngi bym, t esvis rtu al lubb reat b , b , ale Thrap alaimolev pst For,Tl, saoallmuca ymla eer fDsrfoo ciitisisagcirmctorirc ltn u ychirdndsi g io 1diatisnarp ab ilnac lairtohnla pces aaiis tapni ot y s unarir araiteomotmy b eren e b tnluluatprympllerre ovtosunie hc g pmv porniorltpO- a coots o gkni ,tvnot ae e crldv ircti irt gryt, hta eeogoru cevenp prr arteccira o p pna udeaf e p noin nuT p ptlmson yHoc eGnineMehxeyheSrTeLefniuduSub rtaevegverTFyhyhuc ets. e,snaynSdet , ya m -ho no taee ci lehtgao .fne rt-yhsrnvrd eeu itapyohtmlD .ht A5Rtdr mungiarpcou yg s r I e-orryhn aNaor y etPed1Eo yC M5CS esHiDL sXyldodPnenEFMFM 34 5 6ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001evitpenVcor ,aa trf ,onoyme ersr,ce ,faimroes taaap mrwlerreddhc in ae n isn gototseietani mr, nb,re ol oonlsdeterxo,t oh omisineu cii tetrnie ortP, lap su ehths a,dr ah ba,aGvm d,t e dn eloatc epde ud sasis tf ri o araionsipe , Eelen ge npts ra y ehf, nrg,ot elvlpay rhepotd cneD DEl,dalau pe ,psgd syis otcy seiob ylyc C,itici fo ra tl wa o yt aiopcil ecvisGa lihemkecnnand,yhey obimn,m oy rocnnaayhlcno ri no ucin ouemrtem volorysgiai,seate eh ps ht n nitnorh apml ffi nina,td y lameitifomcs er apbaniwi sorbcnerupdytlane nyro mcyni xe,si-c-e ola,Tagosrs ere hoe o ,er , gn lurut ds tnn ar ro lgmalfeer ixnihCtharpcahrc et uaul HiiV d ceiff cRfi udrep ibis d opfs koRrspsn r ur hadziotrepmapiec nia giSh t e b edimxe Trtcaafe t e eiyre y h n rwcaren edo niriotsam ailily aih evgnD s.ogc Ti-sn aGnynitoplaoG.Cf oot I isBesRasBiD F M 78ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 S Bts sNiLesYesYeYhnIR R R R R R D A A A A A A X esa 0esiM50I100001000 045 5 9 87252714141 0 3 1D MDI 2 2 202325213dodo-1coh2ycesa oha laon arc iit noe dl is dl is ne te.fe eseib pioh at ih at rdu ed scefa al ifsenmtcayrh carh adl a e es yxdarcopts ops atissalid obes e lyyeosrahoetb ar eli oet hp li oit hp naelpyrrxer a nhacteihocmneis.muireorpr nonogeorc ouaf paf pno p dualhott ninarbrcifD Hht pnIyhnIyhCyhyhG HysOacedxeSMFMF F F F FM I M OSDI 1 2 3 4 5 6 7ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001p p p p oDeH B TTl>C>4Clpeuadti sptiero *00s mf x 2 7 299d079.1 5 5 o deH AbD DPRPLS gnnm iesoitsapD R Refnet a at o o sX A AiLaeersineN Nel Dsit0 hern73 3T Paf152584 nI13307606etgx r )sAS at ya3 d13702S (1si essEoeae Seci st tyaync cviso sidG B4Nci6 ci1alE into ailxy tpyht tpyht eesee yogstgi paocnl ahh se gDbrna S deeltiaelppoiapptoaan.vr lvuefeudei ailpoh-n gGalE.laEa rgoD ry cii o p it laDuyh tv hp .v hpyPrdPefdem dam Fem hy oli sBaRibp eacDec eecyhenDn erE E dF F F FDI 1 23ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001i,ll9 6 9352523211139 478 55326 048rayleruc aiai EItoterir dr er dr HucariultinaecoeaNvc dynu5 1Apsaefvyh eSdor arp caautr fborSP923015201 005 40107ct ,n ,irnaesroe H eityhtfotI-ellaatotdal a rn sa .hsay5 d.er u pyh en isi Ttoasg oliLerntir adpr ooistmal aB ih eDlviaiat So GMnye. r DXsy co or iDysT-snlg TNlg S -n Amo Shtdo dey nunten .gocnio na SauyEnC N Chml oyl topiseG GtDcSRmoIPC G B Re afganRralre4 5eid auttnu67 8 v eA McAoCATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 egnahcnxeoi atn mesvarletp,nI ypcaitrueehtpaotreohhpT,adetmascaildpnInez2 or1fEhsLerBFATcitu ohaisepalyes ht na saco iseopdatys b ciliah ah xh loryances yis yrtclyy nceg y gcahotycemnsoir epcr psand neyheid es ncr aleoi ncmaeior ncd eipooitDht o.toory eecl h roitp d- io a1p l2f eehcyxife briyci mafe he feeshiha oduaob d c dddla ohHb nmaofythitrneGacen etiali pm rnhIeguo ihvmyndlotiuralt oHnCryFOPDIIM 1 3,4 567 89 O2SATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001nitoib,enimaihTailgnaglasabe esvi as en siodpser-nitoiB01ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 nec hoin ghghghgtediHiHiHirifH egnn a cio ir anitpsyr ten idtaesaaten liot etaerrearvace ee esnatuiu ryfirsLhTi eDrP oofpN NnI c sl Aeirapfu hScat oN xS )etgrsaya 3 9 3 0AtSd(137 0215201Str,nGa oi ,tyNB641alc es ahtybdciy yEinHreluaptht ci htes ehdtetdro ger onxet p ata aeppp nsaalesisy iryhDal liolaelioleae gyr-D dcotaSutipEhppE htaor M. S AmondnpoGpDa.cveec .vpevudy.fg.SNht u eyreRerDne cEDnrEyhPeend Do nCySC yShRLmlXmIotPnE DI 1 2 3 4 5 6ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001214 1191 374217332es, atimtotnce ec nalmy e eaeglelg&n niim mel-ppmoopy na tloiaippDus h h pBhtus2931 3 7831 4 5526408aEiIdHreadnrceyuve doafoSrbrP 5 1005 40107lfato-elaoretIdrnnisa Too mB. lat SG s ia esii ta iDl hv oTNl.ysTi-s Dnaigo lSaGutncnoyiop g cl to s nDafC GiBe aR GeganR rral eteid autnu8v e7A McAoCATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 Discussion

[0155] The genesis of NGS was 13 years of experience in development and implementation of RDGS for SCGD in critically ill (e.g., hospitalized) children, including GTRx management guidance to facilitate rapid translation of results into precision interventions. However, their scope, objectives, indications, reporting, pre-test probability of positives, cost requirements, results recipients, recall, precision, and turnaround time are quite dissimilar (Table 2). The contrasting premise of each creates very different performance requirements: RDGS is ordered to diagnose individual, acutely ill children with suspected SCGD (50% pre-test probability of positive), whereas NGS is designed to screen (e.g., 3.7 million / year), mainly asymptomatic infants (pre-test probability of SCGD <5%). Thus, while RDGS can use very rapid time to result (ideally 1 day) and 100% desired recall, NGS can use very high positive predictive value (PPV) and low cost per subject. Achievement of acceptable recall despite relatively low pre-test probability of SCGD necessitates a very low false positive rate (e.g., type 1 error). Compounding this is Last’s iceberg of underestimated prevalence of mild formes frustes of SCGD associated with variant diplotypes that lack full penetrance or expressivity. Thus NGS false positives encompass both variant diplotypes mis-categorized as P or LP (according to ACMG criteria), those that, while P or LP for SCGD, have insufficient penetrance or expressivity for utility in population screening. These have previously been referred to as late onset pathogenic (LOP) variants. This exemplar of Scylla and Charybdis is much more difficult to navigate than recognized: Untrained NGS.2 screens with 53,855 P and LP variants were positive in 73.5% of UKB470K subjects. Here it was shown, however, that training based on purifying hyperselection of SCGD-causing diplotypes in large diplotype models circumnavigates the underlying problem of both types of false positives. Federated comparisons of NGS disease-gene dyad (genetic) prevalence in ancestry- and age-diverse UKB470K, RCIGM and MCPS genomic datasets with disorder prevalence in ancestry- and age-matched populations (with appropriate corrections) identified 10 (3%) of 342 NGS.2 genes that contributed 97% of positive “raw” screens. By building upon prior allele frequency methods, federated comparison of diplotype prevalence in the three genomic datasets (or genes in these three genomic datasets) with prevalence in ancestry- and age-ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 matched populations (with appropriate corrections) identified 280 (0.5%) of 53,856 NGS variants that contributed 96% of positive UKB470K subjects. Removal (block listing) (e.g., iterative removal) or parameter adjustment of individual disease-gene-variant-inheritance mode tetrads that exceeded credible frequency thresholds reduced positive NGS screens 26-fold to 2.8% of UKB470K subjects. In CF, G6PDD, and DMD, which have well defined prevalence, locus heterogeneity, allelic heterogeneity, penetrance, expressivity, and functional consequence databases, federated training / learning in large diplotype models removed known NSDCC variants, retained disease-causing variants, and yielded genetic prevalence in accord with population prevalence.

[0156] Retrospective cohort evaluations of NGS, with 412 disease-gene dyads and 1,603 associated therapeutic interventions, suggest high potential for clinical utility. NGS was positive in 7.9% of 705 infant deaths in San Diego county, a 7-fold increase over disorders screened by RUSP NBS. Reassuringly this was ~3-fold higher than the positive rate of NGS in UKB470K. An earlier pilot study yielded concordant results: 47 SCGD were identified by GS in 46 (41%) of 112 infant deaths in San Diego County between 2015 and 2020. Of these 5 (4.4%) would have been identified by NGS (CACNAIC-Long QT Syndrome 8 (MIM:618447), COQ2-Primary Coenzyme Q10 Deficiency, MOCS1- Molybdenum Cofactor Deficiency, SCN1A-Developmental and Epileptic Encephalopthy 6B, and TAFFAZIN-Barth Syndrome) (e.g., and an additional two in some implementations of NGS, which is in development NFKB1-common variable immunodeficiency with autoimmunity 12 (MIM:616576) and PPA2-infantile sudden cardiac failure (MIM:617222)). Had these infants received (e.g., were screened by) NGS at birth, together with confirmatory testing and prompt institution of indicated therapeutic interventions, morbidity and mortality may have been reduced substantially. In ten (1.4%) SOMI infant deaths and four (3.6%) of the 112 infant deaths reported\, counterfactual analysis suggested that the NGS-positive disorders were of sufficient severity and indicated therapeutic interventions of sufficient efficacy that death may have been avoidable. These data suggest that NGS has the potential to reduce US infant mortality.

[0157] Retrospective analysis of NGS in 3,118 newborns (e.g., infants) and children in ICUs who received RDGS also suggested that NGS had clinical utility in critically illATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 children. A recent review suggested that RDGS had a diagnostic rate of ~37%, changed outcomes in ~18% of children in ICUs tested, and led to net healthcare cost savings of $14,265 per ICU infant tested. nNGS was positive in 7.2% of ICU infants and children, a 17-fold increase over disorders screened by RUSP NBS. Among eight children with disorders suitable for actual-counterfactual analysis, use of NGS would have shortened time to diagnosis by a median of 121 days (relative to RDGS), and potentially avoided life- threatening disease presentations in four children, and organ damage in six. A similar analysis of NGS in 2,208 of these infants using different methods identified seven children in whom morbidity would likely have been completely avoided, 21 with most morbidity avoided, and 13 with partial avoidance. In the eight children herein, NGS would potentially have avoided healthcare cost of 181 NICU days (at an estimated 2024 average daily spending of $4,332; total, $784,092) and 150 RDGS tests (at an average reimbursement rate under Current Procedural Terminology codes of $8521; total, $1,278,150). Assuming no additional lifetime healthcare cost savings in the 3,118 infants, NGS would be cost neutral at a fee of $661, which is feasible given GS reagent and computation cost of ~$220 per subject. By comparison, current NBS is considered cost effective at an average state fee of $109. Despite Medicaid coverage policies in 12 US States, current ICU utilization of RDGS is less than 5% of those meeting indications for testing. Thus, an early target for NGS implementation may be the 350,000 infants per year (9.5% of births) at time of admission to a neonatal ICU, with more expensive RDGS reserved either for those with rapidly progressing disorders or for those who screen negative (by less expensive reinterpretation of GS generated for NGS). Results of such a pilot trial will be reported elsewhere (ClinicalTrials.gov NCT06276348).

[0158] Numerous clinical trials of gNBS have recently commenced, each with a distinct panel of disease-gene dyads, endpoints, and unique representation of genetic ancestry. For example, an adaptive, multicenter clinical trial was performed to compare clinical utility and cost effectiveness of NGS and RUSP NBS in up to 100,000 newborns, with an initial enrollment emphasis on Hispanic, Amerindian, and African American ancestry (ClinicalTrials.gov ID NCT06306521, the content of which is incorporated by reference herein in its entirety). In toto, gNBS trials world-wide are evaluating a total of ~1,750 genesATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 (e.g., between 100 and 1,000 disorders associated with a total of 1,750 genes according to disparate criteria).

[0159] In summary, evidence to date suggests that gNBS is feasible for hundreds of SCGD with effective therapies. Given the magnitude of changes in knowledge between some implementations of NGS (e.g., from version 1 of NGS to version 2 of NGS), annual re- review of the underpinning structured rare disease molecular and treatment knowledgebase can be desirable and NGS can in some instances remain an adaptive, open platform. Prevalence and diplotype-based distributed learning appear to be a generalizable approach to achievement of acceptable analytic performance for population testing. Supplementary Results Disorder Selection for Version 2 of NGS

[0160] Seventy-seven new gene-disorder dyads and the 388 NGS disorders were evaluated or re-evaluated, respectively, both as in a first implementation of NGS and to address five additional questions for a second implementation of NGS. The additional questions were:

[0161] Did the gene have several phenotype associations that represent a single phenotype spectrum? The expert consensus was to amalgamate such phenotypes if the mode of inheritance and therapy were identical. This remains incomplete and warrants periodic re- evaluation. The following were amalgamated as single phenotype spectra: a. GBA1 (Mendelian Inheritance in Man (reference 49 in manuscript), MIM:606463) associated Gaucher disease type I (MIM:230800), type II (MIM:230900), type III (MIM:231000), and type Mc (MIM:231005). b. IDUA (MIM:252800)-associated Hurler syndrome (MIM:607014), Hurler- Scheie syndrome (MIM:607015), and Shceie syndrome (MIM:607016) were amalgamated as Mucopolysaccharidosis I. c. RAG1 (MIM:179615)-associated alpha / beta T-cell lymphopenia with gamma / delta T-cell expansion, severe cytomegalovirus infection, and autoimmunity (MIM:609889), Combined cellular and humoral immune defects with granulomas (MIM:233650), B cell-negative severe combinedATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 immunodeficiency (SCID, MIM:601457), and Omenn syndrome (MIM:603554) were amalgamated as RAG1-associated SCID. d. RAG2 (MIM:179616)-associated Combined cellular and humoral immune defects with granulomas (MIM:233650), B cell-negative severe combined immunodeficiency (SCID, MIM:601457), and Omenn syndrome (MIM:603554) were amalgamated as RAG2-associated SCID. e. SLC52A3 (MIM: 613350)-associated Fazio-Londe disease (MIM:211500) and Brown-Vialetto-Van Laere syndrome 1 (MIM:211530). f. SMN1 (MIM:600354)-associated Spinal Muscular Atrophy type 1 (MIM:253300), type 2 (MIM:253550), and type 3 (MIM:253400) were amalgamated as Spinal Muscular Atrophy. g. WAS (MIM: 300392)-associated Wiskott-Aldrich syndrome (MIM:301000), X- linked thrombocytopenia (MIM:313900), and X-linked severe congenital neutropenia (MIM:300299. h. FOXA2 (MIM:600288)-associated congenital isolated hyperinsulinism (Orphanet:657) and FOXA2-associated combined pituitary hormone genetic deficiencies (Orphanet:95494). i. ADA (MIM:608958)-associated severe combined immunodeficiency (MIM:102700) and Omenn syndrome (Orphanet:39041).

[0162] Mild, common disorders associated with genes that also had serious, childhood- onset disorders were removed. Thus, HPD (MIM:609695)-associated autosomal dominant (AD) Hawkinsinuria (MIM:140350) was removed, whereas HPD-associated autosomal recessive (AR) tyrosinemia 3 (MIM: 276710) was retained. Likewise, CSF3R (MIM:138971)-associated AD hereditary neutrophilia (MIM:162830) was removed, while AR severe neutropenia 7 (MIM:617014) was retained. Thrombophilia (MIM:188050) associated with F2 (MIM: 176930) and FI3A1 (MIM:134570) were removed. The multiple phenotypes associated with some channelopathy genes were also re-examined. SCN5A (MIM:600163), for example, is associated with 9 heart rhythm disorders. 3 severe childhood SCN5A phenotypes were retained that have distinct treatments (AD progressive familial heart block type 1A (MIM:113900), AD dilated cardiomyopathy 1E (MIM: 601154), and AD long QT syndrome 3 (MIM:603830). Four less severe or adult-onsetATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 phenotypes were removed (AD Brugada syndrome 1, MIM: 601144, AR sick sinus syndrome 1, MIM: 608567, AD atrial fibrillation 10, MIM:614022, and AD nonprogressive heart block, MIM:113900). CACNA1C (MIM:114205)-associated AD Brugada syndrome 3 (MIM: 611875) was also removed. The 2 remaining SCN5A phenotypes are discussed below, as is penetrance of Brugada syndrome and long QT syndrome. SCN4A (MIM:603967) is associated with six phenotypes. Four (AD hyperkalemic periodic paralysis (MIM:170500), AD hypokalemic periodic paralysis type 2 (MIM:613345), AD paramyotonia congenita (MIM:168300), and AR congenital myasthenic syndrome 16 (MIM:614198)) and removed two (potassium-aggravated myotonia (MIM:608390) and normokalemic periodic paralysis (MIM:170600)) were retained.

[0163] Was the gene-disorder association of uncertain significance (GUS)? The expert consensus was that inclusion required each gene-disorder dyad to have been adequately described in >5 independent families. SCN5A -associated AD familial ventricular fibrillation 1 (MIM:603829) which, when distinct from Brugada syndrome, has only been described in 3 adult patients (Table 1) was removed. Four of 19 Diamond-Blackfan anemia (DBA) dyads were removed (TSR2 (MIM:300945)-associated DBA14 (MIM:300946), RPS27 (MIM:603702)-associated DBA17 (MIM: 617409), RPL35 (MIM:618315)- associated DBA19 (MIM:618312), and RPS15A (MIM:603674)-associated DBA20 (MIM: 618313)). Two of 16 Fanconi anemia dyads were removed for the same reason (MAD2L2 (MIM:604094)-associated complementation group V, (MIM:617243) and RFWD3 (MIM:614151)-associated complementation group W (MIM:617784)). UCP2 (MIM:601693)-associated AD hyperinsulinism (ORPHA:276556) was removed since the literature suggests that should be considered a GUS despite descriptions in 11 patients. MYO9A (MIM:604875)-associated AR presynaptic congenital myasthenic syndrome 24 (MIM: 618198) was rejected since it has only been described in three families. CYBB (MIM:300481)-associated immunodeficiency 34 (MIM:300645) was rejected since it has been described in two families and only in adults.

[0164] Can the majority of affected individuals be identified by short-read GS? Sensitivity may be low in ‘dead zone’ exons (such as adjacent to LINE elements), genes with highlyATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 homologous pseudogenes, or difficult variants (such as inversions). The expert consensus, based on comparisons with other screening tests, was that an overall NGS sensitivity of >90% was desired. This implied that only a few NGS gene-disorder dyads should have lower sensitivity. In such disorders, GS identification of variants associated with >50% of affected individuals was required for retention. The CDKN1C (MIM:600856)-associated disorders Beckwith-Wiedemann syndrome (MIM:130650), intrauterine growth retardation, metaphyseal dysplasia, adrenal hypoplasia congenita, and genital anomalies (MIM:614732), and Silver-Russell syndrome 4 (MIM:180860) were removed (Table 1). In most affected individuals, these disorders are attributable to imprinting that cannot be detected natively by short-read GS-by-synthesis.

[0165] Is childhood penetrance sufficient for population screening? The expert consensus was that an average >40% positive predictive value during early childhood was desired for NGS.8 susceptibility factors were removed for atypical hemolytic uremic syndrome which have individual childhood penetrance of ~5% (AHUS1, MIM:23500, associated with CFH (MIM:134370), CFHR3 (MIM:605336), and CFHR1 (MIM:134371), AHUS2, MIM:612922, associated with MCP (MIM:120920), AHUS3, MIM:612923, associated with CF1 (MIM:217030), AHUS4, MIM:612924, associated with CFB (MIM: 138470), AHUS5, MIM:612925, associated with C3 (MIM:120700), AHUS6, MIM:612926, associated with THBD (MIM:188040), AHUS7, MIM:615008, associated with DGKE (MIM:601440), AHUS8, MIM:301110, associated with C1GALT1C1 (MIM:300611)). SCN5A-associated AR susceptibility to Sudden Infant Death Syndrome (MIM:272120), for which the penetrance is unknown, was also removed. RYR1 (MIM:180901)-associated malignant hyperthermia risk (MIM:145600) was retained since it has an FDA-accredited pathogenic variant set, ~40% penetrance, and a simple, effective intervention exists (avoidance of certain general anesthetics). Long QT syndrome (MIM:603830, MIM:192500, and MIM: 618447) was also retained, which has a penetrance of ~40%. As noted above, Brugada syndrome, which has a penetrance of ~46% and low expressiyity in children, was removed.

[0166] Several NGS-eligible phenotypes are restricted to a very small subset of observed variants, such as rare gain of function or dominant negative variants. Other than RYR1-ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 associated malignant hyperthermia risk and several CACNA1C (MIM:114205) Timothy syndrome (MIM:601005) variants (including c.1216G>A, p.Gly406Arg, ClinVar:17632), NGS does not screen for these disorders in some implementations. They include LAMTOR2 (MIM:610389) c.*23C>A (ClinVar:1239)-associated immunodeficiency (MIM:610798), and TCF3 (MIM:147141) c.1663G>A (p.Glu555Lys, ClinVar: 225870)-associated AD agammaglobulinemia 8. In some implementations, these variants can be incorporated in an “Allow List” that permits screening for these disorders and inheritance patterns delimited to these variants. Inheritance mode used in NGS queries

[0167] The patterns of inheritance used in NGS queries of some gene-disorder dyads were also re-evaluated as part of the transition from a NGS research-grade prototype to NGS clinical test by three questions.

[0168] Would the burden of follow-up of asymptomatic or mildly affected positives overwhelm primary care pediatricians? For mild dominant gene-disease dyads disorders that have severe recessive forms, should NGS queries be limited to the severe recessive form? G6PD (MIM:305900) deficiency (G6PDD, MIM:307800) was queried in an X- linked dominant manner in BeginNGS.1. However, G6PDD is highly prevalent in some ancestries, and heterozygote females typically have mild disease. In response, G6PDD queries were changed in BeginNGS.2 to an X-linked recessive model and limited to variants associated with more severe manifestations (WHO classifications II IV).

[0169] For autosomal disorders in which MIM lists inheritance as both AD and AR, which inheritance mode should be employed in NGS queries? For 29 autosomal BeginNGS.1 disorders, MIM lists inheritance as both AD and AR. In some implementations of NGS, these were queried in a dominant manner. For some disorders, this was appropriate, and homozygotes are more severely affected than heterozygotes. For others, while inheritance is almost certainly AR, a few heterozygous affected individuals or families were described in the older literature. It is recently that GS has been more generally used to identify second variants in heterozygous affected individuals, and the older literature is unreliable for ruling out a second, non-coding variant in such individuals. For NGS, an unambiguous singleATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 inheritance pattern was selected based on literature review as well as comparison of the incidence of the disorder and the total frequency of positive genotypes in large diplotype models under AD and AR inheritance (see results). As a result, 15 such disorders associated with 14 genes were queried under a dominant model for NGS: a. ABCC8 (MIM:600509)-associated familial hyperinsulinemic hypoglycemia 1 (MIM:256450) and permanent neonatal diabetes mellitus 3 (MIM:618857), b. CPOX (MIM:612732)-associated hereditary coproporphyria (MIM:121300), c. FAS (MIM:134637)-associated autoimmune lymphoproliferative syndrome (MIM: 601859), d. FOXA2-associated combined genetic pituitary hormone deficiencies (Orphanet:95494), e. GCH1 (MIM:600225)-associated dopa-responsive dystonia (MIM:128230), f. GCK (MIM:138079)-associated familial hyperinsulinemic hypoglycemia 3 (MIM: 602485), g. GLRA1 (MIM:138491)-associated hyperekplexia 1 (MIM: 149400), h HESX1 (MIM:601802)-associated septooptic dysplasia (MIM:182230), i. INS (MIM:176730)-associated permanent neonatal diabetes mellitus 4 (MIM:618858), j. KCNJ11 (MIM:600937)-associated hyperinsulinemic hypoglycemia 2 (MIM:601820), k. LHX4 (MIM:602146)-associated combined pituitary hormone deficiency 4 (MIM: 262700), 1. POU1F1 (MIM:173110)-associated combined pituitary hormone deficiency 1 (MIM:613038), m. SLC2A1 (MIM:138140)-associated GLUT1 deficiency syndrome 1 (MIM:606777), and n. SLC16A1 (MINI: 600682)-associated familial hyperinsulinemic hypoglycemia 7 (MIM: 610021). Fourteen such disorders associated with 12 genes were queried under a recessive model for NGS: a. AHCY (MIM:607826)-associated hypermethioninemia with S-denosylhomocysteine hydrolase deficiency (MIM:613752),ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 b. AQP2 (MIM:107777)-associated autosomal nephrogenic diabetes insipidus (MIM:125800), c. CHAT (MIM:118490)-associated presynaptic congenital myasthenic syndrome 6 (MIM:254210), d. CHRNE (MIM:100725)-associated congenital myasthenic syndrome 4A (MIM:605809), e. COLQ (MIM:603033)-associated congenital myasthenic syndrome (MIM:5603034), f. CYP11A1 (MIM:118485)-associated congenital adrenal insufficiency with partial / complete 46, XY sex reversal (MIM:613743), g. GFPT1 (MIM:138292)-associated congenital myasthenic syndrome 12 (MIM: 610542), h. SLC6A5 (MIM:604159)-associated hyperekplexia 3 (MIM:614618), i. SLC7A9 (MIM:604144)-associated cystinuria (MIM:220100), j. TH (MIM:191290)-associated Segawa syndrome (MIM: 605407), k. FGA (MIM:134820)-associated afribrinogenemia / hypofibrinogenemia (MIM:202400) and dysfibrinogenemialhypodysfibrinogenemia (MIM:616004), and l. FGB (MIM:134830)-associated afribrinogenemia / hypofibrinogenemia (MIM:202400) and dysfibrinogenemialhypodysfibrinogenemia (MIM:616004). Under a dominant model, the genetic prevalence of these disorders was higher than the upper bound of the population prevalence (see below).

[0170] For X-linked recessive (XR) disorders, does Lyonization lead to sufficiently severe disease in female carriers to warrant population screening? NGS screened for these disorders under a recessive model. However, 22% of female carriers of OTC (MIM:300461)-associated ornithine transcarbamylase deficiency (MIM:311250) are clinically affected, with episodic (11%), chronic (7.5%), or neonatal forms of the disease (3.5%). Carrier status was included in NGS for OTC and five other XR disorders (ATP7A (MIM:300011)-associated occipital horn syndrome (MIM:304150), AVPR2 (MIM:300538)-associated nephrogenic inappropriate antidiuresis syndrome (MIM:300539), GLA (MIM:300634)-associated Fabry disease (MIM:301500), SLC6A8 (MIM:300036)-associated creatine transport deficiency (MIM:300352), and XIAP (MIM:300079)-associated lymphoproliferative syndrome 2 (MIM:300635).ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0171] In toto, the pattern of inheritance used in NGS queries of 36 (9%) of 412 gene- disorder dyads was changed. Periodic review of these allocations will be necessary considering new knowledge.

[0172] Additional examples of gene-disorder dyads characterized for potential NSDCC variants include:

[0173] GCH1, for example, is one of 37 loci associated with AD dopa-responsive dystonia (DRD) and accounts for 27% of affected individuals (locus heterogeneity). Twelve variants contributed 569 GCH1-DRD positive UKB470k subjects, resulting in an uncorrected genetic prevalence of 121 / 100,00 (Table 6). The adult UK prevalence of DRD is ~10 / 100,000 (Orphanet). Employing estimates of 50% adult penetrance, 70% adult expressivity and a maximum contribution per diplotype of 50% of GCH1-DRD positives, the corrected UK diplotype frequency was 18 / 100,000. 328 UKB470k subjects had diplotypes containing GCH1 c.671A>G (ACE benign, ClinVar 3 P assertions, 5 LP assertions, 4 VUS assertions) suggesting a penetrance of 2.4% if indeed disease-causing. Based on such evidence, this and two other of 12 GCH1 variants were designated NSDCC. Block listing these variants resulted in a corrected genetic prevalence of GCH1-DRD of 3.4 / 100,000 in UKB470k subjects, which was within the corrected UK prevalence of DRD (8 / 100,000). Likewise, HESX1 is one of 8 loci associated with combined pituitary hormone deficiency (CPHD). Ten variants contributed 320 positive HESX1-CPHD diplotypes in the UKB470k cohort, resulting in a genetic prevalence of 68 / 100,00. The UK adult prevalence of CPHD is 9 per 100,000 (Orphanet). Assuming that HESX1 accounted for 50% of CPHD, had 50% adult penetrance, 70% adult expressivity, and the maximum contribution per diplotype was 50% of HESX1 -CPHD positives, the corrected UK diplotype frequency was 18 / 100,000. Based on such evidence, 3 of 12 HESX1 variants were designated NSDCC, and block listed, resulting in a corrected genetic UKB470K prevalence of 19 per 100,000, which was close to the corrected UK adult prevalence of CPHD of 13 / 100,000. Supplementary Discussion

[0174] The model described herein for identification of NSDCC variants includes (See Methods):ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0175] However, this formula may be individualized for specific genes and disorders. For example, the value of d (diplotype heterogeneity) may need to be adjusted if n, the number of diplotypes represented in the training set of 53,855 ClinVar and Mastermind variants, is lower than the number of known causal variants for that disorder. Likewise, for some genes, P, the sum of the prevalence of the corresponding x genetic disease(s) associated with that gene may need to be only the prevalence of the NGS disorder(s) since well qualified variant sets are available for that disorder. In addition, penetrance and expressivity may differ between the disorders associated with a gene. An example is the age at which penetrance and expressivity reach their maximum values.

[0176] Fig. 7 illustrates a flowchart of a method 700 to identify a misclassified gene, according to an embodiment. In some implementations, method 700 is performed by a processor (e.g., processor 402).

[0177] At 702, a dataset (e.g., dataset 406) comprising a plurality of genetic variants is received (e.g., at compute device 400 and / or memory 404). At 704, a first set of data (e.g., data 426) generated (1) without receiving a first set of genetic data (e.g., genetic data 428), (2) at a first remote compute device (e.g., compute device 420) that has access to the first set of genetic data and not a second set of genetic data (e.g., genetic data 448), and (3) based on the first set of genetic data and not the second set of genetic data is received. At 706, a second set of data (e.g., data 446) generated (1) without receiving the second set of genetic data, (2) at a second remote compute device (e.g., compute device 440) that has access to the second set of genetic data and not the first set of genetic data, and (3) based on the first set of genetic data and not the second set of genetic data is received. At 708, at least one misclassified genetic variant from the plurality of genetic variants is identified, using a model (e.g., model 410), based on the first set of data and the second set of data. At 710, the dataset is updated based on the at least one misclassified genetic variant to generate an updated dataset (e.g., updated dataset 408) that achieves an improved result (e.g., positive predictive value (PPV)) for a genetic screening.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0178] Some implementations of method 700 further include receiving a sample from a subject associated with a genetic variant and analyzing the sample based on the updated dataset to determine whether the genetic variant is a risk factor for a genetic disease.

[0179] In some implementations of method 700, the dataset indicates that a genetic variant from the plurality of genetic variants has a first risk for a genetic disease, and the updated dataset indicates that the genetic variant has a second risk for the genetic different than the first risk.

[0180] In some implementations of method 700, the model is configured to perform an iterative process. The iterative process includes obtaining the first set of data comprising a first plurality of genes. Each gene from the first plurality of genes can be associated with one or more genetic diseases. The iterative process further includes, for each gene from the first plurality of genes, calculating an expected prevalence of the one or more genetic diseases in a first population. The iterative process further includes, for each gene from the first plurality of genes, calculating an observed prevalence of the one or more genetic diseases in a first cohort derived from the first population. The iterative process further includes identifying, to identify a first set of identified genes, genes from the first plurality of genes where the observed prevalence exceeds the expected prevalence beyond a first predetermined acceptable threshold for genes. The iterative process further includes, for each identified gene of the first set of identified genes, selecting a first set of at least one diplotypes associated with the one or more genetic diseases. The iterative process further includes, in a first population, calculating an expected maximum frequency of each diplotype of the first set of at least one diplotypes. The iterative process further includes, in the first cohort, calculating an observed frequency of each diplotype of the first set of at least one diplotypes. The iterative process further includes detecting a first detected diplotype where an the observed frequency of the first detected diplotype exceeds an expected maximum frequency of the first detected diplotype beyond a first predetermined acceptable threshold. The iterative process further includes categorizing genetic variants that contribute to the first detected diplotype into a first set of allowed genetic variants and a first set of blocked genetic variants. The iterative process further includes removing or adjusting the first set of blocked genetic variants to generate an updated first set of geneticATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 data. In some implementations, the iterative process is repeated, replacing the first set of data with the updated first set of data until the observed prevalence is within a predetermined acceptable range of the expected prevalence.

[0181] Some implementations of method 700 further include obtaining the second set of data comprising a second plurality of genes and a second plurality of diplotypes. Each gene from the second plurality of genes is associated with the one or more genetic disease. For each gene from the second plurality of genes, an expected prevalence of the one or more genetic disease in a second population is calculated. For each from the second plurality of genes, an observed prevalence of the one or more genetic diseases in a second cohort derived from the second population is calculated. Genes from the second plurality of genes where the observed prevalence exceeds the expected prevalence beyond a second predetermined acceptable threshold for genes is identified to identify a second set of identified genes. For each identified gene of the second set of identified genes, a second set of at least one diplotypes associated with the one or more genetic diseases is selected. In the second population, an expected maximum frequency of each diplotype of the second set of at least one diplotypes is calculated. A second detected diplotype where an observed frequency of the second detected diplotype exceeds an expected maximum frequency of the second detected diplotype beyond a second predetermined acceptable threshold for diplotypes is detected. Genetic variants that contribute to the second detected diplotype are categorized into a second set of allowed genetic variants and a second set of blocked genetic variants. The second set of blocked genetic variants is removed or adjusted to generate an updated second set of data. The process can repeat, replacing the second set of data with the updated second set of data, until the observed prevalence is within a predetermined acceptable range of the expected prevalence. In some implementations, the at least one misclassified genetic variant is further identified based on the first set of blocked genetic variants and / or the second set of blocked genetic variants.

[0182] In some implementations of method 700, the first set of data is a first set of weights used in a machine learning model and the second set of data is a second set of weights used in the machine learning model.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0183] In some implementations of method 700, the updated dataset indicates at least one of (1) a subset of genetic variants from the plurality of genetic variants that are block-listed or (2) a subset of genetic variants from the plurality of genetic variants that are allow-listed.

[0184] In some implementations of method 700, the first set of data indicates at least one of a diplotype, a zygosity, a gene, a pattern of inheritance (POI), a count, or metadata.

[0185] In some implementations of method 700, the first set of genetic data includes data predetermined as sensitive, and the first set of data does not include the data predetermined as sensitive.

[0186] In some implementations of method 700, the at least one misclassified variant is a false positive for a genetic disease.

[0187] Some implementations of method 700 further include updating the model based on the updated dataset.

[0188] In an embodiment, genetic sequence information of a newborn is obtained (e.g., received). An apparatus configured to perform method 700 analyzes the genetic sequence information, and the risk of the newborn having one or more genetic disorders is determined.

[0189] Fig. 8 illustrates a flowchart of a method 800 to generate a more accurate dataset, according to an embodiment. In some implementations, method 800 is performed by a processor (e.g., processor 402).

[0190] At 802, a dataset (e.g., dataset 406) including a plurality of genetic variants is received. At 804, a plurality of sets of data (e.g., data 426 and 446) is received from a plurality of remote compute devices (e.g., compute device 420 and 440). Each set of data from the plurality of sets of data is generated (1) at a remote compute device from the plurality of remote compute devices different for remaining sets of data from the plurality of sets of data and (2) based on a set of genetic data from a plurality of sets of genetic data (e.g., genetic data 428 and 448) accessible by the remote compute device that generated that set of data and not remaining remote compute devices from the plurality of remoteATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 compute devices. At 806, in response to executing a model (e.g., model 410), at least one misclassified genetic variant from the plurality of genetic variants is identified based on the plurality of sets of data. At 808, the dataset is updated based on the at least one misclassified genetic variant to generate an updated dataset (e.g., updated dataset 408) associated with a lower false positive rate than the dataset and for a genetic screening.

[0191] In some implementations of method 800, a set of data from the plurality of sets of data indicates a diplotype, a zygosity, a gene, a pattern of inheritance (POI), and a count.

[0192] Fig. 9 illustrates a flowchart of a method 900 to determine whether a variant is a risk factor for a genetic disease, according to an embodiment. In some implementations, method 900 is performed by a processor (e.g., processor 402). At 902, a sample from a subject (e.g., newborn) associated with a variant is received (e.g., for a genetic screening). At 904, the sample is analyzed based on a dataset (e.g., updated dataset 408) to determine whether the variant is a risk factor for a genetic disease. The dataset can be generated using a process that includes receiving a preliminary dataset (e.g., dataset 406) including a plurality of genetic variants (where the plurality of genetic variants include the variant). The process can further include receiving a plurality of sets of data (e.g., data 426 and 446) from a plurality of remote compute devices (e.g., compute device 420 and 440). Each set of data from the plurality of sets of data can be generated (1) at a remote compute device from the plurality of remote compute devices different for remaining sets of data from the plurality of sets of data and (2) based on a set of genetic data from a plurality of sets of genetic data accessible by the remote compute device that generated that set of data and not remaining remote compute devices from the plurality of remote compute devices. The process can further include identifying at least one misclassified genetic variant from the plurality of genetic variants based on the plurality of sets of data and using a model (e.g., model 410). The process can further include updating the preliminary dataset based on the at least one misclassified genetic variant to generate the dataset.

[0193] In some implementations of method 900, the plurality of sets of data is a plurality of sets of weights, and the model is trained based on the plurality of sets of weights to identify the at least one misclassified genetic variant.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0194] In some implementations of method 900, each set of genetic data from the plurality of sets of genetic data represents genetic data associated with a population from a plurality of populations different for remaining sets of genetic data from the plurality of sets of genetic data. The model can be configured to receive samples associated with the plurality of populations.

[0195] In some implementations of method 900, identifying the at least one misclassified genetic variant includes repeating a process until an observed prevalence of a gene associated with the plurality of sets of data is within a predetermined acceptable range of an expected prevalence of the gene.

[0196] Combinations of the foregoing concepts and additional concepts discussed here (provided such concepts are not mutually inconsistent) are contemplated as being part of the subject matter disclosed herein. The terminology explicitly employed herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.

[0197] The skilled artisan will understand that the drawings primarily are for illustrative purposes, and are not intended to limit the scope of the subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0198] To address various issues and advance the art, the entirety of this application (including the Cover Page, Title, Headings, Background, Summary, Brief Description of the Drawings, Detailed Description, Embodiments, Abstract, Figures, Appendices, and otherwise) shows, by way of illustration, various embodiments in which the embodiments may be practiced. As such, all examples and / or embodiments are deemed to be non-limiting throughout this disclosure.

[0199] It is to be understood that the logical and / or topological structure of any combination of any program components (a component collection), other componentsATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 and / or any present feature sets as described in the Figures and / or throughout are not limited to a fixed operating order and / or arrangement, but rather, any disclosed order is an example and all equivalents, regardless of order, are contemplated by the disclosure.

[0200] Various concepts may be embodied as one or more methods, of which at least one example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments. Put differently, it is to be understood that such features may not necessarily be limited to a particular order of execution, but rather, any number of threads, processes, services, servers, and / or the like that may execute serially, asynchronously, concurrently, in parallel, simultaneously, synchronously, and / or the like in a manner consistent with the disclosure. As such, some of these features may be mutually contradictory, in that they cannot be simultaneously present in a single embodiment. Similarly, some features are applicable to one aspect of the innovations, and inapplicable to others.

[0201] The indefinite articles “a” and “an,” as used herein in the specification and in the embodiments, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0202] The phrase “and / or,” as used herein in the specification and in the embodiments, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001

[0203] As used herein in the specification and in the embodiments, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of” or “exactly one of,” or, when used in the embodiments, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the embodiments, shall have its ordinary meaning as used in the field of patent law.

[0204] As used herein in the specification and in the embodiments, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non- limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0205] In the embodiments, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to meanATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 including but not limited to. Only the transitional phrases “consisting of” and “consisting essentially of” shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

[0206] Some embodiments described herein relate to a computer storage product with a non-transitory computer-readable medium (also can be referred to as a non-transitory processor-readable medium) having instructions or computer code thereon for performing various computer-implemented operations. The computer-readable medium (or processor- readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) may be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to, magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application- Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a computer program product, which can include, for example, the instructions and / or computer code discussed herein.

[0207] Some embodiments and / or methods described herein can be performed by software (executed on hardware), hardware, or a combination thereof. Hardware modules may include, for example, a processor, a field programmable gate array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can include instructions stored in a memory that is operably coupled to a processor, and can be expressed in a variety of software languages (e.g., computer code), including C, C++, Java™, Ruby, Visual Basic™, and / or other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, butATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher- level instructions that are executed by a computer using an interpreter. For example, embodiments may be implemented using imperative programming languages (e.g., C, Fortran, etc.), functional programming languages (Haskell, Erlang, etc.), logical programming languages (e.g., Prolog), object-oriented programming languages (e.g., Java, C++, etc.) or other suitable programming languages and / or development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.

[0208] The terms “instructions” and “code” should be interpreted broadly to include any type of computer-readable statement(s). For example, the terms “instructions” and “code” may refer to one or more programs, routines, sub-routines, functions, procedures, etc. “Instructions” and “code” may include a single computer-readable statement or many computer-readable statements.

[0209] While specific embodiments of the present disclosure have been outlined above, many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the embodiments set forth herein are intended to be illustrative, not limiting.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 NUMBERED EMBODIMENTS 1. A method for newborn screening comprising: a) performing genome sequencing on a sample from a newborn individual; b) querying the genomic sequence with a set of variants associated with severe childhood genetic disorders, wherein the set of variants has been prequalified by training in large cohorts to achieve high recall and positive predictive value; c) identifying positive findings based on the query results and predefined inheritance patterns; d) generating a report of the positive findings; and e) providing an electronic management guidance system to facilitate the translation of positive screening results into effective therapeutic interventions. 2. The method of embodiment 1, wherein prequalifying the set of variants comprises: (a) calculating an expected prevalence for each disorder based on population data; (b) observing a genetic prevalence for each disorder in one or more large genomic datasets; (c) comparing the observed genetic prevalence to the expected prevalence; and (d) removing or adjusting variants associated with disorders where the observed genetic prevalence significantly exceeds the expected prevalence. 3. The method of embodiment 2, wherein calculating the expected prevalence comprises adjusting for penetrance, expressivity, and locus heterogeneity of the disorder. 4. The method of embodiment 2, wherein observing the genetic prevalence comprises summing frequencies of diplotypes containing variants associated with each disorder. 5. The method of embodiment 2, wherein comparing the observed genetic prevalence to the expected prevalence uses a statistical threshold to identify significant differences. 6. The method of embodiment 1, wherein the large cohorts comprise genomic data from multiple sites and federated learning is used for prequalification. 7. The method of embodiment 1, further comprising supplementing the variant query with identification of novel loss-of-function variants. 8. The method of embodiment 1, further comprising using an artificial intelligence tool to assist in interpreting query results.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 9. The method of embodiment 1, wherein the electronic management guidance comprises information on confirmatory testing, therapeutic interventions, and their efficacy. 10. The method of embodiment 1, wherein the severe childhood genetic disorders comprise at least 300 disorders associated with at least 300 genes. 11. A system for newborn screening of genetic disorders, comprising: a) a sequencing device configured to obtain genomic sequences from newborns; b) a database containing a pre qualified set of variants associated with severe childhood genetic disorders; c) a processor configured to query the genomic sequences with the prequalified set of variants and identify positive findings based on predefined inheritance patterns; d) a report generator configured to produce a report of the positive findings; and e) an electronic management guidance module configured to provide information on therapeutic interventions for positive findings. 12. The system of embodiment 11, further comprising a federated learning module configured to prequalify the set of variants using genomic data from multiple sites. 13. The system of embodiment 11, further comprising an artificial intelligence module configured to assist in interpreting query results. 14. The system of embodiment 11, wherein the database is regularly updated based on new genomic data and clinical knowledge. 15. A method for prequalifying variants for use in newborn genomic screening, comprising: a) obtaining genomic data from multiple large cohorts; b) calculating an expected prevalence for each disorder of interest based on population data; c) observing a genetic prevalence for each disorder in the genomic data; d) comparing the observed genetic prevalence to the expected prevalence; e) identifying variants associated with disorders where the observed genetic prevalence significantly exceeds the expected prevalence; and f) removing or adjusting the identified variants in a screening database. 16. A method for achieving optimal outcomes in individuals with genetic diseases comprising presymptomatic population screening by genomic sequencing comprisingATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 automated interpretation and electronic management guidance to translate positive screens into effective therapeutic interventions. 17. The method of embodiment 16, wherein the automated interpretation is accomplished by querying variants in the genomic sequence of individuals with genetic diseases with a prequalified set of variants, diplotypes, haplotypes, genes, disorders, inheritance patterns or any combination thereof. 18. The method of embodiment 17, wherein the prequalified set of variants, diplotypes, haplotypes, genes, disorders, inheritance patterns or any combination thereof, are subjected to a prequalification by training in genomic sequences of large cohorts of individuals to achieve sufficiently high recall and positive predictive value to be acceptable for presymptomatic population screening. 19. The method of embodiment 17, wherein the prequalified set of variants, diplotypes, haplotypes, genes, disorders, inheritance patterns or any combination thereof are subjected to a prequalification by training in genomic sequences of large cohorts of individuals enables identification of variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns, or any combination thereof, that do not cause severe disease with sufficient penetrance and expressivity to be acceptable for presymptomatic population screening. 20. The method of embodiment 17, wherein the prequalification by training in genomic sequences of large cohorts of individuals leads to removal of variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns, or any combination thereof, that do not cause severe disease in individuals with genetic diseases. 21. The method of embodiment 20, wherein variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns, or any combination thereof, that do not cause severe disease in screened individuals are identified based on having higher observed genetic prevalence in genomic sequences of large cohorts of individuals than an expected prevalence of one or more disorders to be screened in individuals with genetic diseases. 22. The method of embodiment 21, wherein the expected prevalence of one or more disorders to be screened and associated with a gene is determined by the sum of the frequency of the prevalence of each of those disorders in the individuals with genetic diseases.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 23. The method of embodiment 21, wherein the expected prevalence of one or more disorders to be screened and associated with a gene is corrected for the penetrance, expressivity, and locus heterogeneity of those disorders in the individuals with genetic diseases. 24. The method of embodiment 21, wherein the observed genetic prevalence of one or more disorders associated with a gene to be screened is determined by the sum of the observed frequencies of all diplotypes containing many putatively pathogenic variants in combinations that fit a pattern or patterns of inheritance of those disorders in genomic sequences of large cohorts of individuals. 25. The method of embodiment 17, further comprising prequalification of a pattern or patterns of inheritance of one or more disorders associated with a gene to be screened for which the observed genetic prevalence in genomic sequences of large cohorts of individuals exceeds their expected prevalence in a population to be screened is achieved by changing from dominant to recessive. 26. The method of embodiment 21, wherein the expected prevalence of each putatively pathogenic variant associated with one or more disorders associated with a gene to be screened is the product of diplotype heterogeneity in the population and the expected prevalence calculated according to any of the preceding claims. 27. The method of embodiment 21, wherein the expected prevalence in the population to be screened is adjusted under a suitable distribution and one-tailed confidence interval. 28. The method of embodiment 19, wherein the variants, diplotypes, haplotypes, genes, disorders, and inheritance patterns or any combination thereof, that do not cause severe disease in screened individuals are identified or refined by the additional step of identification of other evidence of lack of disease causality. 29. The method of embodiment 16, wherein a plurality of genomic sequences of large cohorts of individuals is used. 30. The method of embodiment 29, wherein the plurality of genomic sequences of large cohorts of individuals are aggregated at a single site. 31. The method of embodiment 30, wherein the plurality of genomic sequences of large cohorts of individuals are at multiple sites and federated or distributed learning is used forATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 prequalification. 1’. A method for newborn screening comprising: a) performing genome sequencing on a sample from a newborn individual; b) interpreting data from the genome sequencing using large diplotype models to identify genetic variants associated with a set of genetic diseases; c) automatically classifying the identified genetic variants based on their potential to cause severe disease in the newborn individual; and d) providing an electronic management guidance system to facilitate the translation of positive screening results into effective therapeutic interventions. 2’. The method of embodiment 1’, wherein the large diplotype models are used to determine the recall, positive predictive value, clinical utility, and cost-effectiveness of the genome sequencing-based newborn population screening. 3’. The method of embodiment 2’, further comprising removing variants, diplotypes, haplotypes, inheritance patterns, and disease-gene dyads from the screening that are determined not to cause severe disease with sufficient penetrance and expressivity. 4’. The method of embodiment 3’, wherein non-severe-disease-causing variants are identified by analyzing their genetic prevalence in the large diplotype models to the corresponding prevalence in the population to be screened. 5’. The method of embodiment 4’, wherein the genetic prevalence of disease-gene dyads is determined by summing the frequency of all positive diplotypes of all potentially pathogenic variants for that locus in the large diplotype models. 6’. The method of embodiment 5’, further comprising correcting the genetic prevalence for the occurrence of more than one positive diplotype per gene.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 7’. The method of embodiment 4’, wherein the disorder population prevalence is established as: the upper bound of prevalence after adjustment for genetic disease heterogeneity, genetic penetrance, and expressivity. 8’. The method of embodiment 7’, wherein genetic heterogeneity is defined as the proportion of disease-affected subjects associated with a specific gene-disease dyad. 9’. The method of embodiment 7’, wherein penetrance is defined as the proportion of subjects who screen positive for a disease-gene dyad and are affected by that disease. 10’. The method of embodiment 7’, wherein expressivity is defined as the proportion of subjects who screen positive and are affected in whom the disease is of sufficient severity to warrant therapeutic intervention. 11’. The method of embodiment 3’, further comprising changing the screened pattern of inheritance of disease-gene dyads with higher genetic prevalence than adjusted population prevalence from dominant to recessive. 12’. The method of embodiment 3’, further comprising identifying and removing individual nonsevere-disease-causing variants in disease-gene dyads from the screen. 13’. The method of embodiment 12’, wherein non-severe-disease-causing variants are identified by comparing their diplotype frequency in a large diplotype model to the maximum credible diplotype frequency in a matched population. 14’. The method of embodiment 13’, wherein the maximum credible population diplotype frequency is calculated by: adjusting the disorder population frequency for diplotype heterogeneity. 15’. The method of embodiment 14’, wherein diplotype heterogeneity is defined as the upper bound of the proportion of genetic prevalence attributable to diplotypes containing a specific variant. 16’. The method of embodiment 15’, wherein the maximum credible population diplotype frequency is adjusted using a suitable distribution and one-tailed confidence interval. 17’. The method of embodiment 13, further comprising identifying non-severe-disease- causing variants in disease-gene dyads by additional evidence of lack of disease causality.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 18’. The method of any one of embodiments 1’-17’, wherein a plurality of large diplotype models is used. 19’. The method of embodiment 18’, wherein the large diplotype models are aggregated at a single site. 20’. The method of embodiment 18’, wherein federated or distributed learning is used to query large diplotype models at multiple sites.

Claims

ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 What Is Claimed:

1. An apparatus, comprising: a memory; and a processor operatively coupled to the memory, the processor configured to: receive a dataset comprising a plurality of genetic variants; receive, without receiving a first set of genetic data, a first set of data generated (1) at a first remote compute device that has access to the first set of genetic data and not a second set of genetic data and (2) based on the first set of genetic data and not the second set of genetic data; receive, without receiving the second set of genetic data, a second set of data generated (1) at a second remote compute device that has access to the second set of genetic data and not the first set of genetic data and (2) based on the first set of genetic data and not the second set of genetic data; identify, using a model, at least one misclassified genetic variant from the plurality of genetic variants based on the first set of data and the second set of data; and update the dataset based on the at least one misclassified genetic variant to generate an updated dataset that achieves an improved result for a genetic screening.

2. The apparatus of claim 1, wherein the processor is further configured to: receive a sample from a subject associated with a genetic variant; and analyze the sample based on the updated dataset to determine whether the genetic variant is a risk factor for a genetic disease.

3. The apparatus of claim 1, wherein the dataset indicates that a genetic variant from the plurality of genetic variants has a first risk for a genetic disease, and the updated dataset indicates that the genetic variant has a second risk for the genetic variant different than the first risk.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 4. The apparatus of claim 1, wherein the model is configured to perform a process that includes: obtaining the first set of data comprising a first plurality of genes and a first plurality of diplotypes, each gene from the first plurality of genes associated with one or more genetic diseases; for each gene from the first plurality of genes, calculating an expected prevalence of the one or more genetic diseases in a first population; for each gene from the first plurality of genes, calculating an observed prevalence of the one or more genetic diseases in a first cohort derived from the first population; identifying, to identify a first set of identified genes, genes from the first plurality of genes where the observed prevalence exceeds the expected prevalence beyond a first predetermined acceptable threshold for genes; for each identified gene of the first set of identified genes, selecting a first set of at least one diplotypes associated with the one or more genetic diseases; in the first population, calculating an expected maximum frequency of each diplotype of the first set of at least one diplotypes; in the first cohort, calculating an observed frequency of each diplotype of the first set of at least one diplotypes; detecting a first detected diplotype where an observed frequency of the first detected diplotype exceeds an expected maximum frequency of the first detected diplotype beyond a first predetermined acceptable threshold; categorizing genetic variants that contribute to the first detected diplotype into a first set of allowed genetic variants and a first set of blocked genetic variants; and removing or adjusting the first set of blocked genetic variants to generate an updated first set of data.

5. The apparatus of claim 4, wherein the processor is further configured to: repeat, replacing the first set of data with the updated first set of data, the process until the observed prevalence is within a predetermined acceptable range of the expected prevalence.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 6. The apparatus of claim 4, wherein the process further includes: obtaining the second set of data comprising a second plurality of genes and a second plurality of diplotypes, each gene from the second plurality of genes associated with the one or more genetic diseases; for each gene from the second plurality of genes, calculating an expected prevalence of the one or more genetic diseases in a second population; for each gene from the second plurality of genes, calculating an observed prevalence of the one or more genetic diseases in a second cohort derived from the second population; identifying, to identify a second set of identified genes, genes from the second plurality of genes where the observed prevalence exceeds the expected prevalence beyond a second predetermined acceptable threshold for genes; for each identified gene of the second set of identified genes, selecting a second set of at least one diplotypes associated with the one or more genetic diseases; in the second population, calculating an expected maximum frequency of each diplotype of the second set of at least one diplotypes; in the second cohort, calculating an observed frequency of each diplotype of the second set of at least one diplotypes; detecting a second detected diplotype where an observed frequency of the second detected diplotype exceeds an expected maximum frequency of the second detected diplotype beyond a second predetermined acceptable threshold for diplotypes; categorizing genetic variants that contribute to the second detected diplotype into a second set of allowed genetic variants and a second set of blocked genetic variants; and removing or adjusting the second set of blocked genetic variants to generate an updated second set of data.

7. The apparatus of claim 6, wherein the processor is further configured to: repeat, replacing the second set of data with the updated second set of data, the process until the observed prevalence is within a predetermined acceptable range of the expected prevalence.ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 8. The apparatus of claim 6, wherein the at least one misclassified genetic variant is further identified based on the first set of blocked genetic variants and the second set of blocked genetic variants.

9. The apparatus of claim 4, wherein the at least one misclassified genetic variant is further identified based on the first set of blocked genetic variants.

10. The apparatus of claim 1, wherein the first set of data is a first set of weights used in a machine learning model and the second set of data is a second set of weights used in the machine learning model.

11. The apparatus of claim 1, wherein the updated dataset indicates at least one of (1) a subset of genetic variants from the plurality of genetic variants that are block-listed or (2) a subset of genetic variants from the plurality of genetic variants that are allow-listed.

12. The apparatus of claim 1, wherein the first set of data indicates at least one of a diplotype, a zygosity, a gene, a pattern of inheritance (POI), a count, or metadata.

13. The apparatus of claim 1, wherein the first set of genetic data includes data predetermined as sensitive, and the first set of data does not include the data predetermined as sensitive.

14. The apparatus of claim 1, wherein the at least one misclassified genetic variant is a false positive for a genetic disease.

15. The apparatus of claim 1, wherein the processor is further configured to update the model based on the updated dataset.

16. The apparatus of claim 1, wherein the improved result for genetic screening includes an improved Positive Predictive Value (PPV).ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 17. The apparatus of claim 2, wherein the subject is a newborn.

18. A method for newborn screening of genetic disorders, comprising: obtaining genetic sequence information of a newborn; using the apparatus of claim 1 to analyze the genetic sequence information; and determining a risk of the newborn having one or more genetic disorders.

19. A method, comprising: receiving a dataset including a plurality of genetic variants; receiving a plurality of sets of data from a plurality of remote compute devices, each set of data from the plurality of sets of data generated (1) at a remote compute device from the plurality of remote compute devices different for remaining sets of data from the plurality of sets of data and (2) based on a set of genetic data from a plurality of sets of genetic data accessible by the remote compute device that generated that set of data and not remaining remote compute devices from the plurality of remote compute devices; identifying, in response to executing a model, at least one misclassified genetic variant from the plurality of genetic variants based on the plurality of sets of data; and updating the dataset based on the at least one misclassified genetic variant to generate an updated dataset associated with a lower false positive rate than the dataset and for a genetic screening.

20. The method of claim 19, wherein a set of data from the plurality of sets of data indicates a diplotype, a zygosity, a gene, a pattern of inheritance (POI), and a count.

21. A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to: receive a sample from a subject associated with a variant for a genetic screening; and analyze the sample based on a dataset to determine whether the variant is a risk factor for a genetic disease, the dataset generated using a process that includes:ATTORNEY DOCKET NO. RCHH-001 / 02WO 350895-2001 receiving a preliminary dataset including a plurality of genetic variants, the plurality of genetic variants including the variant; receiving a plurality of sets of data from a plurality of remote compute devices, each set of data from the plurality of sets of data generated (1) at a remote compute device from the plurality of remote compute devices different for remaining sets of data from the plurality of sets of data and (2) based on a set of genetic data from a plurality of sets of genetic data accessible by the remote compute device that generated that set of data and not remaining remote compute devices from the plurality of remote compute devices; identifying at least one misclassified genetic variant from the plurality of genetic variants based on the plurality of sets of data and using a model; and updating the preliminary dataset based on the at least one misclassified genetic variant to generate the dataset.

22. The non-transitory, processor-readable medium of claim 21, wherein the plurality of sets of data is a plurality of sets of weights, and the model is trained based on the plurality of sets of weights to identify the at least one misclassified genetic variant.

23. The non-transitory, processor-readable medium of claim 21, wherein each set of genetic data from the plurality of sets of genetic data represents genetic data associated with a population from a plurality of populations different for remaining sets of genetic data from the plurality of sets of genetic data, and the model is configured to receive samples associated with the plurality of populations.

24. The non-transitory, processor-readable medium of claim 21, wherein identifying the at least one misclassified genetic variant includes repeating a process until an observed prevalence of a gene associated with the plurality of sets of data is within a predetermined acceptable range of an expected prevalence of the gene.

Citation Information

Patent Citations

  • Screening, Diagnosis and Prognosis of Autism and Other Developmental Disorders

    US20150227681A1

  • Method of characterising a cancer

    WO2022200293A1

  • Method and system for newborn screening for genetic diseases by whole genome sequencing

    WO2023014816A1