A method for predicting drug response or time to death or cancer progression in patients with non-small cell lung cancer (NSCLC) from circulating tumor DNA (ctDNA) using signals from both baseline ctDNA levels and longitudinal changes in ctDNA levels over time.

A predictive model combining baseline ctDNA levels and ctDNA changes using machine learning improves the accuracy of molecular response scoring, allowing for personalized cancer treatment decisions.

JP2025538871APending Publication Date: 2025-12-02GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025527722
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2023-11-10
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing methods for calculating molecular response from circulating tumor DNA (ctDNA) levels are often inaccurate and lack clear interpretation for predicting patient response to treatment or cancer progression, necessitating improved methods for determining meaningful molecular response scores that incorporate baseline ctDNA levels and changes over time.

Method used

A predictive model that integrates baseline ctDNA levels and interactions between linear and non-linear ctDNA changes using machine learning algorithms to determine a composite score for predicting time to death or progression events, taking into account patient demographics and ctDNA metrics.

Benefits of technology

The model provides accurate and clinically significant predictions of patient response to treatment and cancer progression, enabling personalized treatment decisions based on risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538871000001_ABST
    Figure 2025538871000001_ABST
Patent Text Reader

Abstract

Provided herein is a method for determining molecular response score for use in predictive model.Molecular response score can be used to monitor and guide the administration of treatment to subject.The molecular response score can be generated by a method comprising: for at least one variant among a plurality of variants classified into somatic cells, determining the weighted average of the first variant allele fraction (MAF) and the weighted average of the second MAF based on the first MAF and the second MAF.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 383,838, filed November 15, 2022, and U.S. Provisional Patent Application No. 63 / 493,075, filed March 30, 2023, which are incorporated herein by reference in their entireties for all purposes. [Background technology]

[0002] background Molecular response is calculated by measuring changes in circulating tumor DNA (ctDNA) levels observed in samples taken from a subject at various time points. In certain cases, calculations are based on the fraction of somatic variants in the total cell-free DNA (cfDNA) in the sample. In other cases, calculations are based on the concentration of ctDNA in the sample (i.e., normalized for the cfDNA concentration in the sample). A common problem is that existing calculations of molecular response frequently result in inaccurate or unclear molecular response scores. Furthermore, there is little information on how to combine these variables to better interpret ctDNA results and improve prediction of patient response to treatment or patient progression. Predicting patient response to treatment or cancer progression is important information for clinicians, who can use such results to modify patient treatment to more or less aggressive options depending on the results. Therefore, there remains a need for molecular response scores with meaningful interpretations that have clinical significance, such as correlations between response to treatment and changes in ctDNA quantity, and methods for determining absolute baseline (pre-treatment) ctDNA levels that have been shown to be associated with cancer patient prognosis. Summary of the Invention [Means for solving the problem]

[0003] Described herein are models that incorporate the effect of baseline ctDNA levels and interactions between baseline ctDNA levels and linear and non-linear relative ctDNA change to predict the length of time until a patient experiences death or a progression event, taking into account patient demographics and / or baseline ctDNA levels, ctDNA measurements, and metrics of ctDNA change.

[0004] Abstract receiving, by a computer system, genetic information of a subject, the genetic information including data obtained at two or more time points, wherein cancer has been detected in the subject; extracting from the genetic information one or more features identified from a plurality of samples obtained from the subject; generating, by a first classifier implemented in the computer system, a first output indicative of a first classification of the subject using a first machine learning algorithm; generating, by a second classifier implemented in the computer system, a second output indicative of a second classification of the subject using a second machine learning algorithm; identifying, by the computer system, additional subjects from a population having genetic information that matches the genetic information of the subject based on the first classification and the second classification; determining, by the computer system, at least one score for the subject based on the additional subjects having matching genetic information; determining, by the computer system, a composite score using the at least one score; and determining, by a recommender implemented in the computer system, a recommendation indicative of a treatment for the subject based on the composite score. Described herein are methods comprising:

[0005] In another embodiment, the one or more features each comprise a first variant allele fraction (MAF) and a second MAF from two or more time points. In another embodiment, at least one score is based on the first variant allele fraction (MAF) and the second MAF, a weighted average of the first MAF, and a weighted average of the second MAF. In another embodiment, at least one score is based on the ratio of the weighted average of the first MAF to the weighted average of the second MAF, and a confidence interval. In another embodiment, at least one score is based on the first variant allele fraction (MAF) at a first time point and the second MAF at a second time point, a first measure of central tendency of the first MAF, and a second measure of central tendency of the second MAF. In another embodiment, at least one score is based on the ratio of the first measure of central tendency at the first time point to the second measure of central tendency at the second time point. In another embodiment, the measure of central tendency is one or more of the mean, median, or mode. In another embodiment, the method includes comparing a molecular response score for a subject with cancer to a predetermined cutoff point, and identifying the subject as a likely responder to one or more cancer treatments if the molecular response score is below the predetermined cutoff point, or identifying the subject as a likely non-responder to one or more cancer treatments if the molecular response score is at or above the predetermined cutoff point. In another embodiment, the one or more treatments include one or more immunotherapies. In another embodiment, the method includes administering one or more cancer treatments to the subject taking into account the at least one score. In another embodiment, the method includes discontinuing administering one or more cancer treatments to the subject taking into account the at least one score. In another embodiment, the method includes using the at least one score as a prognostic and / or predictive biomarker for the subject. In another embodiment, the method includes calculating the standard deviation of each MAF ratio in the set of MAF ratios using molecular counting. In another embodiment, the method includes propagating the variance through each MAF ratio in the set of MAF ratios.

[0006] In another embodiment, the method includes one or more germline and / or clonal hematopoietic variants when determining mutant allele frequency (MAF). In another embodiment, the first time point includes a pre-treatment time point, and the second time point includes a treatment period or a post-treatment time point. In another embodiment, the method includes generating sequence information from nucleic acid molecules obtained from one or more tissues or cells in a sample. In another embodiment, the method includes generating sequence information from extracellular free nucleic acid (cfNA) in a sample obtained from the subject. In another embodiment, the cfNA includes circulating tumor DNA (ctDNA). In another embodiment, the first and / or second classifiers are each implemented in a machine learning algorithm. In another embodiment, the machine learning algorithm is selected from a neural network, a support vector machine, a hidden Markov model, or a random forest model. In another embodiment, at least one score corresponds to a level of responsiveness to the treatment among a plurality of levels of responsiveness to the treatment.

[0007] The present invention describes a method that includes the steps of: generating genetic information using a genetic analysis device; receiving a training dataset in a computer memory for each of a plurality of individuals with cancer disease, the training dataset including: (1) genetic information from the individual generated at a first time point; and (2) the individual's subsequent treatment response to one or more therapeutic interventions determined at a second time point; and training a computer classifier using the training dataset to obtain a trained computer classifier, wherein the trained computer classifier determines at least one score for the subject based on variant allele frequencies (MAFs) at least at the first time point and the second time point; and determining the amount of change between the MAFs. In another embodiment, the method includes predicting the subject's treatment response based on the amount of change between the MAFs.

[0008] In other embodiments, the method includes selecting a treatment for the subject based on the amount of variation between the MAFs.

[0009] In certain aspects, the present disclosure provides methods for determining a molecular response score at least in part using a computer. The method includes determining a first plurality of sequence reads and a second plurality of sequence reads associated with the subject, wherein the first plurality of sequence reads are determined before administration of the treatment and the second plurality of sequence reads are determined after administration of the treatment; classifying a plurality of variants in the first plurality of sequence reads and the second plurality of sequence reads as somatic or germline; determining a weighted average of a first variant allele fraction (MAF) and a weighted average of a second MAF for at least one variant of the plurality of variants classified as somatic based on the first MAF and the second MAF; determining a ratio of the weighted average of the first MAF to the weighted average of the second MAF for the subject; determining a confidence interval based on the ratio of the weighted average of the first MAF to the weighted average of the second MAF; and outputting the ratio and the confidence interval of the weighted average of the first MAF to the weighted average of the second MAF as a molecular response score.

[0010] In one aspect, the present disclosure provides a method for determining a molecular response score at least partially using a computer. The method includes the steps of: determining a first plurality of sequence reads and a second plurality of sequence reads associated with a subject, wherein the first plurality of sequence reads are determined before administering a treatment, and the second plurality of sequence reads are determined after administering the treatment; classifying a plurality of variants in the first plurality of sequence reads and the second plurality of sequence reads into somatic or germline; determining a variant allele fraction (MAF) ratio for at least one variant among the plurality of variants classified as somatic based on a first MAF and a second MAF; determining a weighted average of the MAF ratios for the subject; determining a confidence interval associated with the weighted average of the MAF ratios based on the weighted average of the MAF ratios; and outputting the weighted average of the MAF ratios and the confidence interval as a molecular response score.

[0011] In certain aspects, the present disclosure provides methods for determining a molecular response score at least in part using a computer. The method includes determining a first plurality of sequence reads and a second plurality of sequence reads associated with the subject, the first plurality of sequence reads being determined before administration of the treatment and the second plurality of sequence reads being determined after administration of the treatment; classifying a plurality of variants in the first plurality of sequence reads as somatic or germline; classifying a plurality of variants in the second plurality of sequence reads as somatic or germline; reclassifying at least one variant of the plurality of variants to resolve a classification discrepancy between the first and second plurality of sequence reads; determining a first variant allele fraction for at least one variant of the plurality of variants classified or reclassified as somatic based at least in part on the first plurality of sequence reads; determining a second variant allele fraction for at least one variant of the plurality of variants classified or reclassified as somatic based at least in part on the second plurality of sequence reads; and determining a molecular response score based on the first variant allele fraction and the second variant allele fraction.

[0012] In certain aspects, the present disclosure provides a method for determining a molecular response score at least in part using a computer, the method including the steps of: determining a first plurality of sequence reads and a second plurality of sequence reads associated with a subject, wherein the first plurality of sequence reads are determined before administration of a treatment and the second plurality of sequence reads are determined after administration of the treatment; classifying a plurality of variants in the first plurality of sequence reads and the second plurality of sequence reads as somatic or germline; and classifying at least one variant of the plurality of variants as associated with Clonal Hematopoiesis of Indeterminate Potential. the plurality of variants includes determining a variant as a somatically classified (somatically classified) Potential (CHIP) variant; removing at least one CHIP variant from the plurality of variants; determining a first variant allele fraction for at least one variant of the plurality of somatically classified variants based at least in part on the first plurality of sequence reads; determining a second variant allele fraction for at least one variant of the plurality of somatically classified variants based at least in part on the second plurality of sequence reads; and determining a molecular response score based on the first variant allele fraction and the second variant allele fraction.

[0013] In certain aspects, the present disclosure provides a method for determining a molecular response score at least partially using a computer, the method comprising: determining a first plurality of sequence reads and a second plurality of sequence reads associated with a subject, the first plurality of sequence reads being determined before administration of a treatment and the second plurality of sequence reads being determined after administration of the treatment; classifying a plurality of variants in the first plurality of sequence reads as somatic or germline; classifying a plurality of variants in the second plurality of sequence reads as somatic or germline; reclassifying at least one variant among the plurality of variants to resolve classification discrepancies between the first and second plurality of sequence reads; determining at least one variant among the plurality of variants as a clonal hematopoietic with undetermined potential (CHIP) variant; removing at least one CHIP variant from the plurality of variants; and classifying the at least one CHIP variant as somatic or germline. determining a first variant allele fraction for at least one variant of the plurality of somatically classified or reclassified variants based at least in part on the first plurality of sequence reads; determining a second variant allele fraction for at least one variant of the plurality of somatically classified or reclassified variants based at least in part on the second plurality of sequence reads; determining a MAF ratio for at least one variant of the plurality of somatically classified or reclassified variants based on the first variant allele fraction and the second variant allele fraction; determining a weighted average of the MAF ratios for the subject; determining a confidence interval associated with the weighted average of the MAF ratios based on the weighted average of the MAF ratios; and outputting the weighted average of the MAF ratios and the confidence interval as a molecular response score.

[0014] In one aspect, the present disclosure provides a method for determining a molecular response score at least partially using a computer, the method comprising: determining a first plurality of sequence reads and a second plurality of sequence reads associated with a subject, the first plurality of sequence reads being determined before administration of a treatment and the second plurality of sequence reads being determined after administration of the treatment; classifying a plurality of variants in the first plurality of sequence reads as somatic or germline; classifying a plurality of variants in the second plurality of sequence reads as somatic or germline; reclassifying at least one variant among the plurality of variants to resolve classification discrepancies between the first plurality of sequence reads and the second plurality of sequence reads; determining at least one variant among the plurality of variants as a clonal hematopoietic with undetermined potential (CHIP) variant; removing at least one CHIP variant from the plurality of variants; and classifying the plurality of variants classified as somatic. determining a first mutant allele fraction (MAF) for at least one variant among the plurality of somatically classified variants based at least in part on the first plurality of sequence reads; determining a second MAF for at least one variant among the plurality of somatically classified variants based at least in part on the second plurality of sequence reads; determining a weighted average of the first MAF and a weighted average of the second MAF for at least one variant among the plurality of somatically classified variants based on the first MAF and the second MAF; determining a ratio of the weighted average of the first MAF to the weighted average of the second MAF for the subject; determining a confidence interval based on the ratio of the weighted average of the first MAF to the weighted average of the second MAF; and outputting the ratio and confidence interval of the weighted average of the first MAF to the weighted average of the second MAF as a molecular response score.

[0015] In one aspect, the present disclosure provides a method for determining a molecular response score at least partially using a computer. The method includes the steps of: determining a first plurality of sequence reads and a second plurality of sequence reads associated with a subject, wherein the first plurality of sequence reads are determined at a first time point before administration of a treatment, and the second plurality of sequence reads are determined at a second time point after administration of the treatment; classifying a plurality of variants in the first plurality of sequence reads and the second plurality of sequence reads into somatic or germline; determining a first measure of central tendency of the first variant allele fraction (MAF) at the first time point and a second measure of central tendency of the second MAF for at least one variant of the plurality of variants classified as somatic, based on the first MAF at the first time point and the second MAF at the second time point; determining a ratio of the first measure of central tendency at the first time point to the second measure of central tendency at the second time point; and outputting the ratio of the first measure of central tendency at the first time point to the second measure of central tendency at the second time point as a molecular response score.

[0016] In one aspect, the present disclosure provides a method for determining a molecular response score for a subject with cancer, at least in part, using a computer. The method includes: (a) determining, by a computer, mutant allele frequencies (MAFs) for a plurality of variants from sequence information generated from targeted nucleic acids associated with one or more cancer types in samples obtained from the subject at a first time point and a second time point, thereby generating a first MAF and a second MAF set for each variant within the plurality of variants. The method also includes: (b) calculating, by a computer, a ratio between the first MAF and the second MAF for each variant within the plurality of variants, thereby generating a set of MAF ratios and a corresponding standard deviation for each MAF ratio within the set of MAF ratios. Furthermore, the method also includes: (c) calculating, by a computer, a weighted average and a confidence interval of the MAF ratios, thereby determining a molecular response score for the subject with cancer.

[0017] In another aspect, the present disclosure provides a method for treating cancer in a subject. The method includes: (a) determining mutant allele frequencies (MAFs) for multiple variants from sequence information generated from targeted nucleic acids associated with one or more cancer types in samples obtained from the subject at a first time point and a second time point, thereby obtaining a first MAF and a second MAF set for each variant in the multiple variants. The method also includes: (b) calculating a ratio between the first MAF and the second MAF for each variant in the multiple variants, thereby obtaining a set of MAF ratios and a corresponding standard deviation for each MAF ratio in the set of MAF ratios. The method also includes: (c) calculating a weighted average and confidence interval of the MAF ratios to determine a molecular response score for the subject. Furthermore, the method also includes: (d) administering one or more therapies to the subject based on at least the molecular response score, thereby treating the cancer in the subject.

[0018] In another aspect, the present disclosure provides a method for treating cancer in a subject. The method includes administering one or more treatments to the subject based on at least the molecular response score of the subject. The molecular response score is calculated by: (a) determining, by a computer, the mutant allele frequencies (MAFs) for multiple variants from sequence information generated from targeted nucleic acids associated with one or more cancer types in samples obtained from the subject at a first time point and a second time point, to obtain a first MAF and a second MAF set for each variant in the multiple variants; (b) calculating, by a computer, the ratio of the first MAF to the second MAF for each variant in the multiple variants, to obtain a set of MAF ratios and corresponding standard deviations for each MAF ratio in the set of MAF ratios; and (c) calculating, by a computer, the weighted average and confidence interval of the MAF ratios to determine the molecular response score for the subject.

[0019] In another aspect, the present disclosure provides a method for identifying clonal hematopoietic variants in a subject with cancer, using at least a computer. The method includes: (a) determining, by a computer, a change in tumor burden (R) relative to a change in tumor fraction P(R) for each of a plurality of variants from sequence information generated from targeted nucleic acids associated with one or more cancer types in samples obtained from the subject at a first time point and a second time point, to produce a set of changes in tumor burden. The method also includes (b) identifying, by a computer, one or more resistance signatures corresponding to one or more clonal hematopoietic variants from the set of changes in tumor burden, thereby identifying clonal hematopoietic variants in the subject with cancer.

[0020] In another aspect, the present disclosure provides a method for identifying clonal hematopoietic variants in a subject with cancer, using at least a computer. The method includes: (a) calculating, by a computer, a probability density function for the change in tumor fraction P(R) for each of a plurality of variants from sequence information generated from targeted nucleic acids associated with one or more cancer types in samples obtained from the subject at a first time point and a second time point. The method also includes: (b) grouping, by a computer, one or more of the variants into one or more clones by P(R); and (c) generating, by a computer, an updated P(R) for each of the clones. Furthermore, the method also includes: (d) identifying, by a computer, one or more clones having a change in fraction at or above a predetermined threshold between the first and second time points, thereby identifying clonal hematopoietic variants in the subject with cancer. In some of these embodiments, the method includes determining the likelihood that a given pair of variants exhibits the same fractional change, merging the most likely pair of variants into a single clone, and updating P(R) for the single clone.

[0021] In another aspect, the present disclosure provides a method for identifying germline variants in a subject with cancer using, at least in part, a computer. The method includes: (a) determining, by a computer, a mutant allele frequency (MAF) for a given variant from sequence information generated from targeted nucleic acids associated with one or more cancer types in a sample obtained from the subject; (b) identifying, by a computer, a given variant as a germline variant if the MAF of the given variant increases the maximum MAF of the sample, the sample contains the maximum fraction of diploid genes (max frac_diploid), and / or if the MAF of the given variant is at least about 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, or greater than one or more other MAFs determined from samples obtained from the subject, thereby identifying a germline variant in the subject with cancer.

[0022] In some embodiments, the methods disclosed herein include comparing a molecular response score for a subject having cancer to a predetermined cutoff point, and identifying the subject as a likely responder to one or more treatments for cancer if the molecular response score is below the predetermined cutoff point, or identifying the subject as a likely non-responder to one or more treatments for cancer if the molecular response score is at or above the predetermined cutoff point. In some embodiments, the one or more treatments include one or more immunotherapies. In some embodiments, the methods disclosed herein include administering one or more treatments for cancer to the subject in consideration of the molecular response score. In some embodiments, the methods disclosed herein include discontinuing administering one or more treatments for cancer to the subject in consideration of the molecular response score. In some embodiments, the methods disclosed herein include recommending one or more treatments. In some embodiments, the methods disclosed herein include recommending discontinuing one or more treatments. In some embodiments, the methods disclosed herein include using the molecular response score as a prognostic and / or predictive biomarker for the subject.

[0023] In some embodiments, the methods disclosed herein include calculating the standard deviation of each MAF ratio in the set of MAF ratios using molecular counting. In some embodiments, the methods disclosed herein include propagating the variance through each MAF ratio in the set of MAF ratios. In some embodiments, the methods disclosed herein include excluding one or more germline and / or clonal hematopoietic variants when determining a mutant allele frequency (MAF) for a plurality of variants. In some embodiments, the plurality of variants includes somatic nucleic acid variants. In some embodiments, the methods disclosed herein include excluding one or more somatic variants having a MAF of less than about 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, or less than 0.9% at both the first and second time points. In some embodiments, the first time point includes a time point before treatment, and the second time point includes a time point during or after treatment.

[0024] In some embodiments, the method disclosed herein comprises generating sequence information from nucleic acid molecules obtained from one or more tissues or cells in a sample.In some embodiments, the method disclosed herein comprises generating sequence information from extracellular free nucleic acid (cfNA) in a sample obtained from a subject.In some embodiments, cfNA comprises circulating tumor DNA (ctDNA).

[0025] In some embodiments, the ratio comprises the second MAF to the first MAF for each variant in the plurality of variants. In some embodiments, the methods disclosed herein comprise calculating a weighted average of the MAF ratios using the formula: Σ[weight × ratio] / Σ[weight], where weight is 1 / range2 for a given variant in the plurality of variants, range is the difference between the value of the first MAF and the value of the second MAF for a given variant in the plurality of variants, and ratio is a given MAF ratio in the set of MAF ratios. In some embodiments, the methods disclosed herein comprise calculating a confidence interval using the formula: weighted average of MAF ratios + / - sqrt[ratio variance], where ratio variance is 1 / Σ[weight].

[0026] In some embodiments, the variants include one or more single nucleotide variants (SNVs), insertion / deletion mutations (indels), gene amplifications, and / or gene fusions. In some embodiments, the methods disclosed herein include determining a molecular response score for a subject with cancer using one or more additional genomic data sources. In some embodiments, the additional genomic data sources include one or more of coverage, off-target coverage, epigenetic signatures, and / or microsatellite instability scores. In some embodiments, the epigenetic signatures include cfNA fragment length, location, and / or endpoint density distribution. In some embodiments, the epigenetic signatures include epigenetic states or statuses represented at one or more epigenetic loci within a given targeted genomic region. In some embodiments, the epigenetic state or status comprises the presence or absence of methylation, hydroxymethylation, acetylation, ubiquitination, phosphorylation, sumoylation, ribosylation, citrullination, and / or histone post-translational modifications or other histone variations.

[0027] The present application discloses methods, computer-readable media, and systems useful for determining a molecular response score for a subject with cancer. Related methods for identifying clonal hematopoietic and / or germline variants are also disclosed. Additional advantages of the disclosed methods, systems, and / or compositions will be set forth in part in the description that follows, and in part will be understood from the description, or may be learned by practice of the disclosed methods and compositions. The advantages of the disclosed methods and compositions will be realized and attained by the elements and combinations particularly pointed out in the appended claims. It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claimed invention.

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the disclosed methods and compositions and, together with the description, serve to explain the principles of the disclosed methods and compositions. [Brief explanation of the drawings]

[0029] [Figure 1] Visualization of a CPH model (utilizing a spline function for MR) to predict whether a patient will have a TTNT event. The model is specified to use the interaction between ctDNA change (using the Guardant molecular response score; MR score = mean treatment-onset VAF / mean baseline VAF) and ctDNA level variables (mean baseline VAF). The plot shows how baseline and ctDNA levels interact with ctDNA change in the CPH regression model to affect response. The plot shows the association between outcome (TTNT) and explanatory variables (ctDNA change interacting with baseline and ctDNA levels), holding other variables constant (weighted comorbidity score = 21, gender female, patient age = 66).

[0030] [Figure 2]Forest plot of the hazard (HR) of death from CPH for molecular responders defined by various thresholds or ctDNA changes (from a 90% reduction at the top to a 20% reduction at the bottom). This figure is for the Dataset 2 ICI cohort, and we selected -60% (60% reduction in ctDNA) as the optimal genomic MR threshold for defining MR responders / non-responders for the ICI-treated patient cohort. All thresholds below were significant for TTNT. The figure shows that the HR for OS was significantly reduced when patients achieved a 50% or greater reduction in ctDNA. In contrast to OS, all thresholds were significant for longer rwTTNT (20% or greater reduction in ctDNA).

[0031] [Figure 3-1]Patients falling into the "responder" and "non-responder" categories based solely on the ctDNA change metric were further subdivided by baseline ctDNA levels (high / low), resulting in four categories that better correlate with both OS and TTNT in real-world clinical practice. Non-responders with high levels of baseline ctDNA consistently had the shortest OS and TTNT in both ICI- and TKI-treated cohorts. A) Patients within the ICI- and TKI-treated cohorts were categorized into four categories based on molecular responder (R) / non-responder (N) and high (H) or low (L) baseline ctDNA levels: responder / low baseline (R / L), responder / high baseline (R / H), non-responder / low baseline (N / L), and non-responder / high baseline (N / H). Patient counts and plots reflect individually optimized thresholds for rwTTNT. B-C) Maximum variant allele fraction (MVAF) between baseline and time points during treatment for each MR / MVAF group within the ICI cohort (B) and TKI cohort (C) is plotted (mean and 95% CI). Kaplan-Meier plots of rwOS (D) and rwTTNT (F) in the ICI cohort and rwOS (E) and rwTTNT (G) in the TKI cohort. [Figure 3-2] Same as above.

[0032] [Figure 4-1]One-year survival probability for rwOS and rwTTNT for the groups in Figure 3. Patients using only the ctDNA change metric can be further subdivided by baseline ctDNA level (high / low). Patients with low baseline ctDNA and declining ctDNA change metrics had a higher percentage of patients with OS and TTNT event-free within 1 year, whereas patients with high baseline ctDNA and increasing ctDNA change metrics had the lowest percentage of patients with OS and TTNT event-free within 1 year. In the ICI cohort (A, B) and TKI cohort (C, D), molecular responders (R, white) and non-responders (N, black), low baseline ctDNA (L, maroon), high baseline ctDNA (H, gray), responder / low baseline (R / L, orange), responder / high baseline (R / H, green), non-responder / low baseline (N / L, red), and non-responder / high baseline (N / H, blue). [Figure 4-2] Same as above. DETAILED DESCRIPTION OF THE INVENTION

[0033] Detailed Description The methods and compositions of the present disclosure may be understood more readily by reference to the following detailed description of specific embodiments and examples included therein, as well as the drawings and accompanying description.

[0034] It is to be understood that the methods and compositions of the present disclosure are not limited to particular synthetic methods, particular analytical techniques, or particular reagents, unless otherwise specified, and as such may vary.

[0035] The present invention relates to compositions and methods for cancer diagnosis, research and treatment, including but not limited to cancer markers. In particular, the present invention relates to predicting the likelihood of patient progression by computer modeling, in which both baseline ctDNA levels and measures of ctDNA changes are incorporated into the same predictive model. Baseline ctDNA levels can be broadly interpreted as a measure of ctDNA at a time point prior to the last time point within a series (two or more) of measurements at various time points. Although this study focuses on ctDNA measurements based on genomic alterations, measures of ctDNA based on DNA methylation can be used in place of genomic alterations, as long as the amount of ctDNA in a sample at a given time point is quantified.

[0036] Previous studies have attempted linear combinations of individual features and their associations with patient outcomes such as PFS and OS. This study differs from previous studies in the successful conception and implementation of a predictive model that incorporates explicit interactions between baseline ctDNA levels and measures of ctDNA change between time points. The significance of ctDNA change within the predictive model may vary depending on the value of the baseline ctDNA level.

[0037] The trained model developed in this study can be used to predict patient outcomes in terms of time to OS event and time to progression event (TTNT) for clinical patients to help clinicians determine the optimal treatment option for a given patient.

[0038] This model allows patients with high risk of OS or progression to switch from their current treatment to an alternative treatment, and the alternative treatment can be a more aggressive treatment regimen, for example, a combination of ICI and chemotherapy instead of ICI monotherapy.Similarly, patients with low risk of OS or progression can switch to a less aggressive treatment option, for example, ICI monotherapy instead of ICI and chemotherapy combination.

[0039] The modeling predictions of the model described in this study can be converted into a report that quantifies the risk of death or progression during cancer patient treatment. There are many different embodiments for how risk can be quantified and communicated from a predictive model. A categorical risk level can be derived from a predictive model by using a threshold value for the predicted time to OS or progression (TTNT) event. For example, patients who are predicted to have OS or TTNT events within the shortest quartile range can be categorized as "high risk", while patients whose predicted time to OS or TTNT events is in the longest quartile can be categorized as "low risk". Clinicians can then directly use this information to optimize the care of patients. Alternatively, the predicted time to OS or TTNT can be provided as an integer score, for example, with a high score associated with low risk. I. Terms and Abbreviations

[0040] It is also understood that the terminology used herein is for the purpose of describing particular embodiments only and is not limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer-readable media, and systems, the following terminology, and grammatical variations thereof, are used in accordance with the definitions set forth below.

[0041] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods, and / or steps of the type described herein and / or that would become apparent to those of skill in the art upon reading this disclosure. Temperatures, concentrations, times, numbers of bases or base pairs, coverage, etc. discussed in this disclosure are implicitly preceded by "about," and therefore, equivalents having slight and insubstantial differences are understood to be within the scope of this disclosure. In this application, the use of the singular includes the plural unless otherwise specified. Additionally, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting.

[0042] About: As used herein, "about" or "approximately," when applied to one or more values ​​or elements of interest, refers to a value or element similar to the stated reference value or element. In certain embodiments, the term "about" or "approximately," unless otherwise specified or otherwise clear from the context, refers to a range of values ​​or elements that are within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or a lower percentage in either direction (greater or less) of the stated reference value or element (except where such number exceeds 100% of the possible value or element).

[0043] CI = confidence interval

[0044] CPH = Cox proportional hazards model

[0045] ctDNA = circulating extracellular free tumor DNA

[0046] G360 = Guardant360, HR = Hazard Ratio

[0047] ICI = immune checkpoint inhibitor

[0048] TKI = tyrosine kinase inhibitor

[0049] rwTTNT = time to next treatment in clinical practice. Time from start of treatment to start of next treatment, death, or loss to follow-up.

[0050] Rw TTD = time to treatment discontinuation. Time from treatment initiation to completion of the treatment regimen, death, or loss to follow-up.

[0051] LOT = Line of Treatment

[0052] AIC = Akaike Information Criterion.

[0053] It will also be understood by those skilled in the art that terms used interchangeably, including "maximum mutant allele frequency," "maximum variant allele frequency," "maximum MAF," "MAX MAF," "maximum VAF," "max-MAF," or "MAX VAF," refer to the largest or greatest MAF of all somatic variants present or observed in a given sample.

[0054] Mutant allele frequency: As used herein, "mutant allele frequency," "variant allele frequency," "mutant allele fraction," "variant allele fraction," "MAF," or "VAF" refers to the frequency of occurrence of a mutant allele in a given population of nucleic acids, such as a sample obtained from a subject. MAF is generally expressed as a fraction or percentage.

[0055] It will also be understood by those skilled in the art that response refers to changes in the frequency, level, or amount of one or more circulating tumor DNA (ctDNA) variant alleles observed between samples obtained from a given subject at different time points.

[0056] Molecular responder: As used herein, "molecular responder" or "responder" refers to a subject having a molecular response score indicating a decrease in one or more circulating tumor DNA (ctDNA) variant allele frequency, level, or amount observed among samples obtained from a subject at various time points. II. Molecular Response Scoring

[0057] In one embodiment, a method for determining a molecular response (MR) score is disclosed. The disclosed method can have a wide variety of uses for the manipulation, preparation, identification, quantification, and / or analysis of extracellularly free nucleic acids. Molecular response is an assessment of the change in circulating tumor DNA (ctDNA) load over the course of treatment (usually 3-10 weeks) compared to a pretreatment baseline. Molecular response correlates with patient response to treatment and long-term outcomes across solid tumors and treatment types. Molecular response can also be used to predict clinical response earlier than radiographic and / or RECIST response. Many methods have been used to calculate molecular response, and there is no consensus as to which method is best.

[0058] A method and system for assessing response to treatment using a molecular response (MR) score is described. In an embodiment, baseline (pre-treatment) gene expression data can be obtained for multiple patients before treatment, and gene expression data during treatment can be obtained for multiple patients during treatment. In an embodiment, the baseline gene expression data (e.g., variant data) and / or the gene expression data during treatment can be analyzed to determine a molecular response (MR) score. The MR score can indicate whether a patient is a responder or non-responder to treatment. In an embodiment, a variant allele fraction (MAF) can be determined as part of the MR score. In an embodiment, the variance of each MAF can be incorporated into the determination of the molecular response score. This ensures that the molecular response score contains accurate variance, resulting in a significant improvement in making correct conclusions from the molecular response score. When the molecular response score is a ratio, this improvement is even more significant because ratios are sensitive to the variance of the denominator. Incorporating variance into the molecular response score can be done either by mathematically deriving the molecular response variance, or by simulating or sampling from the variance distribution of each variant to determine the molecular response variance. a. cfDNA isolation and extraction

[0059] At a first time point T0, baseline cfDNA can be obtained from one or more baseline samples obtained pre-treatment from one or more subjects in step 101, and at a second time point T1, treatment cfDNA can be obtained from one or more treatment samples obtained post-treatment from one or more subjects in step 102. Treatment can be performed / initiated at any time after time point T0. For example, treatment can be performed minutes, hours, days, etc., after time point T0. As a further example, treatment can be performed 30 minutes after time point T0, 1 to 2 hours after time point T0, 1 to 2 days after time point T0, 1 to 2 weeks after time point T0, 1 to 2 months after time point T0, 6 months to 1 year after time point T0, 1 to 2 years after time point T0, etc. The T1 time point can be any amount of time after the T0 time point, such as between 1 hour and 24 hours inclusive, between 1 day and 180 days inclusive, between 1 week and 12 weeks inclusive, between 6 months and 12 months inclusive, etc.

[0060] As described herein, polynucleotides can include any type of nucleic acid, such as DNA and / or RNA. For example, if polynucleotides are DNA, they can be genomic DNA, complementary DNA (cDNA), or any other deoxyribonucleic acid. Polynucleotides can also be extracellular free nucleic acids, such as extracellular free DNA (cfDNA). For example, polynucleotides can be circulating cfDNA. Circulating cfDNA can include DNA released from body cells by apoptosis or necrosis. cfDNA released by apoptosis or necrosis can originate from normal (e.g., healthy) body cells. When there is abnormal tissue growth, such as in cancer, tumor DNA can be released. Circulating cfDNA can include circulating tumor DNA (ctDNA). i. Sample

[0061] Isolation and extraction of extracellularly free polynucleotides can be performed by sampling using various techniques. The sample can be any biological sample isolated from a subject. Samples can include body tissue, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies (e.g., biopsies of known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid (e.g., fluid from the intercellular spaces), gingival exudate, gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Preferably, the sample is a bodily fluid, particularly blood and its fractions, and urine. Such samples contain nucleic acids excreted from tumors. Nucleic acids can include DNA and RNA, and can be in double-stranded or single-stranded forms. Sample can be the form that is originally separated from subject, or can be further processed to remove or add components such as cell, enrich one component for another, or to convert one form of nucleic acid into another form, for example, convert RNA to DNA, or convert single-stranded nucleic acid into double-stranded.Therefore, for example, the body fluid sample for analysis is the plasma or serum that contains extracellular free nucleic acid, for example, extracellular free DNA (cfDNA).

[0062] In some embodiments, the sample volume of bodily fluid obtained from a subject depends on the desired read depth of the region to be sequenced. Exemplary volumes are about 0.4 to 40 ml, about 5 to 20 ml, or about 10 to 20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml, or more milliliters. The volume of sampled blood is typically about 5 ml to about 20 ml.

[0063] Samples can contain varying amounts of nucleic acid. Typically, the amount of nucleic acid in a given sample is equivalent to multiple genome equivalents. For example, a sample of about 30 ng of DNA will contain about 10,000 (10 4 ) haploid human genome equivalents, and in the case of cfDNA, approximately 200 billion (2 x 10 11Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, or in the case of cfDNA, about 600 billion individual molecules.

[0064] In some embodiments, the sample contains nucleic acid from different sources, for example, nucleic acid from cells and nucleic acid from extracellular free sources (such as, for example, blood samples). Typically, the sample contains nucleic acid carrying mutations. For example, the sample optionally contains DNA carrying germline mutations and / or somatic mutations. Typically, the sample contains DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations). In some embodiments of the present disclosure, the extracellular free nucleic acid in a subject may be derived from a tumor. For example, the extracellular free DNA isolated from a subject may include ctDNA.

[0065] Exemplary amounts of extracellularly free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), for example, from about 1 picogram (pg) to about 200 nanograms (ng), from about 1 ng to about 100 ng, or from about 10 ng to about 1000 ng. In some embodiments, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of extracellularly free nucleic acid molecules. Optionally, the amount of extracellularly free nucleic acid molecules is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng. In certain embodiments, the amount of extracellularly free nucleic acid molecules is at most about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng. In some embodiments, the method comprises obtaining between about 1 fg and about 200 ng of extracellularly free nucleic acid molecules from the sample.

[0066] Extracellularly free nucleic acids typically have a size distribution between about 100 and about 500 nucleotides in length, with molecules between about 110 and about 230 nucleotides in length accounting for about 90% of the molecules in a sample, a mode at about 168 nucleotides in length, and a second minor peak in length ranging from about 240 to about 440 nucleotides in length. In certain embodiments, the extracellularly free nucleic acids are between about 160 and about 180 nucleotides in length, or between about 320 and about 360 nucleotides in length, or between about 440 and about 480 nucleotides in length.

[0067] In some embodiments, extracellular-free nucleic acids are isolated from bodily fluids by a partitioning step that separates the extracellular-free nucleic acids found in solution from intact cells and other insoluble components of the bodily fluid. In some of these embodiments, partitioning involves techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid are lysed, and the extracellular-free nucleic acids and intracellular nucleic acids are processed together. Typically, after the addition of a buffer and a washing step, the extracellular-free nucleic acids are precipitated, for example, with alcohol. In certain embodiments, an additional clarification step, such as a silica-based column, is used to remove contaminants or salts. To optimize certain aspects of the exemplary procedure, such as yield, for example, a large amount of nonspecific carrier nucleic acid is optionally added throughout the reaction. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If necessary, the single-stranded DNA and / or single-stranded RNA is converted to a double-stranded form for inclusion in subsequent processing and analysis steps. Additional details regarding cfDNA partitioning and analysis of associated epigenetic modifications, as optionally adapted for use in practicing the methods disclosed herein, are described, for example, in WO2018 / 119452, filed December 22, 2017, which is incorporated by reference. ii. Nucleic acid tags

[0068] In certain embodiments, tags that provide molecular identifiers or barcodes are incorporated or otherwise joined to the adapters by chemical synthesis, ligation, or other methods of overlap extension PCR. In some embodiments, the assignment of unique or non-unique identifiers or molecular barcodes in the reaction follows the methods and utilizes the systems described therein, for example, in U.S. Patent Application Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731, each of which is incorporated by reference.

[0069] The linking (e.g., ligation) of the tag to the sample nucleic acid is performed randomly or non-randomly. In some embodiments, the tags are introduced into the microwells at the expected ratio of identifiers (e.g., combinations of unique and / or non-unique barcodes). For example, the identifiers can be loaded so that about more than 1, more than 2, more than 3, more than 4, more than 5, more than 6, more than 7, more than 8, more than 9, more than 10, more than 20, more than 50, more than 100, more than 500, more than 1000, more than 5000, more than 10000, more than 50,000, more than 100,000, more than 500,000, more than 1,000,000, more than 10,000,000, more than 50,000,000, or more than 1,000,000,000 identifiers are loaded per genome sample. In some embodiments, the identifiers are loaded such that less than about 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 20, less than 50, less than 100, less than 500, less than 1000, less than 5000, less than 10000, less than 50,000, less than 100,000, less than 500,000, less than 1,000,000, less than 10,000,000, less than 50,000,000 or less than 1,000,000,000 identifiers are loaded per genomic sample.In certain embodiments, the average number of identifiers loaded per genomic sample is about less than or more than 1, less than or more than 2, less than or more than 3, less than or more than 4, less than or more than 5, less than or more than 6, less than or more than 7, less than or more than 8, less than or more than 9, less than or more than 10, less than or more than 20, less than or more than 50, less than or more than 100, less than or more than 500, less than or more than 1000, Identifiers may be greater than or equal to 1,000, less than or equal to 5,000, less than or equal to 10,000, less than or equal to 50,000, less than or equal to 100,000, less than or equal to 500,000, less than or equal to 100,000, less than or equal to 500,000, less than or equal to 1,000,000, less than or equal to 10,000,000, less than or equal to 50,000,000, or less than or equal to 1,000,000,000. Identifiers are generally unique or non-unique.

[0070] One exemplary format uses about 2 to about 1,000,000 different tags, or about 5 to about 150 different tags, or about 20 to about 50 different tags ligated to both ends of a target nucleic acid molecule. For 20-50 x 20-50 tags, a total of 400-2500 tags are created. Such a number of tags is typically sufficient to ensure a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) that different molecules with the same start and end points will receive different combinations of tags.

[0071] In some embodiments, the identifier is an oligonucleotide of a predetermined, random, or semi-random sequence. In other embodiments, multiple barcodes can be used, and the barcodes in the multiple barcodes are not necessarily unique to each other. In these embodiments, barcodes are generally attached to individual molecules (e.g., by ligation or PCR amplification) such that the combination of the barcode and the sequence to which it can be attached creates a unique sequence that can be individually tracked. As described herein, detecting a combination of non-uniquely tagged barcodes and sequence data at the beginning (start) and end (end) of a sequence read typically allows for the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequence read is also optionally used to assign a unique identity to a given molecule. As described herein, fragments from a single strand of nucleic acid that have been assigned a unique identity can subsequently identify fragments from the parental and / or complementary strands. iii. Nucleic acid amplification

[0072] The sample nucleic acid flanked by adapters is typically amplified by PCR and other amplification methods using nucleic acid primers that bind to the primer binding sites of the adapters located on both sides of the DNA molecule to be amplified.In some embodiments, the amplification method involves cycles of extension, denaturation and annealing that are brought about by thermocycling, or can be isothermal, for example, in transcription amplification methods.Other exemplary amplification methods that can be used as needed include, among other techniques, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication.

[0073] To introduce a sample index / tag into a nucleic acid molecule using conventional nucleic acid amplification methods, one or more rounds of amplification cycles are generally applied. Amplification is typically carried out in one or more reaction mixtures. In some embodiments, molecular tags and sample index / tags are introduced before and / or after the sequence capture step. In some embodiments, only molecular tags are introduced before probe capture, and sample index / tags are introduced after the sequence capture step. In certain embodiments, both molecular tags and sample index / tags are introduced before the probe-based capture step. In some embodiments, sample index / tags are introduced after the sequence capture step (i.e., nucleic acid enrichment). Typically, sequence capture protocols involve introducing a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, such as a coding sequence in a genomic region, and a mutation in such a region associated with a type of cancer. Typically, the amplification reaction produces a plurality of non-uniquely or uniquely tagged nucleic acid amplicons with molecular tags and sample indexes / tags, with sizes ranging from about 200 nucleotides (nt) to about 700nt, from 250nt to about 350nt, or from about 320nt to about 550nt.In some embodiments, the size of the amplicon is about 300nt.In some embodiments, the size of the amplicon is about 500nt. iv. Nucleic acid enrichment

[0074] In some embodiments, sequences are enriched before nucleic acid sequencing. Enrichment can be performed for specific target regions or non-specifically ("target sequences") as needed. For example, enrichment can be performed non-specifically based on a size selection method that is not sequence-specific but sequence fragment size-specific. In some embodiments, nucleic acid capture probes ("baits") selected for one or more bait set panels can be used to enrich targeted regions of interest using a discriminatory tiling and capture scheme. In discriminatory tiling and capture schemes, bait sets with different relative concentrations are typically used to differentially tile (e.g., at different "resolutions") across the genome sections to which the baits are attached, subject to a set of constraints (e.g., sequencer constraints, e.g., sequencing load, availability of each bait, etc.), and the targeted nucleic acids are captured at a desired level for downstream sequencing. These targeted genome sections of interest optionally contain natural or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotin-labeled beads and probes to one or more sections of interest can be used to capture target sequences, and optionally, the sections can then be amplified to enrich for the region of interest.

[0075] Sequence capture typically involves the use of oligonucleotide probes that hybridize with target nucleic acid sequences.In certain embodiments, probe set strategy involves tiling probes across the target section.Such probes can be, for example, about 60 nucleotides to about 120 nucleotides in length.The depth of the set can be about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x or more.The effectiveness of sequence capture generally depends in part on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe. b. Nucleic acid sequencing

[0076] After cfDNA extraction and isolation from the sample, the cfDNA can be sequenced. Generally, the sample nucleic acid, optionally flanked by adapters, is subjected to sequencing, with or without prior amplification. Optionally, the sequencing method or commercially available format may include, for example, Sanger sequencing, high-throughput sequencing, bisulfite sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, primer walking, PacBio, SOLiD, Ion Torrent, or sequencing using a nanopore platform.Sequencing reactions can be carried out in a variety of sample processing devices, which may include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sample sets substantially simultaneously. The sample processing device may also include multiple sample chambers, allowing multiple runs to be processed simultaneously.

[0077] Sequencing reaction can be performed on one or more nucleic acid fragment types or sections that are known to contain markers for cancer or other diseases.Sequencing reaction can also be performed on any nucleic acid fragment present in a sample.The sequence coverage of genome that is obtained by sequencing reaction can be at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of genome.In other cases, the sequence coverage of genome can be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of genome.

[0078] Simultaneous sequencing reaction can be carried out using multiplex sequencing technology.In some embodiments, the sequencing of extracellular free polynucleotide is carried out at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reaction.In other embodiments, the sequencing of extracellular free polynucleotide is carried out at less than about 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50000 or less than 100,000 sequencing reaction.Sequencing reaction is typically carried out sequentially or simultaneously.Subsequent data analysis is generally carried out on all or part of sequencing reaction. In some embodiments, data analysis is carried out for at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions.In other embodiments, data analysis can be carried out for less than about 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50000 or less than 100,000 sequencing reactions.Exemplary read depths are about 1000 to about 50,000 reads per locus (base position).

[0079] In some embodiments, a nucleic acid population is prepared for sequencing by enzymatically forming blunt ends on double-stranded nucleic acids with single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Exemplary enzymes or catalytic fragments thereof that can be used as needed include Klenow large fragment and T4 polymerase. In the 5' overhang, the enzyme typically extends the recessed 3' end of the opposite strand until it overlaps with the 5' end, creating a blunt end. In the 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposite strand, and sometimes beyond. If this digestion proceeds beyond the 5' end of the opposite strand, the gap can be filled with an enzyme with the same polymerase activity as used for the 5' overhang. Creating blunt ends on double-stranded nucleic acids facilitates, for example, the attachment of adapters and subsequent amplification.

[0080] In some embodiments, the population of nucleic acids is subjected to additional processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA. These forms of nucleic acids are also optionally ligated with adapters and amplified.

[0081] The nucleic acid that is subjected to the above-mentioned blunt-end forming process, with or without prior amplification, and other nucleic acids in sample as needed, can be sequenced to obtain sequenced nucleic acid.Sequenced nucleic acid can refer to either the sequence of nucleic acid (i.e., sequence information) or the nucleic acid that has been sequenced.Sequencing can be carried out so that the sequence data of each nucleic acid molecule in sample is directly or indirectly obtained from the consensus sequence of the amplification product of each nucleic acid molecule in sample.

[0082] In some embodiments, adapters containing barcodes are ligated to both ends of double-stranded nucleic acids with single-stranded overhangs in the sample after blunt end formation, and the nucleic acid sequence and the in-line barcode incorporated by the adapter are determined by sequencing. Blunt-ended DNA molecules are optionally ligated to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Alternatively, complementary nucleotide tails can be added to the blunt ends of the sample nucleic acid and adapter to facilitate ligation (e.g., sticky end ligation).

[0083] Typically, a nucleic acid sample is contacted with a sufficient number of adapters so that the probability that any two copies of the same nucleic acid will receive the same combination of adapter barcodes from adapters ligated to both ends is low (e.g., less than 1% or less than 0.1%). This use of adapters allows for the identification of families of nucleic acid sequences that have the same start and end points in the reference nucleic acid and are ligated with the same combination of barcodes. Such families represent the sequences of the amplification products of the template / parent nucleic acid in the sample before amplification. The sequences of family members can be compiled to derive consensus nucleotide(s) or the complete consensus sequence of the nucleic acid molecules in the original sample modified by blunt-end formation and adapter attachment. In other words, the nucleotide occupying a particular position in the nucleic acid in the sample is determined to be the consensus of the nucleotide occupying the corresponding position in the family member sequences. A family can include sequences from one or both strands of a double-stranded nucleic acid. When a family member comprises the sequence of both strands of double-stranded nucleic acid, all sequences are compiled, and the sequence of one strand is converted into its complement to derive consensus nucleotide(s) or sequence.Some families only comprise a single member sequence.In this case, this sequence can be considered as the sequence of the nucleic acid in the sample before amplification.Alternatively, families that only have a single member sequence can be excluded from subsequent analysis.

[0084] The nucleotide variation in sequenced nucleic acid can be determined by comparing the sequenced nucleic acid with a reference sequence.The reference sequence is often a known sequence, for example, a known whole genome sequence or partial genome sequence from a subject (for example, the whole genome sequence of a human subject).The reference sequence can be, for example, hG19 or hG38.The sequenced nucleic acid can represent the sequence determined directly for the nucleic acid in the sample, or the consensus of the sequence of the amplification product of such nucleic acid as described above.Comparison can be performed with respect to one or more designated positions on the reference sequence.When each sequence is maximally aligned, a subset of sequenced nucleic acid can be identified that includes a position corresponding to the designated position of the reference sequence. Within such a subset, it can determine which sequenced nucleic acids, if any, contain nucleotide variations at designated positions; the length of a given cfDNA fragment based on where the end points (i.e., the 5'-end and 3'-end nucleotides) map to in the reference sequence; the offset of the midpoint of the given cfDNA fragment from the midpoint of the genomic region within the cfDNA fragment; and, if necessary, which contain reference nucleotides (i.e., the same nucleotide as in the reference sequence).If the number of sequenced nucleic acids containing nucleotide variants in a subset exceeds a selected threshold, it can call a variant nucleotide at that designated position.The threshold can be, among other possibilities, a simple number, for example, at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 sequenced nucleic acids in the subset that contain nucleotide variants, or a ratio, for example, at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20 of the sequenced nucleic acids in the subset that contain nucleotide variants.Comparison can be repeated for any designated position of interest within the reference sequence. Sometimes, comparison can be performed for designated positions occupying at least about 20, 100, 200, or 300 contiguous positions on the reference sequence, for example, about 20-500, or about 50-300 contiguous positions.

[0085] Additional details regarding nucleic acid sequencing, including the formats and applications described herein, can be found in, for example, Levy et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364: 1-11 (2012); Voelkerding et al., Clinical Chem., 55: 641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009); Astier et al., J Am Chem Soc., 128 (5): 1705-1000 (2009), each of which is incorporated by reference in its entirety. (2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503. i. Sequencing Panel

[0086] To improve the likelihood of detecting genomic regions of interest and, if necessary, tumors that exhibit mutations, the DNA sections to be sequenced can include a panel of genes or genome sections that include known genomic regions. Selecting limited sections for sequencing (e.g., a limited panel) can reduce the total amount of sequencing required (e.g., the total amount of nucleotides sequenced). A sequencing panel can target multiple different genes or regions, for example, to detect a single cancer, a set of cancers, or all cancers. Alternatively, DNA sequencing can be performed by whole genome sequencing (WGS) or other unbiased sequencing methods without using a sequencing panel. Examples of panels and targets suitable for use in panels can be found in the epigenetic targets described in U.S. Provisional Patent Application No. 62 / 799,637, filed January 31, 2019, the entire contents of which are incorporated by reference.

[0087] In some embodiments, a panel is selected that targets multiple different genes or genomic regions (e.g., transcription factor binding regions, distal regulatory elements (DREs), repetitive elements, intron-exon junctions, transcription start sites (TSSs), and / or the like), such that a predetermined percentage of subjects with cancer exhibit genetic variants or tumor markers for one or more different genes in the panel. The panel can be selected to limit the region so that a fixed number of base pairs are sequenced. The panel can be selected to sequence a desired amount of DNA. Furthermore, the panel can be selected to achieve a desired sequence read depth. The panel can be selected to achieve a desired sequence read depth or sequence read coverage for a certain amount of sequenced base pairs. The panel can be selected to achieve a theoretical sensitivity, specificity, and / or accuracy for detecting one or more genetic variants in a sample.

[0088] The probes for detecting the panel of regions may include probes for detecting genomic regions of interest (hotspot regions) and nucleosome recognition probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of fragment size variations and GC sequence composition, which are influenced by cfDNA coverage and nucleosome binding patterns. As used herein, the term "region" may also include non-hotspot regions that are optimized based on nucleosome positions and GC models. A panel may include multiple subpanels, including a subpanel for identifying tissue of origin (e.g., using published literature to define 50-100 baits representing genes (not necessarily promoters) with the most diverse transcriptional profiles across tissues), a subpanel for identifying whole-genome scaffolds (e.g., to identify ultraconserved genomic content and sparsely tiling across chromosomes using a small number of probes for copy number baseline setting), and a subpanel for identifying transcription start sites (TSSs) / CpG islands (e.g., to capture differentially methylated regions (DMRs) in promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer). In some embodiments, the tissue of origin marker is a tissue-specific epigenetic marker.

[0089] Some example lists of genomic locations of interest can be found in Table 1 and Table 2. In some embodiments, the genomic locations used in the methods of the present disclosure include at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or 97 of the genes in Table 1. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the genes in Table 1. In some embodiments, the genomic locations used in the methods of the present disclosure include at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 1. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the SNVs in Table 1. In some embodiments, the genomic locations used in the methods of the present disclosure include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the CNVs in Table 1. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the CNVs in Table 1. In some embodiments, the genomic locations used in the methods of the present disclosure include at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 1. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the fusions in Table 1. In some embodiments, the genomic locations used in the methods of the present disclosure include at least a portion of at least one, at least two, or three of the indels in Table 1. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the indels in Table 1.In an embodiment, the genomic locations used in the methods of the present disclosure include all of the genes, SNVs, CNVs, fusions, and indels in Table 1. In some embodiments, the genomic locations used in the methods of the present disclosure include at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, or 115 of the genes in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the genes in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the genes in Table 1 and Table 2. In some embodiments, the genomic locations used in the methods of the present disclosure include at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the SNVs in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the SNVs in Table 1 and Table 2. In some embodiments, the genomic locations used in the methods of the present disclosure include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the CNVs in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the CNVs in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all of the CNVs in Table 1 and Table 2.In some embodiments, a genomic location used in a method of the present disclosure comprises at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 2. In an embodiment, a genomic location used in a method of the present disclosure comprises all of the fusions in Table 2. In an embodiment, a genomic location used in a method of the present disclosure comprises all of the fusions in Table 1 and Table 2. In some embodiments, a genomic location used in a method of the present disclosure comprises at least a portion of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 2. In an embodiment, a genomic location used in a method of the present disclosure comprises all of the indels in Table 2. In an embodiment, a genomic location used in a method of the present disclosure comprises all of the indels in Table 1 and Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all genes, SNVs, CNVs, fusions, and indels in Table 2. In an embodiment, the genomic locations used in the methods of the present disclosure include all genes, SNVs, CNVs, fusions, and indels in Tables 1 and 2. For a given bait set panel, each of these genomic locations of interest can be identified as a scaffold region or a hotspot region. Table 1. SNVs, CNVs, fusions and indels [Table 1-1] [Table 1-2] Table 2. SNVs, CNVs, fusions and indels [Table 2-1] [Table 2-2]

[0090] In some embodiments, one or more regions in the panel include one or more loci from one or more genes for detecting residual cancer after surgery. This detection can occur earlier than is possible with existing cancer detection methods. In some embodiments, one or more genomic locations in the panel include one or more loci from one or more genes for detecting cancer in high-risk patient populations. For example, the rate of lung cancer in smokers is much higher than in the general population. In addition, smokers may experience other lung conditions that make cancer detection more difficult, such as the development of irregular nodules in the lungs. In some embodiments, the methods described herein detect a patient's response to cancer treatment (particularly in high-risk patients) earlier than is possible with existing cancer detection methods.

[0091] A genomic location can be selected for inclusion in sequencing panel based on the number of subjects with cancer that have tumor markers in that gene or region.A genomic location can be selected for inclusion in sequencing panel based on the prevalence of subjects with cancer and the tumor markers present in that gene.The presence of tumor markers in a region can indicate subjects with cancer.

[0092] In some cases, a panel can be selected using information from one or more databases. Information about cancer can be derived from cancer tumor biopsies or cfDNA assays. The database can include information describing a population of sequenced tumor samples. The database can include information about mRNA expression in tumor samples. The database can include information about regulatory elements or genomic regions in tumor samples. The information about sequenced tumor samples can include the frequency of various genetic variants, describing the genes or regions in which the genetic variants are present. The genetic variants can be tumor markers. A non-limiting example of such a database is COSMIC. COSMIC is a catalog of somatic mutations found in various cancers. For a particular cancer, COSMIC ranks genes based on the frequency of mutations. Genes can be selected for inclusion in a panel based on the frequency of mutations in a given gene. For example, COSMIC shows that 33% of the population of sequenced breast cancer samples have TP53 mutations, and 22% of the population of sampled breast cancers have KRAS mutations. Other ranked genes, including APC, have mutations found in only about 4% of sequenced breast cancer samples. TP53 and KRAS can be included in a sequencing panel based on their relatively high occurrence in sampled breast cancers (e.g., compared to APC, which is present at a frequency of about 4%). However, COSMIC is provided as a non-limiting example, and any database or set of information linking cancers to tumor markers located in genes or gene regions can be used. In another example, as provided by COSMIC, of ​​1,156 biliary tract cancer samples, 380 samples (33%) carried mutations in TP53. Some other genes, such as APC, have mutations in 4-8% of all samples. Therefore, TP53 can be selected for inclusion in a panel based on its relatively high occurrence in a population of biliary tract cancer samples.

[0093] Genes or genome segments can be selected for a panel if the frequency of tumor markers in sampled tumor tissues or circulating tumor DNA is significantly higher than that found in a given background population. A combination of genome locations can be selected for inclusion in a panel so that at least the majority of subjects with cancer can have tumor markers or genome regions present in at least one of the genome locations or genes in the panel. A combination of genome locations can be selected based on data showing that for a particular cancer or set of cancers, the majority of subjects have one or more tumor markers in one or more of the selected regions. For example, to detect cancer 1, a panel including regions A, B, C, and / or D can be selected based on data showing that 90% of subjects with cancer 1 have tumor markers in regions A, B, C, and / or D of the panel. Alternatively, tumor markers may be shown to be present independently in two or more regions in subjects with cancer, and therefore, when combined, the tumor markers in two or more regions will be present in the majority of subjects with cancer. For example, to detect cancer 2, a panel including regions X, Y, and Z can be selected based on data showing that 90% of subjects have tumor markers in one or more regions, that in 30% of such subjects the tumor marker is detected only in region X, and that in the remaining subjects in which the tumor marker is detected, the tumor marker is detected only in regions Y and / or Z. A tumor marker present at one or more genomic locations previously shown to be associated with one or more cancers can indicate or predict the presence of the target cancer if the tumor marker is detected 50% or more often in one or more of those regions. Computational methods, such as models using the conditional probability of detecting cancer given the cancer frequency of a set of tumor markers in one or more regions, can be used to predict which regions, alone or in combination, may be predictive of cancer.Other approaches to panel selection involve the use of databases containing information from studies using large panels and / or comprehensive genomic profiling of tumors using whole-genome sequencing (WGS, RNA-seq, Chip-seq, bisulfite sequencing, ATAC-seq, and others). Information gleaned from the literature may also describe pathways that are commonly affected and mutated in a particular cancer. The use of genetic ontologies can further inform panel selection.

[0094] The gene included in the panel for sequencing can include the complete transcribed region, promoter region, enhancer region, regulatory element, and / or downstream sequence.To further increase the possibility of detecting tumors that show mutations, only exons can be included in the panel.The panel can include all exons of the selected gene, or can include only one or more exons of the selected gene.The panel can include exons from each of multiple different genes.The panel can include at least one exon from each of multiple different genes.

[0095] In some embodiments, a panel of exons from each of a plurality of different genes is selected such that a predetermined percentage of subjects with cancer exhibit a genetic variant in at least one exon in the panel of exons.

[0096] At least one complete exon from each different gene in a panel of genes can be sequenced. The panel to be sequenced can include exons from multiple genes. The panel can include exons from 2-100 different genes, 2-70 genes, 2-50 genes, 2-30 genes, 2-15 genes, or 2-10 genes.

[0097] The selected panel may contain a varying number of exons. The panel may contain 2-3,000 exons. The panel may contain 2-1,000 exons. The panel may contain 2-500 exons. The panel may contain 2-100 exons. The panel may contain 2-50 exons. The panel may contain 300 or fewer exons. The panel may contain 200 or fewer exons. The panel may contain 100 or fewer exons. The panel may contain 50 or fewer exons. The panel may contain 40 or fewer exons. The panel may contain 30 or fewer exons. The panel may contain 25 or fewer exons. The panel may contain 20 or fewer exons. The panel may contain 15 or fewer exons. The panel may contain 10 or fewer exons. The panel may contain 9 or fewer exons. The panel may contain 8 or fewer exons. The panel may contain 7 or fewer exons.

[0098] A panel may include one or more exons from a plurality of different genes. A panel may include one or more exons from each of a percentage of the plurality of different genes. A panel may include at least two exons from each of at least 25%, 50%, 75%, or 90% of the different genes. A panel may include at least three exons from each of at least 25%, 50%, 75%, or 90% of the different genes. A panel may include at least four exons from each of at least 25%, 50%, 75%, or 90% of the different genes.

[0099] The size of a sequencing panel can vary. A sequencing panel can be larger or smaller (in terms of nucleotide size) depending on several factors, including, for example, the total amount of nucleotides to be sequenced or the number of unique molecules to be sequenced for a particular region within the panel. A sequencing panel can be 5 kb to 50 kb in size. A sequencing panel can be 10 kb to 30 kb in size. A sequencing panel can be 12 kb to 20 kb in size. A sequencing panel can be 12 kb to 60 kb in size. Sequencing panels can be at least 10 kb, 12 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb in size. Sequencing panels can be less than 100 kb, 90 kb, 80 kb, 70 kb, 60 kb, or 50 kb in size.

[0100] The panel selected for sequencing can include at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 genomic locations (e.g., each containing a genomic region of interest).In some cases, the genomic locations in the panel are selected so that the size of the location is relatively small.In some cases, the size of the region in the panel is about 10kb or less, about 8kb or less, about 6kb or less, about 5kb or less, about 4kb or less, about 3kb or less, about 2.5kb or less, about 2kb or less, about 1.5kb or less, or about 1kb or less. In some cases, the size of the genomic locations in the panel is from about 0.5 kb to about 10 kb, from about 0.5 kb to about 6 kb, from about 1 kb to about 11 kb, from about 1 kb to about 15 kb, from about 1 kb to about 20 kb, from about 0.1 kb to about 10 kb, or from about 0.2 kb to about 1 kb. For example, the size of the regions in the panel can be from about 0.1 kb to about 5 kb.

[0101] The panel selected herein may enable deep sequencing sufficient to detect genetic variants with low frequency (e.g., in extracellularly free nucleic acid molecules obtained from a sample). The amount of genetic variants in a sample may be referred to in terms of the mutant allele frequency for a given genetic variant. The mutant allele frequency may refer to the frequency of a mutant allele (e.g., one that is not the most common allele) present in a given group of nucleic acids, such as a sample. Genetic variants with low mutant allele frequency may be present at a relatively low frequency in a sample. In some cases, the panel may enable the detection of genetic variants with a mutant allele frequency of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. The panel may enable the detection of genetic variants with a mutant allele frequency of 0.001% or higher. The panel may enable the detection of genetic variants with a mutant allele frequency of 0.01% or higher. A panel may enable detection of genetic variants present in a sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. A panel may enable detection of tumor markers present in a sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. A panel may enable detection of tumor markers present in a sample at frequencies as low as 1.0%. A panel may enable detection of tumor markers present in a sample at frequencies as low as 0.75%. A panel may make it possible to detect tumor markers occurring as low as 0.5% of a sample. A panel may make it possible to detect tumor markers occurring as low as 0.25% of a sample. A panel may make it possible to detect tumor markers occurring as low as 0.1% of a sample. A panel may make it possible to detect tumor markers occurring as low as 0.075% of a sample.The panel may enable detection of tumor markers occurring at frequencies as low as 0.05% in a sample. The panel may enable detection of tumor markers occurring at frequencies as low as 0.025% in a sample. The panel may enable detection of tumor markers occurring at frequencies as low as 0.01% in a sample. The panel may enable detection of tumor markers occurring at frequencies as low as 0.005% in a sample. The panel may enable detection of tumor markers occurring at frequencies as low as 0.001% in a sample. The panel may enable detection of tumor markers occurring at frequencies as low as 0.0001% in a sample. The panel may enable detection of tumor markers occurring at frequencies as low as 1.0% to 0.0001% in sequenced cfDNA. The panel may enable detection of tumor markers occurring at frequencies as low as 0.01% to 0.0001% in sequenced cfDNA.

[0102] Genetic variants can be represented by percentage relative to the population of subjects with disease (for example, cancer).In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99% of the population with cancer show one or more genetic variants in at least one of the regions in panel.For example, at least 80% of the population with cancer can show one or more genetic variants in at least one of the genome positions in panel.

[0103] A panel may include one or more locations comprising a genomic region of interest from each of one or more genes. In some cases, a panel may include one or more locations comprising a genomic region of interest from each of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, a panel may include one or more locations comprising a genomic region of interest from each of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, a panel may include one or more locations comprising a genomic region of interest from each of about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.

[0104] The position comprising genome region in panel can be selected to detect one or more epigenetic modified regions.One or more epigenetic modified regions can be acetylated, methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated and / or citrullinated.For example, the region in panel can be selected to detect one or more methylated regions.

[0105] The region in panel can be selected to comprise the sequence that is differentially transcribed across one or more tissues.In some cases, the position that comprises genomic region can comprise the sequence that is transcribed at a high level in a certain tissue compared with other tissues.For example, the position that comprises genomic region can comprise the sequence that is transcribed in a certain tissue but not in other tissues.

[0106] The genome position in the panel may comprise coding sequence and / or non-coding sequence.For example, the genome position in the panel may comprise one or more sequences in exon, intron, promoter, 3' untranslated region, 5' untranslated region, regulatory element, transcription start site, and / or splice site.In some cases, the region in the panel may comprise other non-coding sequences, including pseudogene, repeat sequence, transposon, viral element, and telomere.In some cases, the genome position in the panel may comprise the sequence in non-coding RNA, for example, ribosomal RNA, transfer RNA, piwi interacting RNA, and microRNA.

[0107] Genomic locations within the panel can be selected such that cancer is detected (diagnosed) with a desired level of sensitivity (e.g., by detecting one or more genetic variants). For example, regions within the panel can be selected such that cancer (e.g., by detecting one or more genetic variants) is detected with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sensitivity. Genomic locations within the panel can be selected such that cancer is detected with 100% sensitivity.

[0108] The genomic locations within the panel can be selected to detect (diagnose) cancer with a desired level of specificity (e.g., by detecting one or more genetic variants). For example, the genomic locations within the panel can be selected to detect cancer with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity (e.g., by detecting one or more genetic variants). The genomic locations within the panel can be selected to detect one or more genetic variants with 100% specificity.

[0109] Genomic locations within the panel can be selected to detect (diagnose) cancer with a desired positive predictive value. Positive predictive value can be increased by increasing sensitivity (e.g., the likelihood of detecting an actual positive) and / or specificity (e.g., the likelihood of not mistaking an actual negative for a positive). As a non-limiting example, genomic locations within the panel can be selected to detect one or more genetic variants with a positive predictive value of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within the panel can be selected to detect one or more genetic variants with a positive predictive value of 100%.

[0110] The genomic positions within the panel can be selected so that cancer is detected (diagnosed) with a desired accuracy. As used herein, the term "accuracy" can refer to the ability of a test to distinguish between a disease state (e.g., cancer) and a healthy state. Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under the ROC curve, Youden index and / or diagnostic odds ratio.

[0111] Accuracy can be presented as a percentage, referring to the ratio of the number of tests that yield correct results to the total number of tests performed. Regions within the panel can be selected to detect cancer with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Genomic locations within the panel can be selected to detect cancer with 100% accuracy.

[0112] The panel can be selected to have high sensitivity and detect genetic variants with low frequency.For example, the panel can be selected to detect genetic variants or tumor markers that exist in samples with a frequency of 0.01%, 0.05%, or 0.001% with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sensitivity.The genome position in the panel can be selected to detect tumor markers that exist in samples with a frequency of 1% or less with a sensitivity of 70% or higher. A panel can be selected to detect tumor markers occurring in as little as 0.1% of samples with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel can be selected to detect tumor markers occurring in as little as 0.01% of samples with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers with a frequency as low as 0.001% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0113] The panel can be selected to have high specificity and detect genetic variants with low frequency.For example, the panel can be selected to detect genetic variants or tumor markers that exist in samples with a frequency of 0.01%, 0.05% or 0.001% at least with a specificity of 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9%.The genome position in the panel can be selected to detect tumor markers that exist in samples with a frequency of 1% or less with a specificity of 70% or higher. A panel can be selected to detect tumor markers occurring in as little as 0.1% of samples with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel can be selected to detect tumor markers occurring in as little as 0.01% of samples with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers with a frequency as low as 0.001% in the sample with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0114] The panel can be selected to have high accuracy and detect rare genetic variants. The panel can be selected to detect genetic variants or tumor markers present in samples at frequencies as low as 0.01%, 0.05%, or 0.001% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The genomic locations within the panel can be selected to detect tumor markers present in samples at frequencies of 1% or less with 70% or higher accuracy. The panel can be selected to detect tumor markers present in samples at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. A panel can be selected to detect tumor markers occurring as low as 0.01% of samples with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. A panel can be selected to detect tumor markers occurring as low as 0.001% of samples with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy.

[0115] The panel can be selected to be highly predictive and detect rare genetic variants. The panel can be selected so that a genetic variant or tumor marker present in a sample at a frequency as low as 0.01%, 0.05%, or 0.001% can have a positive predictive value of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0116] To capture more nucleic acid molecules in a sample, the concentration of the probe or bait used in the panel can be increased (2 to 6 ng / μL). The concentration of the probe or bait used in the panel can be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or higher. The probe concentration can be about 2 ng / μL to about 3 ng / μL, about 2 ng / μL to about 4 ng / μL, about 2 ng / μL to about 5 ng / μL, or about 2 ng / μL to about 6 ng / μL. The concentration of the probe or bait used in the panel can be 2 ng / μL or higher to 6 ng / μL or less. In some cases, this can enable the analysis of more molecules in an organism, thereby enabling the detection of less frequently occurring alleles.

[0117] In some embodiments, after sequencing, a quality score can be assigned to each sequence read. The quality score can represent the sequence read, indicating whether the sequence read can be useful for subsequent analysis based on a threshold. In some cases, some sequence reads are not of sufficient quality or length to perform the subsequent mapping step. Sequence reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out from the data set of sequence reads. In other cases, sequence reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out from the data set. Sequence reads that meet a specified quality score threshold can be mapped to a reference genome. After mapping alignment, a mapping score can be assigned to each sequence read. The mapping score can represent the sequence read that is mapped to the reference sequence, indicating whether each position is uniquely mappable. Sequence reads with a mapping score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out from the dataset. In other cases, sequence reads assigned a mapping score of less than 90%, less than 95%, less than 99%, less than 99.9%, less than 99.99%, or less than 99.999% can be filtered out from the dataset. c.MAF determination

[0118] After cfDNA sequencing of a sample, one or more variant allele fractions (MAFs) can be determined. Some or all of the MAFs can be determined before variant classification, after variant classification, during variant classification, before variant filtering, after variant filtering, during variant filtering, or a combination thereof. Prior to this step, the cfDNA can be end-repaired, ligated with adapters containing molecular barcodes, amplified, and enriched. Amplification can incorporate a sample index. In some embodiments, MAF values ​​can be determined for all variants or all somatic variants. In some embodiments, MAF values ​​can be determined for fewer than all variants or fewer than all somatic variants. The term variant allele fraction (VAF) is used interchangeably with MAF herein. The variant allele fraction (MAF) represents the number of variant molecules at a particular genomic location divided by the total number of molecules (e.g., molecular coverage):

number

[0119] The maximum MAF can be determined as the maximum or largest MAF of all somatic variants present or observed in a given sample, hi some embodiments, the maximum MAF can be considered the tumor fraction of a given sample.

[0120] The maximum fraction of diploid genes ("max frac_diploid") (minimum allelic imbalance) can be determined. The fraction of diploid genes ("frac_diploid") is a measure of the level of allelic imbalance across samples, as determined by copy number. Samples with high levels of allelic imbalance are more susceptible to germline / somatic misclassification. Therefore, low levels of allelic imbalance (or high frac_diploid) are an indicator of the confidence of somatic classification calls.

[0121] In an embodiment, total coverage profiles can be used to capture fold changes and therefore tumor fractions rather than individual genes. d. Classification of variants

[0122] Sequencing in steps 103 and 104 generates multiple sequence reads. In steps 107 and / or 108, the multiple sequence reads can be analyzed to determine one or more variants and classify the one or more variants. In an embodiment, the classification of some or all variants can be determined before determining the MAF 105 / 106, after determining the MAF 105 / 106, during determining the MAF 105 / 106, or a combination thereof. Variants can include, for example, single nucleotide variants (SNVs), indels, fusions, and copy number variations. Any known technique for variant calling can be used. In an embodiment, multiple sequence reads from a sample can be assembled and / or mapped and aligned to a genomic location based on a reference genome. In some embodiments, the multiple sequence reads (assembled or otherwise) can then be compared to a reference genome to determine how the multiple sequence reads of the subject vary from the sequence reads of the reference genome. Such a process can determine the presence of one or more variants in the multiple sequence reads. In some embodiments, the molecular barcodes and / or start and end genomic positions of the nucleic acid molecules obtained from the multiple sequence reads can be used to identify mutant molecules whose sequence reads belong to molecules that differ from the reference genome. Such a process can determine the presence of one or more variants in the multiple sequence reads.

[0123] In some embodiments, common heterozygous SNPs can be used to model local germline allele counting behavior, and variants can be called somatic variants if they significantly deviate from the observed germline variant allele fraction.Beta-binomial models can be used because they model both the mean and variance of variant allele counts at common SNPs.For example, the beta-binomial model described in PCT / US2018 / 052087, the entire contents of which are incorporated by reference herein, can be used.The use of beta-binomial models is an improvement over the use of simpler methods, such as fixed MAF cutoffs or Poisson models, because these methods cannot adequately represent the variance of molecular counts. e. Variant filtering

[0124] In an embodiment, one or more filtering processes can be applied to sequence reads to exclude them from further analysis. In an embodiment, some or all of the filtering can be applied before determining the MAF, after determining the MAF, during determining the MAF, before variant classification, after variant classification, during variant classification, or a combination thereof.

[0125] In some embodiments, one or more somatic variants having a MAF of less than about 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9% at the first and / or second time points can be excluded from further analysis. In some embodiments, one or more somatic variants having a mutant molecule count of less than 5, 10, 15, 20, 25, or 30 at the first and / or second time points can be excluded from further analysis. In some embodiments, one or more somatic variants having a coverage of less than 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 at the first and / or second time points can be excluded from further analysis.

[0126] In some embodiments, copy number variants can be used to exclude sequence reads from further analysis.Copy number amplification can be determined as known in the art.In method 100, in step 109, the copy number amplification of genes with insufficient probe coverage or insufficient copy number (for example, below 95% detection limit) can be filtered out.

[0127] For example, CNVs can be determined by analyzing sequence reads to generate chromosomal regions of coverage. Chromosomal regions can be divided into windows or bins of various lengths. Read coverage can be determined for each window / bin region. In some embodiments, the quantitative measure of sequencing read coverage is a measure that indicates the number of reads derived from DNA molecules corresponding to a gene locus (e.g., a specific position, base, region, gene, or chromosome from a reference genome). To associate a read with a gene locus, the read can be mapped or aligned to a reference. Software for performing mapping or alignment (e.g., Bowtie, BWA, mrsFAST, BLAST, BLAT) can associate sequencing reads with gene loci. Once sequence read coverage is determined, a probabilistic modeling algorithm can be applied to convert the normalized nucleic acid sequence read coverage for each window / bin region into individual copy number states. In some cases, the algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodologies, and neural networks. The individual copy number states of each window region can be used to identify copy number variations in chromosomal regions. In some cases, all adjacent window / bin regions with the same copy number can be combined into segments to report the presence or absence of copy number variation states. In some cases, various windows / bins can be filtered before being combined with other segments. Copy number variations can be used to report a percentage score indicating how much disease material (or nucleic acids with copy number variations) is present in the extracellular free polynucleotide sample.

[0128] In some embodiments, the presence of CNVs in one or more genes can be used to exclude variants from further analysis. For example, variants with a threshold number of genes reportable by LDT that have a copy number equal to or greater than the gene-specific 95% limit of detection (LoD) in either the T0 or T1 sample. The threshold can be from about 10 to about 30. The threshold can be, for example, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, etc. In some embodiments, the threshold can be 19.

[0129] Copy number variation can indicate the fold change for a given variant. Using a Gaussian model, the ratio of the fold change between T0 and T1 can be determined, which can be used as an estimate of molecular response score.

[0130] In some embodiments, if a subject does not have somatic variants or does not have variants that meet the criteria of the variant filtering process, the subject can be classified as unevaluable. In some embodiments, a subject classified as unevaluable can be further classified as a molecular responder. In some embodiments, a subject with low ctDNA at both T0 and T1 time points can be classified as unevaluable and further classified as a molecular responder. In some embodiments, a subject with low MAF at both T0 and T1 time points can be classified as unevaluable and further classified as a molecular responder. In some embodiments, a subject with low tumor fraction at both T0 and T1 time points can be classified as unevaluable and further classified as a molecular responder. A low MAF or low tumor fraction can refer to a MAF or tumor fraction that is below the detection limit (e.g., below the 95% detection limit) or below the quantification limit. What constitutes low can depend on the design of the panel, but for example, a MAF of 0.1%, 0.2%, or 0.3% can be considered low. i. Germline filter

[0131] In some embodiments, a germline filter can be applied to the sequence reads. Some (e.g., less than all) or all of the steps shown can be performed in any combination and in any order. Samples taken over the course of a subject's treatment (e.g., a sample taken at time T0 and a sample taken at time T1) may have different levels of tumor shedding and allelic imbalance, meaning that the variant classification in steps 107 / 108 may be prone to assigning different somatic classifications to the same variant in the same subject. Because the goal of molecular response is to track somatic variants over the course of treatment, classification discrepancies can be automatically resolved to appropriately remove germline variants from consideration by reclassifying the variants. For example, a variant may be classified as somatic at time T0 and as germline at time T1. For example, a variant may be classified as germline at time T0 and as somatic at time T1. For example, a variant may be classified as germline at time T0 and as unclassified at time T1. For example, a variant may be classified as somatic at time T0 and unclassified at time T1. Germline filter 200 is configured to resolve such discrepancies and reassign the classification of the variant.

[0132] As shown, for at least one variant in sequence reading, whether this variant is the harmful variant (for example, frameshift or nonsense mutation) in tumor suppressor gene (TSG) can be determined.For example, variant can be compared with the database of known TSG.If variant is the harmful variant in TSG, this variant can be classified as somatic regardless of classification result (for example, classification is changed from germline to somatic).

[0133] If variant is not a harmful variant in TSG, germline filter can determine the maximum MAF of variants present in sample, and the maximum fraction of diploid genes for at least one variant in sample.If the maximum fraction of diploid genes for variant (at one of at least two time points) indicates that variant is somatic variant, and the MAF of variant (at one of at least two time points) does not increase the maximum MAF, then regardless of classification result, variant can be classified as somatic (for example, classification is changed from germline to somatic).If the maximum fraction of diploid genes for variant (at one of at least two time points) indicates that variant is germline, and the MAF of variant (at one of at least two time points) increases the maximum MAF, then regardless of classification result, variant can be classified as germline (for example, classification is changed from somatic to germline).

[0134] If the maximum fraction of diploid genes for a variant indicates that the variant is a somatic variant, and the MAF of the variant increases the maximum MAF, or if the maximum fraction of diploid genes for a variant indicates that the variant is a germline, and the MAF of the variant does not increase the maximum MAF, germline filter can determine whether the variant is classified as somatic at less than a threshold percentage in another patient sample (at one of at least two time points).The threshold percentage can be at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, or 9%.If the variant is classified as somatic at less than a threshold percentage in another patient sample, in step 107 / 108, the variant can be classified as somatic regardless of the classification result (for example, the classification is changed from germline to somatic).

[0135] If the variant is not classified as somatic in less than 5% of other patient samples, germline filter can determine whether the MAF of variant (at one of at least two time points) is greater than the other MAF in the sample.For example, germline filter can determine whether the MAF of variant is at least about 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or at least 10 times greater than one or more other MAF in the same sample.The one or more other MAF in the sample can be, for example, the somatic MAF that is the next largest in the sample.If the MAF of variant is greater than the other MAF in the sample, the variant can be classified as germline regardless of classification result (for example, classification is changed from somatic to germline).

[0136] Germline filter can determine whether the MAF of variant (at one of at least two time points) is greater than another MAF in another sample.For example, germline filter can determine whether the MAF of variant is at least about 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or at least 10 times greater than one or more other MAFs in another sample.The one or more other MAFs in another sample can be, for example, the maximum MAF of other samples.If the MAF of variant is greater than another MAF in another sample, the variant can be classified into germline regardless of classification result (for example, classification is changed from somatic to germline).

[0137] If the MAF of the variant is not greater than another MAF in the sample and not greater than another MAF in another sample, the germline filter allows the variant to be classified as germline regardless of the classification outcome (e.g., the classification is changed from somatic to germline).

[0138] Variants classified as germline can be excluded from further analysis, including, for example, MAF determination and / or MR scoring. In some embodiments, if variants are classified as CHIP in at least one patient sample, they are classified as CHIP variants. ii. CHIP filter

[0139] cfDNA can include a collection of cfDNA from any cell type, including tumors, blood cells, etc. Clonal hematopoietic (CHIP) mutations with undetermined potential can also exist in cfDNA. Common methods for CHIP filtering utilize frequently occurring CHIP genes or hotspots curated by large-scale public or internal cohort studies. However, these methods do not address the challenge of identifying random CHIP mutations in plasma-only methods. Residual unfiltered CHIP variants bias the fraction change toward 1 (no change), thus leading to inaccurate subsequent molecular response predictions. To filter informal CHIP variants (e.g., variants that are CHIP but have never been documented or are poorly documented in previous databases of known CHIP variants), mutation measurements between two time points can be used to cluster variants with similar fraction changes. As patients undergo treatment, progression or response results in a fraction of somatic mutations, while CHIP variants remain stable. By clustering mutations into clones, random CHIP variants can be found in clones enriched for the known CHIP list or in stable fraction differential clones.

[0140] Thus, provided herein is an improvement to CHIP filtering that leverages observations between two time points (T0 and T1) to cluster genomic variations into clones with different fractions of alterations. CHIP filtering allows events to be grouped / clustered into clones to estimate % clonal burden changes. The clustering procedure can start with each single event and then be integrated using novel clustering heuristics. Once all events have been used to determine % clonal burden changes, each clone can be examined based on variant composition and % clonal burden changes to determine whether the variant is a CHIP clone.

[0141] In one embodiment, a novel agglomerative hierarchical clustering heuristic is used to cluster genomic mutations / variants. The heuristic quantifies statistical dissimilarity between mutations / variants and clusters using a custom dissimilarity metric. An adjustable stopping rule is used to continue agglomeration until a minimum (or maximum, depending on the metric) acceptable dissimilarity threshold is met. In one embodiment, the custom dissimilarity metric is a modification of the Bhattacharyya distance, and thus a numerical integration (without taking the square root) is performed on the product of the scaled likelihoods of the mutations / variants and / or clusters considered for integration at a given step of the clustering heuristic. The likelihoods are scaled so that when numerically integrated over the support range of the integration, they equal 1. For SNVs and indels, the likelihood is calculated based on a beta-binomial model fit to the observed count data, which informs the determination of the MAF of the variants to be clustered. The variance of the beta-binomial model is set by an adjustable parameter. For CNV, likelihood is calculated based on Gaussian model approximation of the observed fold change estimate of the mutation of interest, and the variability of Gaussian model is also set by adjustable parameters.Mutation aggregation is carried out in a new manner, and therefore in some cases, clustering is carried out stepwise, in which a first set of mutations is clustered until it meets a stopping rule, and then a second set of mutations is introduced, and optionally further aggregation steps are carried out according to the same difference metric and stopping rule.In some cases, a third set of mutations is introduced in the same way, and then clustering heuristic method is applied to the second set of mutations.

[0142] In one embodiment, the CHIP filter generates a scaled likelihood function P for each mutation / variant in the sample. i (R i ) can be estimated, where i=1,...,I mv is a total of I mvFor ease of presentation, we denote the number of observed mutations / variants at time point 1 as the i-th mutation / variant, and the number of observed mutations / variants at time point 2 as the i-th mutation / variant.

number

number

number

number

number

number

number

number

[0143] R i Approximate confidence intervals for i Assuming an improper prior distribution for the value of P i (R i =r i ) is the scaled likelihood i This can be calculated in a variety of ways, including by a technique similar to the highest density interval, which can be considered an approximation of the posterior density of

[0144] The set of mutations / variants is denoted as P i (R i ) can be aggregated pairwise according to all possible pairings {i',i * :i'≠i * ;i',i * =1,2,…,I mv}, P i’ (R i’ ) and P i* (R i* ), the dissimilarity measure between i',i * ) is calculated using the modified Bhattacharyya distance. * ) the larger the value, the more likely it is that the mutation pair {i',i *} are likely to be instantiations from the same underlying distribution of fractional changes. Therefore, pairs of mutations / variants with the largest value D(·,·) can be merged into a single clone, and the P i (R i) can be updated. Pairwise aggregation can continue until a stopping criterion is met or until all mutations / variants are aggregated into a single clone. The threshold can be and / or can include a value ranging from about 0.0005 to 0.005.

[0145] The number of clones and the associated fractional changes between time points can be reported along with confidence intervals. Clones with fractional changes at or above a predetermined threshold between the first and second time points can be identified. If multiple clones are identified, clones with fractional changes close to 1 and / or clones with specific known CHIP variants can be classified as potential CHIP variants. CHIP variants can be excluded from further analysis. In some embodiments, if a variant is classified as CHIP in at least one patient sample, the variants can be classified as CHIP variants.

[0146] An example of the application of the CHIP filter, including an example aggregation procedure, is described herein. In this example, three qualifying variants were identified. The scaled likelihood function (y-axis) for each variant can be displayed across the support range of R (x-axis). The variant corresponding to the first likelihood of the scaled likelihood function for each variant is assumed to be a known CHIP variant. The most similar variant in the left panel is annotated with an asterisk. The center panel displays the aggregated likelihood obtained from integrating the first and second likelihoods of the scaled likelihood function for the clones in the left panel. The third likelihood of the scaled likelihood function for each clone from the left panel has a likelihood function that was not modified by aggregation. The right panel displays the final clonality. Because the composition of the second likelihood clone is 50% CHIP, the second likelihood can be presumptively identified as CHIP. This ensures that the final R value is determined solely by the third likelihood clone.

[0147] Described herein are methods that include determining a change in tumor burden (R) relative to a change in tumor fraction P(R) for each of a plurality of variants from sequence information generated from targeted nucleic acids associated with one or more cancer types in samples obtained from a subject at a first time point and a second time point, resulting in a set of tumor burden changes. Further, the method also includes identifying one or more resistance signatures corresponding to one or more clonal variants from the set of tumor burden changes. f.MR score

[0148] The method may proceed such that an MR score is determined in step 601. In some embodiments, the MR score can be determined using MAF values ​​associated with somatic variants remaining after variant filtering. In some embodiments, the MAF values ​​of all somatic variants can be used. In some embodiments, the MAF values ​​of fewer than all somatic variants can be used. As described in step 601, MAFs can be determined for a plurality of somatic variants from sequence reads generated from targeted nucleic acids associated with one or more cancer types in a sample obtained from a subject at T0 (e.g., before treatment) and T1 (e.g., during treatment) to yield a set of first and second MAFs for somatic variants among the plurality of somatic variants. The MR score can be expressed as a fraction or a percentage. The MR score can be determined according to the method. The method may include, in step 601, determining ratios of the first and second MAFs for somatic variants among the plurality of somatic variants to yield a set of MAF ratios and corresponding standard deviations for the MAF ratios within the set of MAF ratios. In some embodiments, standard deviation can be used as the standard for reporting MR score.For example, the standard deviation of MR score based on the individual standard deviation of at least one variant can be used to determine the confidence interval and the cutoff for the subsequent sample evaluability.In some embodiments, the cutoff can be at least 0.1, 0.15, 0.2, 0.3, 0.4 or 0.5.In step 602, for each subject, the weighted average of MAF ratio can be determined using the following formula:

number

number

number

[0149] In an embodiment, in addition to or instead of using the weighted average of MAF ratios as the MR score, a method is disclosed in which variants are clustered based on MAF ratios, an aggregate MAF ratio is calculated for each cluster, and the MR score is then used as the ratio of a single selected cluster or as the weighted average of cluster ratios.Clustering can be performed by combining pairs of variants with overlapping MAF ratio distributions or by other clustering methods.The single selected cluster can be one that contains known cancer driver variants or one that does not have known clonal hematopoietic variants.The weight of a cluster can also depend on the presence of known cancer driver variants or the maximum VAF or the number of variants in a cluster.

[0150] A. MR score can be determined by the weighted average of the first MAF and the weighted average of the second MAF for somatic variants among multiple somatic variants and the standard deviation of the corresponding weighted MAF ratio.In some embodiments, standard deviation can be used as the basis for reporting MR score.For example, the standard deviation of the MR score based on the individual standard deviation of at least one variant can be used to determine the confidence interval and the subsequent cutoff for the evaluability of the sample.In some embodiments, the cutoff can be at least 0.1, 0.15, 0.2, 0.3, 0.4 or 0.5.In step 1, the ratio of the weighted average of MAF can be determined for the subject.The confidence interval is the variance of the ratio.For example, the confidence interval can be determined using the following formula: R=A / B:var(R)~=var(B) / A^2+var(A)*B^2 / A^4 (where A and B are the weighted average MAF at time points 1 and 2, respectively).

[0151] Clusters can be weighted based on the strength of evidence.For example, max-VAF can indicate which is the main clone, the number of non-CHIP variants can weight clusters with stronger signals, and driver weights can increase the weight of clusters containing drivers of specific cancer types or molecular subtypes, or select the clusters.The weighting applied can be, for example, to apply a greater weight to variants known to be drivers in specific cancer types or molecular subtypes.In some embodiments, weights can be based on max-VAF (any sample), the number of non-CHIP variants, and / or driver weights (tumor type-specific; defined in configuration file).In another embodiment, weighting applied can be, for example, to equally weight somatic variants.

[0152] In some embodiments, classification into molecular responder or molecular non-responder can depend on the VAF of variant and the weight of variant.For example, when MR score is the ratio of average VAF, higher VAF (i.e., more clonality variant) is likely to dominate.When MR score uses the weight of variant, the variant with higher weight (for example, driver variant) may dominate.

[0153] The weighted average of the obtained MAF ratios described or the ratio of the weighted average of MAFs can be the MR score of the subject.For such MR score, the variance of MAF is incorporated into the calculation of molecular response.This ensures that the molecular response score contains accurate variance, which contributes to drawing correct conclusions from molecular response.The MR score can be considered as a "numerically stable" ratio of the average MAF, which is appropriately weighted based on the accuracy of MAF against the change of MAF, and is not prone to overconfidence and erroneous results when MAF fluctuates around the limit of detection (LOD).The MR score can be compared with a threshold value to determine whether the subject responds to treatment or does not respond to treatment.The threshold value can be, for example, from about 25% to about 75% and / or can include it. In some embodiments, the weighting may be based on either the accuracy of the VAF (e.g., location, hotspot regions, depth of coverage, etc.) or existing knowledge about the importance of its variants to the tumor (e.g., known driver or resistance mutations, or variants of uncertain (or unknown) significance).

[0154] To provide a simple example to illustrate the problem addressed by the MR scoring method presented herein, consider a subject in which one variant is detected, the MAF at baseline (T0) is 0.3%, the MAF during treatment (T1) is 0.1%, and the coverage is 3000 molecules at that variant position.Using existing methods, the molecular response score is:

number

[0155] To provide a simple example to illustrate the problem addressed by the MR scoring method presented herein, consider a subject in which two variants (a and b) are detected, and the MAF at baseline (T0) is a=0.1% and b=8.0%, and the MAF during treatment (T1) is a=0.3% and b=2.0%. Using existing methods and considering the average ratio, the molecular response score is calculated as follows:

number

number

[0156] To provide a simple example to illustrate the problem addressed by the MR scoring method presented herein, consider a subject in which two variants (a and b) are detected, and the MAF at baseline (T0) is a=0.3% and b=0.0%, and the MAF during treatment (T1) is a=0.0% and b=0.3%. If existing methods are used to evaluate only variants greater than 0.3% at baseline, the molecular response score will be:

number

number

[0157] Method 100 may include administering one or more treatments to the subject based at least on the molecular response score. Exemplary treatments are further disclosed herein. In some embodiments, method 100 includes comparing the molecular response score for the subject having cancer to a predetermined cutoff point and identifying the subject as likely to be a responder to one or more treatments for cancer (e.g., immunotherapy) if the molecular response score is below the predetermined cutoff point, or identifying the subject as likely to be a non-responder to one or more treatments for cancer if the molecular response score is at or above the predetermined cutoff point. In some embodiments, method 100 includes administering one or more treatments for cancer to the subject taking into account the molecular response score. In some embodiments, method 100 includes discontinuing administering one or more treatments for cancer to the subject taking into account the molecular response score. In some embodiments, method 100 includes using the molecular response score as a prognostic and / or predictive biomarker for the subject.

[0158] In other exemplary embodiments, variance is incorporated into the molecular response calculation by simulating or sampling from the variance distribution of at least one variant to calculate molecular response variance. As further disclosed herein, some applications include weighting variants based on their significance in the tumor or the likelihood of tumor or clonal hematopoiesis. Some embodiments involve integrating multiple genomic data sources to estimate tumor fraction (rather than simply relying on variant (e.g., SNV, indel, and fusion) VAF), coverage (e.g., copy number), off-target coverage, and / or methylation, among other genomic data sources.

[0159] In some embodiments, the method includes determining a molecular response score for a subject with cancer using one or more additional genomic data sources. In some embodiments, the additional genomic data sources include one or more of coverage, off-target coverage, epigenetic signature, tumor mutation burden, and / or microsatellite instability score. For each data source, a tumor fraction can be calculated based on the data source, and the calculated tumor fractions can be combined across the data sources (e.g., using a weighted average to incorporate the reliability of the tumor fraction data source for the particular sample), and then an overall tumor fraction estimate for the sample can be combined to calculate an overall molecular response. In some embodiments, the epigenetic signature includes cfNA fragment length, location, and / or endpoint density distribution. In some embodiments, the epigenetic signature includes the epigenetic state or status represented at one or more epigenetic loci within a given targeted genomic region. In some embodiments, the epigenetic state or status comprises the presence or absence of methylation, hydroxymethylation, acetylation, ubiquitination, phosphorylation, sumoylation, ribosylation, citrullination, and / or histone post-translational modifications or other histone variations.

[0160] Although the present method is described in the context of a first time point T0 and a second time point T1, it should be understood that more than two time points are contemplated, for example, for longitudinal monitoring. At the first time point T0, baseline cfDNA can be obtained from one or more baseline samples obtained from one or more subjects before treatment, and at the second time point T1 or any subsequent time point T nIn the method, the treatment-period cfDNA can be obtained from one or more treatment-period samples obtained from one or more subjects after treatment. The T1 time point can be any amount of time after the T0 time point, for example, between 1 hour and 24 hours, between 1 day and 180 days, between 1 week and 12 weeks, between 1 week and 25 weeks, between 1 week and 30 weeks, inclusive. Further, the method 100 can be performed at T0, T1, ..., T n Any combination of time points can be applied.For example, sample can be obtained at time T1 and time T2, and the sample obtained at both time points is the sample during treatment.In another example, sample can be obtained at time T1 and time T2, and the sample obtained at time T1 represents the sample during treatment, and the sample obtained at time T2 represents the sample when no treatment is performed.

[0161] In some embodiments, the dosage of treatment administered to a subject can be adjusted based on molecular response score.For example, molecular response score can indicate that the subject does not respond to the first treatment, and the dosage of the first treatment can be increased accordingly.In some embodiments, alternative therapies can be identified based on molecular response score.For example, molecular response score can indicate that the subject does not respond to the first treatment, and then the subject can be administered a second treatment instead of or in addition to the first treatment.In some embodiments, the molecular response score of a subject can be determined in a clinical trial, and the molecular response score of the subject receiving a placebo and the subject receiving treatment can be determined.The molecular response scores of the two categories of subjects can be compared to assess treatment.

[0162] In another example, placebo and treatment can be generalized into two arms of a clinical trial comparing different combinations of drugs.The threshold or cutoff can be specific to the use case.Some use cases may require clearance (MR=0), or some use cases may require a certain level of reduction or increase in ctDNA level.

[0163] Examples of practical applications of molecular response scores for patient stratification are described. Patients with advanced cancer may have a baseline MAF determined before treatment at time T0. After 4-10 weeks of treatment, patients with advanced cancer may have a treatment-period MAF determined at time T1. The resulting molecular response score may indicate a decrease in ctDNA in the patient, in which case the patient should continue treatment with the lead study drug. The resulting molecular response score may indicate an increase in ctDNA in the patient, in which case the patient should continue treatment with the lead study drug (or placebo) if the patient is included in the control group. Otherwise, if the patient's ctDNA is increasing, the patient should have one or more therapies added to their treatment regimen, treatment changed, or the dose of the lead study drug changed. Further details can be found in PCT International Application No. PCT / US2022 / 070984. III. Cancer and Other Diseases

[0164] In certain embodiments, methods and aspects disclosed herein are used for longitudinal monitoring of patients with certain disease, disorder or condition.Disclosed methods can be used to track the response of patients to one or more treatments over time.Typically, the disease studied is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, renal clear cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myelomonocytic leukemia (CML), and chronic myelomonocytic leukemia (CML). ML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.

[0165] Non-limiting examples of other genetically based diseases, disorders, or conditions that may be evaluated, if desired, using the methods and systems disclosed herein include achondroplasia, alpha 1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri-a-cat syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down's syndrome, Duane's syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, and fragile X syndrome. These include: Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter's syndrome, Marfan's syndrome, myotonic dystrophy, neurofibromatosis, Noonan's syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland's anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner's syndrome, palatocardiofacial syndrome, WAGR syndrome, and Wilson's disease. IV. Customized Treatment and Related Administration

[0166] In some embodiments, the method disclosed herein relates to identifying the patient with a given disease, disorder or condition and administering treatment.Essentially any cancer treatment (for example, surgery, radiation therapy, chemotherapy, and / or the like) is included as part of these methods.In certain embodiments, the treatment administered to the subject can include at least one chemotherapy drug. In some embodiments, chemotherapeutic agents may include alkylating agents (e.g., but not limited to, chlorambucil, cyclophosphamide, cisplatin, and carboplatin), nitrosoureas (e.g., but not limited to, carmustine and lomustine), antimetabolites (e.g., but not limited to, Fluorauracil, methotrexate, and fludarabine), plant alkaloids and natural products (e.g., but not limited to, vincristine, paclitaxel, and topotecan), antitumor antibiotics (e.g., but not limited to, bleomycin, doxorubicin, and mitoxantrone), hormonal agents (e.g., but not limited to, prednisone, dexamethasone, tamoxifen, and leuprolide), and biological response modifiers (e.g., but not limited to, Herceptin and Avastin, Erbitux, and Rituxan). In some embodiments, the chemotherapy administered to the subject may include FOLFOX or FOLFIRI. In certain embodiments, the subject can be administered a treatment comprising at least one PARP inhibitor. In certain embodiments, PARP inhibitors can include, among others, olaparib, talazoparib, rucaparib, and niraparib (trade name ZEJULA). Typically, the treatment comprises at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of enhancing the immune response to a given type of cancer. In certain embodiments, immunotherapy refers to a method of enhancing the T cell response to tumors or cancer.

[0167] In some embodiments, the immunotherapy or immunotherapeutic agent targets immune checkpoint molecules. Certain tumors can utilize immune checkpoint pathways to evade the immune system. Therefore, targeting immune checkpoints has emerged as an effective approach to counteract tumors' ability to evade the immune system and activate anti-tumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012,12:252-264.

[0168] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that reduces signals involved in T cell responses to antigens. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 limits the activity of T cells in peripheral tissues during inflammatory responses. Furthermore, PD-1 ligands (PD-L1 or PD-L2) are commonly upregulated on the surface of many different tumors, resulting in downregulation of anti-tumor immune responses in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a PD-1 ligand, e.g., PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a CTLA4 ligand, e.g., CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T-cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).

[0169] Antagonists that target these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of an inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4 antibody, an anti-PD-1 antibody, an anti-PD-L1 antibody, or an anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD-L1 antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, a combination of an anti-CTLA4 antibody and an anti-PD-L1 antibody, or a combination of an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain embodiments, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain embodiments, the anti-PD-L1 antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).

[0170] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) against CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments, the antagonist is a soluble version of an inhibitory immune checkpoint molecule, e.g., a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.

[0171] In certain embodiments, the immune checkpoint molecule is a costimulatory molecule that amplifies signals involved in T cell responses to antigens. For example, CD28 is a costimulatory receptor expressed on T cells. When a T cell binds to an antigen through its T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on an antigen-presenting cell, amplifying T cell receptor signaling and promoting T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can counteract or regulate costimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a costimulatory molecule selected from CD28, inducible T cell costimulator (ICOS), CD137, OX40, or CD27. In other embodiments, the immune checkpoint molecule is a ligand for a costimulatory molecule, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.

[0172] Agonists targeting these costimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses to certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of a costimulatory checkpoint molecule. In certain embodiments, the agonist of a costimulatory checkpoint molecule is an agonist antibody, preferably a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS antibody, an anti-CD137 antibody, an anti-OX40 antibody, or an anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80 antibody, an anti-CD86 antibody, an anti-B7RP1 antibody, an anti-B7-H3 antibody, an anti-B7-H4 antibody, an anti-CD137L antibody, an anti-OX40L antibody, or an anti-CD70 antibody.

[0173] Therapeutic options for treating particular genetically based diseases, disorders, or conditions other than cancer are generally well known to those of skill in the art and will become apparent upon consideration of the particular disease, disorder, or condition being considered.

[0174] In certain embodiments, the customized therapy described herein is typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapy (e.g., immunotherapeutic agents, etc.) may be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, urethral, ​​topical, ocular, nasal, and / or auricular administration, and may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc. V. Computer Systems for Processing Real-World Evidence (RWE)

[0175] The method of the present disclosure can be implemented by using or with the assistance of computer system.For example, this method can include: dividing sample into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample comprises the DNA with cytosine modification at a higher rate than the second sub-sample; subjecting the first sub-sample to a procedure that affects the first nucleobase of DNA differently from the second nucleobase of the DNA of the first sub-sample, wherein the first nucleobase is modified nucleobase or unmodified nucleobase, and the second nucleobase is modified nucleobase or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and sequencing the DNA in the first sub-sample and the DNA in the second sub-sample, so that the first nucleobase in the DNA of the first sub-sample is distinguished from the second nucleobase.

[0176] In certain aspects, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing a plurality of sets of target regions from the cfDNA, wherein the plurality of sets of target regions includes a set of sequence variable target regions and a set of epigenetic target regions, thereby resulting in a set of captured cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the set of sequence variable target regions are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the set of epigenetic target regions; obtaining, with a nucleic acid sequencer, a plurality of sequence reads generated by sequencing the captured cfDNA molecules; mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the mapped sequence reads corresponding to the set of epigenetic target regions to determine a likelihood that the subject has cancer.

[0177] The code may be pre-compiled and configured for use on a machine with a processor adapted to execute the code, or it may be compiled at run time. The code may be supplied in a programming language that can be selected to allow the code to be executed in a pre-compiled or compiled-on-demand manner.

[0178] Additional details regarding computer systems and networks, databases, and computer program products are also provided in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety. Further information can be found in PCT Publication No. US2022032250 and U.S. Patent Application No. 17832498.

[0179] According to one or more implementations, a method for generating an integrated data repository containing multiple types of healthcare data is described herein. The architecture may include a data integration and analysis system. The data integration and analysis system may obtain data from a number of data sources and integrate the data from the data sources into the integrated data repository. For example, the data integration and analysis system may obtain data from a health insurance claims data repository. In various examples, the data integration and analysis system and the health insurance claims data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the health insurance claims data repository may be created and maintained by the same entity.

[0180] The data integration and analysis system can be implemented on one or more computing devices. The one or more computing devices can include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In certain implementations, at least a portion of the one or more computing devices can be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices can be implemented in a cloud computing architecture. In scenarios where the computing system used to implement the data integration and analysis system is configured as a distributed computing architecture, processing operations can be performed in parallel by multiple virtual machines. In various examples, multithreading techniques can be implemented in the data integration and analysis system. The implementation of a distributed computing architecture and multithreading techniques allows the data integration and analysis system to utilize fewer computing resources than a computing architecture in which these techniques are not implemented.

[0181] The health insurance claims data repository may store information obtained from one or more health insurance companies corresponding to claims made by subscribers of one or more health insurance companies. The health insurance claims data repository may be organized (e.g., sorted) by patient identifier. The patient identifier may be based on the patient's first name, last name, date of birth, social security number, address, employer, etc. The data stored in the health insurance claims data repository may include structured data arranged in one or more data tables. The one or more data tables in which the structured data is stored may include a number of rows and a number of columns indicating information about health insurance claims made by subscribers of one or more health insurance companies related to procedures and / or treatments the subscribers received from healthcare providers. At least some of the rows and columns of the data tables stored in the health insurance claims data repository may include health insurance codes that may indicate diagnoses of biological conditions and treatments and / or procedures received by subscribers of one or more health insurance companies. In various examples, the health insurance codes may also indicate diagnostic procedures received by individuals related to one or more biological conditions that may exist for the individuals. In one or more examples, a diagnostic procedure may yield information used to detect the presence of a biological condition. A diagnostic procedure may also yield information used to determine the progression of a biological condition. In one or more illustrative examples, a diagnostic procedure may include one or more imaging procedures, one or more assays, one or more laboratory procedures, one or more combinations thereof, etc.

[0182] The data integration and analysis system may also obtain information from a molecular data repository. The molecular data repository may store data related to genomic, genetic, metabolomic, transcriptomic, fragmentomic, immune receptor, methylation, epigenomic, and / or proteomic information for a number of individuals. In one or more examples, the data integration and analysis system and the molecular data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the molecular data repository may be created and maintained by the same entity.

[0183] The genomic information may indicate one or more mutations corresponding to an individual's genes. The mutations in an individual's genes may correspond to differences between the individual's nucleic acid sequence and one or more reference genomes. The reference genome may include a known reference genome, such as hg19. In various examples, the mutations in an individual's genes may correspond to differences between the individual's germline genes and a reference genome. In one or more additional examples, the reference genome may include the individual's germline genome. In one or more further examples, the mutations in an individual's genes may include somatic mutations. The mutations in an individual's genes may be associated with insertions, deletions, single-base variants, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof.

[0184] In one or more illustrative examples, the genomic information stored in the molecular data repository may include a genomic profile of tumor cells present in an individual. In these situations, the genomic information may be derived from an analysis of genetic material from a sample, including, but not limited to, a tissue sample or tumor biopsy, circulating tumor cells (CTCs), exosomes, or efferosomes, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA), or from circulating nucleic acids (e.g., extracellular free DNA) found in the individual's blood sample, which are present due to the degradation of tumor cells present in the individual. In one or more examples, the genomic information of the individual's tumor cells may correspond to one or more target regions. One or more mutations present in one or more target regions may indicate the presence of tumor cells in the individual. The genomic information stored in the molecular data repository may be generated in connection with an assay or other diagnostic test that can determine one or more mutations in one or more target regions of a reference genome.

[0185] In one or more additional examples, the data integration and analysis system may retrieve information from one or more additional data repositories. The one or more additional data repositories may store data related to an individual's electronic medical record for which data is present in at least one of the health insurance claims data repository or the molecular data repository. Furthermore, the one or more additional data repositories may store data related to an individual's pathology report for which data is present in at least one of the health insurance claims data repository or the molecular data repository. In various examples, the one or more additional data repositories may store data related to a biological condition and / or a treatment for a biological condition. In one or more examples, the data integration and analysis system and at least a portion of the one or more additional data repositories may be created and maintained by different entities. In one or more further examples, the data integration and analysis system and at least a portion of the one or more additional data repositories may be created and maintained by the same entity.

[0186] In one or more further implementations, the data integration and analysis system may retrieve information from one or more reference information data repositories. The one or more reference information data repositories may store information including definitions, standards, protocols, vocabularies, one or more combinations thereof, etc. In various examples, the information stored in the one or more reference information data repositories may correspond to biological conditions and / or treatments for biological conditions. In one or more illustrative examples, the one or more reference information data repositories may include RxNorm (which provides normalized names for clinical drugs and links the names to multiple drug lexicons used in medication management and drug interaction software). In one or more examples, the data integration and analysis system and at least a portion of the one or more reference information data repositories may be created and maintained by different entities. In one or more further examples, the data integration and analysis system and at least a portion of the one or more reference information data repositories may be created and maintained by the same entity.

[0187] The data integration and analysis system may obtain data from at least one of the health insurance claims data repository, the molecular data repository, the one or more additional data repositories, or the reference information data repository via one or more communication networks accessible to the data integration and analysis system and accessible to at least one of the health insurance claims data repository, the molecular data repository, the one or more additional data repositories, or the reference information data repository. The data integration and analysis system may also obtain data from at least one of the health insurance claims data repository, the molecular data repository, the one or more additional data repositories, or the reference information data repository via one or more secure communication channels. Additionally, the data integration and analysis system may obtain data from at least one of the health insurance claims data repository, the molecular data repository, the one or more additional data repositories, or the reference information data repository via one or more application programming interface (API) calls.

[0188] The data integration and analysis system may include a data integration system. The data integration system may retrieve data from the health insurance claims data repository and the molecular data repository to generate the integrated data repository. The data integration system may also retrieve data from one or more additional data repositories to generate the integrated data repository. In various examples, one or more natural language processing techniques may be implemented in the data integration system to integrate data from the one or more additional data repositories into the integrated data repository.

[0189] In one or more examples, the data integration system may generate one or more tokens to identify individuals having data stored in a health insurance claims data repository and individuals having data stored in a molecular data repository. In various examples, the data integration system may generate one or more tokens by implementing one or more hash functions. One or more hash functions can be implemented in the data integration system to generate one or more tokens based on information stored in at least one of the health insurance claims data repository or the molecular data repository. For example, the information the data integration system uses to generate individual tokens by implementing a hash function may include at least one of an identifier for each individual, a date of birth for each individual, a zip code for each individual, a date of birth for each individual, or a gender for each individual. In one or more illustrative examples, the identifier for each individual may include a combination of at least a portion of a first name and a last name for each individual. Tokens generated using data from different data repositories may correspond to the same or similar information or the same or similar types stored in the different data repositories. By way of example, a token may be generated using an individual's first and last name, date of birth, at least a portion of their zip code, and a portion of their gender obtained from health insurance claims data repositories and molecular data repositories.

[0190] A data integration system may integrate data from a number of different data sources by analyzing tokens generated by implementing one or more hash functions using data obtained from the number of different data sources. For example, the data integration system may obtain one or more first tokens generated from data stored in a health insurance claims data repository and one or more second tokens generated from data stored in a molecular data repository. The data integration system may analyze the one or more first tokens against the one or more second tokens to determine individual first tokens corresponding to the individual second tokens. In one or more illustrative examples, the data integration system may identify individual first tokens that match individual second tokens. A first token may match a second token if the data of the first token has at least a threshold amount of similarity to the data of the second token. In one or more examples, a first token may match a second token if the data of the first token is the same as the data of the second token. For example, a first token may match a second token if the alphanumeric string of the first token is the same as the alphanumeric string of the second token.

[0191] By determining a first token generated using data stored in the health insurance claims data repository that corresponds to a second token generated using data stored in the molecular data repository, the data integration system may identify individuals that have both data stored in the health insurance claims data repository and data stored in the molecular data repository. In this manner, the data integration system may obtain data from the health insurance claims data repository from a number of individuals and data from the molecular data repository from the same number of individuals and store the health insurance claims data and molecular data for that number of individuals in an integrated data repository.

[0192] The data integration system may integrate data stored in one or more additional data repositories with data from the health insurance claims data repository and the molecular data repository to generate an integrated data repository. By way of example, the data integration system may obtain one or more third tokens generated from data stored in the additional data repository, such as a data repository storing data corresponding to pathology reports. The data integration system may analyze the one or more third tokens against a first token generated using information stored in the health insurance claims data repository and a second token generated using information stored in the molecular data repository to determine respective third tokens corresponding to each of the first tokens and each of the second tokens. In one or more illustrative examples, the data integration system may identify the third tokens generated using one or more hash functions and a common set of information obtained from the health insurance claims data repository, the molecular data repository, and the additional data repository.

[0193] By determining a third token generated using data stored in the additional data repository that corresponds to the first token generated using data stored in the health insurance claims data repository and the second token generated using data stored in the molecular data repository, the data integration system may identify individuals having data stored in the health insurance claims data repository, data stored in the molecular data repository, and data stored in the additional data repository. In this manner, the data integration system may obtain data from the health insurance claims data repository from a number of individuals and data from the molecular data repository and the additional data repository from the same number of individuals and store the health insurance claims data, molecular data, and additional data for the number of individuals in the integrated data repository.

[0194] The data stored in the integrated data repository for the number of individuals may be accessible using the individual's respective identifier. The data integration system may implement several techniques as part of the de-identification process for storing and retrieving an individual's information in the integrated data repository. The individual's identifier may correspond to a key generated using at least one hash function. The individual's identifier may also be generated by implementing one or more salting processes on the key generated using at least one hash function. The token generated using one or more hash functions and a common set of information obtained from the health insurance claims data repository, the molecular data repository, and / or the additional data repository. In one or more illustrative examples, the identifier generated by the data integration system to access information about each individual stored in the integrated data repository may be unique for each individual. In one or more examples, the individual's identifier may be generated using at least a portion of the information used to generate the token associated with the individual. In one or more additional examples, the individual's identifier may be generated using information different from the information used to generate the token associated with the individual.

[0195] The data integration system may also generate an integrated data repository from different combinations of a number of data repositories as well. For example, the data integration system may obtain tokens generated from information stored in the health insurance claims data repository and additional tokens generated from information stored in one or more additional data stores. The data integration system may determine tokens generated from information stored in the health insurance claims data repository that correspond to each additional token generated from information stored in the one or more additional data repositories. By determining tokens generated using data stored in the health insurance claims data repository that correspond to additional tokens generated using data stored in the additional data repositories, the data integration system may identify individuals who have both data stored in the health insurance claims data repository and data stored in the additional data repositories. In this manner, the data integration system may obtain data from the health insurance claims data repository from a number of individuals and data from the additional data repositories from the same number of individuals and store the health insurance claims data and additional data for the number of individuals in the integrated data repository. The health insurance claims data and the data stored in the integrated data repository for the additional number of individuals may be accessible using the individuals' respective identifiers.

[0196] In one or more further examples, the data integration system may obtain tokens generated from information stored in the molecular data repository and tokens generated from information stored in one or more additional data stores. The data integration system may determine tokens generated from information stored in the molecular data repository that correspond to each additional token generated from information stored in the one or more additional data repositories. By determining tokens generated using data stored in the molecular data repository that correspond to additional tokens generated using data stored in the additional data repositories, the data integration system may identify individuals who have both data stored in the molecular data repository and data stored in the additional data repositories. In this manner, the data integration system may obtain data from the molecular data repository from a number of individuals and data from the additional data repositories from the same number of individuals and store molecular data and additional data for the number of individuals in the integrated data repository. The molecular data and data stored in the integrated data repository for the additional number of individuals may be accessible using the individuals' respective identifiers.

[0197] The data stored in the integrated data repository may be stored in accordance with one or more regulatory frameworks that protect privacy and ensure the security of individuals' diagnostic records, health information, and insurance information. For example, data may be stored in the integrated data repository in accordance with one or more government regulatory frameworks that cover the protection of personal information, such as the Health Insurance Portability and Accountability Act (HIPAA) and / or the General Data Protection Regulation (GDPR). The integrated data repository also stores data in a de-identified and anonymized manner to ensure the protection of the privacy of individuals whose data is stored in the integrated data repository. To further ensure the privacy of individuals whose data is stored in the integrated data repository, the data integration system may periodically regenerate the integrated data repository. For example, the data integration system may create the integrated data repository once a quarter. In one or more additional examples, the data integration system may generate the integrated data repository monthly, weekly, or biweekly. By periodically regenerating the integrated data repository, rather than simply refreshing it when new data becomes available, the integrated data repository enhances privacy protection for the data stored in the integrated data repository. That is, in situations where a data repository is simply refreshed with new data, the number of new individuals added at a given time will typically be smaller than the number of existing individuals whose data is already stored in the data repository, so it may be easier to track the individuals associated with the newly added data to the data repository.

[0198] In various examples, the data stored in the integrated data repository may be accessed via a database management system. Additionally, the integrated data repository may store data according to one or more database models. In one or more examples, the integrated data repository may store data according to one or more relational database technologies. For example, the integrated data repository may store data according to a relational database model. In one or more additional examples, the integrated data repository may store data according to an object-oriented database model. In one or more further examples, the integrated data repository may store data according to an extensible markup language (XML) database model. In an additional example, the integrated data repository may store data according to a structured query language (SQL) database model. In yet another example, the integrated data repository may store data according to an image database model.

[0199] A data integration system may generate an integrated data repository by generating a number of data tables and creating links between the data tables. The links may indicate logical connections between the data tables. The data integration system may generate the data tables by extracting specific sets of data from information obtained from the data repositories and storing the data in rows and columns of respective data tables. In various examples, the logical connections between the data tables may include at least one of a one-to-one link, in which a row of information in one data table corresponds to a row of information in another data table, a one-to-many link, in which a row of information in one data table corresponds to multiple rows of information in another data table, or a many-to-many link, in which multiple rows of information in one data table correspond to multiple rows of information in another data table.

[0200] The number of data tables can be arranged according to a data repository schema. In an exemplary example, the data repository schema includes a first data table, a second data table, a third data table, a fourth data table, and a fifth data table. While the exemplary example includes five data tables, in additional implementations, the data repository schema may include more or fewer data tables. The data repository schema may also include links between the data tables. Links between the data tables may indicate that information retrieved from one of the data tables results in the retrieval of additional information stored in one or more additional data tables. Furthermore, not all of the data tables need to link to each of the other data tables. In an exemplary example, the first and second data tables are logically linked by a first link, and the first and fourth data tables are logically linked by a second link. Furthermore, the second and third data tables are logically linked via a third link, and the fourth and fifth data tables are logically linked via a fourth link. Furthermore, the third data table and the fifth data table are logically linked via a fifth link.

[0201] In various examples, as data tables are added and / or removed from the data repository schema, links between additional data tables can be added or removed from the data repository schema. In one or more illustrative examples, the integrated data repository may store data tables according to the data repository schema for at least some of the individuals for whom the data integration system retrieves information from a combination of at least two of the health insurance claims data repository, the molecular data repository, the one or more additional data repositories, and the one or more reference information data repositories. As a result, the integrated data repository may store respective instances of data tables according to the data repository schema for thousands, tens of thousands, hundreds of thousands, or even more individuals.

[0202] The data integration and analysis system may also include a data pipeline system. The data pipeline system may include several algorithms, software codes, scripts, macros, or other computer-executable instruction bundles that process information stored in the integrated data repository to generate additional data sets. The additional data sets may include information obtained from one or more of the data tables. The additional data sets may also include information derived from data obtained from one or more of the data tables. The components of the data pipeline system implemented to generate a first additional data set may be different from the components of the data pipeline system used to generate a second additional data set.

[0203] In one or more examples, the data pipeline system may generate a dataset indicating drug treatments received by a number of individuals. In one or more illustrative examples, the data pipeline system may analyze information stored in at least one of the data tables to determine health insurance codes corresponding to drug treatments received by a number of individuals. The data pipeline system may analyze the health insurance codes corresponding to the drug treatments against a library of data indicating specific drug treatments corresponding to the one or more health insurance codes to determine the names of the drug treatments received by the individuals. In one or more additional examples, the data pipeline system may analyze information stored in an integrated data repository to determine medical procedures received by a number of individuals. By way of example, the data pipeline system may analyze information stored in one of the data tables to determine treatments received by an individual via at least one injection or intravenous infusion. In one or more further examples, the data pipeline system may analyze information stored in the integrated data repository to determine episodes of care for an individual, lines of therapy received by the individual, progression of a biological condition, or time to next treatment. In various examples, the datasets generated by the data pipeline system may differ for different biological conditions. For example, the data pipeline system may generate a first number of datasets relating to a first cancer type, e.g., lung cancer, and a second number of datasets relating to a second cancer type, e.g., colorectal cancer.

[0204] The data pipeline system may also determine one or more confidence levels to assign to information associated with individuals having data stored in the integrated data repository. Each confidence level may correspond to a different measure of accuracy for the information associated with individuals having data stored in the integrated data repository. The information associated with each confidence level may correspond to one or more characteristics of the individuals derived from the data stored in the integrated data repository. Confidence level values ​​for the one or more characteristics may be generated by the data pipeline system in conjunction with generation of one or more datasets from the integrated data repository. In one or more examples, the first confidence level may correspond to a first range of the accuracy scale, the second confidence level may correspond to a second range of the accuracy scale, and the third confidence level may correspond to a third range of the accuracy scale. In one or more additional examples, the second range of the accuracy scale may include values ​​less than the values ​​in the first range of the accuracy scale, and the third range of the accuracy scale may include values ​​less than the values ​​in the second range of the accuracy scale. In one or more illustrative examples, information corresponding to a first confidence level may be referred to as gold standard information, information corresponding to a second confidence level may be referred to as silver standard information, and information corresponding to a third confidence level may be referred to as bronze standard information.

[0205] The data pipeline system may determine a confidence level value for an individual's characteristic based on several factors. For example, each set of information can be used to determine the individual's characteristic. The data pipeline system may determine the confidence level for the individual's characteristic based on the amount of completeness of each set of information used to determine the individual's characteristic. In a situation where one or more pieces of information are missing from a set of information associated with a first number of individuals, the confidence level for the characteristic may be lower than the confidence level for a second number of individuals for which no information is missing from the set of information. In one or more examples, the data pipeline system may use the amount of missing information to determine the confidence level for the individual's characteristic. By way of example, a greater amount of missing information used to determine the individual's characteristic may result in a lower confidence level for the characteristic than a situation where a lesser amount of missing information is used to determine the characteristic. Furthermore, different types of information may correspond to confidence levels for various characteristics. In one or more examples, the presence of a first piece of information used to determine the individual's characteristic may result in a higher confidence level for the characteristic than if a second piece of information used to determine the characteristic were present.

[0206] In one or more illustrative examples, the data pipeline system may determine a number of individuals to be included in a cohort having a primary diagnosis of lung cancer (or other biological condition). The data pipeline system may determine a confidence level for each individual to be classified as having a primary diagnosis of lung cancer. The data pipeline system may use information from a number of columns included in a data table to determine the confidence level for inclusion of an individual in the lung cancer cohort. The number of columns may include health insurance codes associated with a diagnosis of a biological condition and / or a treatment for a biological condition. Further, the number of columns may correspond to a diagnosis date and / or a treatment for a biological condition. The data pipeline system may determine that the confidence level of an individual characterized as being part of the lung cancer cohort is higher in a scenario where information is available for each of the number of columns, or at least a threshold number of columns, than when information is available for fewer than a threshold number of columns. Furthermore, the data pipeline system may determine the confidence level for an individual to be included in the lung cancer cohort based on the type of information associated with one or more columns and the availability of the information. By way of example, in a situation where one or more diagnostic codes are present in association with one or more time periods for a group of individuals and one or more treatment codes are absent, the data pipeline system may determine that the confidence level for inclusion of the group of individuals in the lung cancer cohort is higher than in a situation where at least one of the diagnostic codes is absent and the treatment code used to determine whether the individuals are included in the lung cancer cohort is present.

[0207] The data integration and analysis system may include a data analysis system. The data analysis system may receive integrated data repository requests from one or more computing devices, such as an example computing device. The one or more integrated data repository requests may cause data to be retrieved from the integrated data repository. In various examples, the one or more integrated data repository requests may cause data to be retrieved from one or more datasets generated by the data pipeline system. The integrated data repository requests may specify data to be retrieved from the integrated data repository and / or one or more datasets generated by the data pipeline system. In one or more additional examples, the integrated data repository requests may include one or more pre-constructed queries corresponding to computer-executable instructions that cause specified sets of data to be retrieved from one or more datasets generated by the integrated data repository and / or the data pipeline system.

[0208] In response to the one or more integrated data repository requests, the data analysis system may analyze data retrieved from at least one of the integrated data repository or one or more datasets generated by the data pipeline system to generate data analysis results. The data analysis results may be transmitted to one or more computing devices, such as the exemplary computing device. While the illustrative example shows one or more integrated data repository requests and data analysis results from one computing device being transmitted to another computing device, in one or more additional implementations, the data analysis results may be received by the same computing device that transmits the one or more integrated data repository requests. The data analysis results may be displayed by one or more user interfaces rendered by the computing device or computing device.

[0209] In one or more examples, at least one of one or more machine learning techniques or one or more statistical techniques may be implemented in the data analysis system to analyze the data retrieved in response to one or more integrated data repository requests. In one or more examples, one or more artificial neural networks may be implemented in the data analysis system to analyze the data retrieved in response to one or more integrated data repository requests. By way of example, at least one of one or more convolutional neural networks or one or more residual neural networks may be implemented in the data analysis system to analyze the data retrieved from the integrated data repository in response to one or more integrated data repository requests. In at least some examples, one or more random forest techniques, one or more support vector machines, or one or more hidden Markov models may be implemented in the data analysis system to analyze the data retrieved in response to one or more integrated data repository requests. One or more statistical models may also be implemented to analyze the data retrieved in response to one or more integrated data repository requests to identify at least one correlation or measure of significance between individual characteristics. For example, a log-rank test may be applied to the data retrieved in response to one or more integrated data repository requests. Additionally, a Cox proportional hazards model may be implemented on the data retrieved in response to one or more integrated data repository requests. Additionally, a Wilcoxon signed-rank test may be applied to the data retrieved in response to one or more integrated data repository requests. In yet another example, a z-score analysis may be performed on the data retrieved in response to one or more integrated data repository requests. In yet an additional example, a Kaplan-Meier analysis may be performed on the data retrieved in response to one or more integrated data repository requests.In at least some examples, a combination of one or more machine learning techniques and one or more statistical techniques may be implemented to analyze data retrieved in response to one or more integrated data repository requests.

[0210] In one or more illustrative examples, the data analysis system may determine a survival rate in response to one or more treatments for individuals presenting with lung cancer. In one or more additional illustrative examples, the data analysis system may determine a survival rate in response to one or more treatments for individuals having mutations in one or more genomic regions presenting with lung cancer. In various examples, the data analysis system may generate data analysis results in situations where data retrieved from at least one of the one or more datasets generated by the integrated data repository or data pipeline system meets one or more criteria. For example, the data analysis system may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests meets a confidence level threshold. In situations where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests is below the confidence level threshold, the data analysis system may avoid generating at least a portion of the data analysis results. In scenarios where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests is at least the confidence level threshold, the data analysis system may generate at least a portion of the data analysis results. In various examples, the confidence level threshold may be related to the type of data analysis results generated by the data analysis system.

[0211] In one or more illustrative examples, the data analysis system may receive an integrated data repository request to generate data analysis results indicative of survival rates for one or more individuals. In these cases, the data analysis system may determine whether the data stored in the integrated data repository and / or one or more datasets generated by the data pipeline system meet a confidence level threshold, e.g., a gold standard confidence level. In one or more additional examples, the data analysis system may receive an integrated data repository request to generate data analysis results indicative of treatments received by one or more individuals. In these implementations, the data analysis system may determine whether the data stored in the integrated data repository and / or one or more datasets generated by the data pipeline system meet a lower confidence level threshold, e.g., a bronze standard confidence level.

[0212] In one or more additional illustrative examples, the data analysis system may receive an integrated data repository request to determine individuals who have one or more genomic mutations and have received one or more treatments for a biological condition. Continuing with this example, the data analysis system may determine the survival rate of individuals with one or more genomic mutations in relation to the one or more treatments they have received. The data analysis system may then identify, based on the survival rate of the individuals and the effect of the treatments on the individuals, the association of genomic mutations that may be present in the individuals. In this way, by identifying potential treatments that may be more effective than current treatments provided to individuals for a population of individuals with one or more genomic mutations, health outcomes for the individuals may be improved.

[0213] Described herein is a framework that supports the arrangement of data tables in an integrated data repository according to one or more implementations. In an illustrative example, the framework includes a data repository schema including a first data table, a second data table, a third data table, a fourth data table, a fifth data table, a sixth data table, and a seventh data table. While seven data tables are shown in the illustrative example, in additional implementations, the data repository schema may include more or fewer data tables. The data repository schema may also include links between the data tables. Links between the data tables may indicate that information retrieved from one of the data tables results in the retrieval of additional information stored in one or more additional data tables. Furthermore, not all of the data tables need to link to each of the other data tables. In an illustrative example, the first and second data tables are logically linked by a first link, and the third and second data tables are logically linked by a second link. Additionally, the second data table and the fourth data table are logically linked by a third link, the second data table and the fifth data table are logically linked by a fourth link, and the second data table and the sixth data table are logically linked by a fifth link. Further, the fifth data table and the sixth data table are logically linked by a sixth link, and the sixth data table and the seventh data table are logically linked by a seventh link. Further, the seventh data table and the fourth data table are logically linked by an eighth link. In various examples, as data tables are added and / or removed from the data repository schema, links between additional data tables can be added or removed from the data repository schema.In one or more illustrative examples, the integrated data repository may store data tables according to the data repository schema for at least some of the individuals for whom information is obtained by the data integration system from a combination of at least two of the health insurance claims data repository, the molecular data repository, and the one or more additional data repositories. As a result, the integrated data repository may store respective instances of data tables according to the data repository schema for thousands, tens of thousands, hundreds of thousands, or even more individuals.

[0214] In one or more examples, the first data table may store data corresponding to genomics and genomics tests for an individual. For example, the first data table may include columns containing genomics data, mutations in genomic regions, mutation types, copy numbers of genomic regions, coverage data indicating the number of nucleic acid molecules with one or more mutations identified in the sample, test dates, and information corresponding to the panel used to generate the patient information. The first data table may also include one or more columns containing health insurance data codes that may correspond to one or more diagnostic codes. Additionally, the information in the first data table may include at least one identifier for the individual associated with the instance of the first data table.

[0215] The second data table may store data regarding one or more patient visits by an individual to one or more healthcare providers. The third data table may store information corresponding to each service provided to the individual for one or more patient visits to one or more healthcare providers represented by the second data table. For example, an individual may visit a healthcare provider, and multiple services may be performed on the individual during that visit. The second data table may include columns indicating information about each of the multiple services performed during the patient visit. Multiple third data tables related to the patient visit may be generated, each containing columns indicating a more granular level of information about the patient visit than the information about the patient visit stored in the second data table, for each service provided during the patient visit. For example, the second data table may include multiple columns indicating health insurance codes for different services provided to the individual during the patient visit, and the third data table related to one of the services may include multiple columns for additional health insurance codes corresponding to additional information about the respective service. The second and third data tables related to the patient visit may indicate one or more dates for the services corresponding to the patient visit.

[0216] The fourth data table may include columns that indicate information about the individuals whose information is stored in the integrated data repository. For example, the fourth data table may include columns that indicate information about at least one of the individual's location, the individual's gender, the individual's date of birth, the individual's date of death (if applicable), or one or more keys associated with the individual. In one or more examples, the fourth data table may include one or more columns regarding whether erroneous data has been identified for the individual. In various examples, a single fourth data table may be generated for each individual. Thus, the data repository schema may include numerous instances of the fourth data table, e.g., thousands, tens of thousands, hundreds of thousands, or more.

[0217] The fifth data table may include columns indicating information about the health insurance company or government agency that pays for one or more services provided to each individual. For example, the fifth data table may include one or more payer identifiers. The sixth data table may include columns containing information corresponding to health insurance coverage information for each individual. In one or more examples, the sixth data table may include columns indicating the existence of medical coverage for the individual, the existence of drug coverage for the individual, and the type of health insurance plan for the individual, such as a health maintenance organization (HMO), preferred provider organization (PPO), etc.

[0218] The seventh data table may include columns indicating information regarding medications received by each individual. In one or more examples, the seventh data table may include one or more columns indicating health insurance codes corresponding to medications available through the pharmacy. The health insurance codes may correspond to individual medications. Further, the health insurance codes may indicate a diagnosis of a biological condition for the individual. The seventh data table may also include additional information, such as at least one of dosage, number of days of medication, dispensed quantity, number of refills allowed, date of service, or information regarding the individual receiving the medication.

[0219] In various examples, the data repository schema may provide analysis of information stored in data tables in a more efficient manner than a typical data repository schema. For example, the data repository schema may arrange logical joins between data tables so that related data can be efficiently searched across different data tables. In situations where the data tables are arranged in a sequential manner and / or where a larger number of data tables are logically joined, retrieval of data from one or more of the data tables from the integrated data repository in response to a request for information from the integrated data repository may be less efficient than in situations where a data repository schema is implemented.

[0220] According to one or more implementations, an architecture is described herein for generating one or more datasets incorporating health-related data from several sources from information retrieved from a data repository. The architecture may include a data integration and analysis system and an integrated data repository. Further, the data integration and analysis system may include at least a data pipeline system and a data analysis system. The data pipeline system may include several sets of data processing instructions executable to generate respective datasets that can be analyzed by the data analysis system to generate data analysis results in response to an integrated data repository request.

[0221] The data pipeline system may include first data processing instructions, second data processing instructions, and up to Nth data processing instructions. The data processing instructions may be executable by one or more processing devices to perform several operations to generate respective datasets using information retrieved from the integrated data repository. In one or more illustrative examples, the data processing instructions may include at least one of software code, scripts, API calls, macros, etc. The first data processing instructions may be executable to generate the first dataset. Further, the second data processing instructions may be executable to generate the second dataset. Further, the Nth data processing instructions may be executable to generate the Nth dataset. In various examples, after the data integration and analysis system generates the integrated data repository, the data pipeline system may execute the data processing instructions to generate datasets. In one or more examples, the datasets may be stored in the integrated data repository or in an additional data repository accessible to the data integration and analysis system. At least some of the data processing instructions may analyze health insurance codes to generate at least some of the datasets. Further, at least some of the data processing instructions may analyze genomics data to generate at least some of the datasets.

[0222] In one or more examples, the first data processing instructions may be executable to retrieve data from one or more first data tables stored in the integrated data repository. The first data processing instructions may also be executable to retrieve data from one or more specified columns of the one or more first data tables. In various examples, the first data processing instructions may be executable to identify individuals having health insurance codes stored in one or more column and row combinations corresponding to one or more diagnostic codes. The first data processing instructions may then be executable to analyze the one or more diagnostic codes to determine the biological condition with which the individual has been diagnosed. In one or more illustrative examples, the first data processing instructions may be executable to analyze the one or more diagnostic codes against a library of diagnostic codes indicating one or more biological conditions corresponding to each diagnostic code. The library of diagnostic codes may include hundreds, up to thousands, of diagnostic codes. The first data processing instructions may also be executable to determine individuals diagnosed with a biological condition by analyzing timing information for the individuals, such as date of treatment, date of diagnosis, date of death, one or more combinations thereof, etc.

[0223] The second data processing instructions may be executable to retrieve data from one or more second data tables stored in the integrated data repository. The second data processing instructions may also be executable to retrieve data from one or more specified columns of the one or more second data tables. In various examples, the second data processing instructions may be executable to identify individuals having health insurance codes stored in one or more column and row combinations corresponding to one or more procedure codes. The one or more procedure codes may correspond to procedures obtained from a pharmacy. In one or more additional examples, the one or more procedure codes may correspond to procedures received by medical procedure, such as an injection or intravenous infusion. The second data processing instructions may be executable to determine one or more procedures corresponding to each health insurance code included in the one or more second data tables by analyzing the health insurance code in association with a set of predetermined information. The set of predetermined information may include a data library indicating one or more procedures corresponding to one of hundreds, or up to thousands, of health insurance codes. The second data processing instructions may generate a second dataset indicating each procedure received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to the individuals included in the first dataset. The second dataset may be arranged in rows and columns, with one or more rows corresponding to a single individual and one or more columns indicating the treatment that each individual received.

[0224] The Nth processing instructions (N may be any positive integer) may be executable to generate the Nth dataset by combining information from several previously generated datasets, e.g., a first dataset and a second dataset. Furthermore, the Nth processing instructions may be executable to generate the Nth dataset by retrieving additional information from one or more additional columns of the integrated data repository and incorporating the additional information from the integrated data repository with the information obtained from the first dataset and the second dataset. For example, the Nth processing instructions may be executable to identify individuals included in the first dataset who have been diagnosed with a biological condition and analyze specified columns of one or more additional data tables of the integrated data repository to determine treatment dates indicated in the second dataset that correspond to individuals included in the first dataset. In one or more further examples, the Nth processing instructions may be executable to analyze columns of one or more additional data tables of the integrated data repository to determine dosages of treatments indicated in the second dataset received by individuals included in the first dataset. In this manner, the Nth processing instructions may be executable to generate an episode of care dataset based on information contained in the cohort dataset and the treatment dataset.

[0225] In one or more illustrative examples, in response to receiving the integrated data repository request, the data analysis system may determine one or more datasets corresponding to characteristics of a query related to the integrated data repository request. For example, the data analysis system may determine that information included in a first dataset and a second dataset is applicable to responding to the integrated data repository request. In these scenarios, the data analysis system may analyze at least a portion of the data included in the first dataset and the second dataset to generate a data analysis result. In one or more additional examples, the data analysis system may determine different datasets responsive to different queries included in the integrated data repository request to generate the data analysis result.

[0226] By using a specific set of data processing instructions to generate each dataset, the number of inputs by a user of the data integration and analysis system may be reduced, while also reducing the computational load, such as the amount of processing resources and memory, utilized to process integrated data repository requests. For example, without the specific architecture of the data pipeline system, each integrated data repository request is received and the data utilized to respond to the integrated data repository request is assembled from the data repository. In contrast, when a data pipeline system is implemented to execute data processing instructions to generate datasets, the data required to respond to various integrated data repository requests is already assembled, and the data analysis system can access that data to respond to the integrated data repository requests. Therefore, by implementing a data pipeline system to generate datasets, fewer computing resources are utilized to respond to integrated data repository requests than a typical system that performs information parsing and collection processing for each integrated data repository request. Furthermore, in situations where a data pipeline system is not implemented, a user of a data integration and analysis system may need to submit multiple integrated data repository requests to analyze the information that the user intends to analyze, either due to the inaccuracy of ad-hoc data collection in typical systems to respond to an integrated data repository request, or due to the multiple invocations of a data analysis system to perform analysis of information that could be performed using a single integrated data repository request if a data pipeline system were implemented.

[0227] Described herein is an architecture for generating an integrated data repository containing de-identified health insurance claims data and de-identified genomics data according to one or more implementations. The architecture may include a data integration and analysis system, a health insurance claims data repository, and a molecular data repository. The data integration and analysis system may obtain patient information from the molecular data repository. The patient information may include genomics data about individuals having data stored in the molecular data repository. The genomics data may represent the results of one or more nucleic acid sequencing operations that analyze the sequences of nucleic acid molecules contained in samples obtained from the individuals with respect to one or more target genomic regions. In one or more examples, the samples may be obtained from tissues of one or more individuals. In one or more additional examples, the samples may be obtained from fluids, such as blood or plasma, of one or more individuals. The one or more target genomic regions may correspond to genomic regions corresponding to the presence of one or more biological conditions. For example, the target regions may correspond to genomic regions of a reference genome having mutations present in individuals with the biological conditions. In one or more illustrative examples, the target regions may correspond to genomic regions of a reference human genome in which one or more mutations are present in individuals with one or more forms of cancer. Patient information may also include information indicating personal information about an individual with data stored in a molecular data repository as well as information corresponding to tests and analyses performed on samples provided by the individual.

[0228] The data integration and analysis system may perform an anonymization process to de-identify personal information retrieved from the molecular data repository. The data integration and analysis system may implement one or more computational techniques as part of the anonymization process to de-identify data about individuals stored in the molecular data repository, thus protecting the privacy of the individuals and complying with one or more regulatory privacy frameworks. The anonymization process may include accessing a token. In various examples, the token may include an alphanumeric string. In one or more examples, the token may be generated by the data integration and analysis system. In one or more additional examples, the token may be generated by a third party and retrieved by the data integration and analysis system.

[0229] The token can be generated using one or more hash functions on a subset of the patient's information. By way of example, for individuals whose information is stored in the molecular data repository, the token can be generated using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's date of birth, the individual's gender, and at least a portion of each individual's location identifier. The de-identification process can also include generating identifiers for individuals whose data is stored in the molecular data repository. The identifiers can be generated by the data integration and analysis system using one or more hash functions that are different from the one or more hash functions used to generate the tokens. In one or more illustrative examples, the data integration and analysis system can generate intermediate versions of each identifier using one or more hash functions and then apply one or more salting techniques to the intermediate versions of the identifiers to generate final versions of the identifiers. The salting function includes functionality configured to add at least one random bit to each intermediate identifier to generate the respective final identifier. In various examples, the data integration and analysis system can generate identifiers using at least a portion of the information about each individual stored in the molecular data repository. In one or more illustrative examples, the identifier may be generated based on a patient identifier included in the patient's information. The identifier generated by the data integration and analysis system may be unique to the individual having data stored in the respective molecular data repository.

[0230] During operation, the data integration and analysis system may generate corrected patient information based on the identifiers. The corrected patient information may include genomics data for individuals associated with the molecular data repository and an identifier for each individual. The corrected patient information may have a data structure. The data structure may include a column containing an identifier for each individual associated with the molecular data repository and a number of columns containing genomics data for the individual, e.g., an identifier for one or more genes, a modification of one or more genes, a type of modification of a gene, etc.

[0231] The data integration and analysis system may generate a token file. The token file may include first tokens accessed during operation for individuals having data stored in the respective molecular data repositories. The token file may have a data structure including a number of columns containing information about each individual. The data structure may include a column indicating an identifier generated by each data integration and analysis system and a column indicating one or more first tokens associated with each identifier. The data integration and analysis system may transmit the token file to a health insurance claims data management system coupled to a health insurance claims data repository. The health insurance claims data management system may parse the first tokens into corresponding second tokens. The second tokens may be accessed or generated by the health insurance claims data management system. The second tokens may be generated using the same or similar subset of information about individuals having data stored in the health insurance claims data repository as the subset of patient information. For example, the second token may be generated using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's date of birth, the individual's gender, and at least a portion of each individual's location identifier.

[0232] In various examples, a health insurance claims data management system may search a health insurance claims data repository for health insurance claims data for individuals associated with each second token that matches a corresponding first token. A first token may match a second token if the data of the first token has at least a threshold amount of similarity to the data of the second token. In one or more examples, a first token may match a second token if the data of the first token is the same as the data of the second token.

[0233] In response to identifying health insurance claim data for individuals having respective second tokens corresponding to respective first tokens, the health insurance claim data management system may generate corrected health insurance claim data. The health insurance claim data management system may transmit the corrected health insurance claim data to a data integration and analysis system. In one or more examples, the corrected health insurance claim data may be formatted according to a data structure. The data structure may include a column including a subset of second tokens corresponding to the first tokens and a number of columns including the health insurance claim data.

[0234] During operation, the data integration and analysis system may integrate genomics data and health insurance claim data for individuals common to both the molecular data repository and the health insurance claim data repository. The data integration and analysis system may determine individuals common to both the molecular data repository and the health insurance claim data repository by determining the genomics data and health insurance claim data corresponding to common tokens. The data integration and analysis system may determine that a first token related to a portion of the genomics data corresponds to a second token related to a portion of the health insurance claim data by determining a measure of similarity between the first token and the second token. In scenarios in which the first token has at least a threshold amount of similarity with the second token, the data integration and analysis system may store the corresponding portion of the genomics data and the corresponding portion of the health insurance claim data in association with an identifier for the individual in the integrated data repository, for example, the integrated data repository.

[0235] An implementation of the architecture can implement cryptographic protocols that enable the integration of de-identified information from separate data repositories into a single data repository, thus increasing the security of the data stored in the integrated data repository. Furthermore, the cryptographic protocols implemented in the architecture can enable more efficient searching and accurate analysis of the information stored in the integrated data repository than would be possible without the architecture's cryptographic protocols. For example, by generating a token file containing a first token using cryptographic techniques based on a specified set of information stored in a molecular data repository and utilizing a second token generated using the same or similar cryptographic techniques for a similar or identical set of information stored in a health insurance claims data repository, the data integration and analysis system can match information stored in the separate data repositories that correspond to the same individual. Without implementing the architecture's cryptographic protocols, the probability of information from one data repository being erroneously attributed to one or more individuals increases, thereby reducing the accuracy of the results provided by the data integration and analysis system in response to an integrated data repository request sent to the data integration and analysis system.

[0236] According to one or more implementations, a framework is described herein for generating a dataset based on data stored in an integrated data repository by a data pipeline system. The integrated data repository may store health insurance claims data and genomics data for a group of individuals. For example, the integrated data repository may store information obtained from health insurance claims records for the group of individuals. For each individual in the group of individuals, the integrated data repository may store information obtained from multiple health insurance claims records. In various examples, the information stored in the integrated data repository may include and / or be derived from thousands, tens of thousands, hundreds of thousands, or even millions of health insurance claims records for a number of individuals. Furthermore, each health insurance claim record may include multiple columns. As a result, the integrated data repository can be generated by analyzing millions of columns of health insurance claims data.

[0237] Furthermore, while health insurance claims data can be organized according to a structured data format, health insurance claims data is typically organized to present financial information and insurance code information regarding services provided by health insurance providers to individuals for review by health insurance providers, patients, and healthcare providers. Therefore, health insurance claims data may be available in the context of characteristics of individuals in whom a biological condition exists and is not easily analyzed to obtain insights that may aid in the individual's treatment for the biological condition. An integrated data repository may be generated and organized by analyzing and modifying raw health insurance claims data in a manner that allows further analysis of the data stored in the integrated data repository to determine trends, characteristics, features, and / or insights regarding individuals in whom one or more biological conditions may exist. For example, health insurance codes may be stored in the integrated data repository such that, for a given individual, at least one of a medical procedure, a biological condition, a treatment, a dosage, a drug manufacturer, a drug distributor, or a diagnosis can be determined based on the individual's health insurance claims data. In various examples, the data integration and analysis system may generate and implement one or more tables that correlate health insurance claims data with various treatments, symptoms, or biological conditions corresponding to the health insurance claims data. Additionally, an integrated data repository can be generated using genomics data records for a group of individuals. In various examples, large amounts of health insurance claims data can be matched with genomics data for a group of individuals to generate an integrated data repository.

[0238] By integrating genomics data records with health insurance claims records for a group of individuals, the data integration and analysis system can determine correlations between the presence of one or more biomarkers present in the genomics data records and other individual characteristics indicated by the health insurance claims data records that cannot typically be determined using existing systems. For example, the data integration and analysis system can determine one or more genomic characteristics of the individual corresponding to the treatment the individual received, the timing of the treatment, the dosage of the treatment, the individual's diagnosis, smoking status, the presence of one or more biological conditions, the presence of one or more symptoms of the biological conditions, one or more combinations thereof, etc. Based on the correlations determined by the data integration and analysis system using the integrated data repository, a cohort of individuals who may benefit from one or more treatments can be identified that would not have been identified using existing systems. In one or more examples, the processes and techniques implemented to integrate health insurance claims records and genomics claims records to generate an integrated data repository can be complex, and efficiency-enhancing techniques, systems, and processes can be implemented to minimize the amount of computing resources used to generate the integrated data repository.

[0239] In one or more illustrative examples, the data pipeline system may access information stored in the integrated data repository to generate a dataset including several additional data records containing information about at least a portion of a group of individuals. In an illustrative example, the additional data records include information indicating whether the individual is included in a cohort of individuals in which lung cancer is present. The data pipeline system may execute multiple different sets of data processing instructions to determine the cohort of individuals in which lung cancer is present. In various examples, the additional data records may indicate information used to determine the individual's status with respect to lung cancer, such as one or more insurance transaction identifiers, one or more International Classification of Diseases (ICD) codes, and one or more health insurance transaction dates. In addition to including a column indicating whether the individual is included in a lung cancer cohort, the additional data records may include a column indicating a confidence level of the individual's status with respect to the presence of lung cancer.

[0240] A schematic diagram of a computing architecture 600 for incorporating diagnostic record data into an integrated data repository is described herein. In various examples, at least a portion of the operations of the computing architecture may be performed by the data integration and analysis system of Figures 1, 3, and 4. In one or more examples, at least a portion of the operations of the computing architecture may be controlled, maintained, or implemented by a service provider and performed by one or more additional computing systems that control, maintain, or implement the data integration and analysis system. In one or more additional examples, at least a portion of the operations of the computing architecture may be performed by several servers in a distributed computing environment.

[0241] The computing architecture may include a diagnostic record data repository. The diagnostic record data repository may store diagnostic record data from a number of individuals. The diagnostic record data may include imaging information, laboratory test results, diagnostic examination information, clinical findings, dental hygiene information, healthcare practitioner notes, medical history forms, diagnostic request forms, medical procedure order forms, medical information charts, one or more combinations thereof, etc. In various examples, for a given individual, the diagnostic record data repository may store information obtained from one or more healthcare practitioners related to the individual.

[0242] The computing architecture may perform operations including retrieving a data package from a diagnostic record data repository. In one or more examples, the data package may be retrieved in response to one or more requests sent to the diagnostic record data repository for diagnostic records corresponding to one or more individuals. In one or more additional examples, the data package may be retrieved by the computing architecture using one or more application programming interface (API) calls. In one or more illustrative examples, a first data package, a second data package, up to Nth data package may be retrieved using the computing architecture. Each data package may correspond to a diagnostic record for a respective individual. For example, the first data package may include a diagnostic record for a first individual, the second data package may include a diagnostic record for a second individual, and the Nth data package may include a diagnostic record for a third individual.

[0243] An individual data package may include several components. In one or more examples, an individual data package may include individual components corresponding to diagnostic records from different healthcare providers. In one or more additional examples, an individual data package may include individual components corresponding to different portions of diagnostic records corresponding to one or more healthcare providers. In an illustrative example, a second data package may include a first component, a second component, up to Nth components. In one or more illustrative examples, the first component may include a first portion of the individual's diagnostic record, the second component may include a second portion of the individual's diagnostic record, and the Nth component may include a third portion of the individual's diagnostic record. In various examples, the first component may correspond to a diagnostic record by a first healthcare provider for the individual, the second component may correspond to a diagnostic record by a second healthcare provider for the individual, and the third component may correspond to a diagnostic record by a third healthcare provider for the individual. In one or more additional illustrative examples, the first component may include a first section of the individual's diagnostic record, e.g., one or more forms related to a diagnostic test or procedure, and the second component may include a second section of the individual's diagnostic record, e.g., a pathology report for the individual.

[0244] During operation, the computing architecture may pre-process individual data packages to identify a corpus of information to be analyzed. In one or more examples, pre-processing of data packages retrieved from the diagnostic record data repository may include transforming data included in the data packages. For example, pre-processing of data packages may include converting at least a portion of the data retrieved from the diagnostic record data repository into machine-encoded information. By way of example, pre-processing of data packages may include performing one or more optical character recognition (OCR) operations on at least a portion of the data packages retrieved from the diagnostic record data repository. Converting at least a portion of the data packages retrieved from the diagnostic record data repository into machine-encoded information allows the data packages to be subjected to certain operations that cannot be performed on at least a portion of the data packages retrieved from the diagnostic record data repository, such as one or more parsing operations to identify one or more characters or strings of characters, or one or more editing operations.

[0245] In one or more examples, pre-processing of the individual data packages may include determining information contained in the individual data packages that should be excluded from further analysis by the computing architecture. In various examples, one or more components of the individual data packages may be excluded from the corpus of information to be analyzed. For example, with respect to the second data package, the computing architecture may determine that the first component should be excluded from further analysis by the computing architecture. In one or more examples, the computing architecture may analyze the components {circle around (x)}, {circle around (y)}, and / or {circle around (x)} for one or more keywords to identify at least one of the components {circle around (y)}, {circle around (y)}, and / or {circle around (y)} to exclude from further analysis by the computing architecture. In one or more illustrative examples, the computing architecture may parse the components {circle around (y)}, {circle around (y)}, and / or {circle around (y)} to identify one or more keywords, and in response to identifying one or more keywords in the components {circle around (y)}, and / or {circle around (y)}, the computing architecture may determine that the respective components {circle around (y)} and / or {circle around (y)} to exclude from further analysis by the computing architecture. For example, the computing architecture may determine that the first component of the second data package is a test request form for one or more diagnostic procedures or tests. In these scenarios, the computing architecture may determine that the first component should be excluded from further analysis by the computing architecture. Additionally, the computing architecture may determine at least one of the second components {circumflex over (x)} and / or {circumflex over (x)} that corresponds to one or more pathology reports for the individual based on one or more keywords included in the second component or at least one of the Nth components. In these cases, the computing architecture may determine at least a portion of the second component and / or at least a portion of the Nth component for inclusion in the corpus of information to be further analyzed by the computing architecture.

[0246] Additionally, a subset of the components of the data packages retrieved from the individual diagnostic record data repositories can be included in the corpus of information. In various examples, one or more additional operations can be performed to refine the corpus of information. For example, one or more queries can be applied to the subset of information retrieved from the diagnostic record data repositories. The one or more queries may extract information from the one or more data packages that satisfies the one or more queries. In at least some examples, the one or more queries can be a collection of queries applied to individual components of the data packages. In one or more illustrative examples, the collection of queries can determine information to include in the corpus of information and additional information to exclude from the corpus of information. In one or more additional examples, one or more sections of at least one component of the data package can be excluded from the corpus of information.

[0247] In one or more additional illustrative examples, once it is determined that the first component should be excluded from further analysis by the computing architecture, the computing architecture may then implement one or more queries on at least one of the second component or the Nth component. In these scenarios, the one or more queries may determine that a section of the second component, e.g., a section indicating a family history of one or more biological conditions, should be excluded from the corpus of information. In various examples, the one or more queries may be directed to identifying certain keywords and / or combinations of keywords included in at least one of the second component or the Nth component. In these cases, the computing architecture may exclude from the corpus of information one or more portions of individual components of the data package that include one or more keywords or combinations of keywords. In one or more additional examples, the computing architecture may exclude from the corpus of information a certain number of words, characters, and / or symbols following one or more keywords included in one or more portions of individual components of the data package.

[0248] Further, during operation, the computing architecture may analyze the corpus of information to determine characteristics of individuals. In one or more examples, the computing architecture may analyze the corpus of information to determine individuals having one or more phenotypes. In various examples, the computing architecture may analyze the corpus of information to determine one or more biomarkers indicative of a biological state. For example, the computing architecture may analyze the corpus of information to determine individuals having one or more genetic traits. The one or more genetic traits may include at least one of one or more variants of a genomic region corresponding to a biological state. In one or more illustrative examples, the one or more genetic traits may correspond to one or more variants of a genomic region corresponding to a type of cancer. In one or more additional illustrative examples, the one or more biomarkers may correspond to analyte levels outside of a specified range. By way of example, the computing architecture may analyze the corpus of information to determine individuals having one or more protein levels and / or one or more small molecule levels indicative of a biological state. In these scenarios, the computing architecture may analyze laboratory test results to determine the individual's analyte levels. In one or more additional examples, the computing architecture may analyze the corpus of information to determine individuals in which one or more symptoms indicative of a biological condition are present. In one or more further examples, the computing architecture may analyze imaging information included in the corpus of information to determine individuals in which one or more biomarkers are present.

[0249] In one or more examples, one or more machine learning techniques may be implemented in the computing architecture to analyze the corpus of information. For example, one or more artificial neural networks, such as one or more convolutional neural networks or one or more residual neural networks, may be implemented in the computing architecture to analyze the corpus of information. One or more random forest techniques, one or more hidden Markov models, or one or more support vector machines may also be implemented in the computing architecture to analyze the corpus of information.

[0250] In at least some implementations, the computing architecture may analyze the corpus of information by performing one or more queries on the corpus of information. The one or more queries may correspond to one or more keywords and / or keyword combinations. The one or more keywords and / or keyword combinations may correspond to at least one of characters or symbols corresponding to one or more biological conditions. For example, the keywords may correspond to characters related to mutations in a genomic region, such as HER2. In one or more additional illustrative examples, one or more criteria may be associated with the keyword combination. For example, the criteria corresponding to the keyword combination may include a number of words that exist within a specified distance of each other in a portion of the corpus of information about the individual, e.g., the words fatigue, blood pressure, and swelling that exist within characters of each other. In these cases, the computing architecture may parse the corpus of information for one or more keywords and / or keyword combinations. In various examples, in response to determining the presence of one or more keywords and / or keyword combinations according to one or more criteria, the computing architecture may determine a biological condition that exists for the given individual.

[0251] In one or more additional examples, the one or more queries may be image-based, and the computing architecture may analyze images in the corpus of information against a template image. The template image may be generated based on analyzing a number of images in which the biological state is present and aggregating the number of images into the template image. In these scenarios, the computing architecture may analyze images in the corpus of information against one or more template images to determine a measure of similarity between the images in the corpus of information and the template image. In situations where the measure of similarity for an individual is at least a threshold, the computing architecture may determine that the characteristic of the biological state is present in the individual.

[0252] After the individuals having one or more characteristics are determined, the computing architecture, upon operation, may generate a data structure that stores data about the individuals having the one or more characteristics. In one or more examples, the computing architecture may generate data tables that indicate individuals having individual characteristics and / or groups of characteristics. For example, the computing architecture may generate a first data table and a second data table. The first data table may indicate individuals having one or more first characteristics, and the second data table may indicate individuals having one or more second characteristics. In one or more illustrative examples, the first data table may indicate individuals having one or more first biomarkers for a biological state, and the second data table may indicate individuals having one or more second biomarkers for the biological state. The one or more first biomarkers may correspond to one or more first genomic variants associated with the biological state, and the one or more second biomarkers may correspond to one or more second genomic variants associated with the biological state. In various examples, the data tables may indicate whether one or more characteristics associated with the individual data tables are present for the individual individuals. By way of example, the first data table may include a first indicator for individuals in which the one or more first genomic variants are present and a second indicator for individuals in which the one or more first genomic variants are absent. In one or more additional examples, the first data table may indicate the smoking status of individuals, and the second data table may indicate whether each individual has received one or more treatments for a biological condition.

[0253] In one or more illustrative examples, the first data table and the second data table may have rows corresponding to individual individuals. In at least some examples, an individual identifier may be present in each row. The individual identifier may include at least one alphanumeric character or code corresponding to the individual. In various examples, an individual identifier may be present in a data package corresponding to the individual. The columns of the first data table and the second data table may indicate the status of the individual with respect to one or more characteristics. For example, the columns of the data tables may include an identifier including at least one alphanumeric character or code indicating the presence or absence of one or more characteristics for a given individual. Furthermore, while the illustrative examples include a first data table and a second data table, the computing architecture may generate more or fewer data tables.

[0254] During operation, the computing architecture may store data structures in the additional data repository. For example, the computing architecture may store at least the first data table and / or the second data table in the intermediate data repository. In various examples, the first data table and the second data table may be temporarily stored in the intermediate data repository. In one or more illustrative examples, the first data table and the second data table may be stored in the intermediate data repository before being added to the integrated data repository. In one or more examples, the integrated data repository may be periodically generated and / or updated. In these scenarios, the data structures generated by the computing architecture based on analysis of the corpus of information may be stored in the intermediate data repository until at least one of the integrated data repository is generated or updated.

[0255] Prior to adding the data structures stored in the intermediate data repository to the integrated data repository, the computing architecture may perform one or more de-identification processes during operation. The data structures stored in the intermediate data repository may be de-identified to protect the privacy of individuals. The one or more de-identification processes may include applying one or more electronically implemented cryptographic techniques to information about individuals included in the data structures stored in the intermediate data repository. In one or more examples, the computing architecture may generate tokens corresponding to each individual having information stored in the data structures of the intermediate data repository. The tokens may be generated by applying one or more hash functions to information about each individual. In one or more examples, the one or more de-identification processes may include applying a salt function to information corresponding to each individual to generate a token for each individual. In various examples, the one or more cryptographic techniques applied to de-identify the data structures stored in the intermediate data repository may be the same as or similar to those applied to information retrieved from the health insurance claims data repository.

[0256] During operation, the computing architecture may store the de-identified data structure together with the integrated data repository. For example, information stored in the intermediate data repository about a given individual may be stored together with additional information about the given individual in the integrated data repository. By way of example, the integrated data repository may store at least two of information about the given individual obtained from a molecular data repository, information obtained from a health insurance claims data repository, and information obtained from the intermediate data repository. In this manner, information about a given individual obtained from a number of separate data repositories may be stored in the integrated data repository. As a result, information about an individual obtained from different data repositories may be analyzed together, rather than separately as in many existing systems.

[0257] In various examples, the information stored in the intermediate data repository can be used to validate one or more decisions made by the data integration and analysis system. For example, the data integration and analysis system may analyze information retrieved from the health insurance claims data repository and the molecular data repository to determine characteristics of an individual. The data integration and analysis system may then analyze the information retrieved from the intermediate data repository to determine whether predicted characteristics identified from the information retrieved from the health insurance claims data repository and from the molecular data repository correspond to characteristics for the same individual for the information stored in the intermediate data repository.

[0258] The one or more cryptographic techniques applied to anonymize the data structures stored in the intermediate data repository may utilize the same or similar information as that used to generate at least one of the first token or the second token. For example, one or more cryptographic techniques may be implemented to anonymize the data structures in the intermediate data repository using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's date of birth, the individual's gender, and at least a portion of each individual's location identifier. By utilizing the same or similar cryptographic techniques and a subset of the same or similar information as that used to generate at least one of the first token or the second token to anonymize the data structures stored in the intermediate data repository, the information stored in the intermediate data repository can be synchronized with information about the same individual whose information is stored in the integrated data repository. Both the integrated data repository and the intermediate data repository may store information for thousands, tens of thousands, or even millions of individuals. Thus, without the use of the specified cryptographic protocols described herein to synchronize individuals with records stored in the integrated data repository and the intermediate data repository, the data structures of the integrated data repository and the intermediate data repository associated with the same individual may not be stored in a manner that allows the information stored in the integrated data repository and the information stored in the intermediate data repository for a given individual to be searched together, which may lead to inaccurate information being provided by the data integration and analysis system. The absence of the specified cryptographic protocols described herein may also lead to the use of greater computing resources to determine the information stored in the integrated data repository and the information stored in the intermediate data repository from other data sources that corresponds to a given individual. Figures 7 and 8 illustrate example processes for generating an integrated data repository and generating datasets used to analyze the information stored in the integrated data repository.An example process is illustrated as a collection of blocks in a logical flow graph, which represents a sequence of operations that can be implemented in hardware, software, or a combination thereof. The blocks are referenced by numbers. In the software context, the blocks represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processing units (e.g., hardware microprocessors), perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described blocks can be combined in any order and / or in parallel to implement a process.

[0259] A data flow diagram of an example process for generating an integrated data repository storing health insurance claims data and genomics data according to one or more implementations is described herein. During operation, the process may include generating a data file including tokens generated using a first hash function. An individual token may correspond to each individual of a group of individuals having data stored in the molecular data repository. In one or more examples, an individual having data stored in the molecular data repository may be associated with one or more tokens. The tokens may be generated by applying one or more first hash functions to a subset of information corresponding to the group of individuals stored in the genomics data repository. In various examples, the individual tokens may be generated by applying one or more first hash functions to one or more combinations of at least a portion of the first name of each individual of the group of individuals, at least a portion of the last name of each individual of the group of individuals, a location identifier of each individual of the group of individuals, a gender of each individual of the group of individuals, and a date of birth of each individual of the group of individuals. In one or more illustrative examples, the tokens may be generated by a data integration and analysis system coupled with the genomics data repository. In one or more additional illustrative examples, the token may be generated by a third-party system and accessed by a data integration and analysis system coupled with the molecular data repository. During operation, the process may also include transmitting the data file to a health insurance claims data management system. The health insurance claims data management system may match the token included in the data file with a second token accessed by the health insurance data management system and generated based on information stored in the health insurance claims data repository.

[0260] Further, during operation, the process may include obtaining first data corresponding to the group of individuals from a health insurance claims data management system in response to the data file, where the first data includes health insurance claims data. In some implementations, obtaining affirmative consent from members of the group of individuals to migrate their data from the health insurance claims data management system. In one or more examples, the data is migrated in a de-identified format, so that the data cannot be traced back to individual members. The health insurance claims data management system may be coupled to a health insurance claims data repository that stores health insurance claims information for a number of individuals. In one or more examples, the health insurance claims data management system may parse tokens in the data file against additional tokens generated by the health insurance claims data management system. The additional tokens may be generated based on the same set of information included in the data file that was used to generate the tokens. However, the identity of the individual cannot be determined based on the tokens. In various examples, the health insurance claims data management system may match tokens included in the data file with additional tokens generated based on information stored in the health insurance claims data repository to determine individuals having information stored in the health insurance claims data repository that also have information stored in a genomics data repository. The technology disclosed herein complies with legal and best practice privacy standards such as HIPAA and GDPR.

[0261] During operation, the process may include generating a number of identifiers using a second hash function different from the first hash function. In one or more examples, each identifier may correspond to one or more tokens for each individual in the group of individuals. The identifiers may be unique to a given individual in the group of individuals and may be anonymized. Furthermore, the identifiers may be generated using information stored in the genomics data repository for the group of individuals that differs from the information stored in the genomics data repository used to generate the tokens. In various examples, intermediate identifiers may be generated by applying a second hash function to the information for each group of individuals, and final versions of the identifiers may be generated by applying one or more salting techniques to the intermediate identifiers. The information stored in the genomics data repository for each individual may be stored in association with an identifier, such that at least a portion of the information for a given individual stored in the genomics data repository can be accessed using the given individual's respective identifier.

[0262] Furthermore, the process may, when operated, include obtaining data from a second molecular data repository for the group of individuals using the number of identifiers, and the process may, when operated, include determining, for the group of individuals, respective portions of the first data that correspond to respective portions of the second data. For example, for a given individual, first data corresponding to health insurance claim data for the given individual can be identified in addition to second data, e.g., genomics data, that corresponds to the given individual's molecular data. In this manner, for a given individual, both health insurance claim data and molecular data can be identified.

[0263] The process, when operated, may include generating an integrated data repository that stores a respective portion of the first data and a respective portion of the second data in association with a respective identifier of the number of identifiers. For example, the integrated data repository may store health insurance claims data and genomics claims data for a given individual in association with an identifier that can be used to access the health insurance claims data and genomics claims data for the given individual. The information stored in the integrated data repository may be organized according to a data repository schema. For example, the integrated data repository may store health insurance claims data and genomics data for a group of individuals in a number of data tables. In one or more examples, the information stored in the number of data tables may be linked. For example, information about a given individual stored in a first data table of the data repository schema may be linked with additional information about the given individual stored in a second data table of the data repository schema. In this manner, information accessed in one data table of the data repository schema may provide access to additional information stored in another data table of the data repository schema.

[0264] In one or more illustrative examples, the data repository schema may include a first data table storing genomics data for a group of individuals. For example, the first data table may store information corresponding to the panel used to generate the genomics data, mutations in genomic regions, the type of mutation, the copy number of the genomic region, coverage data indicating the number of nucleic acid molecules with one or more mutations identified in the sample, the test date, and the patient information. The data repository schema may also include a second data table storing data regarding one or more patient visits by the individual to one or more healthcare providers, and a third data table storing information corresponding to each service provided to the individual for the one or more patient visits to one or more healthcare providers indicated by the second data table. Furthermore, the data repository schema may include a fourth data table storing personal information for the group of individuals and a fifth data table storing information regarding health insurance companies or government agencies that pay for services provided to the group of individuals. Furthermore, the data repository schema may include a sixth data table storing health insurance coverage information for the group of individuals, e.g., information corresponding to the type of health insurance plan for the group of individuals. The data repository schema may also include a seventh data table that stores information about medications received by the group of individuals.

[0265] In one or more examples, the integrated data repository may also store diagnostic records corresponding to at least a portion of the group of individuals. In these examples, the diagnostic records may be retrieved from one or more data repositories in which the diagnostic records are stored. One or more optical character recognition (OCR) operations may be performed on the diagnostic records. Additionally, the diagnostic records may be analyzed to determine one or more pieces of additional information to remove to result in a corpus of information. In various examples, the corpus of information may be analyzed to determine a portion of a subset of the group of additional individuals that corresponds to one or more biomarkers.

[0266] From the corpus of information, one or more data structures can be generated that store identifiers for a portion of the subset of the additional group of individuals and that store an indication that the portion of the subset of the additional group of individuals corresponds to one or more biomarkers. The one or more data structures can be stored in an intermediate data repository. One or more de-identification operations can be performed on the identifiers for the portion of the subset of the additional group of individuals before modifying the integrated data repository to store at least some of the additional information for the diagnostic records of the portion of the subset of the additional group of individuals in association with the number of identifiers. After de-identification of the information stored in the one or more data structures, the information stored in the integrated data repository can be added to the integrated data repository. In at least some examples, the de-identified diagnostic record information can be added to the integrated data repository in addition to or instead of health insurance claims data. In various examples, the one or more data structures that store the de-identified diagnostic record information with respect to biomarker data can have one or more logical connections with other data structures stored in the integrated data repository.By way of example, the one or more data structures storing de-identified diagnostic record information with respect to biomarker data may have one or more logical connections with at least one of: a first data table that may store information corresponding to genomics data, mutations in genomic regions, types of mutations, copy numbers of genomic regions, coverage data indicating the number of nucleic acid molecules with one or more mutations identified in a sample, test dates, and panels used to generate the patient information; a second data table that stores data regarding one or more patient visits by an individual to one or more healthcare providers; a third data table that stores information corresponding to each service provided to an individual for the one or more patient visits to one or more healthcare providers indicated by the second data table; a fourth data table that stores personal information of a group of individuals; a fifth data table that stores information regarding health insurance companies or government agencies that pay for services provided to the group of individuals; a sixth data table that stores health insurance coverage information for the group of individuals, e.g., information corresponding to the types of health insurance plans for the group of individuals; or a seventh data table that stores information regarding medications obtained by the group of individuals.

[0267] In various examples, diagnostic record data can be added to an integrated data repository by generating a data file including first tokens generated using a first hash function. Each first token can correspond to a respective individual in a group of individuals having data stored in a molecular data repository. Furthermore, the data file can be transmitted to a diagnostic record data management system, and diagnostic record data corresponding to the group of individuals can be retrieved from the diagnostic record data management system in response to the data file. Furthermore, a number of identifiers can be generated using a second hash function different from the first hash function. Each identifier can correspond to one or more tokens for each individual in the group of individuals. For the group of individuals, second data can be retrieved from the molecular data repository using the number of identifiers. In various examples, for the group of individuals, respective portions of the first data corresponding to respective portions of the second data can be determined. In this manner, an integrated data repository can be generated in which respective portions of the first data and respective portions of the second data are stored in association with respective identifiers in the number of identifiers.

[0268] After an integrated data repository containing the diagnostic record data has been created, a request can be received to determine data for a number of individuals having data stored in the integrated data repository. The request can include one or more search criteria. In one or more examples, a subset of the number of individuals having one or more characteristics corresponding to the one or more search criteria can be determined, and information for the subset of the number of individuals can be analyzed to determine a measure of significance of the one or more characteristics with respect to the biological condition.

[0269] In one or more exemplary examples, one or more genomic mutations present in a subset of the number of individuals can be determined, and multiple treatments provided to the subset of the number of individuals can also be determined. In various examples, the respective survival rates for the subset of the number of individuals, for example, survival rates in clinical practice, can be determined. In at least some examples, the significance measure can correspond to the survival rate for a treatment among the multiple treatments and a genomic mutation among the one or more genomic mutations. Based on the significance measure, the effect of the treatment on the subset of the number of individuals can be determined. In one or more examples, the individuals within the subset of the number of individuals who have not received the treatment can be determined. One or more therapeutically effective doses of the treatment can be administered to the individuals within the subset of the number of individuals who have not received the treatment.

[0270] Described herein is a data flow diagram of an example process for generating a number of datasets to be used to analyze information stored in an integrated data repository, where health insurance claims data and genomics data are stored, according to one or more implementations. When operated, the process may include determining a first set of executable data processing instructions in association with first data stored in the integrated data repository. The integrated data repository may store health insurance claims data and molecular data for a common group of individuals. In one or more examples, the first set of data processing instructions may be included in multiple sets of data processing instructions that are part of a data processing pipeline. Each of the sets of data processing instructions in the data processing pipeline may be executed to generate a respective analysis-ready dataset. For example, the set of data processing instructions in an individual data processing pipeline may be executable to generate a dataset including a specified portion and / or combination of information stored in the integrated data repository. In one or more additional examples, the set of data processing instructions in an individual data processing pipeline may be executable to analyze and modify a portion of the information stored in the integrated data repository to generate a respective dataset. Furthermore, each set of data processing instructions may be executable with respect to a respective subset of the information stored in the integrated data repository.

[0271] The process, when operated, may also include executing a first set of data processing instructions to generate a first dataset. The first dataset may represent a subset of a group of individuals in which a biological condition exists. The first set of data processing instructions may be executed to analyze the data stored in the integrated data repository to identify a cohort of individuals in which a biological condition exists. In one or more illustrative examples, the biological condition may include cancer. By way of example, the first set of data processing instructions may be executed to analyze the data stored in the integrated data repository to identify a cohort of individuals in which lung cancer exists. In various examples, the data processing pipeline may include multiple sets of data processing instructions to identify cohorts of individuals in which different biological conditions exist.

[0272] In one or more examples, the first set of data processing instructions can be executed to analyze at least one of health insurance claim data or molecular data to determine a cohort of individuals in which a biological condition exists. For example, the first set of data processing instructions can be executed to identify individuals with one or more health insurance codes present in the health insurance claim data to determine a group of individuals in which a biological condition exists. Furthermore, the first set of data processing instructions can be executed to identify individuals in which one or more mutations exist in a genomic region of a nucleic acid molecule derived from a sample obtained from the individual to determine a group of individuals in which a biological condition exists.

[0273] Additionally, the process, when operated, may include determining a second set of executable data processing instructions in association with second data stored in the second integrated data repository. The second set of data stored in the integrated data repository may be different from the first set of data stored in the integrated data repository and may be analyzed in association with the first set of data processing instructions. For example, the first data may correspond to a first column of one or more first data tables stored in the integrated data repository, and the second data may correspond to a second column of one or more second data tables stored in the integrated data repository.

[0274] During operation, the process may include executing a second set of data processing instructions to generate a second dataset indicating one or more treatments provided to a second subset of the group of individuals. The second dataset may indicate a subset of the group of individuals that received one or more treatments. One or more treatments may be provided to individuals with one or more biological conditions. In one or more examples, the second set of data processing instructions may be executed to analyze data stored in the integrated data repository to identify a cohort of individuals that received one or more treatments. By way of example, the second set of data processing instructions may be executed to analyze at least one health insurance claims data or genomics data to determine a cohort of individuals that received one or more treatments. In one or more illustrative examples, the second set of data processing instructions may be executed to identify individuals having one or more health insurance codes present in the health insurance claims data to determine a group of individuals that received one or more treatments.

[0275] Furthermore, the process may, upon operation, include determining a third subset of the group of individuals that includes a portion of the first subset of the group of individuals that overlaps with a portion of the second subset of the group of individuals. As a result, the third subset of the group of individuals corresponds to individuals for whom the biological condition exists and to whom one or more treatments have been provided. The process may include analyzing the first dataset and the second dataset with respect to the third subset of the group of individuals to determine a measure of significance of a trait for the third subset of the group of individuals. In one or more examples, one or more machine learning or statistical techniques may be applied to information contained in at least one of the first dataset and the second dataset with respect to the third subset of the group of individuals. The measure of significance may correspond to a measure of statistical significance for the trait. In one or more additional examples, the measure of significance may correspond to a probability of the trait being present in individuals for whom the biological condition exists.

[0276] In one or more illustrative examples, the characteristic may include one or more treatments provided to an individual in whom a biological condition exists. In one or more additional illustrative examples, the characteristic may include the presence of a mutation in a genomic region of a nucleic acid molecule derived from a sample obtained from an individual in whom a biological condition exists. In various examples, information included in at least one of the first dataset or the second dataset can be analyzed to determine the effect of the characteristic on one or more metrics. In one or more examples, information included in at least one of the first dataset or the second dataset can be analyzed to determine the amount of effect of a treatment on the survival rate of an individual in whom a biological condition exists. In one or more further examples, information included in at least one of the first dataset or the second dataset can be analyzed to determine the amount of effect of a mutation in a genomic region on the survival rate of an individual in whom a biological condition exists. Furthermore, information included in the first dataset and the second dataset can be analyzed to determine the amount of effect of one or more treatments on an individual in whom a biological condition exists and one or more genomic mutations are also present.

[0277] According to an example, according to an example implementation, a machine in the form of a computer system is described herein that can execute a set of instructions to cause the machine to perform any one or more of the methodologies discussed herein. For example, the machine in the example form of a computer system can execute instructions (e.g., software, programs, applications, applets, apps, or other executable code) to cause the machine to perform any one or more of the methodologies discussed herein. For example, the instructions can cause the machine to implement the architecture and framework described above and perform the methods described in connection with the foregoing.

[0278] The instructions transform a general non-programmed machine into a specific programmed machine that performs the described and illustrated functions in the described manner. In alternative implementations, the machine may operate as a stand-alone device or may be coupled (e.g., networked) with other machines. In a networked deployment, the machine may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a mobile phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing, sequentially or otherwise, instructions that specify operations to be performed by the machine. Furthermore, although only a single machine is illustrated, the term "machine" may be construed to include a collection of machines that individually or collectively execute instructions to perform any one or more of the methodologies discussed herein.

[0279] Examples of computing devices may include logic, one or more components, circuits (e.g., modules), or mechanisms. A circuit is a tangible entity configured to perform certain operations. In one example, a circuit may be arranged in a predetermined manner (e.g., internally or in relation to external entities such as other circuits). In one example, one or more computer systems (e.g., stand-alone, client, or server computer systems) or one or more hardware processors (processors) may be configured as circuitry that operates with software (e.g., instructions, portions of an application, or applications) to perform certain operations described herein. In one example, the software may reside (1) on a non-transitory machine-readable medium or (2) as a transmission signal. In one example, the software, when executed by the hardware underlying the circuit, causes the circuit to perform certain operations.

[0280] In one example, a circuit may be implemented mechanically or electronically. For example, a circuit may include dedicated circuitry or logic specifically configured to perform one or more techniques, such as those described above, including, for example, a special-purpose processor, a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In one example, a circuit may include programmable logic (e.g., circuitry contained in a general-purpose processor or other programmable processor) that can be temporarily configured (e.g., by software) to perform certain operations. It will be appreciated that the decision to implement a circuit mechanically (e.g., in dedicated, permanently configured circuitry) or in temporarily configured circuitry (e.g., configured by software) may be determined by cost and time considerations.

[0281] Thus, the term "circuitry" is understood to encompass a tangible entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily (e.g., transiently) configured (e.g., programmed) to operate in a specified manner or to perform specified operations. In one example, considering multiple temporarily configured circuits, each of the circuits need not be configured or instantiated at any one time. For example, if a circuit includes a general-purpose processor configured via software, the general-purpose processor may be configured at different times as each different circuit. Thus, the software may, for example, configure the processor so that a particular circuit is configured at one time and so that a different circuit is configured at a different time.

[0282] In one example, a circuit can provide information to and receive information from other circuits. In this example, a circuit can be considered to be communicatively connected to one or more other circuits. When multiple such circuits exist simultaneously, communication can be achieved through signal transmission (e.g., via appropriate circuits and buses) coupling the circuits. In implementations in which multiple circuits are configured or instantiated at different times, communication between such circuits can be achieved, for example, through storage and retrieval of information in a memory structure accessed by the multiple circuits. For example, one circuit can perform an operation and store the output of that operation in a memory device to which the circuit is communicatively coupled. Another circuit can then access the memory device at a later time to retrieve and process the stored output. In one example, a circuit can be configured to initiate or receive communication with an input or output device and can operate on a resource (e.g., a set of information).

[0283] Various operations of the example methods described herein may be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented circuitry that operates to perform one or more operations or functions. In one example, circuitry referred to herein may include processor-implemented circuitry.

[0284] Similarly, the methods described herein may be at least partially processor-implemented. For example, at least some or all of the operations of a method may be performed by one or more processors or processor-implemented circuitry. The performance of certain operations may be distributed among one or more processors, and the processors may reside within a single machine or may be deployed across several machines. In one example, the processor(s) may be located at a single location (e.g., in a home environment, a work environment, or a server farm), while in other examples, the processors may be distributed across several locations.

[0285] The one or more processors may also operate to support implementation of associated operations in a "cloud computing" environment or as "software as a service" (SaaS).

[0286] For example, at least some of the operations may be performed by a group of computers (as an example of a machine that includes a processor), and these operations may be accessible over a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application program interfaces (APIs)).

[0287] An example implementation (e.g., a device, system, or method) can be implemented in digital electronic circuitry, computer hardware, firmware, software, or any combination of these. An example implementation can be implemented using a computer program product (e.g., a computer program tangibly embodied in an information carrier or machine-readable medium for execution by or to control operations by a data processing apparatus, e.g., a programmable processor, a computer, or multiple computers).

[0288] A computer program may be written in any type of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program, or as a software module, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be executed on one computer or on multiple computers, at one site or distributed across multiple sites interconnected by a communication network.

[0289] In one example, the operations may be performed by one or more programmable processors executing a computer program to perform functions by manipulating input data and generating output. Example method operations may also be performed by, and example apparatus may be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)).

[0290] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. It will be understood that in implementations deploying a programmable computing system, both hardware and software architectures need to be considered. In particular, it will be understood that the choice of whether to implement particular functionality in permanently configured hardware (e.g., an ASIC), temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently and temporarily configured hardware may be a design choice. Described below are hardware (e.g., computing devices) and software architectures that may be deployed in example implementations.

[0291] In one example, a computing device may operate as a stand-alone device, or the computing device may be coupled (eg, networked) with other machines.

[0292] In a networked deployment, a computing device may operate as either a server or a client machine in a server-client network environment. In one example, a computing device may act as a peer machine in a peer-to-peer (or other distributed) network environment. A computing device may be a personal computer (PC), a tablet PC, a set-top box (STB), a mobile phone, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing instructions (sequential or otherwise) that specify operations to be taken (e.g., performed) by the computing device. Furthermore, while only a single computing device is illustrated, the term "computing device" shall be taken to include any collection of machines that individually or collectively execute a set (or sets) of instructions to implement any one or more of the methodologies discussed herein.

[0293] An example of a computing device may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory, and a static memory, some or all of which may communicate with each other via a bus. The computing device may further include a display unit, an alphanumeric input device (e.g., a keyboard), and a user interface (UI) navigation device (e.g., a mouse). In one example, the display unit, the input device, and the UI navigation device may be a touchscreen display. The computing device may further include a storage device (e.g., a drive unit), a signal generating device (e.g., a speaker), a network interface device, and one or more sensors, for example, a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor.

[0294] A storage device may include a machine-readable medium on which is stored one or more data structures or sets of instructions (e.g., software) that embody or are utilized in any one or more of the methodologies or functions described herein. The instructions may also reside, completely or at least partially, in main memory, static memory, or within the processor during execution by the computing device. In one example, one or any combination of the processor, main memory, static memory, or storage device may constitute a machine-readable medium.

[0295] While the machine-readable medium is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) configured to store one or more instructions. The term "machine-readable medium" may also be interpreted to encompass any tangible medium capable of storing, encoding, or carrying instructions for execution by a machine, causing a machine to perform any one or more of the methodologies of the present disclosure, or capable of storing, encoding, or carrying data structures utilized by or associated with such instructions. Accordingly, the term "machine-readable medium" may be interpreted to include, but is not limited to, solid-state memory, optical media, and magnetic media. Specific examples of machine-readable media may include, by way of example, semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) and flash memory devices); magnetic disks, e.g., internal hard disks and removable disks; magneto-optical disks; and non-volatile memory, including CD-ROM and DVD-ROM disks.

[0296] The instructions may further be transmitted or received over a communications network 828 using a transmission medium via the network interface device 822, utilizing any one of several transport protocols (e.g., Frame Relay, IP, TCP, UDP, HTTP, etc.). Examples of communications networks include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile phone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., the IEEE 802.11 family of standards known as Wi-Fi®, the IEEE 802.16 family of standards known as WiMax®), peer-to-peer (P2P) networks, among others. The term “transmission medium” shall be taken to encompass any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, including digital or analog communications signals or other intangible media that facilitate communication of such software.

[0297] As used herein, a component may refer to a device, a physical entity, or logic with boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that result in the division or modularization of specific processing or control functions. Components can be combined with other components through their interfaces to perform machine processes. A component may also be a packaged functional hardware unit designed to be used with other components, and part of a program that performs a specific function, usually related functionality. A component may constitute either a software component (e.g., code embodied in a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that can perform certain operations and can be configured or arranged in a certain physical manner. In various example implementations, one or more computer systems (e.g., stand-alone computer systems, client computer systems, or server computer systems), or one or more hardware components of a computer system (e.g., a processor or group of processors), may be configured by software (e.g., an application or part of an application) as hardware components that operate to perform certain operations described herein. [Example]

[0298] Example 1 method Following the techniques described herein, two rounds of dataset generation were used. Both datasets included patients with molecular algorithm results consisting of a comparison of the mean VAF from an initial Guardant 360 test with the mean VAF from a genomic test (e.g., Guardant G360) at a second time point (cit). For the first dataset, we queried a real-world evidence database containing aggregated medical claims and de-identified records from commercial insurance companies for over 225,000 individuals who underwent comprehensive ctDNA testing. Specifically, we retrospectively evaluated patients with NSCLC treated with immune checkpoint inhibitors (ICIs) (monotherapy or combination therapy) who underwent ctDNA testing within 15 weeks before treatment initiation and a second test 3–15 weeks after treatment initiation (Dataset 1). Example 2 Additional Methods

[0299] Next, we generated a second dataset containing additional, more recent records from the Real-World Evidence Database. Specifically, we retrospectively evaluated NSCLC patients treated with immune checkpoint inhibitors (ICIs) (monotherapy or combination therapy) who underwent ctDNA testing within 15 weeks before treatment initiation and a second test between 3 and 15 weeks after treatment initiation (Dataset 2; ICI-treated cohort). Furthermore, we retrospectively evaluated NSCLC patients treated with tyrosine kinase inhibitors (including immune checkpoint inhibitors, gefitinib, erlotinib, afatinib, and osimertinib) (monotherapy or combination therapy) who underwent ctDNA testing within 15 weeks before treatment initiation and a second test between 3 and 15 weeks after treatment initiation (Dataset 2; TKI-treated cohort).

[0300] Cox proportional hazards (CPH) was used to analyze time to next treatment (TTNT) and time to treatment discontinuation (TTD) in RW. Gender, age, line of treatment (LOT), and comorbidities were included as covariates. Example 3 Modeling and Analysis

[0301] For the first dataset, ctDNA change from baseline to treatment was modeled as a continuous variable using the regularized cubic spline function shown in Table 1 or a similar s...

Claims

1. receiving, by a computer system, genetic information of a subject, the genetic information including data obtained at two or more time points, wherein cancer has been detected in the subject; extracting from the genetic information one or more features identified from a plurality of samples obtained from the subject; generating, by a first classifier implemented on the computer system, a first machine learning algorithm to generate a first output indicative of a first classification of the object; generating, by a second classifier implemented on the computer system, a second output indicative of a second classification of the subject using a second machine learning algorithm; identifying, by the computer system, additional subjects from the population having genetic information that matches the genetic information of the subject based on the first classification and the second classification; determining, by the computer system, at least one score for the subject based on additional subjects with matching genetic information; determining, by the computer system, a composite score using the at least one score; determining, by a recommender implemented in the computer system, a treatment recommendation for the subject based on the composite score; A method comprising:

2. 2. The method of claim 1, wherein the one or more features comprise a first variant allele fraction (MAF) and a second MAF from two or more time points, respectively.

3. 3. The method of claim 2, wherein the at least one score is based on a first variant allele fraction (MAF) and a second MAF, a weighted average of the first MAF, and a weighted average of the second MAF.

4. The method of claim 3 , wherein the at least one score is based on a ratio of a weighted average of the first MAF to a weighted average of the second MAF, and a confidence interval.

5. 3. The method of claim 2, wherein the at least one score is based on a first variant allele fraction (MAF) at the first time point and a second MAF at the second time point, a first measure of central tendency for the first MAF, and a second measure of central tendency for the second MAF.

6. 6. The method of claim 5, wherein the at least one score is based on a ratio of the first measure of central tendency at the first time point to the second measure of central tendency at the second time point.

7. 7. The method of claim 6, wherein the measure of central tendency is one or more of the mean, median, or mode.

8. 10. The method of any one of the preceding claims, further comprising the step of comparing the molecular response score for the subject with the cancer to a predetermined cutoff point, and identifying the subject as a likely responder to one or more treatments for the cancer if the molecular response score is below the predetermined cutoff point, or identifying the subject as a likely non-responder to one or more treatments for the cancer if the molecular response score is at or above the predetermined cutoff point.

9. 10. The method of any one of the preceding claims, wherein the one or more treatments include one or more immunotherapies.

10. 10. The method of any one of the preceding claims, further comprising administering to the subject one or more treatments for the cancer taking into account the at least one score.

11. 10. The method of any one of the preceding claims, further comprising discontinuing administering one or more treatments for the cancer to the subject in consideration of the at least one score.

12. 10. The method of any one of the preceding claims, comprising using said at least one score as a prognostic and / or predictive biomarker for said subject.

13. 10. The method of any one of the preceding claims, comprising calculating the standard deviation of each MAF ratio in the set of MAF ratios using molecular counting.

14. 10. The method of any one of the preceding claims, comprising propagating the dispersion through each MAF ratio in the set of MAF ratios.

15. 10. The method of any one of the preceding claims, further comprising excluding one or more germline and / or clonal hematopoietic variants when determining the mutant allele frequency (MAF).

16. 10. The method of any one of the preceding claims, wherein the first time point comprises a time point before treatment and the second time point comprises a time point during or after treatment.

17. 10. The method of any one of the preceding claims, comprising generating sequence information from nucleic acid molecules obtained from one or more tissues or cells in the sample.

18. 10. The method of any one of the preceding claims, comprising generating sequence information from extracellular free nucleic acid (cfNA) in the sample obtained from the subject.

19. 10. The method of any one of the preceding claims, wherein the cfNA comprises circulating tumor DNA (ctDNA).

20. 10. The method according to any one of the preceding claims, wherein the first classifier and / or the second classifier are each implemented in a machine learning algorithm.

21. 10. The method of any one of the preceding claims, wherein the machine learning algorithm is selected from a neural network, a support vector machine, a hidden Markov model, or a random forest model.

22. 10. The method of any one of the preceding claims, wherein the at least one score corresponds to a level of responsiveness to treatment among a plurality of levels of responsiveness to treatment.

23. generating genetic information using a genetic analysis device; receiving into a computer memory, for each of a plurality of individuals having a cancer disease, a training dataset including: (1) genetic information from the individual generated at a first time point, and (2) a subsequent treatment response of the individual to one or more therapeutic interventions determined at a second time point; training a computer classifier using the training data set to obtain a trained computer classifier, the trained computer classifier comprising: determining at least one score for the subject based on the variant allele frequency (MAF) at least at the first time point and the second time point; Determine the amount of change between MAFs The steps are structured as follows: A method comprising:

24. 24. The method of claim 23, further comprising predicting the subject's treatment response based on the amount of variation between the MAFs.

25. 24. The method of claim 23, comprising selecting a treatment for the subject based on the amount of variation between the MAFs.

26. 10. Apparatus configured to perform the method according to any of the preceding claims.

27. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform the method of any of the preceding claims.

28. A system configured to perform the method according to any of the preceding claims.