Joint modeling of longitudinal data and time-to-event data to predict patient survival

Through the joint modeling of longitudinal data and time-to-event data, combined with hierarchical random effect model and proportional hazards model, the problem of early cancer detection is solved, accurate monitoring of cancer development and treatment response prediction is achieved, and diagnostic accuracy and sensitivity are improved.

CN120604294APending Publication Date: 2025-09-05GUARDANT HEALTH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202480007445.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-01-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively use liquid biopsy technology to detect cancer in the early stage, especially in the early stage of malignant diseases. There are a small number of distortions, confusions and a lack of understanding of the meaning of driving changes, resulting in difficulty in diagnosis.

Method used

A combined modeling method of longitudinal data and time-to-event data was used, combined with the time change analysis of biomarkers, and a hierarchical random effect model and proportional hazard model were used to measure the allelic frequency and tumor score of circulating tumor DNA (ctDNA), and a cubic spline model was generated, combining medical records and insurance record databases to predict patient responses.

Benefits of technology

Improves the accuracy and sensitivity of early cancer detection, enables monitoring of cancer development and treatment response, supports personalized treatment options, and predicts patient survival.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604294A_ABST
    Figure CN120604294A_ABST
Patent Text Reader

Abstract

Changes in ctDNA levels may significantly fluctuate across patients over time, and the results may be difficult to interpret. Methods and techniques are described herein that are capable of capturing these complexities while taking into account a distinct set of patient traits. Furthermore, an observation dataset consisting of a patient suffering from cancer undergoing treatment is analyzed, the analysis results being graphically rendered. These results demonstrate utility of the described methods and techniques in obtaining a comprehensive understanding of how response patterns evolve and how different patient characteristics affect these evolutions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 479,470 filed on January 11, 2023, U.S. Provisional Application Serial No. 63 / 496,765 filed on April 18, 2023, and U.S. Provisional Application Serial No. 63 / 612,218 filed on December 19, 2023, each of which is incorporated herein by reference in its entirety.

[0003] background

[0004] Today, knowledge of the molecular pathogenesis of cancer is increasing, and with the advent of next-generation sequencing technologies, the potential to study early molecular alterations in cancer development is also increasing. This includes liquid biopsies in body fluids. Genetic and epigenetic alterations associated with cancer development can be found in cell-free DNA (cfDNA), such as those in plasma, serum, urine, etc., with the potential to be used as diagnostic biomarkers. Non-invasive sampling methods promote patient compliance because they are easier, faster, and more economical to perform.

[0005] Such liquid biopsy technology supports the characterization of the genomic composition of different tissues in the subject. Although usually released by all types of cells, cfDNA can be derived from necrotic cells or apoptotic cells for the identification of specific tumor-related changes, such as mutations, methylation, and copy number variations (CNVs). In view of the need to distinguish between signals derived from disease tissues (such as cancer) and signals derived from germline cells that release cfDNA in a wider range of tissues (such as healthy tissues and white blood cells undergoing hematopoiesis), the improved characterization of this circulating tumor DNA (ctDNA) is challenging. Signals can be enriched by identifying variant alleles with an allele fraction that does not comply with the exemplary 1: 1 ratio of heterozygous alleles in the germline.

[0006] Despite these advances, the use of cfDNA for diagnostics has largely focused on advanced tumor stages, with little understanding of the characteristics of early malignant disease stages. However, there are several obstacles to early detection, including a smaller number of aberrations, confounding phenomena such as clonal non-tumor tissue expansion, adventitious cancer-associated mutations, and a lack of understanding of the significance of driver changes. Therefore, there is a great need in the art for improved technologies for characterizing early disease stages to support the development of cfDNA- and ctDNA-related diagnostics.

[0007] Described herein is the use of detection measurements that can include multiple parameters, including longitudinal data and time-to-event data, to support understanding how temporal changes in biomarkers are related to time-to-event responses and patient outcomes. For example, the methods and techniques described herein combine longitudinal data support and time-to-event data support to decipher temporal changes in biomarkers associated with time-to-event responses. In addition, the methods and techniques described herein allow for assessment of patient characteristics, such as age, sex, etc., in the analysis. Repeated measurements via liquid biopsy provide an opportunity to assess patient outcomes. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 Distribution of allele frequencies and tumor fractions. Figure 1 A. Depiction of EGFR L858R allele frequency and tumor fraction. Figure 1 B. Depiction of allele frequencies and logarithmic transformations of EGFR L858R, KRAS G12D, and KRAS G12V.

[0010] Figure 2 .Spaghetti plot of allele frequency and tumor percentage. Figure 2 A. Spaghetti plot of EGFR L858R allele frequency and tumor percentage. Figure 2 B. Spaghetti plot of allele frequencies and log-transformation of EGFR L858R, KRAS G12D, and KRAS G12V.

[0011] Figure 3 . GLMM results based on fitted cubic splines for log-transformed biomarkers of EGFR L858R.

[0012] Figure 4 Biomarker evolution and corresponding survival curves: The biomarker EGFR L858R evolution of patients at 300 days, 600 days, and 900 days is shown.

[0013] Figure 5 Biomarker evolution and corresponding survival curves: The biomarker EGFR L858R evolution of patients at 300 days, 600 days, and 900 days is shown.

[0014] Figure 6 Biomarker evolution and corresponding survival curves. The biomarker KRAS G12V evolution of patients at 300 days, 600 days, and 900 days is shown.

[0015] Figure 7 Time-to-event submodel: Overall survival. Delineation of EGFR L858R, KRAS G12D, and KRAS G12V.

[0016] Figure 8 Random-effects modeling. Delineation of EGFR L858R, KRAS G12D, and KRAS G12V.

[0017] Figure 9 Distribution of ctDNA levels and logit-transformed ctDNA levels and corresponding spaghetti plots.

[0018] Figure 10 Unconditional model fits with and without data points for the non-small cell lung cancer (NSCLC) cohort. Unconditional model fits with and without data points for the NSCLC cohort. The black curve represents the response pattern for the cohort, while each black dot represents a ctDNA level value. The purple area represents the 95% confidence band for the estimated trajectory.

[0019] Figure 11 Response patterns for different baseline age values ​​and ELIX scores: For female non-smokers receiving first-line anti-EGFR therapy, Figure 11 A. Surviving patients and Figure 11 B. Deceased patients.

[0020] Figure 12 . For the following velocity (IRC) graphs with different baseline age values ​​and ELIX scores: For female non-smokers receiving first-line anti-EGFR therapy, Figure 12 A. Surviving patients and Figure 12 B. Deceased patients. SUMMARY OF THE INVENTION

[0022] Described herein is a method for determining a patient response in at least one patient, comprising: obtaining nucleic acid sequence information from the at least one patient, including a measurement of temporal variation in a biomarker; and determining a patient response in the at least one patient. In various embodiments, the biomarker comprises ctDNA. In various embodiments, the biomarker comprises allele frequency and tumor fraction. In various embodiments, the method comprises determining a patient response in the at least one patient, comprising the use of a database. In various embodiments, the method comprises a database comprising medical records and / or insurance records. In various embodiments, the method comprises the use of a database, comprising the application of a model. In various embodiments, the model is a hierarchical model. In various embodiments, the model is an effect model. In various embodiments, the model is a regression model. In various embodiments, the model is a joint model. In various embodiments, the hierarchical model is a hierarchical random effects model. In various embodiments, the model comprises cubic splines. In various embodiments, the model comprises a regression model. In various embodiments, the hierarchical random effects model comprises generating data from nucleic acid sequence information, wherein the nucleic acid sequence information comprises temporal variation in a biomarker comprising circulating tumor DNA (ctDNA) from at least one subject of more than one subject. In various embodiments, the generation of data includes generating cubic splines for at least one subject in more than one subject. In various embodiments, the generation of data includes generating a response parameter including one or more covariates. In various embodiments, the generation of data includes generating a response parameter without covariates. In various embodiments, the response parameter applies a multivariate normal distribution. In various embodiments, the method includes determining a patient response for at least one patient, including generating a velocity map. In various embodiments, the method includes determining a patient response for at least one patient, including comparing with a model. In various embodiments, the joint model includes at least two models. In various embodiments, the joint model includes an association factor between at least two models. In various embodiments, the joint model includes cubic splines and a proportional hazards model. In various embodiments, biomarkers are measured using next-generation DNA sequencing. In various embodiments, next-generation DNA sequencing includes attaching a non-unique barcode to ctDNA. In various embodiments, next-generation DNA sequencing includes attaching a unique barcode to ctDNA. In various embodiments, next-generation DNA sequencing includes attaching a non-unique barcode to ctDNA fragments, wherein the non-unique barcode is present in at least 20x, at least 30x, at least 50x, or at least 100x molar excess.

[0023] A system comprising a machine comprising at least one processor and memory comprising instructions capable of performing any of the foregoing methods. A computer-readable medium comprising instructions capable of performing any of the foregoing methods.

[0024] Described herein is a method for determining a patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from at least one patient, the method comprising measuring the time variation of a biomarker comprising circulating tumor DNA (ctDNA); and determining the patient response of at least one patient, comprising using a database comprising medical records and / or insurance records from more than one subject, wherein the use of the database comprises applying a hierarchical random effects model. In various embodiments, the hierarchical random effects model comprises generating data from nucleic acid sequence information, the nucleic acid sequence information comprising the time variation of ctDNA from at least one subject in more than one subject. In various embodiments, the hierarchical random effects model comprises generating cubic splines for at least one subject in more than one subject. In various embodiments, the hierarchical random effects model comprises a response parameter comprising one or more covariates of at least one subject in more than one subject. In various embodiments, the database comprises medical records and / or insurance records of more than one subject. Described herein is a system comprising a machine comprising at least one processor and memory comprising instructions for executing a method for determining a patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from at least one patient, comprising measuring temporal changes in biomarkers comprising circulating tumor DNA (ctDNA); and determining a patient response in at least one patient, comprising using a database comprising medical records and / or insurance records from more than one subject, wherein the use of the database comprises applying a hierarchical random effects model. In various embodiments, the hierarchical random effects model comprises generating data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject in more than one subject. In various embodiments, the hierarchical random effects model comprises generating cubic splines for at least one subject in more than one subject. In various embodiments, the hierarchical random effects model comprises a response parameter comprising one or more covariates of at least one subject in more than one subject. In various embodiments, the database comprises medical records and / or insurance records of more than one subject. Described herein is a computer-readable medium comprising instructions capable of executing a method for determining a patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from the at least one patient, comprising measuring temporal variation of a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response in the at least one patient, comprising using a database comprising medical records and / or insurance records from more than one subject, wherein using the database comprises applying a hierarchical random effects model. In various embodiments, the hierarchical random effects model comprises generating data from nucleic acid sequence information comprising temporal variation of ctDNA from at least one of the more than one subject.In various embodiments, the hierarchical random effects model comprises generating a cubic spline for at least one of the more than one subjects. In various embodiments, the hierarchical random effects model comprises a response parameter comprising one or more covariates for at least one of the more than one subjects. In various embodiments, the database comprises medical records and / or insurance records for the more than one subject.

[0025] Described herein is a method for determining a patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from at least one patient, comprising measuring temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response in at least one patient, comprising using a database comprising medical records and / or insurance records from more than one subject, wherein using the database comprises applying a joint model comprising a cubic spline and a proportional hazards model generated from data of nucleic acid sequence information from at least one subject in the more than one subject. In various embodiments, the database comprises medical records and / or insurance records from more than one subject. Described herein is a system comprising a machine comprising at least one processor and a memory comprising instructions capable of executing a method for determining a patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from at least one patient, comprising measuring temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response in at least one patient, comprising using a database comprising medical records and / or insurance records from more than one subject, wherein using the database comprises applying a joint model comprising a cubic spline and a proportional hazards model generated from data of nucleic acid sequence information from at least one subject in the more than one subject. In various embodiments, the database includes medical records and / or insurance records of more than one subject. A computer-readable medium is described herein, comprising instructions for executing a method for determining a patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from at least one patient, comprising measuring temporal changes in biomarkers comprising circulating tumor DNA (ctDNA); and determining a patient response in at least one patient, comprising using a database comprising medical records and / or insurance records from more than one subject, wherein the use of the database comprises the application of a joint model comprising a cubic spline and a proportional hazards model generated from data of nucleic acid sequence information from at least one of the more than one subjects. In various embodiments, the database includes medical records and / or insurance records of more than one subject.

[0026] Details

[0027] analyze

[0028] The method of the present invention can be used to diagnose the presence of a condition, particularly cancer, in a subject, to characterize the condition (e.g., to stage the cancer or determine the heterogeneity of the cancer), monitor the response of the condition to treatment, and achieve a prognosis for the risk of the condition developing or the subsequent course of the condition. The present disclosure can also be used to determine the efficacy of a specific treatment option. If the treatment is successful, the successful treatment option may increase the amount of copy number variation or rare mutations detected in the subject's blood as more cancers may die and DNA is shed. In other examples, this may not occur. In another example, perhaps certain treatment options may be associated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if it is observed that the cancer is in remission after treatment, the method of the present invention can be used to monitor residual disease or the recurrence of the disease.

[0029] The types and numbers of cancers that can be detected may include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc. The type and / or stage of cancer can be detected based on genetic variations, including mutations, rare mutations, insertions / deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosome fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0030] Genetic and other analyte data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data can allow characterization of specific subtypes of cancer, which can be important in diagnosing or treating that specific subtype. This information can also provide clues to the prognosis of a specific type of cancer to a subject or practitioner, and allow the subject or practitioner to adjust treatment options based on the progression of the disease. Some cancers can progress and become more aggressive and genetically unstable. Other cancers can remain benign, inactive, or dormant. The systems and methods of the present disclosure can be used to determine disease progression.

[0031] The analysis of the present invention can also be used to determine the effectiveness of a specific treatment selection. If the treatment is successful, the successful treatment selection may increase the amount of copy number variation or rare mutation detected in the subject's blood as more cancer may die and DNA is shed. In other examples, this may not happen. In another example, perhaps certain treatment options may be related to the genetic profile of the cancer over time. This correlation can be used to select therapy. Additionally, if it is observed that the cancer is in remission after treatment, the method of the present invention can be used to monitor residual disease or the recurrence of the disease.

[0032] The method of the present invention can also be used to detect genetic variations in conditions other than cancer. After certain diseases occur, immune cells, such as B cells, can undergo rapid clonal expansion. Copy number variation detection can be used to monitor clonal expansion, and certain immune states can be monitored. In this example, copy number variation analysis can be performed over time to produce a spectrum of how a specific disease may progress. Copy number variation or even rare mutation detection can be used to determine how pathogen populations change during the course of infection. This may be particularly important during chronic infection (such as HIV / AID or hepatitis infection), whereby the virus can change life cycle state and / or mutate into a more virulent form during the course of infection. When immune cells attempt to destroy transplanted tissue, the method of the present invention can be used to determine or analyze the rejection activity of the host body, to monitor the state of transplanted tissue and to change the process of rejection treatment or prevention.

[0033] For example, when an individual is experiencing stress, many types of dysfunction and abnormalities that commonly occur in the cardiovascular system (which are not diagnosed or treated) will gradually reduce the body's ability to supply enough oxygen to meet coronary oxygen demand. The progressive decline in the cardiovascular system's ability to supply oxygen under stress conditions will ultimately culminate in a heart attack, i.e., a myocardial infarction event caused by an interruption in blood flow through the heart, leading to oxygen starvation of the myocardial tissue (i.e., the myocardium). In many cases, permanent damage will occur to the cells that make up the myocardium, which will subsequently predispose the individual to additional myocardial infarction events.

[0034] The methods of the present disclosure can characterize dysfunction and abnormalities associated with myocardial and valvular tissue (e.g., hypertrophy), a reduction in blood flow and oxygen supply to the heart that is typically a secondary symptom of weakening and / or deterioration of the blood flow and supply systems caused by physical and biochemical stress. Examples of cardiovascular diseases that are directly affected by these types of stress include atherosclerosis, coronary artery disease, peripheral vascular disease, and peripheral arterial disease, as well as various heart diseases and arrhythmias that may represent other forms of disease and dysfunction.

[0035] In addition, the methods of the present disclosure can be used to characterize the heterogeneity of abnormal conditions in a subject. Such methods can include, for example, generating a genetic profile of extracellular polynucleotides derived from a subject, wherein the genetic profile includes more than one data point obtained by copy number variation and rare mutation analysis. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can be a condition that leads to a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of cancer. In other examples, heterogeneity can include multiple lesions of the disease. Again, in the example of cancer, there can be multiple tumor lesions, perhaps one or more of which are the result of metastasis that has spread from the primary site.

[0036] The methods of the present invention can be used to generate or analyze a fingerprint or dataset that is the sum of genetic information from different cells in a heterogeneous disease. The dataset can include copy number variation and mutation analysis alone or in combination.

[0037] The methods of the present invention can be used to diagnose, prognose, monitor or observe cancer or other diseases. In some embodiments, the methods herein do not involve diagnosis, prognosis or monitoring of a fetus and, therefore, do not involve non-invasive prenatal testing. In other embodiments, these methods can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in unborn subjects, whose DNA and other polynucleotides can co-circulate with maternal molecules.

[0038] Modified Nucleic Acid Analysis Methods

[0039] The present disclosure provides an optional method for analyzing modified nucleic acids (e.g., methylated, connected to histones, and other modifications discussed above). In some such methods, a nucleic acid population with varying degrees of modification (e.g., each nucleic acid molecule has 0, 1, 2, 3, 4, 5, or more methyl groups) is contacted with an adapter, and then the population is graded according to the degree of modification. The adapter is attached to one or both ends of the nucleic acid molecule in the population. Preferably, the adapter comprises a sufficient number of different tags so that the number of tag combinations causes the probability of two nucleic acids with the same starting and ending points receiving different tag combinations is high, e.g., 95%, 99%, or 99.9%. After attaching the adapter, the nucleic acid is amplified from the primers that bind to the primer binding sites in the adapter. Regardless of whether the adapters with the same or different tags can comprise the same or different primer binding sites, it is preferred that the adapters comprise the same primer binding sites. After amplification, the nucleic acid is contacted with an agent (such as the agent previously described) that is preferably bound to the nucleic acid with the modification. The nucleic acid is divided into at least two partitions, and the difference between at least two partitions is that the degree of binding of the nucleic acid with the modification to the agent is different. For example, if the agent has affinity to the nucleic acid with modification, then the modification is preferentially combined with the agent by the nucleic acid (compared with the median representation in the colony) that is overrepresented, and the nucleic acid that modification is not fully represented does not bind the agent or is more easily eluted from the agent.After separation, different partitions can then undergo other processing steps, which generally include parallel but independent other amplification and sequence analysis.Then the sequence data from different partitions can be compared.

[0040] Both ends of the nucleic acid can be connected to a Y-shaped adapter containing a primer binding site and a tag. The molecule is amplified. The amplified molecule is then partitioned by contact with an antibody that preferentially binds to 5-methylcytosine to produce two partitions. One partition contains the original molecule that lacks methylation and the amplified copy that has lost methylation. The other partition contains the original DNA molecule with methylation. The two partitions are then processed and sequenced separately, and the methylated partition is further amplified. The sequence data of the two partitions can then be compared. In this example, the tag is not used to distinguish between methylated DNA and unmethylated DNA, but to distinguish between different molecules in these partitions, so that people can determine whether reads with the same start and end points are based on the same or different molecules.

[0041] The present disclosure also provides methods for analyzing nucleic acid populations, wherein at least some of the nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, the nucleic acid population is contacted with an adapter comprising one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably, all cytosine residues in such an adapter are also modified, or all such cytosines in the primer binding region of the adapter are modified. The adapter is attached to both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags so that the number of tag combinations results in a high probability, for example, 95%, 99% or 99.9%, that two nucleic acids with the same starting and ending points receive different tag combinations. The primer binding sites in such an adapter can be the same or different, but preferably the same. After the adapter is attached, the nucleic acid is amplified by primers that bind to the primer binding sites of the adapter. The amplified nucleic acid is divided into a first aliquot and a second aliquot. Sequence data is determined for the first aliquot with or without further processing. Determine thus the sequence data of the molecule in the first aliquot sample and regardless of the initial methylation state of the nucleic acid molecule. The nucleic acid molecule in the second aliquot sample is treated with bisulfite. This process converts unmodified cytosine into uracil. The nucleic acid experience amplification of bisulfite treatment then, and this amplification is caused by the primer for the original primer binding site of the adapter connected to the nucleic acid. Now only the nucleic acid molecule (different from its amplification product) that is initially connected to the adapter is amplified, because these nucleic acids retain cytosine in the primer binding site of the adapter, and the amplification product has lost the methylation of these cytosine residues, and these cytosine residues that lose methylation have been converted into uracil in bisulfite treatment. Therefore, only the original molecule in the colony (at least some of them are methylated) is experienced amplification. After amplification, these nucleic acid experiences sequence analysis. The comparison of the sequence determined from the first aliquot sample and the second aliquot sample can especially indicate which cytosine has experienced methylation in the nucleic acid colony.

[0042] Partitioning a sample into more than one subsample; aspects of a sample; analysis of epigenetic signatures

[0043] In certain embodiments described herein, different forms of nucleic acid populations (e.g., hypermethylated DNA and hypomethylated DNA in a sample, such as a capture group of cfDNA as described herein) can be physically partitioned based on one or more features of the nucleic acid and then further analyzed, for example, differentially modified or separated nucleobases, labeled, and / or sequenced. This method can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, hypermethylated variable epigenetic target regions are analyzed to determine whether they show the hypermethylation characteristics of tumor cells, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they show the hypomethylation characteristics of tumor cells. In addition, by partitioning heterogeneous nucleic acid populations, one can increase rare signals, for example, by enriching for rare nucleic acid molecules that are more prevalent in a fraction (or a partition) of a population. For example, by partitioning a sample into hypermethylated and hypomethylated nucleic acid molecules, genetic variations that are present in hypermethylated DNA but are less (or absent) in hypomethylated DNA can be more easily detected. By analyzing more than one fraction of a sample, multidimensional analysis of individual loci or nucleic acid species of the genome can be performed and, therefore, greater sensitivity can be achieved.

[0044] In some cases, a heterogeneous nucleic acid sample is partitioned into two or more partitions (e.g., at least 3, 4, 5, 6, or 7 partitions). In some embodiments, each partition is differentially labeled. The labeled partitions can then be pooled together for collective sample preparation and / or sequencing. The partition-labeling-pooling step can occur more than once, with each round of partitioning occurring based on a different feature (examples provided herein) and labeled using a differential label that is distinct from other partitions and partitioning methods.

[0045] In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention.In some embodiments, the partition of the present invention can be used to partition the nucleic acid of the present invention. Alternatively or additionally, the heterogeneous nucleic acid population can be partitioned into the nucleic acid molecules relevant to nucleosomes and the nucleic acid molecules not containing nucleosomes. Alternatively or additionally, the heterogeneous nucleic acid population can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively or additionally, the heterogeneous nucleic acid population can be partitioned based on nucleic acid length (for example, a molecule of maximum 160bp and a molecule with a length greater than 160bp).

[0046] In some cases, each partition (representing different nucleic acid forms) difference labeling is performed, and the partitions are brought together and then sequenced. In other cases, different forms are sequenced separately. In some embodiments, different nucleic acid populations are partitioned into two or more different partitions. Each partition represents different nucleic acid forms, and the first partition (also referred to as subsample) includes a DNA modified with cytosine that has a larger ratio than the second subsample. Each partition is labeled differently. The first subsample is subjected to a program that differently affects the first core base in the DNA of the first subsample and the second core base in the DNA, wherein the first core base is a modified or unmodified core base, and the second core base is a modified or unmodified core base different from the first core base, and the first core base and the second core base have identical base pairing specificity. The labeled nucleic acids are brought together and then sequenced. Sequence reads are obtained and analyzed, including computer simulation (in silico) distinguishing the first core base and the second core base in the DNA of the first subsample. Labels are used to sort the reads from different partitions. Analyze to detect genetic variation at the level of each partition and at the whole nucleic acid population level. For example, analysis can include computer simulation analysis to determine genetic variation, such as CNV, SNV, insertion / deletion, fusion in the nucleic acid of each partition. In some cases, computer simulation analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the positioning of nucleosomes in chromatin. Higher coverage can be associated with higher nucleosome occupancy in the genomic region, while lower coverage can be associated with lower nucleosome occupancy or nucleosome depleted region (NDR).

[0047] The sample may include nucleic acids of various modifications, including post-replicative modifications to the nucleotides and association (usually non-covalently) with one or more proteins.

[0048] In embodiments, nucleic acid colony is the nucleic acid colony obtained from serum, plasma or blood sample of the experimenter suspected of having vegetation, tumor or cancer or previously diagnosed as having vegetation, tumor or cancer.Nucleic acid colony includes nucleic acid with different methylation levels.Methylation can be by any one or more replication or post-transcription modification occurs.Replication modification includes modification to nucleotide cytosine, particularly at the 5-position of core base, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine.Affinity agent can be an antibody with desired specificity, natural binding partner or its variant (Bock et al., Nat Biotech 28:1106-1114 (2010); Song et al., Nat Biotech 29:68-72 (2011)), or artificial peptides with specificity to a particular target, such as selected by phage display.

[0049] Examples of capture moieties contemplated herein include methyl binding domains (MBDs) and methyl binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies that preferentially bind to 5-methylcytosine. Similarly, partitioning of nucleic acids of different forms can be performed using histone binding proteins that can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. For some affinity agents and modifications, although binding to the agent can occur in a substantially all-or-nothing manner depending on whether the nucleic acid is modified, separation can be to a certain extent. In such cases, nucleic acids that are overrepresented in a modification bind to the agent to a greater extent than nucleic acids that are underrepresented in the modification. Alternatively, nucleic acids with modifications can be bound in an all-or-nothing manner. However, then, various levels of modification can be sequentially eluted from the binding agent.

[0050] For example, in some embodiments, partition can be binary or based on the degree / level of modification.For example, methyl binding domain protein (such as MethylMiner methylated DNA enrichment kit (ThermoFisherScientific)) can be used to partition all methylated fragments with unmethylated fragments. Subsequently, other partitions can include eluting fragments with different methylation levels by adjusting the salt concentration of the solution containing methyl binding domain and binding fragment. Along with salt concentration increases, the fragment with larger methylation level is eluted. In some cases, final partition represents the nucleic acid with varying degrees of modification (modification of excessive representation (over representative) or insufficient representation (underrepresentative)). Excessive representation and insufficient representation can be defined by the median of the modification of the number of the modification that nucleic acid is carried with respect to every chain in the colony. For example, if the median of 5-methylcytosine residues is 2 in the nucleic acid in the sample, then the modification comprising more than two 5-methylcytosine residues is excessive representation, and the nucleic acid with 1 or 0 5-methylcytosine residues is insufficient representation. The effect of affinity separation is to enrich nucleic acids whose modifications are over-represented in the bound phase and nucleic acids whose modifications are under-represented in the unbound phase (i.e., in solution). The nucleic acids in the bound phase can be eluted before subsequent processing.

[0051] When using MethylMiner methylated DNA enrichment test kit (ThermoFisher Scientific), sequential elution can be used by the methylation partition of different levels.For example, by making nucleic acid colony contact with the MBD being attached to magnetic bead from test kit, low methylation partition (for example, without methylation) is separated from methylation partition.Pearl is used for separating methylated nucleic acid from unmethylated nucleic acid.Subsequently, one or more elution steps are carried out in sequence to elute the nucleic acid with different methylation levels.For example, first group of methylated nucleic acid can be eluted at 160mM or higher salt concentration, for example, at least 150mM, at least 200mM, at least 300mM, at least 400mM, at least 500mM, at least 600mM, at least 700mM, at least 800mM, at least 900mM, at least 1000mM or at least 2000mM.After such methylated nucleic acid is eluted, magnetic separation is used again for the methylated nucleic acid of higher level and the nucleic acid separation with lower methylation level. The elution and magnetic separation steps themselves can be repeated to generate various partitions, such as a hypomethylated partition (representing no methylation), a methylated partition (representing low methylation levels), and a hypermethylated partition (representing high methylation levels).

[0052] In some methods, the nucleic acid bound to the agent for affinity separation undergoes a washing step. The washing step washes away the nucleic acid weakly bound to the affinity agent. Such nucleic acid can be enriched with a modified nucleic acid close to the mean or median (that is, the intermediate value between the nucleic acid bound to the solid phase and the nucleic acid not bound to the solid phase when the sample is initially in contact with the agent). Affinity separation results in at least two and sometimes three or more partitions with nucleic acids of different modification degrees. When the partition is still separated, the nucleic acid of at least one partition and usually two or three (or more) partitions is connected to a nucleic acid tag, which is generally provided as a component of an adapter, and the nucleic acid in different partitions receives different tags that distinguish the member of a partition from the member of another partition. The tags connected to the nucleic acid molecules of the same partition can be the same or different from each other. But if different from each other, the tag can have a part of the common coding, so that the molecules to which they are attached are identified as belonging to a specific partition. For more details about partitioning nucleic acid samples based on features such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some embodiments, nucleic acid molecules can be fractionated into different partitions based on nucleic acid molecules that bind to a specific protein or fragment thereof and nucleic acid molecules that do not bind to the specific protein or fragment thereof.

[0053] Nucleic acid molecules can be separated by fractionation based on DNA-protein combination. Protein-DNA complex can be separated by fractionation based on the specific characteristics of protein. The example of such characteristic includes various epitopes, modification (for example, histone methylation or acetylation) or enzymatic activity. The example of the protein that can be used as the basis for fractionation in conjunction with DNA can include but is not limited to protein A and protein G. Any suitable method can be used for separating nucleic acid molecules by fractionation based on protein binding region. The example for separating nucleic acid molecules by fractionation based on protein binding region includes but is not limited to SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography and asymmetric field flow fractionation (AF4).

[0054] In some embodiments, nucleic acid partitioning is performed by contacting the nucleic acid with a methylation binding domain ("MBD") of a methylation binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is attached to paramagnetic beads (such as Partitioning into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentrations.

[0055] An exemplary method for molecular tag identification of MBD bead-partitioned libraries by NGS is as follows:

[0056] Extracted DNA samples (eg, plasma DNA extracted from human samples) are physically partitioned using a methyl-binding domain protein-bead purification kit, retaining all eluates from the process for downstream processing.

[0057] Differential molecular tags and NGS actionable adapter sequences were applied to each partition in parallel. For example, a high methylation partition, a residual methylation ('wash') partition, and a low methylation partition were ligated with NGS adapters carrying molecular tags.

[0058] All molecularly tagged partitions were reassembled and subsequently amplified using adapter-specific DNA primer sequences.

[0059] The recombined and amplified total library is enriched / hybridized to target genomic regions of interest (e.g., cancer-specific genetic variations and differentially methylated regions).

[0060] The enriched total DNA library is re-amplified, sample tags are added, and different samples are pooled and multiplexed on an NGS instrument.

[0061] Bioinformatics analysis of NGS data uses molecular signatures to identify unique molecules and deconvolute samples into molecules with differential MBD partitioning. This analysis can generate information on relative 5-methylcytosine levels across genomic regions simultaneously with standard sequencing / variant detection.

[0062] Examples of MBPs contemplated herein include, but are not limited to:

[0063] (a) MeCP2, a protein that preferentially binds to 5-methylcytosine over unmodified cytosine.

[0064] (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethyl-cytosine compared to unmodified cytosine.

[0065] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferentially bind to 5-formyl-cytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14:R119 (2013)).

[0066] (d) Antibodies specific for one or more methylated nucleotide bases.

[0067] Typically, elution varies with the number of methylation sites per molecule, with more methylated molecules eluting at increasing salt concentrations. To elute DNA into different populations based on the degree of methylation, one can use a series of elution buffers with increasing NaCl concentrations. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, this process results in three (3) partitions. The molecules are contacted with a solution of a first salt concentration, and the solution contains molecules containing a methyl binding domain that can be attached to a capture moiety such as streptavidin. At the first salt concentration, one population of molecules will bind to the MBD, and one population will remain unbound. The unbound population can be separated into a "hypomethylated" population. For example, a first partition representing a hypomethylated form of DNA is a partition that remains unbound at a low salt concentration (e.g., 100 mM or 160 mM). A second partition representing moderately methylated DNA is eluted using a moderate salt concentration (e.g., a concentration between 100 mM and 2000 mM). This is also separated from the sample. A third fraction, representing hypermethylated forms of DNA, is eluted using a high salt concentration (eg, at least about 2000 mM).

[0068] The present disclosure also provides a method for analyzing a nucleic acid population, wherein at least some of the nucleic acids comprise one or more modified cytosine residues, such as 5-methylcytosine and any other modification described previously. In these methods, after partitioning, a nucleic acid subsample is contacted with an adapter comprising one or more cytosine residues modified at the 5C position (such as 5-methylcytosine). Preferably, all cytosine residues in such an adapter are also modified, or all such cytosines in the primer binding region of the adapter are modified. The adapter is attached to both ends of the nucleic acid molecule in the population. Preferably, the adapter comprises a sufficient number of different tags so that the number of tag combinations results in a high probability, such as 95%, 99% or 99.9%, of two nucleic acids having the same starting and ending points receiving different tag combinations. The primer binding sites in such an adapter can be the same or different, but preferably the same. After the adapter is attached, the nucleic acid is amplified by a primer that is bound to the primer binding site of the adapter. The amplified nucleic acid is divided into a first aliquot and a second aliquot. In the case of carrying out or not carrying out further processing, sequence data is measured to the first aliquot. Thus determine the sequence data of molecules in the first aliquot without regard to the initial methylation state of nucleic acid molecules. The nucleic acid molecules in the second aliquot experience differently affect the program of the first core base in DNA and the second core base in DNA, wherein the first core base includes the cytosine modified at position 5, and the second core base includes unmodified cytosine. This program can be another program of bisulfite treatment or unmodified cytosine converted into uracil. The nucleic acid then undergoes this program with a primer amplification for the original primer binding site of the adapter connected to nucleic acid. Now only the nucleic acid molecules (different from its amplification product) initially connected to the adapter are amplified, because these nucleic acids retain cytosine in the primer binding site of the adapter, and the amplification product has lost the methylation of these cytosine residues, and these cytosine residues that lose methylation have undergone conversion into uracil in bisulfite treatment. Therefore, only the original molecule in the colony (at least some of them are methylated) experiences amplification. After amplification, these nucleic acid experiences sequence analysis. Comparison of the sequences determined from the first aliquot and the second aliquot can indicate, among other things, which cytosines in the nucleic acid population have undergone methylation.

[0069] Such analysis can be carried out using the following exemplary program. After partitioning, the two ends of the methylated DNA are connected to the Y-shaped adapter containing primer binding site and label. The cytosine in the adapter is modified (e.g., 5-methylated) at position 5. The modification of the adapter is used to protect the primer binding site in subsequent conversion steps (e.g., bisulfite treatment, TAP conversion or any other conversion that does not affect modified cytosine but affects unmodified cytosine). After the adapter is attached, the DNA molecule is amplified. The amplified product is divided into two aliquots for conversion and sequencing without conversion. The aliquots that have not undergone conversion can be subjected to sequence analysis when further processed or not. Another aliquot experience differently affects the program of the first core base in DNA and the second core base in DNA, wherein the first core base includes the cytosine modified at position 5, and the second core base includes unmodified cytosine. The program can be another program of bisulfite treatment or unmodified cytosine converted into uracil. When contacted with primers specific to the original primer binding site, only primer binding sites protected by cytosine modification can support amplification. Therefore, only the original molecule, rather than the copy from the first amplification, undergoes further amplification. The further amplified molecule is then subjected to sequence analysis. The sequences from the two aliquots can then be compared. As in the separation scheme discussed above, the nucleic acid tags in the adapter are not used to distinguish between methylated DNA and unmethylated DNA, but are used to distinguish nucleic acid molecules within the same partition.

[0070] Subjecting the first subsample to a process that differentially affects a first nucleobase in the DNA and a second nucleobase in the DNA of the first subsample Nucleobase Program

[0071] Methods disclosed herein include subjecting a first subsample to a procedure that differently affects a first nucleobase in the DNA of the first subsample and a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, if the first nucleobase is modified or unmodified adenine, the second nucleobase is modified or unmodified adenine; if the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine; if the first nucleobase is modified or unmodified guanine, the second nucleobase is modified or unmodified guanine; if the first nucleobase is modified or unmodified thymine, the second nucleobase is modified or unmodified thymine (wherein for the purposes of this step, modified and unmodified uracil are included in modified thymine).

[0072] In some embodiments, the first core base is a modified or unmodified cytosine, and then the second core base is a modified or unmodified cytosine. For example, the first core base can include unmodified cytosine (C), and the second core base can include one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second core base can include C, and the first core base can include one or more of mC and hmC. Other combinations are also possible, as indicated in the above overview and the following discussion, such as wherein the first core base and the second core base include mC and another include hmC.

[0073] In some embodiments, the program of the first core base in the DNA of the first subsample and the second core base in the DNA comprises bisulfite conversion. Bisulfite treatment converts unmodified cytosine and the cytosine nucleotides of certain modifications (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) into uracil, while other modified cytosines (e.g., 5-methylcytosine and 5-hydroxymethylcytosine) are not converted. Therefore, when bisulfite conversion is used, the first core base comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine or other cytosine forms affected by bisulfite, and the second core base can comprise one or more of mC and hmC, such as mC and optionally hmC. The position of cytosine read as is identified as mC position or hmC position by sequencing of bisulfite-treated DNA. At the same time, positions that read as T are identified as T or bisulfite-susceptible forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Thus, bisulfite conversion of the first subsample as described herein facilitates identification of positions containing mC or hmC using sequence reads obtained from the first subsample. For an exemplary description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.

[0074] In some embodiments, the program of the first core base in the DNA of the first subsample and the second core base in the DNA differently affects comprises oxidation bisulfite (Ox-BS) conversion.In some embodiments, the program of the first core base in the DNA of the first subsample and the second core base in the DNA differently affects comprises TET auxiliary bisulfite (TAB) conversion.In some embodiments, the program of the first core base in the DNA of the first subsample and the second core base in the DNA differently affects comprises Tet auxiliary substituted borane reducing agent conversion, optionally wherein substituted borane reducing agent is 2-picoline borane, pyridine borane, tert-butylamine borane or ammonia borane.In some embodiments, the program of the first core base in the DNA of the first subsample and the second core base in the DNA differently affects comprises chemical auxiliary substituted borane reducing agent conversion, optionally wherein substituted borane reducing agent is 2-picoline borane, pyridine borane, tert-butylamine borane or ammonia borane. In some embodiments, the procedure that differentially affects a first nucleobase in the DNA and a second nucleobase in the DNA of the first subsample comprises APOBEC-coupled epigenetic (ACE) conversion.

[0075] In some embodiments, the procedure that differentially affects the first nucleobase in the DNA of the first subsample and the second nucleobase in the DNA comprises an enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA.bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosine, converting it to uracil.

[0076] In some embodiments, the procedure that differentially affects a first nucleobase in the DNA and a second nucleobase in the DNA of the first subsample comprises separating DNA that initially comprises the first nucleobase from DNA that initially does not comprise the first nucleobase.

[0077] In some embodiments, the first core base is a modified or unmodified adenine, and the second core base is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA) or N6-formyladenine (fA).

[0078] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases (such as mA) from other DNA. See, for example, Kumar et al., Frontiers Genet.2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific for mA are described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies against various modified nucleobases (such as thymine / uracil forms, including halogenated forms, such as 5-bromouracil) are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can be produced by deamination and is read as G in sequencing. See, for example, U.S. Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, Chapter 14, “Mutation, Repair, and Recombination.”

[0079] Enrichment / capture step; amplification; adapters; barcodes

[0080] In some embodiments, the methods disclosed herein include the step of capturing one or more target region groups of DNA such as cfDNA. Any suitable method known in the art can be used for capture. In some embodiments, capture includes contacting the DNA to be captured with a target-specific probe group. The target-specific probe group can have any feature of the target-specific probe group described herein, including but not limited to the features in the embodiments set forth above and the parts related to the probe below. Capture can be performed on one or more subsamples prepared during the method disclosed herein. In some embodiments, DNA is captured from at least the first subsample or the second subsample, for example, at least the first subsample and the second subsample. In the case where the first subsample undergoes a separation step (e.g., separating DNA that initially includes the first nuclear base (e.g., hmC) and DNA that initially does not include the first nuclear base, such as hmC-seal), any one, any two, or all of the DNA that initially includes the first nuclear base (e.g., hmC), the DNA that initially does not include the first nuclear base, and the second subsample can be captured. In some embodiments, differential labeling (e.g., as described herein) is performed on the subsamples and then pooled before undergoing capture.

[0081] The capture step can be performed using conditions suitable for specific nucleic acid hybridization, which conditions generally depend to some extent on the characteristics of the probe, such as length, base composition, etc. Appropriate conditions will be familiar to those skilled in the art given the general knowledge of nucleic acid hybridization in the art. In some embodiments, a complex of the target-specific probe and the DNA is formed.

[0082] In some embodiments, the methods described herein include capturing more than one target area group of cfDNA obtained from a test subject. The target area includes an epigenetic target area that can show differences in methylation levels and / or fragmentation patterns, depending on whether they are derived from tumor cells or from healthy cells. The target area also includes a sequence variable target area that can show sequence differences, depending on whether they are derived from tumor cells or from healthy cells. The capture step produces a capture group of cfDNA molecules, and in the capture group of cfDNA molecules, the cfDNA molecules corresponding to the sequence variable target area group are captured with a capture yield greater than that of the cfDNA molecules corresponding to the epigenetic target area group. For further discussion of capture steps, capture yields, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.

[0083] In some embodiments, the methods described herein include contacting cfDNA obtained from a test subject with a target-specific probe set, wherein the target-specific probe set is configured to capture cfDNA corresponding to a set of sequence variable target regions at a greater capture yield than cfDNA corresponding to a set of epigenetic target regions.

[0084] It is beneficial to capture cfDNA corresponding to a set of sequence variable target regions at a greater capture yield than cfDNA corresponding to a set of epigenetic target regions because the sequencing depth that may be required to analyze sequence variable target regions with sufficient confidence or accuracy is greater than the sequencing depth that may be required to analyze epigenetic target regions. The amount of data required to determine fragmentation patterns (e.g., testing perturbations of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in high-methylation and low-methylation partitions) is generally less than the amount of data required to determine the presence or absence of cancer-associated sequence mutations. Capturing target region groups at different yields can facilitate sequencing target regions to different sequencing depths in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing pool).

[0085] In various embodiments, the method further comprises sequencing the captured cfDNA to, for example, different degrees of sequencing depth for the epigenetic target region group and the sequence variable target region group, consistent with the discussion herein. In some embodiments, the complex of the target-specific probe and the DNA is separated from the DNA that is not bound to the target-specific probe. For example, where the target-specific probe is covalently or non-covalently bound to a solid support, a washing or aspiration step can be used to separate the unbound material. Alternatively, chromatography can be used where the complex has chromatographic properties different from those of the unbound material (e.g., where the probe comprises a ligand that binds to a chromatography resin).

[0086] As discussed in detail elsewhere herein, target-specific probe sets can include more than one set, such as probes for a sequence variable target set and probes for an epigenetic target set. In some such embodiments, the capture step is performed simultaneously using probes for sequence variable target regions and probes for epigenetic target regions in the same container, for example, the probes for sequence variable target regions and epigenetic target regions are in the same composition. This method provides a relatively more efficient workflow. In some embodiments, the concentration of probes for the sequence variable target region set is greater than the concentration of probes for the epigenetic target region set.

[0087] Alternatively, the capture step is performed in a first container with a sequence variable target probe set and in a second container with an epigenetic target probe set, or the contacting step is performed at a first time and in a first container with a sequence variable target probe set and at a second time, either before or after the first time, with an epigenetic target probe set. The method allows for the preparation of separate first and second compositions comprising captured DNA corresponding to the sequence variable target set and captured DNA corresponding to the epigenetic target set. The compositions can be processed separately as desired (e.g., fractionated based on methylation, as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis, such as sequencing.

[0088] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capture step. In some embodiments, the amplification is performed after the capture step.

[0089] In some embodiments, the DNA includes an adapter. This can be performed simultaneously with the amplification procedure, for example, by providing an adapter in the 5' portion of the primer, for example, as described above. Alternatively, the adapter can be added by other methods such as ligation.

[0090] In some embodiments, the DNA comprises a tag, which can be a barcode or comprises a barcode. The tag can help identify the source of the nucleic acid. For example, a barcode can be used to allow the source of the DNA to be identified after more than one sample is pooled for parallel sequencing, for example, a subject. This can be performed simultaneously with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer, for example, as described above. In some embodiments, the adapter and the tag / barcode are provided by the same primer or primer set. For example, the barcode can be located 3' of the adapter and 5' of the target hybridization portion of the primer. Alternatively, the barcode can be added by other methods, such as ligation, optionally in the same ligation substrate together with the adapter.

[0091] Additional details regarding amplification, tagging, and barcoding are discussed below in the "General Features of the Method" section, which details may be combined, to the extent practicable, with any of the preceding embodiments and the embodiments set forth in the "Introduction and Overview" section.

[0092] Computer systems, processing of real-world evidence (RWE)

[0093] The methods of the present disclosure can be implemented using or with the aid of a computer system. For example, such methods can include partitioning a sample into more than one subsample, the more than one subsample including a first subsample and a second subsample, wherein the first subsample comprises DNA having a greater proportion of cytosine modifications than the second subsample; subjecting the first subsample to a procedure that differentially affects a first nucleobase in the DNA of the first subsample and a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and sequencing the DNA in the first subsample and the DNA in the second subsample in a manner that distinguishes the first nucleobase from the second nucleobase in the DNA of the first subsample.

[0094] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing more than one target region set from the cfDNA, wherein the more than one target region set comprises a sequence variable target region set and an epigenetic target region set, thereby generating a set of captured cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining more than one sequence reads generated by a nucleic acid sequencer by sequencing the captured cfDNA molecules; mapping the more than one sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence variable target region set and the epigenetic target region set to determine a likelihood that the subject has cancer.

[0095] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or may be compiled during runtime. The code may be provided in a programming language that may be selected so that the code can be executed in a precompiled or as-compiled manner.

[0096] Additional details related to computer systems and networks, databases, and computer program products are also provided, for example, in: Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), all of which are incorporated by reference in their entireties. Additional information can be found in PCT Publication No. US2022032250 and US Application No. 17832498.

[0097] This document describes a method for generating an integrated data repository and / or analysis system that includes multiple types of healthcare data, according to one or more implementations. The architecture may include a data integration and / or analysis system. The data integration and analysis system may obtain data from multiple data sources and integrate the data from the data sources into the integrated data repository. For example, the data integration and analysis system may obtain data from a health insurance claims data repository. In various examples, the data integration and analysis system and the health insurance claims data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the health insurance claims data repository may be created and maintained by the same entity.

[0098] The data integration and analysis system can be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In certain implementations, at least a portion of the one or more computing devices can be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices can be implemented in a cloud computing architecture. In a scenario where the computing system for implementing the data integration and analysis system is configured in a distributed computing architecture, processing operations can be performed concurrently by multiple virtual machines. In various examples, the data integration and analysis system can implement multi-threading technology. The implementation of the distributed computing architecture and multi-threading technology enables the data integration and analysis system to utilize fewer computing resources relative to a computing architecture that does not implement these technologies.

[0099] A health insurance claims data repository may store information obtained from one or more health insurance companies, corresponding to insurance claims filed by subscribers of one or more health insurance companies. The health insurance claims data repository may be arranged (e.g., sorted) by patient identifier. The patient identifier may be based on the patient's first name, last name, date of birth, social security number, address, employer, etc. The data stored by the health insurance claims data repository may include structured data arranged in one or more data tables. The one or more data tables storing the structured data may include multiple rows and columns indicating information regarding health insurance claims filed by subscribers of one or more health insurance companies related to procedures and / or treatments received by the subscribers from healthcare providers. At least a portion of the rows and columns of the data tables stored by the health insurance claims data repository may include health insurance codes, which may indicate diagnoses and treatments and / or procedures for biological conditions obtained by subscribers of one or more health insurance companies. In various examples, the health insurance codes may also indicate diagnostic procedures obtained by an individual related to one or more biological conditions that may be present in the individual. In one or more examples, a diagnostic procedure may provide information for detecting the presence of a biological condition. A diagnostic procedure may also provide information for determining the progression of a biological condition. In one or more illustrative examples, a diagnostic procedure may include one or more imaging procedures, one or more assays, one or more laboratory procedures, one or more combinations thereof, and the like.

[0100] The data integration and analysis system can also obtain information from the molecular data repository. The molecular data repository can store data related to genomic information, genetic information, metabolome (metabolomic) information, transcriptome information, fragment group (fragmentomic) information, immune receptor (immune receptor) information, methylation (methylation) information, epigenome (epigenomic) information and / or proteome information of multiple individuals. In one or more examples, the data integration and analysis system and the molecular data repository can be created and maintained by different entities. In one or more other examples, the data integration and analysis system and the molecular data repository can be created and maintained by the same entity.

[0101] The genomic and / or epigenomic information can indicate one or more mutations corresponding to an individual's gene. The mutation of an individual's gene may correspond to the difference between the individual's nucleic acid sequence and one or more reference genomes. The reference genome may include a known reference genome, such as hg19. In various examples, the mutation of an individual's gene may correspond to the difference of an individual's germline gene relative to the reference genome. In one or more additional examples, the reference genome may include an individual's germline genome. In one or more further examples, the mutation of an individual's gene may include somatic mutations. The mutation of an individual's gene may be relevant to insertions, deletions, single nucleotide variations, loss of heterozygosity, doublings, amplifications, translocations, fusion genes, or one or more combinations thereof.

[0102] In one or more illustrative examples, the genome and / or epigenomic information stored by the molecular data repository can include the genome and / or epigenomic spectrum of the tumor cells present in the individual. In these cases, genome and / or epigenomic information can be derived from the analysis of genetic material, and this genetic material is such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA) from sample, and this sample includes but is not limited to tissue sample or tumor biopsy, circulating tumor cell (CTC), exosome (exosome) or cell burial body (efferosome), or from the circulating nucleic acid (for example, cell-free DNA) found in the blood sample of individual, it exists due to the degradation of the tumor cells present in the individual. In one or more instances, the genome and / or epigenomic information of individual tumor cells can correspond to one or more target regions. One or more mutations existing about one or more target regions can indicate the presence of tumor cells in individual. The genome and / or epigenomic information stored by the molecular data repository can be generated in relation to assay or other diagnostic tests, and this assay or other diagnostic tests can determine one or more mutations about one or more target regions of reference genome.

[0103] Multiple data tables can be arranged according to a data repository schema. In an illustrative example, the data repository schema includes a first data table, a second data table, a third data table, a fourth data table, and a fifth data table. Although the illustrative example includes five data tables, in other implementations, the data repository schema may include more data tables or fewer data tables. The data repository schema may also include links between the data tables. Links between the data tables may indicate that information retrieved from one of the data tables causes additional information stored by one or more additional data tables to be retrieved. Furthermore, not all data tables may be linked to every other data table. In an illustrative example, the first data table is logically coupled to the second data table via a first link, and the first data table is logically coupled to the fourth data table via a second link. Furthermore, the second data table is logically coupled to the third data table via a third link, and the fourth data table is logically coupled to the fifth data table via a fourth link. Furthermore, the third data table is logically coupled to the fifth data table via a fifth link.

[0104] In various examples, as data tables are added to and / or removed from the data repository schema, additional links between data tables may be added to or removed from the data repository schema. In one or more illustrative examples, the integrated data repository may store data tables according to the data repository schema for at least a portion of individuals for whom the data integration system obtains information from a combination of at least two of a health insurance claims data repository, a molecular data repository, one or more additional data repositories, and one or more reference information data repositories. As a result, the integrated data repository may store corresponding instances of data tables for thousands, tens of thousands, up to hundreds of thousands, or more individuals according to the data repository schema.

[0105] The data integration and analysis system may also include a data pipeline system. The data pipeline system may include multiple algorithms, software codes, scripts, macros, or other computer-executable instruction packages that process information stored by the integrated data repository to generate additional data sets. The additional data sets may include information obtained from one or more data tables. The additional data sets may also include information derived from data obtained from one or more data tables. The components of the data pipeline system implemented to generate the first additional data set may be different from the components of the data pipeline system used to generate the second additional data set.

[0106] In one or more examples, a data pipeline system can generate a data set indicating medications received by multiple individuals. In one or more illustrative examples, the data pipeline system can analyze information stored in at least one of the data tables to determine health insurance codes corresponding to medications received by the multiple individuals. The data pipeline system can analyze the health insurance codes corresponding to the medications against a library indicating designated medications corresponding to the one or more health insurance codes to determine the names of the medications received by the individuals. In one or more additional examples, the data pipeline system can analyze information stored by an integrated data repository to determine medical procedures received by the multiple individuals. For illustration, the data pipeline system can analyze information stored by one of the data tables to determine treatments received by the individuals via at least one injection or intravenously. In one or more further examples, the data pipeline system can analyze information stored by the integrated data repository to determine episodes of care for the individuals, lines of treatment received by the individuals, progression of a biological condition, or the time to next treatment. In various examples, the data sets generated by the data pipeline system can be different for different biological conditions. For example, a data pipeline system may generate a first number of data sets regarding a first type of cancer (eg, lung cancer), and a second number of data sets regarding a second type of cancer (eg, colorectal cancer).

[0107] The data pipeline system may also determine one or more confidence levels to assign to information associated with an individual having data stored by the integrated data repository. The respective confidence levels may correspond to different accuracy measures for the information associated with the individual having data stored by the integrated data repository. The information associated with the respective confidence levels may correspond to one or more characteristics of the individual derived from the data stored in the integrated data repository. The data pipeline system may generate confidence level values ​​for the one or more characteristics in conjunction with generating one or more data sets based on the integrated data repository. In one or more examples, a first confidence level may correspond to a first range of the accuracy measure, a second confidence level may correspond to a second range of the accuracy measure, and a third confidence level may correspond to a third range of the accuracy measure. In one or more additional examples, the second range of the accuracy measure may include values ​​smaller than the first range of values ​​for the accuracy measure, and the third range of the accuracy measure may include values ​​smaller than the second range of values ​​for the accuracy measure. In one or more illustrative examples, information corresponding to the first confidence level may be referred to as gold standard information, information corresponding to the second confidence level may be referred to as silver standard information, and information corresponding to the third confidence level may be referred to as bronze standard information.

[0108] The data pipeline system can determine the value of the confidence level of an individual's characteristic based on many factors. For example, a corresponding information set can be used to determine the individual's characteristic. The data pipeline system can determine the confidence level of an individual's characteristic based on the amount of completeness of the corresponding information set used to determine the individual's characteristic. In the case where one or more pieces of information are missing from the information set associated with a first number of individuals, the confidence level of the characteristic can be lower than the confidence level of a second number of individuals in the information set where no information is missing. In one or more examples, the data pipeline system can use the amount of missing information to determine the confidence level of an individual's characteristic. To illustrate, a larger amount of missing information used to determine a characteristic may result in a lower confidence level for the characteristic than a lower amount of missing information used to determine the individual's characteristic. In addition, different types of information can correspond to various confidence levels for a characteristic. In one or more examples, the presence of a first piece of information used to determine the characteristic may result in a higher confidence level for the characteristic than the presence of a second piece of information used to determine the individual's characteristic.

[0109] In one or more illustrative examples, a data pipeline system can determine a plurality of individuals included in a group with a preliminary diagnosis of lung cancer (or other biological conditions). The data pipeline system can determine a confidence level for the respective individuals regarding being classified as having a preliminary diagnosis of lung cancer. The data pipeline system can use information from multiple columns included in a data table to determine a confidence level that an individual is included in the lung cancer group. The multiple columns can include health insurance codes associated with the diagnosis of the biological condition and / or the treatment of the biological condition. In addition, the multiple columns can correspond to dates of diagnosis and / or treatment of the biological condition. The data pipeline system can determine that, if information for each column or at least a threshold number of columns in the multiple columns is available, the confidence level that the individual is characterized as part of the lung cancer group is higher than the confidence level if information for less than a threshold number of columns is available. In addition, the data pipeline system can determine a confidence level for an individual included in the lung cancer group based on the type of information and the availability of information associated with one or more columns. To illustrate, where one or more diagnosis codes are present and one or more treatment codes are not present for one or more time periods for a group of individuals, the data pipeline system may determine that the confidence level for including the group of individuals in a lung cancer cohort is greater than the confidence level for including the group of individuals in a lung cancer cohort where at least one diagnosis code is not present and a treatment code used to determine whether an individual is included in the lung cancer cohort is present.

[0110] The data analysis system can receive an integrated data repository request from one or more computing devices (e.g., an example computing device). The one or more integrated data repository requests can result in data being retrieved from the integrated data repository. In various examples, the one or more integrated data repository requests can result in data being retrieved from one or more datasets generated by the data pipeline system. The integrated data repository request can specify data to be retrieved from the integrated data repository and / or one or more datasets generated by the data pipeline system. In one or more additional examples, the integrated data repository request can include one or more pre-built queries corresponding to computer-executable instructions for retrieving a specified dataset from the integrated data repository and / or one or more datasets generated by the data pipeline system.

[0111] In response to one or more integrated data repository requests, the data analysis system can analyze data retrieved from the integrated data repository or at least one of the one or more data sets generated by the data pipeline system to generate data analysis results. The data analysis results can be sent to one or more computing devices, such as the example computing device. Although the illustrative example shows that one or more integrated data repository requests come from one computing device and the data analysis results are sent to another computing device, in one or more other implementations, the data analysis results can be received by the same computing device that sent the one or more integrated data repository requests. The data analysis results can be displayed by one or more user interfaces presented by the computing device or the computing device.

[0112] This paper describes the method for analyzing nucleic acid sequence information.In various embodiments, analytical method includes one or more models, and each in one or more models includes one or more of the survival, sub-modeling, disease node determination and identification (for example, driving mutation), disease association, disease subtyping, recurrence, transfer, the time to next treatment etc. as a separate component.In various embodiments, model includes hierarchical model (for example, nested model, multilevel model), mixed model (for example, regression, such as logistic regression and Poisson regression, pooling, random effect, fixed effect, mixed effect, linear mixed effect, generalized linear mixed effect), risk model, odds ratio model and / or repeated samples (for example, repeated measurement, such as ANOVA).In various embodiments, model is hierarchy random effects model.In various embodiments, model is hierarchy cubic spline random effects model.In various embodiments, model is cubic spline model.In various embodiments, model is generalized linear effects model.In various embodiments, model is linear effects model.In various embodiments, model is Cox proportional hazards model.In various embodiments, analytical method includes assembling the models together.In various embodiments, assembling includes the generation of associated parameters. In one or more embodiments, analytical method comprises patient survival information and patient genetic information.As an example, model assembly together can comprise the different models of the different types of cancer (comprising subtype) represented in patient survival information.Each in different models can be configured to determine the correlation between the survival time of the patient of the cancer of the respective type that they are configured to assess and genetic factors and are diagnosed with.For example, can recommend the genetic factors that are determined to have strong correlation with cancer survival time (for example, relatively short survival time and / or relatively long survival time) as potential therapeutic target.

[0113] In various embodiments, analysis can include one or more of survival, submodeling, disease node determination and identification (e.g., driver mutation), disease association, disease subtyping, recurrence, metastasis, the time to next treatment, etc. as separate components. For example, modeling can be beneficial to application as mentioned above, such as patient survival information and patient genetic information. In various embodiments, the submodeling component can determine the subset of patient survival information and patient genetic information for generating different patient groups associated with different types of cancer and cancer subtypes. In various embodiments, the submodel includes a hierarchical model (e.g., a nested model, a multilevel model), a mixed model (e.g., regression, such as logistic regression and Poisson regression, pooling, random effects, fixed effects, mixed effects, linear mixed effects, generalized linear mixed effects), a risk model, an odds ratio model, and / or repeated samples (e.g., repeated measurements, such as ANOVA). In various embodiments, the submodel is a hierarchical random effects model. In various embodiments, the submodel is a hierarchical cubic spline random effects model. In various embodiments, the submodel is a cubic spline model. In various embodiments, the submodel is a generalized linear effects model. In various embodiments, the submodel is a linear effects model. In various embodiments, the submodel is a Cox proportional hazards model. Each subset of patient survival information and patient genetic information can include information of patients diagnosed with different types of cancer and cancer subtypes. For example, the submodeling component can also apply the subset of patient survival information and patient genetic information to corresponding individual survival models developed for different cancer types (including subtypes). In various embodiments, the information generated for the analysis method can be stored in memory (e.g., as model data). In various embodiments, the information generated for the analysis method generates one or more survival models for individual subjects.

[0114] In various embodiments, a survival model is used to analyze patient survival information and patient genetic information, including a disease node determination and identification component, to identify disease nodes included in the patient genetic information for each type of cancer, the disease nodes being involved in the genetic mechanisms used by the corresponding cancer type for proliferation. In various embodiments, the disease node component identifies disease nodes based on observed correlations between genetic factors and cancer survival times provided in the patient survival information. For example, genetic factors that are often observed to be associated with short survival times of a particular type of cancer and less often observed to be associated with long survival times of the particular type of cancer can be identified as active genetic factors that have an active role in the genetic mechanisms of the particular type of cancer (including subtypes).

[0115] In various embodiments, disease node determination and identification include disease association parameters about the association between different cancer types, to promote identification of active genetic factors associated with different cancer types. For example, highly associated cancer types can share one or more common key potential genetic factors. As readily understood by those of ordinary skill, the model (e.g., survival model) of the associated cancer type dialectically exchanges information to determine and / or identify active genetic factors across cancer types (including subtypes). In various embodiments, disease association parameters applied by disease node determination and identification are promoted by modeling. In various embodiments, the generation of individual survival models can adopt one or more machine learning algorithms to promote survival, modeling, disease nodes associated with specific types of cancer (including subtypes) based on patient genetic information and disease association parameters.

[0116] In some embodiments, the node determination and identification of cancer types (including subtypes) are associated with a scoring system that includes determining disease nodes. For example, the scoring of disease nodes for a particular type of cancer (including subtypes) reflects the association of the survival time of the disease node with the particular type of cancer (including subtypes). In various embodiments, scoring can be based on the frequency of direct or indirect identification of specific genetic factors for patients diagnosed with a particular cancer type. In various embodiments, analysis includes survival, submodeling, disease node determination and identification (for example, driver mutations), disease association, disease subtyping, recurrence, metastasis, time to the next treatment, etc. mentioned above, which can be related to a threshold value less than a definition, greater than a definition. For example, the greater the score associated with a disease node and a cancer type (including subtypes), the greater the contribution of the disease node to survival time. In various embodiments, information about the disease node of the corresponding type of cancer (including subtypes) can be collated in a data structure such as a database, and the score determined for active genetic factors.

[0117] This paper describes the analytical method comprising effect modeling. In various embodiments, effect modeling includes random effects, fixed effects, mixed effects, linear mixed effects and generalized linear mixed effects. In various embodiments, effect includes cubic splines. In various embodiments, effect modeling includes regression. In various embodiments, effect modeling includes logistic regression and Poisson regression. In various embodiments, model does not include covariates. In various embodiments, model includes covariates. In various embodiments, covariates are information from medical records (including laboratory test records, such as genome, epigenome, nucleic acid and other analyte results), insurance records, etc. Example includes age, treatment line, smoking status (yes / no), sex and multiple scoring and / or staging systems for specific cancer disease patients, wherein illustrative examples include age (in years), anti-EGFR treatment line, smoking status (yes / no), sex (female / male) and Van Walraven Elixhauser comorbidity (ELIX) score (expressed as a weighted measure across multiple common comorbidities) specific to lung cancer patients. A skilled artisan will readily appreciate that covariates can include any number of data elements for individuals and individuals in a population, such as data elements from medical records (including laboratory test records, such as genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc.

[0118] In various embodiments, the analysis method includes generating a hierarchy comprising at least one first-order equation. In various embodiments, the first-order equation includes a truncated cubic spline. In various embodiments, the truncated cubic spline includes longitudinal data. This includes, for example, direct or indirect measurements of ctDNA levels, allele fractions, tumor fractions. In various embodiments, additional first-order equations include covariates. In various embodiments, covariates are information about individuals or groups extracted and / or stored from medical records (including laboratory test records, such as genome, epigenome, nucleic acid and other analyte results), insurance records, etc. Examples include age, treatment line, smoking status (yes / no), sex, and various scoring and / or staging systems that have been used for patients with specific cancer diseases. In various embodiments, a velocity map is generated. In various embodiments, the velocity map is a derivative or one or more equations, such as at least one first-order equation. In various embodiments, the analysis method includes one or more of equations (1), (2), and (3) described in the examples.

[0119] This article describes an analysis method that includes jointly solving different analysis components, including one or more of survival, modeling and sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. In various embodiments, the analysis method includes jointly solving one or more different models of different cancer types under a joint model framework. For example, the analysis method may include jointly solving one or more different survival models of different cancer types under a joint model framework. In various embodiments, the method includes determining association parameters. In various embodiments, association parameters include, for example, the relationship between the estimated current value of the patient's survival and the patient's biomarker, the relationship between the patient's survival and the patient's estimate of the biomarker current change over time. In various embodiments, this includes slope, and the relationship between the current estimated area under the longitudinal trajectory of the subject as a surrogate for the cumulative effect of the biomarker. It is readily understood by those skilled in the art that association parameters can take various forms and can also be combined. For example, the relationship between the estimated current value of the total survival and the estimated current slope of the patient's longitudinal trajectory can be examined.

[0120] In one or more examples, the data analysis system may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze data retrieved in response to one or more integrated data repository requests. In one or more examples, the data analysis system may implement one or more artificial neural networks to analyze data retrieved in response to one or more integrated data repository requests. For illustration, the data analysis system may implement at least one of one or more convolutional neural networks or one or more residual neural networks to analyze data retrieved from the integrated data repository in response to one or more integrated data repository requests. In at least some examples, the data analysis system may implement one or more random forest techniques, one or more support vector machines, or one or more hidden Markov models to analyze data retrieved in response to one or more integrated data repository requests. One or more statistical models may also be implemented to analyze the data retrieved in response to one or more integrated data repository requests to identify at least one of a correlation or a significance measure between individual features. For example, a log-rank test may be applied to the data retrieved in response to one or more integrated data repository requests. In addition, a Cox proportional hazards model can be implemented on the data retrieved in response to one or more integrated data repository requests. In addition, a Wilcoxon signed rank test can be applied to the data retrieved in response to one or more integrated data repository requests. In other examples, a z-score analysis can be performed on the data retrieved in response to one or more integrated data repository requests. In another example, a Kaplan Meier analysis can be performed on the data retrieved in response to one or more integrated data repository requests. In at least some examples, one or more machine learning techniques can be implemented in combination with one or more statistical techniques to analyze the data retrieved in response to one or more integrated data repository requests.

[0121] In one or more illustrative examples, a data analysis system may determine the survival rate of individuals with lung cancer in response to one or more treatments. In one or more additional illustrative examples, the data analysis system may determine the survival rate of individuals with mutations in one or more genomic and / or epigenomic regions in which lung cancer is present in response to one or more treatments. In various examples, the data analysis system may generate a data analysis result if data retrieved from at least one of the integrated data repository or one or more datasets generated by the data pipeline system meets one or more criteria. For example, the data analysis system may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests meets a threshold confidence level. If the confidence level of at least a portion of the data retrieved in response to the one or more integrated data repository requests is less than the threshold confidence level, the data analysis system may refrain from generating at least a portion of the data analysis result. If the confidence level of at least a portion of the data retrieved in response to the one or more integrated data repository requests is at least the threshold confidence level, the data analysis system may generate at least a portion of the data analysis result. In various examples, the threshold confidence level may be associated with the type of data analysis result generated by the data analysis system.

[0122] In one or more illustrative examples, a data analysis system may receive an integrated data repository request to generate a data analysis result indicating survival rates for one or more individuals. In these cases, the data analysis system may determine whether the data stored by the integrated data repository and / or one or more datasets generated by the data pipeline system meets a threshold confidence level, such as a gold standard confidence level. In one or more additional examples, the data analysis system may receive an integrated data repository request to generate a data analysis result indicating treatments received by one or more individuals. In these implementations, the data analysis system may determine whether the data stored by the integrated data repository and / or one or more datasets generated by the data pipeline system meets a lower threshold confidence level, such as a copper standard confidence level.

[0123] In one or more additional illustrative examples, a data analysis system may receive an integrated data repository request to identify individuals who have one or more genomic and / or epigenomic mutations and who have received one or more treatments for a biological condition. Continuing with this example, the data analysis system may determine the survival rate of individuals who have one or more genomic and / or epigenomic mutations with respect to the one or more treatments received by the individuals. The data analysis system may then identify the effectiveness of treatments for the individuals that are associated with genomic and / or epigenomic mutations that may be present in the individuals based on the survival rates of the individuals. In this way, health outcomes for individuals may be improved by identifying prospective treatments that are more effective for a population of individuals who have one or more genomic and / or epigenomic mutations than current treatments provided to the individuals.

[0124] The data pipeline system may include a first data processing instruction, a second data processing instruction, and up to an Nth data processing instruction. The data processing instructions may be executed by one or more processing units to perform multiple operations to generate corresponding data sets using information obtained from the integrated data repository. In one or more illustrative examples, the data processing instructions may include at least one of software code, scripts, API calls, macros, etc. The first data processing instruction may be executed to generate a first data set. Furthermore, the second data processing instruction may be executed to generate a second data set. Furthermore, the Nth data processing instruction may be executed to generate an Nth data set. In various examples, after the data integration and analysis system generates the integrated data repository, the data pipeline system may cause the data processing instructions to be executed to generate the data sets. In one or more examples, the data sets may be stored in the integrated data repository or in an additional data repository accessible to the data integration and analysis system. At least a portion of the data processing instructions may analyze health insurance codes to generate at least a portion of the data set. Furthermore, at least a portion of the data processing instructions may analyze genomic data to generate at least a portion of the data set.

[0125] In one or more examples, the first data processing instruction may be executed to retrieve data from one or more first data tables stored in an integrated data repository. The first data processing instruction may also be executed to retrieve data from one or more specified columns of the one or more first data tables. In various examples, the first data processing instruction may be executed to identify individuals having health insurance codes stored in one or more column and row combinations corresponding to one or more diagnostic codes. The first data processing instruction may then be executed to analyze the one or more diagnostic codes to determine the biological condition that the individual has been diagnosed with. In one or more illustrative examples, the first data processing instruction may be executed to analyze the one or more diagnostic codes with respect to a diagnostic code library indicating one or more biological conditions (corresponding to the corresponding diagnostic codes). The diagnostic code library may include hundreds to thousands of diagnostic codes. The first data processing instruction may also be executed to determine individuals diagnosed with a biological condition by analyzing the individual's timing information (such as treatment date, diagnosis date, death date, one or more combinations thereof, etc.).

[0126] The second data processing instructions may be executed to retrieve data from one or more second data tables stored in the integrated data repository. The second data processing instructions may also be executed to retrieve data from one or more specified columns of the one or more second data tables. In various examples, the second data processing instructions may be executed to identify individuals having health insurance codes stored in one or more column and row combinations corresponding to one or more treatment codes. The one or more treatment codes may correspond to treatments obtained from a pharmacy. In one or more additional examples, the one or more treatment codes may correspond to treatments received through medical procedures (such as injections or intravenous injections). The second data processing instructions may be executed to determine one or more treatments corresponding to the corresponding health insurance codes included in the one or more second data tables by analyzing the health insurance codes in relation to a predetermined information set. The predetermined information set may include a database indicating one or more treatments corresponding to one of hundreds to thousands of health insurance codes. The second data processing instructions may generate a second data set to indicate the corresponding treatments received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to the individuals included in the first data set. The second data set may be arranged into rows and columns, where one or more rows correspond to a single individual and one or more columns indicate a treatment received by the corresponding individual.

[0127] An Nth processing instruction (where N can be any positive integer) can be executed to generate an Nth dataset by combining information from multiple previously generated datasets (such as a first dataset and a second dataset). Furthermore, the Nth processing instruction can be executed to generate the Nth dataset by retrieving additional information from one or more additional columns of the integrated data repository and merging the additional information from the integrated data repository with information obtained from the first dataset and the second dataset. For example, the Nth processing instruction can be executed to identify individuals included in the first dataset who have been diagnosed with a biological condition and analyze designated columns of one or more additional data tables of the integrated data repository to determine dates corresponding to treatments indicated in the second dataset for the individuals included in the first dataset. In one or more further examples, the Nth processing instruction can be executed to analyze columns of one or more additional data tables of the integrated data repository to determine the dosage of treatments indicated in the second dataset received by the individuals included in the first dataset. In this manner, the Nth processing instruction can be executed to generate a care event dataset based on information included in the group dataset and the treatment dataset.

[0128] In one or more illustrative examples, in response to receiving an integrated data repository request, a data analysis system may determine one or more data sets corresponding to characteristics of a query associated with the integrated data repository request. For example, the data analysis system may determine that information included in a first data set and a second data set is suitable for responding to the integrated data repository request. In these scenarios, the data analysis system may analyze at least a portion of the data included in the first data set and the second data set to generate a data analysis result. In one or more additional examples, the data analysis system may determine different data sets to respond to different queries included in the integrated data repository request in order to generate the data analysis result.

[0129] Using a specific set of data processing instructions to generate corresponding data sets can reduce the amount of input from users of the data integration and analysis system and reduce the computational load, such as the amount of processing resources and memory, required to process integration data repository requests. For example, without a specific architecture for a data pipeline system, data used to respond to an integration data repository request is gathered from the data repository each time an integration data repository request is received. In contrast, by implementing a data pipeline system to execute data processing instructions to generate data sets, the data required to respond to various integration data repository requests is already gathered and accessible to the data analysis system in response to the integration data repository requests. Consequently, the computational resources used to generate data sets in response to integration data repository requests by implementing a data pipeline system are less than in a typical system that performs information parsing and collection processes for each integration data repository request. Furthermore, if a data pipeline system has not yet been implemented, users of the data integration and analysis system may need to submit multiple integration data repository requests in order to analyze the information they intend to analyze because the ad hoc collection of data in response to integration data repository requests in a typical system is inaccurate, or because the data analysis system in a typical system is invoked multiple times to perform information analysis that could be performed using a single integration data repository request if a data pipeline system were implemented.

[0130] In operation, the data integration and analysis system can integrate genomic data and health insurance claims data for an individual that is common to both the molecular data repository and the health insurance claims data repository. The data integration and analysis system can determine that the individual is common to both the molecular data repository and the health insurance claims data repository by determining that the genomic data and health insurance claims data correspond to a common token. The data integration and analysis system can determine that a first token corresponds to a second token by determining a similarity measure between a first token associated with a portion of the genomic data and a second token associated with a portion of the health insurance claims data. If the first token has at least a threshold amount of similarity relative to the second token, the data integration and analysis system can store the corresponding portion of the genomic data and the corresponding portion of the health insurance claims data in an integrated data repository, such as the integrated data repository, with respect to an identifier of the individual.

[0131] The implementation of the architecture can implement an encryption protocol that enables de-identified information from different data repositories to be integrated into a single data repository. In this way, the security of the data stored by the integrated data repository is increased. Furthermore, the encryption protocol implemented by the architecture can enable more efficient retrieval and accurate analysis of the information stored by the integrated data repository compared to a situation where the encryption protocol of the architecture is not utilized. For example, by generating a token file including a first token based on a specified set of information stored by a molecular data repository using encryption technology, and utilizing a second token generated using the same or similar encryption technology for a similar or identical set of information stored by a health insurance claims data repository, the data integration and analysis system can match information stored by different data repositories corresponding to the same individual. Without implementing the encryption protocol of the architecture, the probability of misattributing information from one data repository to one or more individuals increases, which reduces the accuracy of the results provided by the data integration and analysis system in response to integrated data repository requests sent to the data integration and analysis system.

[0132] This document describes a framework for generating a dataset based on data stored in an integrated data repository via a data pipeline system, according to one or more implementations. The integrated data repository can store health insurance claim data and genomic data for a group of individuals. For example, the integrated data repository can store information obtained from health insurance claim records for a group of individuals. For each individual included in the group, the integrated data repository can store information obtained from multiple health insurance claim records. In various examples, the information stored by the integrated data repository can include and / or be derived from thousands, tens of thousands, hundreds of thousands, or even millions of health insurance claim records for multiple individuals. Furthermore, each health insurance claim record can include multiple columns. As a result, the integrated data repository can be generated by analyzing millions of columns of health insurance claim data.

[0133] Furthermore, while health insurance claims data can be organized according to a structured data format, it is typically designed to be viewed by health insurance providers, patients, and healthcare providers to display financial information and insurance code information related to services provided by healthcare providers to individuals. Consequently, it is not easy to analyze health insurance claims data to obtain usable insights related to characteristics of individuals with biological conditions that could aid in treating those conditions. An integrated data repository can be generated and organized by analyzing and modifying raw health insurance claims data in a manner that enables further analysis of the data stored in the integrated data repository to identify trends, characteristics, traits, and / or insights about individuals who may have one or more biological conditions. For example, health insurance codes can be stored in the integrated data repository in such a manner that at least one of a medical procedure, biological condition, treatment, dosage, drug manufacturer, drug distributor, or diagnosis can be determined for a given individual based on the individual's health insurance claims data. In various examples, the data integration and analysis system can generate and implement one or more tables indicating correlations between health insurance claims data and various treatments, symptoms, or biological conditions corresponding to the health insurance claims data. Furthermore, an integrated data repository can be generated using a set of individual genomic data records. In various examples, a large amount of health insurance claims data can be matched with a set of individual genomic data to generate an integrated data repository.

[0134] By integrating a set of individual genomic data records with health insurance claims records, a data integration and analysis system can determine correlations between the presence of one or more biomarkers present in the genomic data records and other characteristics of the individuals indicated by the health insurance claims data records, which are typically not determined by existing systems. For example, the data integration and analysis system can determine one or more genomic and / or epigenomic characteristics of an individual that correspond to treatments received by the individual, the timing of treatment, the dosage of treatment, the individual's diagnosis, smoking status, the presence of one or more biological conditions, the presence of one or more symptoms of a biological condition, one or more combinations thereof, and the like. Based on the correlations determined by the data integration and analysis system using the integrated data repository, groups of individuals who may benefit from one or more treatments can be identified that would not be identified using existing systems. In one or more examples, the processes and techniques implemented to integrate health insurance claims records and genomic data records to generate the integrated data repository can be complex, and efficiency-enhancing techniques, systems, and processes are implemented to minimize the amount of computing resources used to generate the integrated data repository.

[0135] In one or more illustrative examples, a data pipeline system can access information stored by an integrated data repository to generate a data set including a plurality of additional data records that include information related to at least a portion of a group of individuals. In an illustrative example, the additional data records include information indicating whether the individual is included in a group of individuals in which lung cancer is present. The data pipeline system can execute more than one set of different data processing instructions to determine a group of individuals in which lung cancer is present. In various examples, the additional data records can indicate information used to determine the status of an individual with respect to lung cancer, such as one or more transaction insurance identifiers, one or more International Classification of Diseases (ICD) codes, and one or more health insurance transaction dates. In addition to including a column indicating whether an individual is included in a group of lung cancer, the additional data records can also include a column indicating a confidence level of the individual's status with respect to the presence of lung cancer.

[0136] This document describes a computing architecture for merging medical record data into an integrated data repository. In various examples, at least a portion of the operations of the computing architecture can be performed by a data integration and analysis system. In one or more examples, at least a portion of the operations of the computing architecture can be performed by one or more additional computing systems, with at least one of control, maintenance, or implementation of these additional computing systems performed by a service provider that also performs at least one of control, maintenance, or implementation of the data integration and analysis system. In one or more additional examples, at least a portion of the operations of the computing architecture can be performed by multiple servers in a distributed computing environment.

[0137] The computing architecture may include a medical record data repository. The medical record data repository may store medical record data from a plurality of individuals. The medical record data may include imaging information, laboratory test results, diagnostic test information, clinical observations, dental health information, healthcare practitioner notes, medical history forms, diagnostic request forms, medical procedure order forms, medical information charts, one or more combinations thereof, and the like. In various examples, for a given individual, the medical record data repository may store information related to the individual obtained from one or more healthcare practitioners.

[0138] The computing architecture can perform operations that include obtaining data packets from a medical records data repository. In one or more examples, the data packets can be obtained in response to one or more requests sent to the medical records data repository for medical records corresponding to one or more individuals. In one or more additional examples, the computing architecture can obtain the data packets using one or more application programming interface (API) calls. In one or more illustrative examples, the computing architecture can be used to obtain a first data packet, a second data packet, and so on. The individual data packets can correspond to medical records for respective individuals. For example, a first data packet can include medical records for a first individual, a second data packet can include medical records for a second individual, and an Nth data packet can include medical records for an Nth individual.

[0139] A separate data packet may include multiple components. In one or more examples, a separate data packet may include components corresponding to medical records from different healthcare providers. In one or more additional examples, a separate data packet may include components corresponding to different portions of the medical records of one or more healthcare providers. In an illustrative example, a second data packet may include a first component, a second component, and so on. In one or more illustrative examples, the first component may include a first portion of an individual's medical record, the second component may include a second portion of the individual's medical record, and the Nth component may include the Nth portion of the individual's medical record. In various examples, the first component may correspond to a medical record for an individual from a first healthcare provider, the second component may correspond to a medical record for an individual from a second healthcare provider, and the third component may correspond to a medical record for an individual from a third healthcare provider. In one or more additional illustrative examples, the first component may include a first section of an individual's medical record, such as one or more forms related to a diagnostic test or procedure, and the second component may include a second section of an individual's medical record, such as an individual's pathology report.

[0140] In operation, the computing architecture may pre-process individual data packets to identify a corpus of information to be analyzed. In one or more examples, pre-processing the data packets obtained from the medical records data repository may include converting data included in the data packets. For example, pre-processing the data packets may include converting at least a portion of the data obtained from the medical records data repository into machine-encoded information. To illustrate, pre-processing the data packets may include performing one or more optical character recognition (OCR) operations on at least a portion of the data packets obtained from the medical records data repository. By converting at least a portion of the data packets obtained from the medical records data repository into machine-encoded information, the data packets may be subjected to a plurality of operations, such as one or more parsing operations for identifying one or more characters or strings, or one or more editing operations that cannot be performed on at least a portion of the data packets obtained from the medical records data repository.

[0141] In one or more examples, pre-processing of an individual data packet may include determining information included in the individual data packet to be excluded from further analysis by the computing architecture. In various examples, one or more components of the individual data packet may be excluded from the corpus of information to be analyzed. For example, with respect to a second data packet, the computing architecture may determine that a first component is to be excluded from further analysis by the computing architecture. In one or more examples, the computing architecture may analyze the components with respect to one or more keywords to identify at least one of the components for exclusion from further analysis by the computing architecture. In one or more illustrative examples, the computing architecture may parse the components to identify one or more keywords, and in response to identifying the one or more keywords in the components, the computing architecture may determine to exclude the corresponding component from further analysis by the computing architecture. For example, the computing architecture may determine that a first component of the second data packet is a test request form for one or more diagnostic procedures or tests. In these scenarios, the computing architecture may determine that the first component is to be excluded from further analysis by the computing architecture. Furthermore, the computing architecture may determine, based on one or more keywords included in at least one of the second component or the Nth component, that at least one of the second components corresponds to one or more pathology reports for an individual. In these cases, the computing architecture may determine that at least a portion of the second component and / or at least a portion of the Nth component are to be included in the information corpus to be further analyzed by the computing architecture.

[0142] In addition, a subset of the components of each data packet obtained from the medical records data repository may be included in the information corpus. In various examples, one or more additional operations may be performed to narrow the information corpus. For example, one or more queries may be applied to a subset of the information obtained from the medical records data repository. The one or more queries may extract information from one or more data packets that satisfy the one or more queries. In at least some examples, the one or more queries may be a set of queries applied to the components of the data packets. In one or more illustrative examples, the set of queries may determine information to be included in the information corpus and additional information to be excluded from the information corpus. In one or more additional examples, one or more segments of at least one component of the data packet may be excluded from the information corpus.

[0143] In one or more additional illustrative examples, after determining that a first component is to be excluded from further analysis by the computing architecture, the computing architecture can then cause one or more queries to be implemented with respect to at least one of the second component or the Nth component. In these scenarios, the one or more queries can determine that a segment of the second component (such as a segment indicating a family history of one or more biological conditions) is to be excluded from the information corpus. In various examples, the one or more queries can involve identifying multiple keywords and / or combinations of keywords included in at least one of the second component or the Nth component. In these cases, the computing architecture can exclude from the information corpus one or more portions of the components of the data packet that include the one or more keywords or combinations of keywords. In one or more additional examples, the computing architecture can exclude from the information corpus multiple words, multiple characters, and / or multiple symbols that follow one or more keywords in one or more portions of the components of the data packet.

[0144] Furthermore, during operation, the computing architecture may analyze the information corpus to determine characteristics of an individual. In one or more examples, the computing architecture may analyze the information corpus to identify individuals with one or more phenotypes. In various examples, the computing architecture may analyze the information corpus to identify one or more biomarkers indicative of a biological condition. For example, the computing architecture may analyze the information corpus to identify individuals with one or more genetic signatures. The one or more genetic signatures may include at least one of one or more variants in genomic and / or epigenomic regions corresponding to the biological condition. In one or more illustrative examples, the one or more genetic signatures may correspond to one or more variants in genomic and / or epigenomic regions corresponding to a type of cancer. In one or more additional illustrative examples, one or more biomarkers may correspond to analyte levels outside a specified range. For illustration, the computing architecture may analyze the information corpus to identify individuals with levels of one or more proteins and / or one or more small molecules indicative of the biological condition. In these scenarios, the computing architecture may analyze laboratory test results to determine the analyte levels in the individual. In one or more additional examples, the computing architecture can analyze the information corpus to identify individuals with the presence of one or more symptoms indicative of a biological condition. In one or more further examples, the computing architecture can analyze imaging information included in the information corpus to identify individuals with the presence of one or more biomarkers.

[0145] In one or more examples, the computing architecture can implement one or more machine learning techniques to analyze the information corpus. For example, the computing architecture can implement one or more artificial neural networks, such as at least one of one or more convolutional neural networks or one or more residual neural networks, to analyze the information corpus. The computing architecture can also implement at least one of one or more random forest techniques, one or more hidden Markov models, or one or more support vector machines to analyze the information corpus.

[0146] In at least some implementations, a computing architecture can analyze an information corpus by executing one or more queries about the information corpus. The one or more queries can correspond to one or more keywords and / or combinations of keywords. The one or more keywords and / or combinations of keywords can correspond to at least one of characters or symbols corresponding to one or more biological conditions. For illustration, a keyword can correspond to a character associated with a mutation in a genomic and / or epigenomic region, such as HER2. In one or more additional illustrative examples, one or more criteria can be associated with the combination of keywords. For illustration, a criterion corresponding to the combination of keywords can include multiple words that are no more than a specified distance apart from each other in a portion of the individual's information corpus, such as the words "fatigue," "blood pressure," and "swelling" that appear no more than 100 characters apart from each other. In these cases, the computing architecture can parse the information corpus for the one or more keywords and / or combinations of keywords. In various examples, in response to determining the presence of one or more keywords and / or combinations of keywords according to the one or more criteria, the computing architecture can determine that a biological condition exists for a given individual.

[0147] In one or more additional examples, one or more queries can be image-based, and the computing architecture can analyze images included in the information corpus relative to a template image. The template image can be generated based on analyzing multiple images in which a biological condition is present and aggregating the multiple images into a template image. In these scenarios, the computing architecture can analyze images included in the information corpus relative to the one or more template images to determine a similarity measure between the images included in the information corpus and the template images. If the similarity measure for an individual is at least a threshold, the computing architecture can determine that a characteristic of the biological condition is present in the individual.

[0148] After determining individuals having one or more characteristics, the computing architecture can, when operated, generate a data structure that stores data about individuals having the one or more characteristics. In one or more examples, the computing architecture can generate a data table that indicates individuals having individual characteristics and / or individuals having a set of characteristics. For example, the computing architecture can generate a first data table and a second data table. The first data table can indicate individuals having one or more first characteristics, and the second data table can indicate individuals having one or more second characteristics. In one or more illustrative examples, the first data table can indicate individuals having one or more first biomarkers for a biological condition, and the second data table can indicate individuals having one or more second biomarkers for the biological condition. The one or more first biomarkers can correspond to one or more first genomic and / or epigenomic variants associated with the biological condition, and the one or more second biomarkers can correspond to one or more second genomic and / or epigenomic variants associated with the biological condition.

[0149] One or more data structures can be generated from the information corpus, storing identifiers for a portion of the additional set of individuals' subsets and an indication that the portion of the additional set of individuals' subsets corresponds to one or more biomarkers. The one or more data structures can be stored by the intermediate data repository. Before modifying the integrated data repository to store at least a portion of the additional information regarding the medical records of the portion of the additional set of individuals' subsets in association with multiple identifiers, one or more de-identification operations can be performed on the identifiers for the portion of the additional set of individuals' subsets. After de-identifying the information stored by the one or more data structures, the information stored by the integrated data repository can be added to the integrated data repository. In at least some examples, the de-identified medical record information can be added to the integrated data repository in addition to or in lieu of health insurance claims data. In various examples, the one or more data structures storing the de-identified medical record information regarding the biomarker data can have one or more logical connections to other data structures stored in the integrated data repository. To illustrate, one or more data structures storing de-identified medical record information regarding biomarker data can have one or more logical connections to at least one of the following data tables: a first data table that can store information corresponding to a panel used to generate genomic data, mutations in genomic and / or epigenomic regions, mutation types, copy numbers of genomic and / or epigenomic regions, coverage data indicating the number of nucleic acid molecules identified in a sample having one or more mutations, testing dates, and patient information; a second data table that stores data related to one or more patient visits of an individual to one or more healthcare providers; a third data table that stores information corresponding to corresponding services provided to the individual regarding the one or more patient visits to the one or more healthcare providers indicated by the second data table; a fourth data table that stores personal information for a group of individuals; a fifth data table that stores information related to a health insurance company or government entity that paid for services provided to the group of individuals; a sixth data table that stores information corresponding to health insurance coverage information for a group of individuals (such as the type of health insurance plan associated with the group of individuals); or a seventh data table that stores information related to medications obtained by a group of individuals.

[0150] This document describes a machine in the form of a computer system according to an example implementation, in which a set of instructions can be executed within the machine to cause the machine to perform any one or more of the methods discussed herein. For example, in the example form of a computer system, instructions (e.g., software, programs, applications, applet, apps, or other executable code) for causing the machine to perform any one or more of the methods discussed herein can be executed within the computer system. For example, the instructions can cause the machine to implement the previously described architecture and framework and perform the previously described methods. For example, one or more machine-executable components embodied in one or more machines (e.g., embodied in one or more computer-readable storage media associated with one or more machines). Such components, when executed by one or more machines (e.g., processors, computers, computing devices, virtual machines, etc.), can cause the one or more machines to perform the operations described by the instructions. For example, the machine can include a computing device having an analysis component. The analysis can include survival, modeling, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments, the analysis component is embodied in a machine-executable component within a system that includes a variety of electronic data sources and data structures that include information that can be used with the analysis component. Non-limiting examples include data sources and structures such as survival information, genetic information, model data, sub-models, disease node identification and characterization, disease association information, disease subtyping, recurrence, metastasis, time to next treatment, etc.

[0151] The computing device may include or be operatively coupled to at least one memory and at least one processor. The at least one memory stores executable instructions for performing the analysis when executed by the at least one processor. In some embodiments, the memory may also store various data sources and / or structures of the system. In other embodiments, the various data sources and structures of the system may be stored in other memories accessible to the computing device (e.g., at a remote device or system).

[0152] Instruction converts a general, unprogrammed machine, such as a computing device, into a specific machine that is programmed to perform the described and illustrated functions in the described manner. In alternative implementations, the machine operates as an independent device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine can operate with the capacity of a server machine or a client machine in a server-client network environment, or operate as a peer machine in a peer-to-peer (or distributed) network environment. The machine can include but is not limited to a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), an intelligent home device (e.g., an intelligent appliance), other intelligent devices, a web device (appliance), a network router, a network switch, a bridge, or any machine that can sequentially or otherwise execute an instruction, and the instruction specifies the action that the machine will take. Technicians understand that the machine includes individually or jointly executing instructions to perform the set of machines of any one or more methods discussed herein.

[0153] Examples of computing devices may include logic, one or more components, circuits (e.g., modules), or mechanisms. A circuit is a tangible entity that is configured to perform certain operations. In an example, the circuit may be arranged in a specified manner (e.g., internally or relative to an external entity such as other circuits). In an example, one or more computer systems (e.g., stand-alone client or server computer systems) or one or more hardware processors (processors) may be configured by software (e.g., instructions, application portions, or applications) to operate as a circuit that performs certain operations described herein. In an example, the software may (1) reside on a non-transitory machine-readable medium, or (2) reside in a transmission signal. In an example, the software, when executed by the circuit's underlying hardware, causes the circuit to perform certain operations.

[0154] The various operations of the method examples described herein may be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily configured or permanently configured, such processors may constitute processor-implemented circuits that operate to perform one or more operations or functions. In an example, the circuits mentioned herein may include processor-implemented circuits.

[0155] Similarly, the method described herein can be processor-implemented at least in part. For example, at least some or all operations of the method can be performed by one or more processors or a circuit implemented by a processor. The execution of certain operations can be distributed in one or more processors, which can reside not only in a single machine, but also can be deployed on multiple machines. In an example, one or more processors can be located in a single location (for example, in a home environment, an office environment, or as a server farm), and in other examples, the processor can be distributed in multiple locations.

[0156] The one or more processors may also operate to support execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some operations may be performed by a group of computers (as an example of a machine including a processor), which may be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application program interfaces (APIs)).

[0157] Example implementations (e.g., apparatus, systems, or methods) may be implemented in digital electronic circuitry, computer hardware, firmware, software, or any combination thereof. Example implementations may be implemented using a computer program product (e.g., a computer program tangibly embodied in an information carrier or machine-readable medium for execution by, or to control the operation of, data processing apparatus such as a programmable processor, a computer, or multiple computers).

[0158] A computer program may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be executed on one computer or on multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.

[0159] In an example, the operations may be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Examples of method operations may also be performed by, and example apparatus may be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0160] The computing system may include a client and a server. The client and the server are typically remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on corresponding computers and have a client-server relationship with each other. In the implementation of a programmable computing system, it should be understood that both hardware and software architectures need to be considered. Specifically, it should be understood that it may be a design choice to select whether to implement certain functions in permanently configured hardware (e.g., ASIC), temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently configured hardware and temporarily configured hardware. The hardware (e.g., computing device) and software architecture that can be deployed in an example implementation are described below.

[0161] In examples, the computing device may operate as a standalone device, or the computing device may be connected (eg, networked) to other machines.

[0162] In a networked deployment, a computing device can operate in the capacity of a server or client machine in a server-client network environment. In an example, a computing device can act as a peer machine in a peer-to-peer (or other distributed) network environment. The computing device can be a personal computer (PC), a tablet PC, a set-top box (STB), a mobile phone, a web device, a network router, a switch or a bridge, or any machine capable of (sequentially or otherwise) executing an instruction that specifies the action to be taken (e.g., executed) by the computing device. In addition, although only a single computing device is shown, the term "computing device" should also be understood to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more methods discussed herein.

[0163] The computing device may also include a storage device (e.g., a drive unit), a signal generating device (e.g., a speaker), a network interface device, and one or more sensors, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor. The storage device may include a machine-readable medium on which one or more data structures or instructions (e.g., software) are stored, which embody or are utilized by any one or more of the methods or functions described herein. During execution of the instructions by the computing device, the instructions may also reside, in whole or in part, in main memory, in static memory, or in a processor. In an example, one or any combination of the processor, main memory, static memory, or storage device may constitute a machine-readable medium.

[0164] Although the machine-readable medium is shown as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions. The term "machine-readable medium" may also be understood to include any tangible medium that can store, encode, or carry instructions for execution by a machine and cause the machine to perform any one or more of the methods of the present disclosure, or can store, encode, or carry data structures utilized by or associated with these instructions.

[0165] As used herein, a component may refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques for partitioning or modularization that provide specific processing or control functions. Components can be combined with other components via their interfaces to perform machine processes. A component can be a packaged functional hardware unit designed to be used with other components, or it can be part of a program that typically performs a specific function among related functions. A component can constitute a software component (e.g., code contained on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that is capable of performing certain operations and can be configured or arranged in a certain physical manner. In various example implementations, one or more computer systems (e.g., an independent computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured as a hardware component by software (e.g., an application or an application portion) that operates to perform certain operations described herein.

[0166] disease

[0167] The method of the present invention can be used to diagnose the presence of a condition in a subject, to characterize the condition, monitor the response of the condition to treatment, and realize the prognosis of the risk of developing a condition or the subsequent course of the condition. The present disclosure can also be used to determine the efficacy of a specific treatment option. If the treatment is successful, the successful treatment option can increase the amount of nucleic acid (such as cell-free nucleic acid) detected in the blood of the subject, because the patient dies from illness and dysfunction and sheds DNA or otherwise exhibits chronic and acute signs of inflammation. In other examples, this may not occur. In another example, perhaps certain treatment options may be associated with the genetic profile of the disease type and subtype over time. This correlation can be used to select therapy.

[0168] In some embodiments, the methods and systems disclosed herein can be used to identify customized or targeted therapies to treat a patient's specific disease or condition based on classifying nucleic acid variations as somatic or germline in origin. Typically, the disease under consideration is cancer.

[0169] In addition, the method of the present disclosure can be used to characterize the heterogeneity of abnormal conditions in subjects. Such methods can include, for example, generating genomic and epigenomic profiles of extracellular polynucleotides derived from a subject, wherein the genetic profile includes more than one data that can characterize dysfunction and abnormalities (e.g., hypertrophy) associated with myocardial and valvular tissue, and the reduction in blood flow and oxygen supply to the heart is typically a secondary symptom of weakness and / or deterioration of the blood flow and supply system caused by physical and biochemical stress. Examples of cardiovascular diseases directly affected by these types of stress include atherosclerosis, coronary artery disease, peripheral vascular disease, and peripheral arterial disease, as well as various heart diseases and arrhythmias that may represent other forms of disease and functional abnormalities. The method of the present invention can be used to generate or analyze a fingerprint map or data set that is the sum of genetic information from different cells in a heterogeneous disease. The data set can include separate or combined copy number variation, epigenetic variation, and mutation analysis.

[0170] The methods of the present invention can be used to diagnose, prognose, monitor or observe cancer or other diseases. In some embodiments, the methods herein do not involve diagnosis, prognosis or monitoring of a fetus and, therefore, do not involve non-invasive prenatal testing. In other embodiments, these methods can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in unborn subjects, whose DNA and other polynucleotides can co-circulate with maternal molecules.

[0171] Non-limiting examples of other genetically based diseases, disorders, or conditions that are optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), Cry-a-cat syndrome, Crohn's disease, cystic fibrosis, Dercum disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson disease, etc.

[0172] Treatment and related administration

[0173] In certain embodiments, methods disclosed herein relate to identifying customized therapy and administering customized therapy to patients in view of the state that nucleic acid variation is somatic cell origin or germline origin. In some embodiments, substantially any cancer therapy (e.g., surgical therapy, radiotherapy, chemotherapy and / or similar therapy) can be included as a part of these methods. Typically, customized therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method for enhancing the immune response for a given cancer type. In certain embodiments, immunotherapy refers to a method for enhancing the T cell response for a tumor or cancer.

[0174] In certain embodiments, the nucleic acid variation of the sample from the subject is a state of somatic cell origin or germline origin and can be compared with a database of comparator results from a reference population to identify a customized or targeted therapy for the subject. Typically, the reference population includes patients with the same cancer or disease type as the subject being tested and / or patients who are receiving or have received the same therapy as the subject being tested. When the nucleic acid variation and comparator results meet certain classification criteria (e.g., basic or approximate matching), customized or targeted therapy (or more than one treatment) can be identified.

[0175] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutics are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutics, etc.) can also be administered by methods such as, for example, buccal, sublingual, rectal, vaginal, intraurethral, ​​topical, intraocular, intranasal, and / or intraaural, which may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.

[0176] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. It is not intended that the present invention be limited to the specific examples provided in this specification. Although the present invention has been described with reference to the description mentioned above, the description and illustration of the embodiments herein are not intended to be interpreted in a restrictive sense. A person skilled in the art will now appreciate various variations, changes, and substitutions without departing from the present invention. In addition, it will be understood that all aspects of the present invention are not limited to the specific description, configuration, or relative proportions set forth herein according to various conditions and variables. It will be understood that various alternatives to the embodiments of the present disclosure described herein may be adopted in practicing the present invention. It is therefore contemplated that the present disclosure should also encompass any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the present invention and thus encompass methods and structures within the scope of these claims and their equivalents.

[0177] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent to those skilled in the art upon reading this disclosure that various changes in form and detail may be made without departing from the true scope of the disclosure and that the present invention may be practiced within the scope of the appended claims. For example, all methods, systems, computer-readable media, and / or component features, steps, elements, or other aspects may be used in various combinations.

[0178] Biomarkers

[0179] The present disclosure provides methods for diagnosing, prognosing, and selecting a therapy for a subject suffering from a disease (e.g., heart failure, cardiovascular disease, cancer, etc.) using biomarkers. A biomarker can be any gene or gene variant whose presence, mutation, deletion, substitution, copy number, or translation (i.e., translation into protein) is an indicator of a disease state. Biomarkers of the present disclosure may include presence, mutation, deletion, substitution, copy number, or translation in any one or more of EGFR, KRAS, MET, BRAF, MYC, NRAS, ERBB2, ALK, Notch, PIK3CA, APC, and SMO.

[0180] Biomarkers are genetic variants. Biomarkers can be identified using any of several resources or methods. Biomarkers can be previously discovered or discovered de novo using experimental or epidemiological techniques. When a biomarker is highly correlated with a disease, detection of the biomarker can indicate the disease. When a biomarker in a region or gene occurs at a frequency greater than that in a given background population or dataset, detection of the biomarker can indicate cancer.

[0181] Publicly available resources such as scientific literature and databases can describe genetic variants in detail.Scientific literature can describe the experiment or genome-wide association study (GWAS) of one or more genetic variants associated.Databases can gather information collected from sources such as scientific literature to provide a more comprehensive resource for determining one or more biomarkers.Non-limiting examples of databases include FANTOM, GTex, GEO, Body Atlas, INSiGHT, OMIM (online human Mendelian inheritance, omim.org), cBioPortal (cbioportal.org), CIViC (clinical explanation of cancer variants, civic.genome.wustl.edu), DOCM (database of selected mutations, docm.genome.wustl.edu) and ICGC data portal (dcc.icgc.org).In another example, COSMIC (somatic mutation catalogue in cancer) database allows biomarkers to be searched by cancer, gene or mutation type.It is also possible to determine biomarkers from the beginning by conducting experiments such as case control or association (for example, genome-wide association study) research.

[0182] One or more biomarkers can be detected in a sequencing panel. A biomarker can be one or more genetic variants. A biomarker can be selected from single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (e.g., insertions / deletions), gene fusions, and inversions. A biomarker can affect the level of a protein. A biomarker can be in a promoter or enhancer and can change the transcription of a gene. A biomarker can affect the transcription and / or translation efficiency of a gene. A biomarker can affect the stability of the transcribed mRNA. A biomarker can cause changes in the amino acid sequence of a translated protein. A biomarker can affect splicing, can change the amino acid encoded by a specific codon, can cause frameshifting, or can cause premature termination codons. A biomarker can cause conservative substitutions of an amino acid. One or more biomarkers can cause conservative substitutions of an amino acid. One or more biomarkers can cause non-conservative substitutions of an amino acid.

[0183] The frequency of a biomarker can be as low as 0.001%. The frequency of a biomarker can be as low as 0.005%. The frequency of a biomarker can be as low as 0.01%. The frequency of a biomarker can be as low as 0.02%. The frequency of a biomarker can be as low as 0.03%. The frequency of a biomarker can be as low as 0.05%. The frequency of a biomarker can be as low as 0.1%. The frequency of a biomarker can be as low as 1%.

[0184] A single biomarker may not be present in more than 50% of subjects with cancer. A single biomarker may not be present in more than 40% of subjects with cancer. A single biomarker may not be present in more than 30% of subjects with cancer. A single biomarker may not be present in more than 20% of subjects with cancer. A single biomarker may not be present in more than 10% of subjects with cancer. A single biomarker may not be present in more than 5% of subjects with cancer. A single biomarker may be present in 0.001% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 30% of subjects with cancer. A single biomarker may be present in 0.01% to 20% of subjects with cancer. A single biomarker may be present in 0.01% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 5% of subjects with cancer.

[0185] Genetic analysis

[0186] Genetic analysis includes detecting nucleotide sequence variants and copy number variation.Genetic variants can be determined by order-checking.Sequencing method can be large-scale parallel sequencing, i.e., simultaneously (or rapidly and continuously) sequencing at least 100,000, 1,000,000, 10,000,000, 100,000,000 or 1,000,000,000 polynucleotide molecules any one.Sequencing method can include but is not limited to: high-throughput sequencing, pyrosequencing, synthesis sequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, connection sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next generation sequencing, single molecule synthesis sequencing (SMSS) (Helicos), large-scale parallel sequencing, cloned single molecule array (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger order-checking, primer walking, use PacBio, SOLiD, Ion Torrent or nanopore platform order-checking and any other sequencing method known in the art.

[0187] Sequencing can be made more efficient by performing sequence capture, i.e., enriching a sample for target sequences of interest, e.g., sequences comprising KRAS and / or EGFR genes or portions thereof containing sequence variant biomarkers. Sequence capture can be performed using immobilized probes that hybridize to the target of interest.

[0188] Cell-free DNA can include small amounts of tumor DNA mixed with germline DNA. Sequencing methods that increase the sensitivity and specificity of detecting tumor DNA, and in particular genetic sequence variants and copy number variations, can be used in the methods of the present invention. Such methods are described, for example, in WO 2014 / 039556. These methods can not only detect molecules with a sensitivity of up to or greater than 0.1%, but can also distinguish these signals from the noise typical in current sequencing methods. Various methods can be used to achieve increased sensitivity and specificity from blood-based cfDNA samples. One method includes efficiently labeling DNA molecules in the sample, for example, labeling at least 50%, 75% or 90% of the polynucleotides in the sample. This increases the likelihood that low-abundance target molecules in the sample will be labeled and subsequently sequenced, and significantly increases the sensitivity of target molecule detection.

[0189] Another approach involves molecular tracking, which identifies sequence reads that have been redundantly generated from the original parental molecules and assigns the most likely identity of the base at each locus or position in the parental molecule. This significantly increases the specificity of detection by reducing the noise generated by amplification and sequencing errors, which reduces the frequency of false positives.

[0190] The methods of the present disclosure can be used to detect genetic variation in initial starting genetic material (e.g., rare DNA) that is non-uniquely tagged with a specificity of at least 99%, 99.9%, 99.999%, 99.9999%, or 99.99999% at a concentration of less than 5%, 1%, 0.5%, 0.1%, 0.05%, or 0.01%. Sequence reads of the tagged polynucleotides can then be tracked to generate a consensus sequence for the polynucleotides with an error rate of no more than 2%, 1%, 0.1%, or 0.01%.

[0191] In other examples, the gene of interest can be amplified using primers that identify the gene of interest. Primers can hybridize with genes upstream and / or downstream (for example, upstream of the mutation site) of a specific region of interest. Detection probes can hybridize with amplified products. Detection probes can hybridize with wild-type sequences or with mutation / variant sequence specificity. Detection probes can be labeled with a detectable marker (for example, with a fluorophore). Detection of wild-type or mutant sequences can be carried out by detecting a detectable marker (for example, fluorescence imaging). In the example of copy number variation, the gene of interest can be compared with a reference gene. The copy number difference between the gene of interest and the reference gene can indicate amplification or deletion / truncation of a gene. Examples of platforms suitable for carrying out methods described herein include digital PCR platforms, such as, for example, Fluidigm digital arrays.

[0192] This paper describes the method for analyzing nucleic acid sequence information.In various embodiments, analytical method includes one or more models, and each in one or more models includes one or more of the survival, sub-modeling, disease node determination and identification (for example, driving mutation), disease association, disease subtyping, recurrence, transfer, the time to next treatment etc. as a separate component.In various embodiments, model includes hierarchical model (for example, nested model, multilevel model), mixed model (for example, regression, such as logistic regression and Poisson regression, pooling, random effect, fixed effect, mixed effect, linear mixed effect, generalized linear mixed effect), risk model, odds ratio model and / or repeated samples (for example, repeated measurement, such as ANOVA).In various embodiments, model is hierarchy random effects model.In various embodiments, model is hierarchy cubic spline random effects model.In various embodiments, model is cubic spline model.In various embodiments, model is generalized linear effects model.In various embodiments, model is linear effects model.In various embodiments, model is Cox proportional hazards model.In various embodiments, analytical method includes assembling the models together.In various embodiments, assembling includes the generation of associated parameters. In one or more embodiments, analytical method comprises patient survival information and patient genetic information.As an example, model assembly together can comprise the different models of the different types of cancer (comprising subtype) represented in patient survival information.Each in different models can be configured to determine the correlation between the survival time of the patient of the cancer of the respective type that they are configured to assess and genetic factors and are diagnosed with.For example, can recommend the genetic factors that are determined to have strong correlation with cancer survival time (for example, relatively short survival time and / or relatively long survival time) as potential therapeutic target.

[0193] In various embodiments, analysis can include one or more of survival, submodeling, disease node determination and identification (e.g., driver mutation), disease association, disease subtyping, recurrence, metastasis, the time to next treatment, etc. as separate components. For example, modeling can be beneficial to application as mentioned above, such as patient survival information and patient genetic information. In various embodiments, the submodeling component can determine the subset of patient survival information and patient genetic information for generating different patient groups associated with different types of cancer and cancer subtypes. In various embodiments, the submodel includes a hierarchical model (e.g., a nested model, a multilevel model), a mixed model (e.g., regression, such as logistic regression and Poisson regression, pooling, random effects, fixed effects, mixed effects, linear mixed effects, generalized linear mixed effects), a risk model, an odds ratio model, and / or repeated samples (e.g., repeated measurements, such as ANOVA). In various embodiments, the submodel is a hierarchical random effects model. In various embodiments, the submodel is a hierarchical cubic spline random effects model. In various embodiments, the submodel is a cubic spline model. In various embodiments, the submodel is a generalized linear effects model. In various embodiments, the submodel is a linear effects model. In various embodiments, the submodel is a Cox proportional hazards model. Each subset of patient survival information and patient genetic information can include information of patients diagnosed with different types of cancer and cancer subtypes. For example, the submodeling component can also apply the subset of patient survival information and patient genetic information to corresponding individual survival models developed for different cancer types (including subtypes). In various embodiments, the information generated for the analysis method can be stored in memory (e.g., as model data). In various embodiments, the information generated for the analysis method generates one or more survival models for individual subjects.

[0194] In various embodiments, a survival model is used to analyze patient survival information and patient genetic information, including a disease node determination and identification component, to identify disease nodes included in the patient genetic information for each type of cancer, the disease nodes being involved in the genetic mechanisms used by the corresponding cancer type for proliferation. In various embodiments, the disease node component identifies disease nodes based on observed correlations between genetic factors and cancer survival times provided in the patient survival information. For example, genetic factors that are often observed to be associated with short survival times of a particular type of cancer and less often observed to be associated with long survival times of the particular type of cancer can be identified as active genetic factors that have an active role in the genetic mechanisms of the particular type of cancer (including subtypes).

[0195] In various embodiments, disease node determination and identification include disease association parameters about the association between different cancer types, to promote identification of active genetic factors associated with different cancer types. For example, highly associated cancer types can share one or more common key potential genetic factors. As readily understood by those of ordinary skill, the model (e.g., survival model) of the associated cancer type dialectically exchanges information to determine and / or identify active genetic factors across cancer types (including subtypes). In various embodiments, disease association parameters applied by disease node determination and identification are promoted by modeling. In various embodiments, the generation of individual survival models can adopt one or more machine learning algorithms to promote survival, modeling, disease nodes associated with specific types of cancer (including subtypes) based on patient genetic information and disease association parameters.

[0196] In some embodiments, the node determination and identification of cancer types (including subtypes) are associated with a scoring system that includes determining disease nodes. For example, the scoring of disease nodes for a particular type of cancer (including subtypes) reflects the association of the survival time of the disease node with the particular type of cancer (including subtypes). In various embodiments, scoring can be based on the frequency of direct or indirect identification of specific genetic factors for patients diagnosed with a particular cancer type. In various embodiments, analysis includes survival, submodeling, disease node determination and identification (for example, driver mutations), disease association, disease subtyping, recurrence, metastasis, time to the next treatment, etc. mentioned above, which can be related to a threshold value less than a definition, greater than a definition. For example, the greater the score associated with a disease node and a cancer type (including subtypes), the greater the contribution of the disease node to survival time. In various embodiments, information about the disease node of the corresponding type of cancer (including subtypes) can be collated in a data structure such as a database, and the score determined for active genetic factors.

[0197] This paper describes the analytical method comprising effect modeling. In various embodiments, effect modeling includes random effects, fixed effects, mixed effects, linear mixed effects and generalized linear mixed effects. In various embodiments, effect includes cubic splines. In various embodiments, effect modeling includes regression. In various embodiments, effect modeling includes logistic regression and Poisson regression. In various embodiments, model does not include covariates. In various embodiments, model includes covariates. In various embodiments, covariates are information from medical records (including laboratory test records, such as genome, epigenome, nucleic acid and other analyte results), insurance records, etc. Example includes age, treatment line, smoking status (yes / no), sex and multiple scoring and / or staging systems for specific cancer disease patients, wherein illustrative examples include age (in years), anti-EGFR treatment line, smoking status (yes / no), sex (female / male) and Van Walraven Elixhauser comorbidity (ELIX) score (expressed as a weighted measure across multiple common comorbidities) specific to lung cancer patients. A skilled artisan will readily appreciate that covariates can include any number of data elements for individuals and individuals in a population, such as data elements from medical records (including laboratory test records, such as genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc.

[0198] In various embodiments, the analysis method includes generating a hierarchy comprising at least one first-order equation. In various embodiments, the first-order equation includes a truncated cubic spline. In various embodiments, the truncated cubic spline includes longitudinal data. This includes, for example, direct or indirect measurements of ctDNA levels, allele fractions, tumor fractions. In various embodiments, additional first-order equations include covariates. In various embodiments, covariates are information about individuals or groups extracted and / or stored from medical records (including laboratory test records, such as genome, epigenome, nucleic acid and other analyte results), insurance records, etc. Examples include age, treatment line, smoking status (yes / no), sex, and various scoring and / or staging systems that have been used for patients with specific cancer diseases. In various embodiments, a velocity map is generated. In various embodiments, the velocity map is a derivative or one or more equations, such as at least one first-order equation. In various embodiments, the analysis method includes one or more of equations (1), (2), and (3) described in the examples.

[0199] This article describes an analysis method that includes jointly solving different analysis components, including one or more of survival, modeling and sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. In various embodiments, the analysis method includes jointly solving one or more different models of different cancer types under a joint model framework. For example, the analysis method may include jointly solving one or more different survival models of different cancer types under a joint model framework. In various embodiments, the method includes determining association parameters. In various embodiments, association parameters include, for example, the relationship between the estimated current value of the patient's survival and the patient's biomarker, the relationship between the patient's survival and the patient's estimate of the biomarker current change over time. In various embodiments, this includes slope, and the relationship between the current estimated area under the longitudinal trajectory of the subject as a surrogate for the cumulative effect of the biomarker. It is readily understood by those skilled in the art that association parameters can take various forms and can also be combined. For example, the relationship between the estimated current value of the total survival and the estimated current slope of the patient's longitudinal trajectory can be examined.

[0200] Example 1 - Joint Modeling

[0201] The inventors applied joint modeling (JM) of longitudinal data and time-to-event data in combination with next-generation sequencing (NGS) genetic testing to demonstrate the ability to detect biomarkers (or several biomarkers) that are associated with a specific patient's probability of survival over time. Detecting and characterizing genomic biomarkers using this method and technology illustrates how the evolution of such biomarkers can be associated with and predict patient survival. As an example, this real-world application of joint modeling has produced a patient-level monitoring system designed to enhance clinician decision-making capabilities.

[0202] Notably, JM includes the ability to appropriately accommodate endogenous time-varying covariates. Since most biomarkers fall into this category, this leads to a reduction in parameter estimation bias, improved statistical inference, and the ability to perform dynamic patient-level predictions, where predictions are based on a partial or complete biomarker history. Joint modeling is flexible because both frequentist and Bayesian approaches have been developed. Here, for computational efficiency, the inventors employed a Bayesian approach based on a Markov Chain Monte Carlo sampling algorithm.

[0203] Example 2 - Genetic Testing via Next Generation Sequencing

[0204] The inventors selected a patient cohort from a real-world evidence database comprising real-world outcomes, anonymized genomic data, and structured payer claims data for >240,000 patients. For indicative purposes, the different target populations in this dataset included patients individually diagnosed with non-small cell lung cancer (NSCLC) with EGFR L858R mutation, colorectal cancer (CRC) with KRAS G12D and KRAS12V. Due to the longitudinal component of this study, only patients with at least three time measurements were included. After meeting these criteria, the resulting cohort consisted of 252 patients. The biomarkers of interest, i.e., longitudinal outcomes, were the patient's mutant allele frequency (AF) and tumor fraction (TF), where we intended to correlate the progression of these biomarkers over time with patient survival.

[0205] Example 3 - Method

[0206] The joint modeling framework was divided into two sub-models for evaluation, wherein, after the sub-models were analyzed, the information from these sub-models was combined with the goal of determining whether there was an association between the two. More specifically, the first sub-model focused on providing adequate representation of longitudinal data (at the patient level), and the second sub-model evaluated patient survival. In this study, generalized linear mixed models (GLMMs) were used to evaluate the temporal progression of each biomarker, while Cox proportional hazards (CPH) models examined patient survival. It is important to note that because the distribution of each biomarker was highly skewed, in order to be consistent with the GLMM normality assumption, the analysis was based on the logarithmic transformation of both AF and TP. In addition, because several patients showed complex biomarker progression, a cubic spline model was used to describe the patient-level response. In addition, because factors such as age and sex are often confounded with survival, these factors were included in the CPH model to serve as statistical controls.

[0207] As readily understood by those of ordinary skill, the methods and techniques described herein support determining the association between longitudinal data and time-to-event data referred to as an association structure. Examples of association structures include, but are not limited to, the relationship between the estimated current value of a patient's survival and the patient's biomarker, the relationship between the current change (e.g., slope) of the patient's survival and the patient's estimate of a biomarker over time, and the relationship between the current estimated area under the patient's longitudinal trajectory, which is typically used as a substitute for the cumulative effect of a biomarker. Association structures can take various forms and can also be combined. For example, the relationship between the estimated current slope of the total survival and the estimated current value plus the patient's longitudinal trajectory can be examined. Here, the inventors have described association structures of current value, slope, and their combination, but it will be appreciated by those skilled in the art that a large number of association structures (many of which are not explicitly mentioned above but readily known to those skilled in the art) can be used for exploration. After establishing suitable JMs for each biomarker, these JMs will then be used to inform dynamic predictions. That is, overall survival for each patient is predicted depending on the nature of the association structure between the longitudinal data and the time-to-event data, or more specifically, the survival of a given patient is predicted using the measures captured up to a given time point, and as additional measures are collected, the patient survival prediction is adjusted accordingly, hence the term "dynamic prediction."

[0208] Example 4 - Statistical Analysis

[0209] All statistical analyses were performed using R version 4.1.3, with the JMBayes2 package performing joint modeling. As previously mentioned, due to the longitudinal component of the study, each patient had at least three time measurements, with the first measurement coinciding with the patient's initial Guardant360 test and the remaining measurements modeled accordingly. A total of 252 patients met these criteria, resulting in 909 measurements collected for AF and TP, respectively, spanning from November 19, 2014, to September 30, 2022. The distribution of each biomarker was Figure 1 The results are given in Table 1, followed by the associated summary statistics (see Table 1). The JM results indicate that the recent change of each biomarker over time is associated with patient survival (AF: p value = 0.0139; TF: p value = 0.0332). Through these associations, a graphical representation of the patient-level survival curve can be displayed to evaluate clinical outcomes based on the patient's unique biomarker evolution.

[0210] Table 1. Summary statistics of allele frequencies and tumor fractions

[0211] Biomarkers Minimum 1Q median average value 3Q Maximum Standard Deviation Allele frequency 0.03 0.60 2.90 11.88 15.50 93.20 18.88 Tumor score 0.04 1.00 4.00 9.44 14.80 84.10 12.28

[0212] Example 5 - Results

[0213] The distribution of allele frequencies and tumor fractions is shown in Figure 1 To supplement the descriptive statistics, patient-level longitudinal data for each biomarker were presented in Figure 2 The results are shown as spaghetti plots in , which illustrate the complexity of the patient-level longitudinal progression of each biomarker and reinforce the skewed nature of the data. To adhere to the normality assumption required by the GLMM, a logarithmic transformation was applied to each biomarker, and to compensate for the complexity observed in the patient-level evolution, natural cubic splines were used to model the longitudinal characteristics of each patient within the GLMM structure. The results of the fitted GLMM for each patient are shown in Figure 3 , and the fixed and random effects of the GLMM for each biomarker are depicted in Figure 8 Since the biomarkers were collected on the same group of patients, only a single CPH model needed to be fitted. Of the 252 patients, 99 experienced an event (death), while the remaining observations were censored. Since both age and sex are often confounded with survival, the initial CPH model included these covariates as statistical controls. However, analysis of the initial model revealed that both age (p value = 0.519) and sex (p value = 0.310) were statistically insignificant at the 0.05 level. Similarly, models that included age and sex alone produced similar results (age, p value = 0.56) and (sex, p value = 0.33). Subsequently, a null CPH model (a model without covariates) was used for joint modeling.

[0214] Example 6-Fitting

[0215] The results of the fitted cubic spline-based GLMM for log-transformed biomarkers are shown in Figure 3 Three JMs were analyzed for each biomarker, each matching the aforementioned association structures (and their combinations). Since the analysis was performed under a Bayesian paradigm, care was taken to ensure that the model parameters were accurately estimated. In doing so, each model consisted of two chains, each with 9000 burn-in iterations followed by 90,000 iterations, and a dilution factor of 3 was implemented to account for potential autocorrelation issues. Similarly, inspection of the trace plots provides a visual construct that the model parameters have fully converged. Tables 2 and 3 summarize the joint modeling results for each corresponding biomarker (numbered 1-3). Due to the Bayesian approach, 95% belief intervals are reported rather than frequentist confidence intervals.

[0216] The results in Tables 2 and 3 reveal a second JM for each biomarker, showing promise as evidenced by the corresponding p-values ​​(0.0139 and 0.0332), indicating an association between current slope and patient survival. Further information can also be extracted from these tables. That is, hazard ratios corresponding to their respective association structures can be calculated. For example, referring to the mean values ​​in Table 2, if the current allele frequency change rate increases by 10% over 100 days, the resulting hazard ratio is 1.19, meaning that the risk of death associated with this increase increases by 19%. Similar calculations can be performed for the maximum tumor percentage.

[0217] Table 2. Joint modeling results of logarithmic transformation of allele frequencies

[0218]

[0219] 1. Inspection of the density plots shows that even when convergence is demonstrated, many posterior distributions are skewed, which means that the deviance information criterion (DIC) may not be suitable for model comparison because the distribution of the combined density plot is not multivariate normal. However, DIC is included above because it is often reported.

[0220] 2.SD stands for standard deviation.

[0221] Table 3. Joint modeling results of logarithmic transformation of tumor percentage

[0222]

[0223] 1. Inspection of the density plots shows that even when convergence is demonstrated, many posterior distributions are skewed, which means that the deviance information criterion (DIC) may not be suitable for model comparison because the distribution of the combined density plot is not multivariate normal. However, DIC is included above because it is often reported.

[0224] 2.SD stands for standard deviation.

[0225] Example 7 - Dynamic Prediction

[0226] HR is a good indicator of overall trends, but from the perspective of precision medicine, the real advantage of the JM method is to produce dynamic predictions. Since the concept of dynamic prediction is best understood through visual representation, Figure 4 and Figure 5 A graphical description of the process is provided in .

[0227] Figure 4The top graph in depicts a longitudinal trajectory associated with a patient's biomarker evolution (as seen from the blue line), where the trajectory adjusts accordingly as additional metrics are captured. It is important to note that the focus is on the current slope of the trajectory, as the JM used to create dynamic predictions is built on this correlation structure. In this example, we examined the time range spanning from 0 days to 300 days, 600 days, and 900 days, respectively. Directly below each trajectory, i.e., the bottom graph, is the matched survival curve. Note that each curve is updated as new biomarker information becomes available. For example, from 0 days to 300 days, the trajectory of patient 106 is as follows: Figure 4 The indicated decrease. By examining the corresponding survival curve, if we extrapolate, for example, 1000 days, i.e., assess survival at 1300 days, the patient's probability of survival is approximately 0.71 or 71%. Similarly, at 600 days, additional biomarker values ​​are captured, which changes the trajectory, wherein even if the trend remains downward, the slope is now increasing. Evaluating outside 1000 days (at 1600 days), we see that the patient's estimated survival has dropped by 6%, from 71% to 65%. Such a result is expected because, in general, as the slope increases, survival decreases. Finally, due to the slight increase in the slope caused by the last set of measurements collected for up to 800 days, the patient's estimated survival has dropped slightly, from 65% to 64%. A 1000-day prediction was used here; However, regardless of the expected time frame, the survival trend is still relatively comparable.

[0228] In contrast to patient 106, patient 94 (see Figure 5 The slope of the trajectory for the HR remains fairly consistent over the time span considered, although a slight increase in slope is observed. Therefore, we should expect to see minimal changes in the survival probability. If we extrapolate to 1000 days as described above, the expected survival probabilities are 71%, 70%, and 69%, respectively, which is consistent with expectations. Similar dynamic predictions can be made based on the maximum tumor percentage, as with the HR calculation.

[0229] Using the methods and techniques described herein, JM results showed that the recent change in each biomarker over time was associated with patient survival (AF: p value = 0.0139; TF: p value = 0.0332). Through these associations, a graphical representation of patient-level survival curves can be displayed to assess clinical outcomes based on the patient's unique biomarker evolution.

[0230] Example 8 - Discussion

[0231] In addition to the many available JM options, dynamic prediction capabilities are particularly beneficial because they are very suitable for enhancing the decision-making ability of clinicians. This is because, in real medical environments, patient conditions are constantly changing, and therefore, making informed decisions using the latest available data is usually in the best interests of the patient. As shown, JM essentially captures the ever-changing patient landscape, and as changes occur, JM adapts accordingly. Therefore, by utilizing JM's ability to link the latest information with patient survival, clinicians can modify and / or adjust treatment plans with the ultimate goal of improving patient survival. In addition, the application of methods such as JM supports the generation of large amounts of genetic data. Those skilled in the art will understand that there are many biomarkers, cancer types, and mutations that can be used for research, because the analysis performed here can be applied to other cancer types and mutations, and in the process, additional relevant biomarkers can be identified. This approach supports the creation of patient-specific monitoring systems that are customized for both specific cancer types and mutation combinations.

[0232] Example 9 - Hierarchical cubic spline random effects model

[0233] This paper describes the use of a hierarchical cubic spline random effects model (HCSREM) applied to a retrospective real-world cohort of patients diagnosed with advanced non-small cell lung cancer (NSCLC). Here, the interest is in ctDNA levels, as measured by the maximum variant allele fraction of all somatic variants detected by liquid biopsy, although those of ordinary skill understand that the proposed framework can be applied to longitudinal biomarkers, combinations of biomarkers. A major advantage of this approach is the ability to incorporate patient information, taking into account several relevant covariates. Finally, to enhance interpretation, the model results are graphically presented in the form of estimated longitudinal predictions, each based on a set of different traits of the patient. In the process, patient-level predictions are directly compared, with comparisons enhanced by subsequently defined velocity plots.

[0234] Example 10 - Data Sources and Patient Groups

[0235] The cohort used to illustrate the utility of this approach is based on observational data and is derived from the Real-World Evidence Anonymous Clinicogenomic Database, which includes structured commercial payer claims collected from inpatient and outpatient facilities in both academic and community settings.

[0236] Patients selected for this cohort were diagnosed with advanced non-small cell lung cancer (NSCLC) and had at least three genomic liquid biopsy tests in the United States between June 1, 2014, and June 30, 2023. Only patients receiving EGFR mutation-targeted therapy were included, and the following treatments were considered: osimertinib, afatinib, dacomitinib, erlotinib, gefitinib, and ervantumab. All patients were required to have at least three blood samples while on a specific anti-EGFR treatment line, or within 30 days before and 30 days after the start of a treatment line. Patients who had their first genomic test on a treatment line more than 120 days after the start of a treatment line were excluded. For patients with multiple treatment lines that met these criteria, the earliest treatment line was selected for inclusion in the study. Finally, patients with suspected germline mutations were removed from the cohort.

[0237] Example 11 - Response Variables and Study Covariates

[0238] The response variable, i.e., the ctDNA measurement results captured over time, are reported as percentages. In the case where the sample contains ctDNA levels below the detection limit of the assay, the value is replaced by a ctDNA level of 0.04% (the lowest value in the group and consistent with the detection limit of the test). All covariates except death are captured at baseline, where the baseline period is defined as six months before the index date (i.e., the date of the patient's first genomic test). Baseline covariates include age (in years), anti-EGFR treatment line, smoking status (yes / no), sex (female / male), and the Van Walraven Elixhauser comorbidity (ELIX) score specific to lung cancer patients (expressed as a weighted measure across multiple common comorbidities). Because the group is based on real data, it is impossible to directly align the treatment start date with the patient's first genomic test as can be achieved in a prospective study. Therefore, the number of days between the first genomic test and the start of treatment is added as a covariate to serve as a statistical control and is set to zero days in the analysis to simulate the post-treatment situation. Also included is the mortality rate of patients captured as survival and death within the study time frame.

[0239] Example 12 - Exemplary Statistical Model

[0240] This paper describes the mathematical details of the HCSREM, which is sufficiently scalable to capture variable nonlinear trends and allows for the direct incorporation of patient characteristics in the form of covariates. In addition to these properties, the model can also provide unique temporal ctDNA patterns for each combination of covariate values. It is the ability to provide this type of patient-specific information that makes this approach attractive in targeted oncology efforts.

[0241] The model is partitioned into first-order and second-order equations, which create a hierarchical structure. The first-order equation takes the form of a truncated cubic spline and captures how the ctDNA level of a particular patient changes over time (see equation (1)). At higher levels, this is achieved by creating a function that is divided into individual segments across the horizontal axis. In each segment, a cubic polynomial is used to fit the data, where the ends of consecutive cubic polynomials are connected by knots. Although there are "automated" methods for determining the amount and placement of knots, the knot locations and the number of knots can be strategically designed based on data inspection. Ultimately, the cubic spline model combines the separated segments to form a single uniform function to represent the data.

[0242]

[0243] in

[0244] In Equation (1), the ctDNA measurement (or its transformation) captured over time is given by Y ij 's, where i is used to index the patient and j is the measurement occasion. The time point captured within the patient is represented by t ij Given, ∈ is the value of the kth knot, π ri ′s are r response parameters, where each, i.e., π 0i ,π 1i ,…,π (k+3)i varies between patients, i.e., a random effect, and ε ij is the error term and is assumed to have mean 0 and variance σ 2 The response parameters are particularly important because they collectively control the shape of each patient's unique longitudinal ctDNA trajectory and serve to bridge the first- and second-level equations.

[0245] The significance of the second-level equations is that they contain information about individual patient characteristics and relate these characteristics to the response parameter itself. The second-level equations are given below.

[0246]

[0247] Among them, X ci represents the desired patient characteristics, β rc Capturing the linear relationship between the response parameter and patient characteristics, β r0 is each corresponding π ri The intercept, and e ri represents the random component and is assumed to follow the following multivariate normal distribution:

[0248]

[0249] When a model includes covariates, it is called a conditional model, otherwise it is an unconditional model. The unconditional model provides group-level results, while the conditional model is responsible for producing patient-level results.

[0250] Furthermore, the velocity plots described are of interest when it is useful to examine the direction and speed of change in ctDNA levels at a given time point, i.e., the instantaneous rate of change (IRC). Each model-generated patient trajectory has a cubic spline at its center. An advantageous property of cubic splines is that they are twice differentiable, and therefore, the IRC at a given time point can be calculated. In the case of the spline model employed, this is equivalent to taking the first derivative of Equation (1) with respect to time, yielding:

[0251]

[0252] The value of the IRC is given by the slope of a line tangent to the patient's trajectory, where positive values ​​correspond to increasing IRC, negative values ​​correspond to decreasing IRC, and an IRC value of zero indicates reaching a peak or trough, or a flattening of the trajectory. The further the IRC value is from zero, the more extreme the rate of change.

[0253] Example 13 - Statistical Analysis and Results

[0254] Data were extracted using SAS software package 9.4 (SAS Institute, Cary, NC, USA), and all statistical analyses for HCSREM were performed using R version 4.1.3. A total of 400 patients with advanced NSCLC who had undergone at least three G360 tests were identified from the GuardantINFORM database. 73 patients were excluded because their first test was more than 120 days after the start of treatment, and 5 patients were excluded due to germline mutations. Of the remaining patients, 163 received anti-EGFR treatment with a total of 561 ctDNA longitudinal measurements, of which these 163 patients defined the cohort used in the analysis. The average age of these patients was 62 years, 66% of them were female, the average line of anti-EGFR treatment was 1, and the average time between the G360 test and the start of treatment was 0 days (-115 days to 30 days range) (Table 4).

[0255] Table 4. Summary of Patient Characteristics

[0256] Characteristics (total N = 163) N / average % / standard deviation Age (years) 61.18 10.88 female 108 66% ELIX Rating 1.89 1.86 Current or former smokers 123 75% Anti-EGFR treatment line 1.44 0.99 Time between G360 testing and treatment initiation (days) 0.29 31.98 ctDNA (%)* 5.66 10.59 Death at the end of the study period 55 33%

[0257] *ctDNA values ​​were extracted and summarized from each test, thus including multiple ctDNA values ​​per patient

[0258] like Figure 9As shown in , the inventors developed an unconditional model fit to the transformed data using knots placed at 50, 125, 250, 500, 750, 1000, and 1250 days. To ensure consistency, other knot orientations were explored, although the different orientations did not change the results much. The results are presented graphically because the spline model parameter estimates are difficult to interpret, although parameter estimates and related outputs are provided in the Supplementary Information for reference. The graphical representation of the unconditional model, called the response pattern, is presented in Figure 10 Here, the black curve represents the response pattern of the cohort, while each black point represents the ctDNA level value. The purple area represents the 95% confidence band of the estimated trajectory.

[0259] The response pattern shows that ctDNA levels drop sharply between the first G360 test and 30 days, then rise rapidly until 150 days, at which time ctDNA levels drop slightly and rise again around 300 days, although the rate is less extreme. In addition, ctDNA levels drop from 550 days to 1000 days and then rise again from 1000 days to 1600 days. As the number of data points decreases, the corresponding 95% confidence bands expand over time. The flexibility built into the unconditional model reveals details hidden in the data that simpler models cannot detect. Nevertheless, the unconditional model only estimates the response pattern of the group and does not account for the unexpected event that patients with different characteristics may show different response patterns. To assess the impact of incorporating patient characteristics, a conditional model incorporating all baseline covariates was fit to the data. As is typical in hierarchical models, all numerical covariates are centered around their corresponding means.

[0260] Example 14 - Age and health status, response pattern

[0261] here, Figure 11 Shown is how baseline age and health status affect the response pattern of female non-smokers receiving their first-line EGFR-TKI treatment as measured by ELIX score. The results are separated by surviving patients and deceased patients. Since the data becomes sparse after 400 days, we only examine the first 400 days. The examples presented above reveal that patients with different characteristics have different response patterns. In the upper left figure, the response curves of the average ELIX score of 30 years old and 80 years old are compared.

[0262] These results show that the 80-year-old patient did not show an initial post-treatment decline in ctDNA levels compared to the 30-year-old patient who showed a rapid decrease followed by a rapid increase. The top middle panel indicates that patients with an average age and a maximum ELIX score of 13 exhibited a very different response pattern compared to the same patients with a minimum ELIX score of 0, suggesting that patients with many comorbidities exhibited a delayed response to treatment. In the top right panel, the response patterns of elderly patients with a high burden of comorbidities and otherwise healthy younger patients are shown, illustrating how age / health status combinations can amplify differences in response patterns. Although not shown, over 400 days, a trend of decreasing ctDNA values ​​was observed for patients who were still alive at the end of the study, while this trend increased for patients who died before the end of the study.

[0263] Example 15 - Speed ​​Map

[0264] To focus on the behavior of the response pattern, a velocity map showing the IRC corresponding to the response pattern was generated ( Figure 12 ). The information presented in the velocity plots can be gleaned from the response patterns themselves, but the differences in response patterns are exacerbated when the response patterns are examined through the IRC lens. Therefore, comparing velocity plots based on IRC values ​​can provide additional clues about where the response patterns are similar and where they diverge. Another advantage of using velocity plots arises when the baseline values ​​between the response patterns are different, and therefore the differences between the response patterns may be due to the fact that the biomarker values ​​were different at the beginning. In these cases, using velocity plots for comparison may be more appropriate because the IRC is constant for the baseline values ​​of the biomarkers.

[0265] Interpreting the velocity graphs is understandable to one of ordinary skill in the art. Here, one can focus on the leftmost graph. Over the first 100 days, the velocity graphs (red curves) for the 80-year-old patients who survived and died showed different patterns. For the survivors, the IRC was initially positive, but slowed to zero around day 20 (indicating a peak in the corresponding response curve referenced by the dotted line), and then declined, with the fastest rate of decline (-0.026 logits per day) occurring around day 43. Beyond day 43, the IRC continued to decline and remained relatively flat through day 100. In contrast, the velocity graphs for the 80-year-old patients who died showed almost the opposite pattern.

[0266] Example 16 - Discussion

[0267] This article describes methods and techniques adapted for the analysis of complex longitudinal genomic data. As shown, the inventors analyzed observational data and demonstrated applications in diverse data settings, including hypothesis generation, statistical inference, and patient monitoring. Here, the inventors utilized 95% confidence bands, without retaining their traditional inferential meaning, but rather as a "guide" for identifying differences in response patterns. This supported the generation of thousands of response patterns.

[0268] Those of ordinary skill will readily appreciate that the described framework can also be applied to representative cohorts. If statistical inference is the goal, since there is the potential to generate and compare many response patterns, the number of comparisons should be minimized based on a priori assumptions, and common considerations such as controlling for type I error should be made. Assumptions can include comparing response patterns between groups of patients with predetermined covariate values ​​(where other study covariates can be used as statistical controls), but can also include assumptions about the nature of the relationship between response pattern behavior and the covariate values ​​themselves.

[0269] Another embodiment includes patient monitoring. The general idea is that each response pattern is a reasonable description of the patient, as described by his or her own unique set of characteristics, and in this way, the same response pattern can be used as a reference for new patients who share these characteristics. In addition, if survival status (dead or alive) is incorporated into the model, a reference response pattern for survivors and non-survivors can be created. Therefore, if the response pattern of the new patient is consistent with that of the survivor, intervention is unnecessary, but if the response pattern reflects that of the non-survivor, intervention may be required. Using velocity maps to compare response patterns can also enhance this process, especially if the baseline values ​​between the response patterns are different. In order to ensure reliable classification, such a monitoring system should undergo internal validation and external validation. Internal validation can be achieved by creating a training data set and a test data set, and then applying, for example, k-fold cross validation to evaluate classification accuracy. If an acceptable level of accuracy is achieved, external validation can be completed if the new patient (i.e., not participating in cross validation) is also classified with a high degree of accuracy.

[0270] As described, changes in ctDNA levels can fluctuate significantly between patients over time.Here, the methods and techniques mentioned above generate patient-level results, wherein such results reveal ctDNA dynamics for clinical decision making.

Claims

1. A method of determining a patient response in at least one patient, the method comprising: Nucleic acid sequence information is obtained from at least one patient, including: Measurement of temporal changes in biomarkers; and A patient response is determined for the at least one patient.

2. The method of any one of the preceding claims, wherein the biomarker comprises circulating tumor DNA (ctDNA).

3. The method of any one of the preceding claims, wherein the biomarkers include allele frequency and tumor fraction.

4. The method according to any of the preceding claims, wherein determining the patient response of the at least one patient comprises the use of a database.

5. A method according to any preceding claim, wherein the database comprises medical and / or insurance records.

6. A method according to any preceding claim, wherein the use of the database comprises the application of a model.

7. A method according to any one of the preceding claims, wherein the model is a hierarchical model.

8. The method according to any one of the preceding claims, wherein the model is an effect model.

9. The method according to any one of the preceding claims, wherein the model is a regression model.

10. A method according to any preceding claim, wherein the model is a joint model.

11. The method according to any one of the preceding claims x, wherein the hierarchical model is a hierarchical random effects model.

12. A method according to any preceding claim, wherein the model comprises cubic splines.

13. A method according to any preceding claim, wherein the model comprises a regression model.

14. The method of any of the preceding claims, wherein the hierarchical random effects model comprises generating data from nucleic acid sequence information comprising temporal changes in biomarkers comprising circulating tumor DNA (ctDNA) from at least one of more than one subject.

15. The method of any one of the preceding claims, wherein the generating of the data comprises generating a cubic spline for at least one of the more than one subjects.

16. A method according to any one of the preceding claims, wherein the generating of the data comprises generating a response parameter comprising one or more covariates.

17. A method according to any one of the preceding claims, wherein the generating of the data comprises generating a response parameter without covariates.

18. The method according to any one of the preceding claims, wherein the response parameter applies a multivariate normal distribution.

19. The method of any of the preceding claims, wherein said determining a patient response of said at least one patient comprises generating a velocity map.

20. The method of any one of the preceding claims, wherein said determining the patient response of said at least one patient comprises comparing with said model.

21. The method according to any of the preceding claims, wherein the joint model comprises at least two models.

22. The method according to any one of the preceding claims, wherein the joint model comprises a correlation factor between the at least two models.

23. The method of any one of the preceding claims, wherein the joint model comprises cubic splines and a proportional hazards model.

24. The method of any one of the preceding claims, wherein the biomarkers are measured using next generation DNA sequencing.

25. The method of any one of the preceding claims, wherein next generation DNA sequencing comprises attaching a non-unique barcode to the ctDNA.

26. The method of any one of the preceding claims, wherein next generation DNA sequencing comprises attaching a unique barcode to the ctDNA.

27. The method of any one of the preceding claims, wherein next-generation DNA sequencing comprises attaching non-unique barcodes to ctDNA fragments, wherein the non-unique barcodes are present in at least 20x, at least 30x, at least 50x, or at least 100x molar excess.

28. A method of determining a patient response in at least one patient, the method comprising: Nucleic acid sequence information is obtained from at least one patient, including: Measurement of temporal changes in biomarkers including circulating tumor DNA (ctDNA); and Determining the patient response of the at least one patient comprises using a database comprising medical records and / or insurance records from more than one subject, wherein using the database comprises applying a hierarchical random effects model.

29. The method of any of the preceding claims, wherein the hierarchical random effects model comprises generating data from nucleic acid sequence information comprising temporal variation in ctDNA from at least one of more than one subject.

30. The method of any one of the preceding claims, wherein the hierarchical random effects model comprises generating a cubic spline for at least one of the more than one subjects.

31. The method of claim x, wherein the hierarchical random effects model includes a response parameter comprising one or more covariates for at least one subject among the more than one subjects.

32. The method of any preceding claim, wherein the database comprises medical records and / or insurance records of the more than one subjects.

33. A method of determining a patient response in at least one patient, the method comprising: Nucleic acid sequence information is obtained from at least one patient, including: Measurement of temporal changes in biomarkers including circulating tumor DNA (ctDNA); and Determining the patient response of the at least one patient comprises using a database comprising medical records and / or insurance records from more than one subject, wherein use of the database comprises application of a joint model comprising cubic splines and a proportional hazards model generated from data of nucleic acid sequence information from at least one of the more than one subjects.

34. The method of any preceding claim, wherein the database comprises medical records and / or insurance records of the more than one subjects.

35. A system comprising a machine comprising at least one processor and memory, the at least one processor and memory comprising instructions capable of performing any of the foregoing methods.

36. A computer readable medium comprising instructions capable of performing any of the foregoing methods.

Citation Information

Patent Citations

  • Process and apparatus for reacting feed with a fluidized catalyst over a temperature profile

    US20220032250A1

  • Methods for accurate sequence data and modified base position determination

    US8486630B2

  • Systems and methods to detect rare mutations and copy number variation

    WO2014039556A1

  • Methods and systems for analyzing nucleic acid molecules

    WO2018119452A2

  • Compositions and methods for isolating cell-free DNA

    WO2020160414A1