Methods for determining mutant frequencies and monitoring disease progression

By aligning sequencing reads with both reference and variant sequences to label gene variants, the method addresses the inefficiencies in detecting gene variants, enabling accurate disease progression monitoring and personalized treatment strategies.

JP2026049017APending Publication Date: 2026-03-17FOUNDATION MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing genomic testing methods face challenges in efficiently detecting gene variants and determining their frequencies, especially in limited samples, which hinders effective monitoring of disease progression and treatment strategies.

Method used

A method involving sequencing reads alignment with both reference and variant sequences to label reads as having or not having gene variants, using techniques like next-generation sequencing, and generating match scores to determine variant frequencies, allowing for efficient detection and monitoring of disease progression.

Benefits of technology

This approach enables accurate and efficient detection of gene variants and monitoring of disease progression, even in low nucleic acid samples, by reducing false positives and computational costs, thereby informing personalized treatment adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026049017000004
    Figure 2026049017000004
  • Figure 2026049017000005
    Figure 2026049017000005
  • Figure 2026049017000006
    Figure 2026049017000006
Patent Text Reader

Abstract

To provide a method for determining mutant frequencies and monitoring disease progression. [Solution] Methods for determining the frequency of variants in a test sample derived from a subject, and methods for labeling sequencing reads as having or not having variants are described herein. Exemplary methods include generating a reference match score and a variant match score by aligning sequencing reads to corresponding variant sequences and corresponding reference sequences, and labeling the sequencing reads as having or not having variants based on the determined match scores. Methods for monitoring disease progression and methods for treating subjects having the disease are also described herein. Devices and systems for implementing such methods are further described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 082,939, filed Sep. 24, 2020, which is hereby incorporated by reference in its entirety.

[0002] Methods and systems for identifying variants, methods and systems for determining variant frequencies in test samples, methods for monitoring disease progression (such as cancer progression), and methods for treating subjects having a disease (such as cancer) are described herein.

Background Art

[0003] Genomic testing has shown great promise in developing a better understanding of cancer and guiding more effective treatment approaches. Genomic testing involves sequencing the genome or a portion thereof of a patient's biological sample (which can contain cancer cells or cell - free nucleic acid products of cancer cells), and identifying any gene variants (e.g., mutations that can be associated with a tumor) in the sample relative to a reference gene sequence. Gene variants can include, for example, insertions, deletions, substitutions, rearrangements, or any combination thereof. Identifying and understanding these gene variants (e.g., mutations) found in a particular patient's cancer can also help in developing better treatment methods and can be useful in identifying (or excluding ineffective approaches) the best approach for treating specific cancer variants using genomic information.

[0004] Generally, biological samples are processed in a laboratory using a variety of possible techniques, with the ultimate goal of extracting and isolating the DNA they contain. The isolated DNA is then sequenced, yielding a data structure representation (which can be electronic) of the DNA from the patient sample. Often, this data structure representation is in the form of thousands or more "reads" (e.g., tens of thousands, hundreds of thousands, millions, tens of millions, or hundreds of millions of reads). A single read typically contains a relatively short (e.g., 50–150 base pairs) subsequence of the patient's DNA. In contrast, the entire human genome is approximately 3 billion base pairs long, and the subregions of interest for the purposes of this application can be tens of thousands of base pairs long.

[0005] The progression of certain diseases, such as cancer and clonal hematopoiesis, can be monitored in patients by determining the frequency of variants among nucleic acid molecules in samples taken from the patient. Cancer severity generally correlates with the number of variants in the tumor genome or the relative frequency of those variants appearing in the sample. For example, cell-free DNA is generally a mixture of genomic DNA and circulating tumor DNA. As cancer severity increases, a larger proportion of cell-free DNA is thought to be cancer-related. Disease progression can be monitored by tracking the relative frequency of variants exhibiting tumor genome characteristics.

[0006] Mutant calling pipelines generally require a threshold number of sequencing reads to be identified as containing mutants before positive mutant calling is performed. Detecting a sufficient number of sequencing reads often requires considerable sequencing depth, which can be impossible if a limited amount of disease-related nucleic acid is available. There is still a need for efficient mutant calling methods with low detection limits that can be used to track disease progression. [Overview of the project] [Means for solving the problem]

[0007] This specification describes methods for labeling sequencing reads from subject-derived test samples as having or not having gene variants, and methods for determining the frequency of variants in subject-derived test samples. Methods for monitoring disease progression and methods for treating subjects with disease are also described herein. Electronic devices and systems for performing such methods are further described.

[0008] In some embodiments, a method for detecting gene variants in a test sample derived from a subject or determining the frequency of a variant allele includes: (a) selecting a gene variant at a variant locus from a variant panel; (b) obtaining one or more sequencing reads related to the test sample that overlap with the variant locus; (c) generating a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene variant; and (d) generating a variant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding variant sequence, wherein the corresponding variant sequence contains the gene variant. (e) To generate labeled sequencing reads based on a reference match score and a variant match score, each of one or more sequencing reads is labeled as having a gene variant, not having a gene variant, or being a null read, wherein the sequencing read is labeled as having a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled as not having a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0009] This method may include sequencing nucleic acid molecules obtained from test samples from subjects to generate one or more sequencing reads.

[0010] The sequencing of nucleic acid molecules may involve the use of massively parallel sequencing (MPS) techniques (e.g., next-generation sequencing (NGS), whole-genome sequencing (WGS), whole-exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing techniques).

[0011] For example, in some implementations of this method, a method for detecting gene variants or determining the frequency of variant alleles in a test sample derived from a subject comprises: providing a plurality of nucleic acid molecules obtained from a test sample derived from a subject; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing the captured nucleic acid molecules with a sequencer to obtain a plurality of sequence reads representing the captured nucleic acid molecules, wherein one or more of the plurality of sequence reads overlap with a variant gene locus within a subgenome section in the sample; receiving one or more sequencing reads corresponding to a reference sequence and a variant sequence in one or more processors; receiving a reference sequence from memory in one or more processors; and aligning each sequencing read to the corresponding reference sequence in one or more processors, thereby determining the frequency of one or more sequencing reads. The process includes generating a reference match score for each; receiving a variant sequence from memory in one or more processors; generating a variant match score for each of the one or more sequencing reads by aligning each sequencing read to the corresponding variant sequence in one or more processors; and labeling each of the one or more sequencing reads in one or more processors, based on the reference match score and the variant match score, as either having a gene variant, not having a gene variant, or being a null read, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence; the sequencing read is labeled not to have a gene variant if the reference match score and the variant match score are equal,It is labeled as a null read. One or more adapters may include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. Captured nucleic acid molecules can be captured from amplified nucleic acid molecules by hybridization to one or more bait molecules. One or more bait molecules may each contain one or more nucleic acid molecules with a region complementary to the region of the captured nucleic acid molecule. Amplifying nucleic acid molecules may include performing polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques.

[0012] In some embodiments, the method further includes calling for the presence of a gene variant in a test sample based on one or more labeled sequencing reads.

[0013] In some embodiments, the corresponding reference sequence and the corresponding mutant sequence include a mutant locus, a 5' flanking region, and a 3' flanking region. In some embodiments, the 5' flanking region and the 3' flanking region are each approximately 5 nucleotides to approximately 5000 nucleotides in length.

[0014] In some embodiments, the method further includes generating a corresponding reference sequence or a corresponding variant sequence.

[0015] In some embodiments, the corresponding reference sequence and the corresponding variant sequence are identical except for the gene variant.

[0016] In some embodiments, the method includes calling for the presence of a gene variant in a test sample based on one or more labeled sequencing reads. In some embodiments, the one or more sequencing reads include multiple sequencing reads that overlap with the variant locus, and the method further includes determining the number of sequencing reads from the multiple sequencing reads that have the gene variant or the number of sequencing reads from the multiple sequencing reads that do not have the gene variant. In some embodiments, the method includes determining the mutant allele frequency for the gene variant using the number of sequencing reads that have the gene variant and the number of sequencing reads that do not have the gene variant.

[0017] In some embodiments, the method involves labeling one or more sequencing reads related to a test sample for multiple gene variants at different variant loci selected from a variant panel.

[0018] In some embodiments, the method includes determining the disease status of a subject. In some embodiments, the disease status is a value proportional to the percentage of circulating tumor DNA (ctDNA) compared to total cell-free DNA (cfDNA) in the test sample. In some embodiments, the disease status is the largest somatic allele fraction of cfDNA. In some embodiments, the disease status includes qualitative factors indicating cancer recurrence in the subject, the presence of cancer resistant to the treatment mode in the subject, or the presence of cancer that can be treated with a particular treatment mode.

[0019] In some embodiments of the methods described herein, the test sample is derived from a liquid biopsy sample from a subject. For example, the liquid biopsy sample may include blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. In some implementations, the liquid biopsy sample contains circulating tumor cells (CTCs). In some implementations, the sample contains a liquid biopsy sample, where tumor nucleic acid molecules are derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and non-tumor nucleic acid molecules are derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample. In some embodiments, the test sample contains cfDNA. In some implementations, the test sample contains a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some implementations, tumor nucleic acid molecules are derived from the tumor portion of a heterogeneous tissue biopsy sample, and non-tumor nucleic acid molecules are derived from the normal portion of the heterogeneous tissue biopsy sample. In some embodiments of the methods described herein, the test sample is derived from a solid tissue biopsy sample from a subject. Optionally, the method may further include obtaining a test sample from a subject.

[0020] In some embodiments, the reference match score and variant match score are determined using a sequence alignment algorithm. In some embodiments, the sequence alignment algorithm is the Smith-Waterman alignment algorithm, the Striped Smith-Waterman alignment algorithm, or the Needleman-Wunsch alignment algorithm.

[0021] In some embodiments, gene variants include single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels, or rearrangement conjugates. In some embodiments, the variant panel is determined by sequencing nucleic acid molecules in a previous test sample taken from a subject and calling one or more gene variants. In some embodiments, the subject received an intervention treatment for the disease between the previously taken test sample and the test sample to be taken.

[0022] In some embodiments, the disease is cancer. In some embodiments, cancer is B-cell carcinoma (multiple myeloma), melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendiceal cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, hematological cancer, adenocarcinoma, inflammatory myofibroblast Cell tumors, gastrointestinal stromal tumors (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorders (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma Osteosarcoma, chordoma, angiosarcoma, endosarcoma, lymphangiosarcoma, lymphangioendosarcoma, synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, liver cancer, cholangiocarcinoma, choriocarcinoma, seminocarcinoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal carcinoma These include glomyocyte tumors, glioblastomas, acoustic neuroblastomas, oligodendrogliomas, meningiomas, neuroblastomas, retinoblastomas, follicular lymphomas, diffuse large B-cell lymphomas, mantle cell lymphomas, hepatocellular carcinomas, thyroid cancers, gastric cancers, head and neck cancers, small cell carcinomas, essential thrombocythemia, aplastic myelogenesis, eosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinomas, or carcinoid tumors.

[0023] In some embodiments, the method further includes adjusting the treatment based on the difference between the subject's disease state determined using a test sample and the subject's previous disease state based on a previous test sample. Adjusting the disease therapy may include, for example, adjusting the dosage of the disease therapy or selecting a different disease therapy in accordance with the progression of the disease. The method may further include administering the adjusted disease therapy to the subject. In some implementations, a first sample is taken from the subject before the subject is administered the disease therapy, and a second sample is taken from the subject after the subject is administered the disease therapy. The disease therapy may include, for example, chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery.

[0024] In some implementations of this method, the detected gene variants or determined variant allele frequencies are used as a basis for enrolling subjects in clinical trials of selected disease treatments (e.g., anti-cancer therapy).

[0025] A method of monitoring disease progression, comprising: sequencing nucleic acid molecules in a first test sample obtained from a subject having a disease to generate a first sequencing read; generating an individualized variant panel for the subject; sequencing nucleic acid molecules in a second test sample obtained from the subject at a time point later than the first test sample to generate a second sequencing read; using the second sequencing read to detect a gene variant or using the second sequencing read to determine a variant allele frequency using one of the methods described above, is also described herein. In some embodiments, the method includes administering a disease therapy to the subject after the first test sample is obtained from the subject and before the second test sample is obtained from the subject. In some embodiments, the method includes generating a first disease state based on some of the first sequencing reads having variants included in the variant panel and generating a second disease state based on some of the second sequencing reads having variants from within the variant panel. In some embodiments, the method further includes determining disease progression by comparing the first disease state and the second disease state. In some embodiments, the method includes administering a disease therapy to the subject after the first test sample is obtained from the subject and before the second test sample is obtained from the subject and adjusting the disease therapy based on the determined disease progression.

[0026] A method of treating a subject having a disease (such as cancer), comprising obtaining a first test sample from the subject, sequencing nucleic acid molecules in the first test sample to generate a first sequencing read, using the first sequencing read to determine a first disease state, generating an individualized variant panel for the subject, administering a disease therapy to the subject, obtaining a second test sample from the subject after the disease therapy has been administered to the subject, sequencing nucleic acid molecules in the second test sample to generate a second sequencing read, using the second sequencing read to detect gene variants or using the second sequencing read to determine variant allele frequencies using one of the methods described above, using the labeled second sequencing read to determine a second disease state, determining disease progression by comparing the first disease state and the second disease state, adjusting the disease therapy administered to the subject based on the disease progression, and administering the adjusted disease therapy to the subject, is also described herein. In some embodiments, the disease is cancer.

[0027] In some embodiments of the above method, the method includes (1) generating or updating a report including identification information for the subject and (2) a call for the presence or absence of gene variants or a call for variant allele frequencies for the gene variants. In some embodiments, the method includes transmitting the report to the subject or a healthcare provider for the subject. In some implementations, the report is transmitted via a computer network or a peer-to-peer connection.

[0028] A computer implementation method for detecting gene variants in a test sample derived from a subject or determining the frequency of a variant allele, comprising an electronic device comprising one or more processors and a memory for storing a reference sequence that does not contain a gene variant and a variant sequence that contains a gene variant at a mutation locus, wherein one or more processors receive one or more sequencing reads related to the test sample corresponding to the reference sequence and the variant sequence; one or more processors receive the reference sequence from the memory; one or more processors generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to the corresponding reference sequence; one or more processors receive the variant sequence from the memory; one or more processors align each sequencing read to the corresponding variant sequence, A computer implementation method is also described herein, comprising generating a variant match score for each of one or more sequencing reads, and labeling each of the one or more sequencing reads in one or more processors as having a gene variant, not having a gene variant, or being a null read, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled not to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0029] In some embodiments of the computer implementation method, the method includes storing in memory an identifier associated with each sequence determination read.

[0030] In some embodiments of the computer implementation method, the method includes using one or more processors to recall the presence or absence of gene variants in a test sample based on one or more labeled sequence reads, and storing the recall of gene variants in memory.

[0031] In some embodiments of the computer implementation method, the method includes using one or more processors to determine the mutant allele frequencies of a gene variant in a test sample based on one or more labeled sequence reads, and storing the mutant allele frequencies in memory.

[0032] In some embodiments of the computer implementation method, the corresponding reference sequence and the corresponding variant sequence include a variant locus, a 5' flanking region, and a 3' flanking region. In some embodiments, the 5' flanking region and the 3' flanking region are each approximately 5 nucleotides to approximately 5000 nucleotides in length.

[0033] In some embodiments of the computer implementation method, the method includes using one or more processors to select a gene variant from a variant panel stored in memory, using one or more processors to generate a reference sequence or variant sequence, and storing the reference sequence or variant sequence in memory.

[0034] In some embodiments of the computer implementation method, the corresponding reference sequence and the corresponding variant sequence are identical except for the gene variant.

[0035] In some embodiments of the computer implementation method, one or more sequencing reads include multiple sequencing reads that overlap with a variant gene locus, and the method further includes using one or more processors to determine some sequencing reads from multiple sequencing reads having gene variants or some sequencing reads from multiple sequencing reads not having gene variants.

[0036] In some embodiments of the computer implementation method, the method includes using one or more processors to label one or more sequence reads related to a test sample for multiple gene variants at different variant loci selected from a variant panel.

[0037] In some embodiments of the computer implementation method, the method includes determining the disease status of a subject using one or more processors. In some embodiments, the disease status is a value proportional to the percentage of circulating tumor DNA (ctDNA) compared to total cell-free DNA (cfDNA) in the test sample. In some embodiments, the disease status is the largest somatic allele fraction of cfDNA. In some embodiments, the disease status includes qualitative factors indicating cancer recurrence in the subject, the presence of cancer resistant to a particular mode of treatment in the subject, or the presence of cancer that can be treated with a specific mode of treatment.

[0038] In some embodiments of the computer implementation method, the test sample includes cfDNA.

[0039] In some embodiments of the computer implementation method, the reference match score and variant match score are determined using a sequence alignment algorithm. In some embodiments, the sequence alignment algorithm is the Smith-Waterman alignment algorithm, the Striped Smith-Waterman alignment algorithm, or the Needleman-Wunsch alignment algorithm.

[0040] In some embodiments of the computer implementation method, the gene variants include single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels, or rearrangement conjugates.

[0041] In some embodiments of the computer implementation method, the variant panel is determined by sequencing nucleic acid molecules in a previous test sample taken from a subject and calling one or more gene variants. In some embodiments, the subject received an intervention treatment for a disease between the previously taken test sample and the test sample to be taken. In some embodiments, the disease is cancer.

[0042] In some embodiments of the computer implementation method, the test sample is derived from a liquid biopsy sample from a subject. In some embodiments of the computer implementation method, the test sample is derived from a solid tissue biopsy sample from a subject.

[0043] In some embodiments of the computer implementation method, the method further includes using one or more processors to generate a report that includes (1) subject identification information and (2) a call for the presence or absence of gene variants or a call for the variant allele frequency. In some embodiments, the method includes transmitting the report to a second electronic device. In some implementations, the report is transmitted via a computer network or peer-to-peer connection.

[0044] In some embodiments of any of the above methods, the mutant is a somatic mutation.

[0045] In some embodiments of any of the above methods, the mutant is a germline mutation.

[0046] The method may further include generating a genomic profile of a subject using one or more labeled sequencing reads, or detected gene variants or determined variant allele frequencies. The subject's genomic profile may include results from comprehensive genomic profiling (CGP) tests, gene expression profiling tests, cancer hotspot panel tests, DNA methylation tests, DNA fragmentation tests, RNA fragmentation tests, or any combination thereof. In some implementations of the method, the method may further include selecting an anticancer drug, administering an anticancer drug, or applying an anticancer treatment to the subject based on the generated genomic profile. In some implementations of the method, the genomic profile is used as a basis for enrolling the subject in a clinical trial for a selected disease treatment (e.g., anticancer therapy).

[0047] In some implementations of this method, the method further includes selecting an anticancer therapy to administer to a subject based on the detection of a gene variant or a determined mutant allele frequency. For example, the detection of a gene variant or the determination of an allele frequency in a test sample can be used when making a decision on a proposed treatment for a subject. In some implementations of this method, the detected gene variant or the determined mutant allele frequency is used as a basis for enrolling a subject in a clinical trial of the selected disease treatment (e.g., the selected anticancer therapy). In some embodiments, the method further includes administering the selected anticancer therapy to the subject. For example, the selected anticancer therapy may include chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery.

[0048] The detection of gene variants or determined mutant allele frequencies can be used to diagnose or confirm a disease in a subject. Accordingly, a method for diagnosing a disease is also provided herein, which may include diagnosing a subject as having a disease based on the detection of gene variants or determined mutant allele frequencies, wherein the gene variants or mutant allele frequencies are determined according to any of the methods described above.

[0049] A method for identifying a patient as eligible for a clinical trial of a disease treatment based on the detection of a gene variant or determined mutant allele frequency, wherein the detected gene variant or determined mutant allele frequency is determined according to one of the methods described above. The method may further include enrolling the patient in the clinical trial. In some implementations, the method may include administering the disease treatment to the patient.

[0050] Subjects of any of the methods described herein may have cancer, be at risk of having cancer, be routinely screened for cancer, or be suspected of having cancer. In some implementations, cancer is a solid tumor. In other implementations, cancer is a blood cancer.

[0051] An electronic device comprising one or more processors, memory, and one or more programs, wherein one or more programs are stored in memory and configured to be executed by one or more processors, and one or more programs perform the following functions: (a) selecting gene variants at mutant loci from a mutant panel; (b) obtaining one or more sequencing reads related to a test sample that overlaps with the mutant locus; (c) generating a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene variant; and (d) generating a mutant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding mutant sequence. Electronic devices are also described herein that include instructions for generating, (e) generating a corresponding variant sequence containing a gene variant, and (f) labeling each of one or more sequencing reads, based on a reference match score and a variant match score, as either having a gene variant, not having a gene variant, or being a null read, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled to not have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0052] Further described herein is a non-temporary computer-readable storage medium for storing one or more programs, wherein when one or more programs are executed by one or more processors of an electronic device having a display, the electronic device performs the following actions: (a) selecting gene variants at mutant loci from a mutant panel; (b) obtaining one or more sequencing reads related to a test sample overlapping with the mutant loci; (c) aligning each sequencing read to a corresponding reference sequence to generate a reference match score for each of the one or more sequencing reads, wherein the corresponding reference sequence does not contain the gene variant; and (d) aligning each sequencing read to a corresponding mutant sequence to generate a mutant match score for each of the one or more sequencing reads. (e) a command to generate a corresponding variant sequence containing a gene variant, and to label each of one or more sequencing reads as either having a gene variant, not having a gene variant, or being a null read, based on the reference match score and the variant match score, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled not to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal. [Brief explanation of the drawing]

[0053] [Figure 1] This document illustrates an exemplary embodiment of a method for labeling sequencing reads.

[0054] [Figure 2]This document provides an exemplary method for determining the frequency of variants in test samples derived from subjects.

[0055] [Figure 3] This document illustrates exemplary methods for monitoring disease progression.

[0056] [Figure 4] This document illustrates an exemplary computer implementation method for determining the frequency of variants in test samples derived from subjects.

[0057] [Figure 5A] An example of a computing device according to one embodiment is shown.

[0058] [Figure 5B] An example of a computing system according to one embodiment is shown.

[0059] [Figure 6A] As further described in the examples, the mutant distribution of the mutants in the Sample 1 panel is shown.

[0060] [Figure 6B] As further described in the examples, the mutant distribution of the mutants in the Sample 2 panel is shown.

[0061] [Figure 7A] A plot of the number of mutant reads detected using the exemplary methods described herein (y-axis) versus the number of mutant reads detected using a standard mutation calling protocol (x-axis) is shown for Sample 1 on a logarithmic scale (left) and normalized scale (right), as described in the Examples.

[0062] [Figure 7B]A plot (y-axis) of the mutant locus depth at each mutant locus for the total number of sequencing reads labeled as having or not having a mutant (i.e., excluding null reads) using the exemplary method described herein, against the mutant locus depth at each mutant locus for the total number of sequencing reads from the initial pool of sequencing reads overlapping with the mutant locus (x-axis), is shown for Sample 1 on a logarithmic scale (left) and normalized (right), as described in the Examples.

[0063] [Figure 8A] A plot of the number of mutant reads detected using the exemplary methods described herein (y-axis) versus the number of mutant reads detected using a standard mutation calling protocol (x-axis) is shown for Sample 2 on a logarithmic scale (left) and normalized scale (right), as described in the Examples.

[0064] [Figure 8B] A plot (y-axis) of the mutant locus depth at each mutant locus for the total number of sequencing reads labeled as having or not having a mutant (i.e., excluding null reads) using the exemplary method described herein, against the mutant locus depth at each mutant locus for the total number of sequencing reads from the initial pool of sequencing reads overlapping with the mutant locus (x-axis), is shown for Sample 2 on a logarithmic scale (left) and normalized (right), as described in the Examples.

[0065] [Figure 9A] A plot of the number of mutant reads detected using another exemplary method described herein (y-axis) versus the number of mutant reads detected using a standard mutation calling protocol (x-axis) is shown for Sample 1 on a logarithmic scale (left) and normalized scale (right), as described in the examples.

[0066] [Figure 9B]A plot (y-axis) of the mutant locus depth at each mutant locus for the total number of sequencing reads labeled as having or not having a mutant (i.e., excluding null reads) using another exemplary method described herein, against the mutant locus depth at each mutant locus for the total number of sequencing reads from the initial pool of sequencing reads overlapping with the mutant locus (x-axis), is shown for Sample 1 on a logarithmic scale (left) and normalized (right), as described in the examples.

[0067] [Figure 10A] A plot of the number of mutant reads detected using another exemplary method described herein (y-axis) versus the number of mutant reads detected using a standard mutation calling protocol (x-axis) is shown for Sample 2 on a logarithmic scale (left) and normalized scale (right), as described in the Examples.

[0068] [Figure 10B] A plot (y-axis) of the mutant locus depth at each mutant locus for the total number of sequencing reads labeled as having or not having a mutant (i.e., excluding null reads) using another exemplary method described herein, against the mutant locus depth at each mutant locus for the total number of sequencing reads from the initial pool of sequencing reads overlapping with the mutant locus (x-axis), is shown for Sample 2 on a logarithmic scale (left) and normalized (right), as described in the Examples. [Modes for carrying out the invention]

[0069] This document describes methods for treating a disease, including methods for determining the frequency of mutant alleles in test samples from subjects, or for detecting the presence or absence of mutants, for monitoring disease progression, for detecting the presence of tumors, for profiling the immune repertoire in subjects, for identifying tumor clones, viral strains, or bacterial strains, for detecting clonal hematopoiesis, and for monitoring disease progression and adjusting treatment therapies based on disease progression. Mutant allele frequency determination or mutant detection can utilize an individualized mutant panel established for subjects using initial samples. The individualized mutant panel contains disease-indicating gene variants. The mutant panel can then be used to rapidly label most sequencing reads from subjects as either having or not having mutant sequences. The labeled sequencing reads can then be used to determine the disease status based on the mutant frequency.

[0070] For clinical decisions to be made when treating a subject, treating physicians need confidence in the diagnostic tools used to evaluate the subject. Sequencing of nucleic acid molecules in a subject and de novo variant calling provide useful information that can be used to characterize a disease. However, nucleic acid sequencing is generally exposed to substantial noise due to mutations introduced during PCR amplification, errors made during nucleotide detection during sequencing, and other anomalies that may be introduced during the sequencing process. For this reason, many sequencing pipelines require a threshold number of unique sequencing reads that have the same variant before the variant can be confidently called. Sequencing at sufficiently deep depths can overcome this obstacle, but it can be expensive and impossible when limited tumor nucleic acids are available (e.g., in the case of circulating tumors (ctDNA), they are excreted from small tumor clones). Furthermore, a specific true variant may be detected but not actively called because the number of detected sequencing reads with the variant does not meet the calling threshold. However, using the method described herein, sequencing reads labeled as having the variant from a given variant panel reduce the limits of detection because the possibility of false-positive variant calling from the pre-panel is less likely to be due to random chance.

[0071] Furthermore, de novo variant calling is computationally expensive. The method described herein streamlines the variant calling process to generate more efficient variant calling and more efficient measurements of allele frequencies of a given variant. For example, the method described herein can be limited to the analysis of a selected number of loci.

[0072] In some embodiments, a method for detecting gene variants in a test sample derived from a subject or determining the frequency of a variant allele is to (a) select a gene variant at a variant locus from a variant panel; (b) obtain one or more sequencing reads related to the test sample that overlap with the variant locus; (c) generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene variant; (d) generate a variant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding variant sequence, wherein the corresponding variant sequence contains the gene variant; and (e) reference match The instructions include labeling each of one or more sequencing reads as either having a gene variant, not having a gene variant, or being a null read, based on the reference match score and the variant match score, wherein the sequencing read is labeled as having a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the variant sequence more closely than the reference sequence; the sequencing read is labeled as not having a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the reference sequence more closely than the variant sequence; and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal, the labeled sequencing reads can then be used to determine the disease status of a subject.

[0073] Methods for determining mutant allele frequencies can be used to monitor disease progression. For example, a method for monitoring disease progression may include sequencing nucleic acid molecules in a first test sample obtained from a subject with the disease to generate a first sequencing read, generating an individualized mutant panel for the subject, sequencing nucleic acid molecules in a second test sample obtained from the subject at a later time than the first test sample to generate a second sequencing read, and labeling the second sequencing read using the method described herein. The labeled sequencing read can then be used to determine the subject's disease status, which can be compared to a previously determined disease status (e.g., the disease status associated with the subject at the time the first test sample was obtained from the subject) to monitor disease progression.

[0074] Disease status monitoring can be further used to treat subjects with a disease, for example, by adjusting disease therapy based on monitored disease progression. For example, in some embodiments, a method for treating a subject with a disease may include: obtaining a first test sample from the subject; sequencing nucleic acid molecules in the first test sample to generate a first sequencing read; generating a personalized variant panel for the subject; administering disease therapy to the subject; obtaining a second test sample from the subject after the disease therapy has been administered to the subject; sequencing nucleic acid molecules in the second test sample to generate a second sequencing read; labeling the second sequencing read using a method described herein; determining disease progression by comparing a first disease state with a second disease state; adjusting disease therapy administered to the subject based on disease progression; and administering the adjusted disease therapy to the subject.

[0075] In some embodiments, the disease is cancer. definition

[0076] As used herein, the singular forms "a," "an," and "the" include plural references unless otherwise explicitly indicated by the context.

[0077] References to values ​​or parameters "about" in this specification include (and describe) variations relating to the value or parameter itself. For example, the statement "about X" includes the statement "X".

[0078] The terms “allele frequency” and “allele fraction” are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular allele relative to the total number of sequence reads for a genomic locus. The terms “mutant allele frequency” and “mutant allele fraction” are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular mutant allele relative to the total number of sequence reads for a genomic locus.

[0079] The terms "individual," "patient," and "subject" are used synonymously and refer to animals such as humans.

[0080] The “reference” sequence is any sequence used for comparison with the test or subject sequence (e.g., sequencing reads) and can be a standardized reference sequence (e.g., a sequence from a standardized reference assembly such as GRCh38 or an alternative reference assembly from the Genome Reference Consortium) or a personalized reference sequence (e.g., a sequence from the subject’s healthy tissue).

[0081] A "subgenome segment" refers to a portion of a genome or exome sequence. A subgenome segment can be, for example, a single nucleotide location or two or more nucleotide locations (e.g., at least 2, 5, 10, 50, 100, 150, or 250 nucleotides in length). A subgenome segment may include an entire gene or a pre-selected portion thereof (e.g., a coding region (or part thereof), a pre-selected intron (or part thereof), or an exon (or part thereof)).

[0082] The term “variant” refers to any sequence difference between a subject sequence and a reference sequence compared to the subject sequence. Therefore, the term “variant” includes differences between sequences from healthy individuals and reference sequences used to identify population variations, or between sequences from diseased tissue (e.g., tumor tissue) and sequences from healthy tissue (i.e., mutations).

[0083] It is understood that the embodiments and variations of the invention described herein include "consisting of" and / or "essentially consisting of."

[0084] Where a range of values ​​is provided, it should be understood that each intermediate value between the upper and lower limits of that range, and any other stated or intermediate values ​​within that range, are included within the scope of this disclosure. If the stated range includes an upper or lower limit, the range excluding either of those limits is also included within this disclosure.

[0085] Some of the analytical methods described herein involve mapping a sequence to a reference sequence, determining sequence information, and / or analyzing sequence information. Complementary sequences can be readily determined and / or analyzed, and it is well understood in the art that the descriptions provided herein encompass analytical methods performed with respect to complementary sequences.

[0086] Section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described herein. Descriptions are presented to enable those skilled in the art to manufacture and use the invention and are provided in the context of a patent application and its requirements. Various modifications to the embodiments described herein will be readily apparent to those skilled in the art, and the general principles herein may apply to other embodiments. Accordingly, the invention is not intended to be limited to the embodiments shown, but rather should be given the broadest scope consistent with the principles and features described herein.

[0087] The diagram illustrates the process according to various embodiments. In the exemplary process, some blocks are arbitrarily combined, the order of some blocks is arbitrarily changed, and some blocks are arbitrarily omitted. In some examples, additional steps can be performed in combination with the exemplary process. Therefore, the actions illustrated (and described in more detail below) are illustrative in nature and should not be considered limiting.

[0088] All publications, patents, and patent application disclosures referenced herein are incorporated herein by reference in their entirety. In the event of any conflict between the references incorporated herein and this disclosure, this disclosure shall prevail. Mutant panel

[0089] Certain methods described herein utilize a variant panel comprising one or more gene variants of interest. The gene variants may be, for example, variants associated with a specific disease (e.g., cancer or cancer clone) or disease condition (e.g., metastasis). In some embodiments, the variant panel is a personalized variant panel. In some embodiments, the variant panel is a disease patient population variant panel based on variants detected in a population of subjects having a specific disease.

[0090] The variant panel can be of any size. The variants are associated with the reference sequence and the variant sequence. Therefore, the reference sequence and the variant sequence can be easily constructed as long as the targeted variant is known in advance. The variants in the variant panel may include, for example, one or more single nucleotide variants (SNVs), one or more multinucleotide variants (MNVs), rearrangement conjugates, and / or one or more indels. An MNV may include consecutive nucleotide variants of two or more nucleotide variants queried using the constructed reference sequence or variant sequence. In some embodiments, the variant panel includes one or more fusion variants or other rearrangement variants (e.g., inversion or deletion events). The variants in the variant panel may include the variant relative to the reference sequence and / or the locus of the variant. As just one example, an SNP variant may include a locus (e.g., gene name and base location within the gene, or base location within the genome) and a variant (e.g., a C→G mutation).

[0091] The variant panel may include any number of variants associated with the disease, or 1 or more, 2 or more, 5 or more, 10 or more, 25 or more, 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, 10,000 or more, 20,000 or more, 50,000 or more, or 100,000 or more, or about 1 to about 10, about 10 to about 25, about 25 to about 100, about 100 to about 500, about 500 to about 1000, about 1000 to about 5000, about 5000 to about 10,000, about 10,000 to about 20,000, about 20,000 to about 50,000, or about 50,000 to about 100,000.

[0092] The variant panel or subject variant may, in some embodiments, include rearrangement conjugates. Rearrangement variants, such as insertions, deletions, or inversion generation, may generate two rearrangement conjugates (or more in the case of complex rearrangements) relative to the reference sequence. Conjugates can be detected, for example, by using a variant sequence containing at least one conjugate, using the methods described herein.

[0093] In some embodiments, the mutant panel is a personalized mutant panel generated for a specific subject. A sample can be obtained for the subject, and nucleic acid molecules (e.g., DNA, RNA, or both) in the sample are sequenced to generate sequencing reads. In some embodiments, the RNA molecule is reverse transcribed to form a corresponding cDNA molecule. A known mutant calling method can then be used to call mutants from the generated sequencing reads.

[0094] Samples obtained from subjects may include nucleic acid molecules derived from diseased tissue, or a mixture of nucleic acid molecules derived from diseased tissue and nucleic acid molecules derived from healthy tissue (or two separate samples may be analyzed, one using nucleic acid molecules derived from diseased tissue and the other from healthy tissue). For example, a sample may include cell-free DNA (cfDNA) containing circulating tumor DNA (ctDNA, i.e., DNA naturally occurring in tumor tissue) and genomic cell-free DNA (i.e., cfDNA, naturally occurring in healthy tissue). The cfDNA can be sequenced and may include a variant panel of tumor-related variants (referred to by genomic cell-free DNA or by some other reference genome) and one or more of the tumor variants. In some embodiments, samples may be derived from diseased tissue (e.g., solid tumor biopsy samples or hematological tumor biopsy samples) or tissue biopsy samples (e.g., solid tissue samples or hematological tissue samples) to obtain healthy tissue. Nucleic acid samples may be derived from tissue samples and may be used to generate sequencing reads.

[0095] In some embodiments, the mutant panel is generated by invoking mutants between nucleic acid molecules obtained from diseased tissue (e.g., tumor tissue) and healthy tissue. For example, mutants can be invoked using a matched normal tumor sample.

[0096] In some embodiments, the mutant panel is generated by calling for mutants between nucleic acid molecules obtained from plasma (e.g., cfDNA) and nucleic acid molecules obtained from peripheral blood mononuclear cells (PBMCs).

[0097] In some embodiments, the sample used to obtain nucleic acid molecules may be blood, serum, saliva, tissue (e.g., solid or hematological tissue), cerebrospinal fluid, amniotic fluid, peritoneal fluid, interstitial fluid, or embryonic tissue. In some embodiments, the tissue is fresh tissue (i.e., not frozen or preserved). In some embodiments, the tissue is frozen or preserved tissue (e.g., formaldehyde-fixed paraffin-embedded (FFPE) or paraformaldehyde-fixed paraffin-embedded (PFPE) tissue).

[0098] In some embodiments, the samples used to generate the personalized variant panel are obtained from the subject before the initiation of disease therapy. In some embodiments, the samples used to generate the personalized variant panel are obtained from the subject after the initiation of disease therapy.

[0099] A personalized variant panel can be generated for a subject with a disease using a personalized reference genome or sequence (i.e., the subject's non-disease genome sequence) or a standard reference genome or sequence (i.e., a reference genome or a reference sequence assembled from one or more other individuals, e.g., a standard or publicly available reference sequence, e.g., the Genome Reference Consortium's Human Genome Build 37 (GRCh37), or another suitable reference genome). Differences between nucleic acid molecules derived from diseased tissue can be compared to the reference to identify variants.

[0100] In some embodiments, the variants in the variant panel include one or more variants known to be associated with a specific disease (such as a specific cancer) or a population of subjects having a specific disease (such as a specific cancer). For example, the variant panel may include one or more variants curated from the literature.

[0101] The mutants in the mutant panel are associated with a corresponding reference sequence and a corresponding mutant sequence that includes the mutant's locus having left and right flanking regions (i.e., a 5' flanking region and a 3' flanking region). The left and right flanking regions of the mutant locus provide the mutant's context and are the same for both the corresponding reference sequence and the corresponding mutant sequence. Thus, the corresponding reference sequence and the corresponding mutant sequence are identical except for the mutant itself. The corresponding mutant sequence includes the mutant, and the corresponding reference sequence does not include the mutant (i.e., it includes a reference or "wild-type" sequence at the mutant's position). In some embodiments, the flanking regions each include approximately 5 or more bases, approximately 10 or more bases, approximately 15 or more bases, approximately 20 or more bases, approximately 25 or more bases, approximately 30 or more bases, approximately 50 or more bases, approximately 75 or more bases, approximately 100 or more bases, approximately 150 or more bases, approximately 200 or more bases, approximately 250 or more bases, approximately 300 or more bases, approximately 400 or more bases, or approximately 500 or more bases. In some embodiments, the flanking regions each contain about 5 to about 5000 bases, for example, about 5 to about 10 bases, about 10 to about 20 bases, about 20 to about 50 bases, about 50 to about 100 bases, about 100 to about 200 bases, about 200 to about 500 bases, about 500 to about 1000 bases, about 1000 to about 2500 bases, or about 2500 to about 5000 bases. In some embodiments, the left and right adjacent regions have the same number of bases, and in some embodiments, the left and right adjacent regions have different numbers of bases.

[0102] The corresponding reference sequence and corresponding mutant sequence can be generated, for example, using a reference sequence (which can be a personalized reference sequence or a standard reference sequence) used to identify the mutant. To generate the corresponding mutant sequence, a mutant is selected, and left and right flanking sequences are added to the mutant using the reference sequence. To generate the corresponding reference sequence, the reference sequence is used with the same base positions as the corresponding mutant sequence. Therefore, in some embodiments, the corresponding reference sequence and the corresponding mutant sequence are identical except for the gene mutant.

[0103] The variant panel may be stored in non-temporary computer-readable memory, or it may be a list stored in a table or file (e.g., a variant calling format (VCF) file or other suitable file format) that can be accessed by one or more processors to perform one or more of the methods described herein. In some embodiments, the corresponding reference sequence and the corresponding variant sequence are stored in the same table or file as the variant panel, and in some embodiments, the corresponding reference sequence and the corresponding variant sequence are stored in a different table or file than the variant panel.

[0104] The variant panel may be a variant panel associated with a disease (such as cancer) or a personalized variant panel associated with the subject's disease (such as cancer). Exemplary diseases include, but are not limited to, B-cell cancers such as multiple myeloma, melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, and pancreatic cancer, as well as gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or adnexal cancer, salivary gland cancer, and thyroid cancer. Cancer, adrenal adenocarcinoma, osteosarcoma, chondrosarcoma, hematological cancers, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma NHL tumor, soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endosarcoma synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonic carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma This includes craniopharyngioma, ependymoma, pineal glandoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenous myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, the familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, and cancerous tumors.

[0105] In some embodiments, the variants in the variant panel are not disease-related. For example, the variant panel can be used to support previous or presumptive calls. Whole-genome sequencing and other sequencing methods can result in calls made with low certainty. The methods described herein can be used to support specific calls (either positive or negative) to provide higher sequence reliability.

[0106] In some embodiments, the variant panel includes one or more variants (e.g., SNPs, MNPs, rearrangement conjugates, or indels) of any of the following genes: ABCB1, ABCC2, ABCC4, ABCG2, ABL1, ABL2, AKT1, AKT2, AKT3, ALK, APC, AR, ARAF, ARFRP1, ARID1A, ATM, ATR, AURKA, AURKB, BCL2, BCL2A1, BCL2L1, BCL2L2, BCL6, BRAF, BRCA1, BRCA2, C1orf144, CARD11, CBL, CCND1, C CND2, CCND3, CCNE1, CDH1, CDH2, CDH20, CDH5, CDK4, CDK6, CDK8, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CRKL, CRLF2, CTNNB1, CYP1B1, CYP2C19, CYP2C8, CYP2D6, CYP3A4, CYP3A5, DNMT3A, DOT1L, DPYD, EGFR, EPHA3, EPHA5, EPHA6, EPHA7, EPHB1, EPHB4, EPHB6, ERBB2, ERBB3, ERBB4, ERCC2, ERG, ESR1 , ESR2, ETV1, ETV4, ETV5, ETV6, EWSR1, EZH2, FANCA, FBXW7, FCGR3A, FGFR1, FGFR2, FGFR3, FGFR4, FLT1, FLT3, FLT4, FOXP4, GATA1, GNA11, GNAQ, GNAS, GP R124, GSTP1, GUCY1A2, HOXA3, HRAS, HSP90AA1, IDH1, IDH2, IGF1R, IGF2R, IKBKE, IKZF1, INHBA, IRS2, ITPA, JAK1, JAK2, JAK3, JUN, KDR, KIT, KRAS, LRP1 B, LRP2, LTK, MAN1B1, MAP2K1, MAP2K2, MAP2K4, MCL1, MDM2, MDM4, MEN1, MET, MITF, MLH1, MLL, MPL, MRE11A, MSH2, MSH6, MTHFR, MTOR, MUTYH, MYC, MYCL1, MYCN, NF1, NF2, NKX2-1, NOTCH1, NPM1, NQO1, NRAS, NRP2, NTRK1, NTRK3, PAK3, PAX5, PDGFRA, PDGFRB, PIK3CA, PIK3R1, PKHD1, PLCG1, PRKDC, PTCH1, PTEN,PTPN11, PTPRD, RAF1, RARA, RB1, RET, RICTOR, RPTOR, RUNX1, SLC19A1, SLC22A2, SLCO1B3, SMAD2, SMAD3, SMAD4, SMARCA4, SMARCB1, SMO, SOD 2, SOX10, SOX2, SRC, STK11, SULT1A1, TBX22, TET2, TGFBR2, TMPRSS2, TOP1, TP53, TPMT, TSC1, TSC2, TYMS, UGT1A1, UMPS, USP9X, VHL, and WT1. ,

[0107] In some embodiments, the variant is a mutation, such as a tumor-related mutation. In some embodiments, the variant is a somatic mutation. In some embodiments, the variant is a germline mutation. Labeling sequence determination reads

[0108] A sequencing read can be labeled as containing a gene variant or not containing a gene variant (or as a “null read” indicating that the sequencing read cannot be labeled as containing a variant or not containing a variant). A sequencing read can be mapped to a location within a reference sequence, and the mapped location is used to select a gene variant from a variant panel associated with the locus. Once a variant and a sequencing read are associated, the sequencing read is claimed to be a reference sequence (i.e., the corresponding sequence without the variant) for generating a reference match score and a variant sequence (i.e., the corresponding sequence containing the variant) for generating a variant match score. A sequencing read can be labeled as containing a variant if the reference match score and the variant match score indicate that the sequencing read matches the variant sequence more closely than the reference sequence, or as not containing a variant if the reference match score and the variant match score indicate that the sequencing read matches the reference sequence more closely. In some embodiments, a sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0109] In some embodiments, a method for detecting the presence or absence of a mutant in a test sample derived from a subject, or for determining the frequency of a mutant allele, includes: (a) selecting a gene mutant at a mutant locus from a mutant panel; (b) obtaining one or more sequencing reads related to the test sample that overlap with the mutant locus; (c) generating a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene mutant; and (d) generating a mutant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding mutant sequence, wherein the corresponding mutant (e) generating a sequence containing a gene variant, and labeling each of one or more sequencing reads based on a reference match score and a variant match score as either having a gene variant, not having a gene variant, or being a null read, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the variant sequence more closely than the reference sequence, the sequencing read is labeled to not have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the reference sequence more closely than the variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0110] Sequence reads are aligned to a reference sequence to determine the location of the sequence reads within the reference genome. Alignment can be used to generate a sequence alignment map file (e.g., a SAM or BAM file) containing the mapping locations for the reads. A variant panel can then be accessed to select a gene variant, and one or more sequence reads that overlap with the locus of the variant can be obtained (e.g., by accessing the sequence alignment map file). The overlap may be at one or more base positions of the variant form (e.g., if the variant is a polynucleotide variant). In some embodiments, sequence reads that overlap with the same single base (e.g., the first base) of the variant are used. The corresponding reference sequence and corresponding variant sequence are also selected and associated with the selected variant.

[0111] A reference match score for any given sequencing read is generated by aligning the sequencing read to its corresponding reference sequence, and a variant match score is generated by aligning the sequencing read to its corresponding variant sequence. Both the reference and variant match scores are generated using the same alignment algorithm, such that the reference and variant match scores are comparable. The match score provides a value indicating how closely the query sequence (i.e., the sequencing read) matches the corresponding variant sequence or corresponding reference sequence. An exemplary alignment algorithm is the Smith-Waterman algorithm (SWA) (e.g., Striped This includes the Smith-Waterman algorithm or the Needleman-Wunsch algorithm (NWA). In some embodiments, the reference match score and variant match score are generated using the Smith-Waterman algorithm. In some embodiments, the reference match score and variant match score are generated using the Striped Smith-Waterman algorithm. In some embodiments, the reference match score and variant match score are generated using the Needleman-Wunsch algorithm.

[0112] Sequencing reads are labeled by comparing the mutant match score with the reference match score. For example, if the reference match score and the mutant match score indicate that the sequencing read matches the mutant sequence more closely than the reference sequence, the sequencing read is labeled as having a gene mutant. If the reference match score and the mutant match score indicate that the sequencing read matches the reference sequence more closely than the mutant sequence, the sequencing read is labeled as not having a gene mutant. In some cases, the reference match score and the mutant match score are equal. In this case, the sequencing read can be labeled as a null read. In some embodiments, sequencing reads labeled as null reads are excluded from further analysis.

[0113] Sequence reads can be obtained by sequencing nucleic acid molecules in a test sample derived from a subject. Targeted sequencing methods, such as selective capture and / or selective amplification of targeted subgenomic regions, can be used. Nucleic acid molecules (e.g., a mixture of tumor and non-tumor nucleic acid molecules) can be extracted from a test sample obtained from a subject. One or more adapters can be ligated to the nucleic acid molecules extracted from the sample. The adapters may include, for example, one or more of the following: amplification primer hybridization sites, flow cell adapter sequences, substrate adapter sequences, sample index sequences, or unique molecular identifiers. Nucleic acid molecules can be amplified before sequencing (e.g., using polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques). Targeted nucleic acid molecules can be captured from the amplified nucleic acid molecules (e.g., by hybridization to one or more bait molecules, each containing one or more nucleic acid molecules, each containing a region complementary to the region of the captured nucleic acid molecule). Nucleic acid molecules (or library proxies derived therefrom) extracted from a sample can be sequenced using, for example, a next-generation (e.g., large-scale parallel) sequencer, using, for example, next-generation (e.g., large-scale parallel) sequencing technology, whole-genome sequencing (WGS) technology, whole-exome sequencing technology, targeted sequencing technology, direct sequencing technology, or Sanger sequencing technology. The assay results can be generated, displayed, transmitted, and / or delivered as a report (e.g., an electronic report, a web-based report, or a paper report) to the subject (or patient), caregiver, healthcare provider, physician, oncologist, electronic medical record system, hospital, clinic, third-party payer, insurance company, or government agency. In some cases, the report includes output from the methods described herein. In some cases, all or part of the report can be displayed in a graphical user interface of an online or web-based healthcare portal. In some cases, the report is transmitted over a computer network or peer-to-peer connection.

[0114] In some examples, the disclosed method includes the steps of (i) obtaining a sample from a subject (e.g., a subject suspected of having cancer or determined to have cancer), (ii) extracting nucleic acid molecules (e.g., a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules) from the sample, (iii) ligating one or more adapters (e.g., one or more amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences) to the nucleic acid molecules extracted from the sample, (iv) amplifying the nucleic acid molecules (e.g., using polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques), and (v) hygroscopically The method may further include one or more of the following steps: (vi) capturing nucleic acid molecules from amplified nucleic acid molecules (by bridization); (vi) sequencing nucleic acid molecules (or library proxies derived therefrom) extracted from a sample using, for example, a next-generation (e.g., massively parallel) sequencer, using, for example, next-generation (ultra-parallel) sequencing technology, whole-genome sequencing (WGS) technology, whole-exome sequencing technology, targeted sequencing technology, direct sequencing technology, or Sanger sequencing technology; and (vii) generating, displaying, transmitting, and / or delivering a report (e.g., an electronic report, a web-based report, or a paper report) to a subject (or patient), caregiver, healthcare provider, physician, oncologist, electronic medical record system, hospital, clinic, practice, third-party payer, insurance company, or government agency. In some cases, the report includes output from the method described herein. In some cases, all or part of the report may be displayed in a graphical user interface of an online or web-based healthcare portal. In some cases, the report may be transmitted over a computer network or peer-to-peer connection.

[0115] In some embodiments, the test sample is the same type of sample used to determine the gene variant in the individualized variant panel. Exemplary test samples include, but are not limited to, blood, serum, saliva, tissue (e.g., solid or hematological tissue), cerebrospinal fluid, amniotic fluid, peritoneal fluid, interstitial fluid, or embryonic tissue. In some embodiments, the tissue is fresh tissue (i.e., not frozen or preserved). In some embodiments, the tissue is frozen or preserved tissue (e.g., formaldehyde-fixed paraffin-embedded (FFPE) or paraformaldehyde-fixed paraffin-embedded (PFPE) tissue).

[0116] A subject may have cancer, be at risk of having cancer, be routinely screened for cancer, or be suspected of having cancer. As further described herein, the results of gene variant detection or variant allele frequency determination methods may be used to diagnose or confirm cancer, or to select a treatment for cancer.

[0117] In some embodiments, the test sample is derived from a liquid biopsy sample (e.g., plasma, peripheral blood, etc.). In some embodiments, the liquid biopsy sample is blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the liquid biopsy sample contains circulating tumor cells (CTCs). In some embodiments, the liquid biopsy sample contains cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof. The liquid biopsy can be divided into two or more matched samples or sample components. For example, the sample may contain a plasma component (which may contain cfDNA) and a peripheral blood mononuclear cell (PBMC) component. The individual components can be analyzed separately to determine the differences between the genetic profiles of each component. This can be used, for example, to identify somatic mutations or clonal hematopoiesis.

[0118] In some embodiments, the sample is derived from a solid tissue biopsy sample. The tissue biopsy may contain cancerous cells, non-cancerous (i.e., healthy) cells, or a mixture thereof. In some embodiments, the tissue biopsy sample is fresh tissue (i.e., not frozen or preserved). In some embodiments, the tissue is frozen or preserved tissue (e.g., formaldehyde-fixed paraffin-embedded (FFPE) or paraformaldehyde-fixed paraffin-embedded (PFPE) tissue).

[0119] In some cases, nucleic acid molecules extracted from a sample may include a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some cases, tumor nucleic acid molecules may originate from the tumor portion of the xenotissue biopsy sample, and non-tumor nucleic acid molecules may originate from the normal portion of the xenotissue biopsy sample. In some cases, the sample may include a liquid biopsy sample, and tumor nucleic acid molecules may originate from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and non-tumor nucleic acid molecules may originate from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample.

[0120] The nucleic acid molecules in the test sample may be DNA, RNA, or a mixture thereof. In some embodiments, RNA molecules are reverse transcribed to form the corresponding cDNA molecules. The test sample obtained from a subject may include nucleic acid molecules derived from diseased tissue, or a mixture of nucleic acid molecules derived from diseased tissue and nucleic acid molecules derived from healthy tissue. For example, the sample may include cell-free DNA (cfDNA) containing circulating tumor DNA (ctDNA, i.e., DNA naturally occurring in tumor tissue) and genomic cell-free DNA (i.e., cfDNA, naturally occurring in healthy tissue). In some embodiments, the sample may be derived from diseased tissue (e.g., a solid tumor biopsy sample or a hematological tumor biopsy sample) or a tissue biopsy sample to obtain healthy tissue (e.g., a solid tissue sample or a hematological tissue sample). The nucleic acid sample may be derived from a tissue sample and may be used to generate sequencing reads.

[0121] The described method for labeling sequencing reads can be repeated for any number of variants using different gene variants at different loci selected from a gene variant panel.

[0122] In some embodiments, labeled sequencing reads are used to invoke the presence of gene variants in a sample from a subject. For example, if one or more sequencing reads (or one or more unique sequencing reads) are labeled as having a gene variant, the presence of that gene variant can be invoked. A threshold set for invoking the presence of a gene variant can be set as desired, depending on the desired level of confidence for making the invoke. For example, in some embodiments, the threshold for invoking the presence of a gene variant can be invoked as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more sequencing reads (or unique sequencing reads) labeled as having a gene variant, where the presence of a gene variant is invoked if the number of sequencing reads (or unique sequencing reads) labeled as having a gene variant satisfies or exceeds the threshold.

[0123] In some embodiments, labeled sequencing reads are used to determine the mutant allele frequency for the mutant in the sample. The mutant allele frequency (F) at locus i of the test sample is calculated as follows: i )teeth,

number

[0124] The methods described herein can be used to determine the frequency of mutant alleles in a sample, two or more different tissues or samples, or two or more different components of the same sample. For example, a blood sample can be divided into plasma (containing cfDNA) and peripheral blood mononuclear cells (PBMCs). The frequency of a first mutant allele can be determined for a first sample or a first sample component (e.g., plasma), and the frequency of a second mutant allele can be determined for a second sample or a second sample component (e.g., PBMCs). For example, the difference in mutant allele frequencies between nucleic acid molecules derived from plasma and nucleic acid molecules derived from PBMCs is useful for subjects with clonal hematopoiesis or clonal hematopoiesis with an uncertain potential (CHIP).

[0125] Figure 1 shows an exemplary embodiment of a method for labeling sequencing reads. In step 100, a gene variant panel (i.e., baseline changes) is generated by sequencing an initial sample taken from a subject. The gene variant panel may include information about each gene variant in the panel, e.g., a subject identifier, the gene containing the variant, the locus of the variant, and / or the variant's changes (compared to the reference). In the corresponding sequencing module 102, the corresponding reference sequence 104 and the corresponding variant sequence read 106 are generated using the variants from the variant panel and the reference sequence used to provide the context for the variants. The corresponding reference sequence 104 and the corresponding variant sequence read 106 are identical except for the variant locus, and an A→G SNP is present (indicated by underlining). Sequencing reads obtained by sequencing a second test sample taken from a subject are aligned to the reference sequence, and the mapped sequencing reads are included in the alignment map file 108. The alignment map file 108 includes the sequences from the sequencing reads, along with locus information for the sequencing reads. Optionally, the alignment map file 108 may include additional information such as subject information, the time when the sample was taken, and / or other sample information. A variant is selected from the variant table, and sequencing reads that overlap with the locus of the variant read are retrieved from the alignment map file 108 in the sequencing read acquisition module 110. In the example shown in Figure 1, sequencing reads 112, 114, 116, and 118 represent sequencing reads that overlap with the locus of the selected variant. In the alignment module 120, sequencing reads 112, 114, 116, and 118 are aligned with their respective reference sequences 104 to generate reference match scores 122, and the corresponding variant sequence read 106 generates a variant match score 124. The reference match score 122 and the variant match score 124 can be generated using an alignment algorithm such as the Smith-Waterman algorithm or the Needleman-Wunsch algorithm.In the classification module 126, for each sequencing read, the reference match score and the mutant match score are compared to label the sequencing read as having a mutant, not having a mutant, or being a null read. In the example shown in Figure 1, sequencing reads 112 and 114 are labeled as not having a mutant because their reference match scores are greater than their mutant match scores (i.e., the sequencing reads more closely match the corresponding reference sequence than the corresponding mutant sequence). Sequencing read 116 is labeled as having a mutant because its mutant match score is greater than its reference match score (i.e., the sequencing reads more closely match the corresponding mutant sequence than the corresponding reference sequence). Sequencing read 118 is labeled as a null read because its mutant match score is equal to its reference match score.

[0126] Figure 2 shows an exemplary method for determining the frequency of variants in a test sample derived from a subject. In step 202, gene variants at the variant locus are selected from the variant panel. In some embodiments, the variant panel is a personalized variant panel. In step 204, sequencing reads that overlap with the variant locus and are relevant to the test sample are obtained. A reference match score for each sequencing read is obtained in step 206 by aligning the sequencing read to the corresponding reference sequence, and a variant match score for each sequencing read is generated in step 208 by aligning the sequencing read to the corresponding variant sequence. Using the reference match score and the variant match score, the sequencing reads are labeled in step 210 as having a variant, not having a variant, or a null read. In step 212, the frequency of gene variants is determined using the number of sequencing reads labeled as having a variant and the number of sequencing reads labeled as not having a variant.

[0127] In some embodiments, the method includes generating or updating a report (such as a printed report or an electronic medical record). The report may include a call for the presence or absence of a gene variant, a call for the frequency of the variant allele, and / or one or more disease conditions. The report may also include subject identification information (e.g., name, identification number). The report may be stored or sent to another person or entity, such as a patient or healthcare provider (e.g., a doctor, nurse, caregiver, hospital, clinic, etc.). Monitoring of disease status and disease progression or recurrence

[0128] The disease status can be determined using the frequency of variants in a test sample at one or more variant loci. In some embodiments, an increase in variant frequency indicates an increase in disease severity. In some embodiments, sequencing reads labeled as having a gene variant are attributed to diseased tissue. In some embodiments, sequencing reads labeled as not having a gene variant are attributed to non-disease tissue. In some embodiments, sequencing reads labeled as having a gene variant are attributed to diseased tissue, and sequencing reads labeled as not having a gene variant are attributed to non-disease tissue. In some embodiments, sequencing reads labeled as having a gene variant are attributed to a first diseased tissue, and sequencing reads labeled as not having a gene variant are attributed to a second diseased tissue and / or non-disease tissue.

[0129] In some embodiments, one or more gene variants are used to characterize a disease or cancer. For example, the presence of one or more gene variants can be used to track the underlying cause of a disease (e.g., primary cancer). In some embodiments, the detection of one or more gene variants can be used to characterize treatment-resistant cancer or cancer as being particularly sensitive to a particular treatment. The variant panel used to characterize a disease may be based on known variants, e.g., controlled variants from the literature.

[0130] In some embodiments, the disease state is determined by the state of each variant. In some embodiments, the disease state is determined using multiple variants from a variant panel. For example, in some embodiments, the disease state (DS) is:

number

[0131] Disease progression can be monitored by determining the disease state at two or more time points. The disease state can be indicated by the frequency of variants in a test sample. For example, a first test sample may be taken from a subject at a first time point, and a second test sample may be taken from a subject at a second time point. In some embodiments, the first test sample is used to generate a variant panel and to determine the disease state at the first time point, and the second test sample uses the generated variant panel to determine the disease state at the second time point.

[0132] Subjects can receive treatment for the disease between the first and second test samples (i.e., interventional treatment). Therefore, by monitoring the progression of the disease, it can be determined whether the treatment is effective in treating the disease. The treatment can be further adjusted according to the progression of the disease. For example, if the disease worsens or does not improve, the therapeutic dose can be increased or an alternative treatment can be used.

[0133] The period between the first time point and the second time point can be a desired frequency to effectively monitor the subjects. In some embodiments, the interval between the first and second time points is approximately one week or longer, approximately two weeks or longer, approximately four weeks or longer, approximately eight weeks or longer, approximately twelve weeks or longer, approximately sixteen weeks or longer, approximately six months or longer, approximately one year or longer, or approximately two years or longer.

[0134] In some embodiments, monitoring a subject for disease progression includes monitoring the subject for disease relapse. For example, a subject considered to be in remission may have a minimal amount of residual disease with some risk of relapse. A test sample of the subject may be taken from time to time to determine the disease status in order to find out if the disease has relapsed. If the disease status has relapsed, the subject may be treated for relapsing disease.

[0135] In some embodiments, a method for monitoring disease progression includes: sequencing nucleic acid molecules in a first test sample obtained from a subject with the disease to generate a first sequencing read; generating a personalized variant panel for the subject; sequencing nucleic acid molecules in a second test sample obtained from the subject at a later time than the first test sample to generate a second sequencing read; and labeling the second sequencing read. The sequencing reads include, for example, (a) selecting a gene variant at a variant locus from the personalized variant panel; (b) obtaining one or more sequencing reads associated with a test sample overlapping with the variant locus; (c) generating a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene variant; (d) generating a variant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding variant sequence, wherein the corresponding variant sequence contains the gene variant; and (e) generating the reference match score and variant Each of one or more sequencing reads can be labeled as either having a gene variant, not having a gene variant, or being a null read, based on the variant match score. A sequencing read is labeled as having a gene variant if the reference match score and variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence. A sequencing read is labeled as not having a gene variant if the reference match score and variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence. A sequencing read is labeled as a null read if the reference match score and variant match score are equal.

[0136] Figure 3 shows an exemplary method for monitoring disease progression. The method includes, in step 302, sequencing nucleic acid molecules in a first test sample obtained from a subject with the disease to generate a first sequencing read. From the first sequencing read, a personalized variant panel is generated for the subject. Optionally, the subject's disease status can be determined, which indicates the severity of the subject's disease. The disease status can be represented, for example, by the variant frequency determined for the subject. After a certain period, a second test sample can be obtained from the subject. In step 306, nucleic acid molecules in the second test sample are sequenced. In step 308, gene variants at variant loci are selected from the personalized variant panel. In step 310, sequencing reads that overlap with the variant loci and are associated with the test sample are obtained. A reference match score for each sequencing read is obtained in step 312 by aligning the sequencing read to the corresponding reference sequence, and a variant match score for each sequencing read is generated in step 314 by aligning the sequencing read to the corresponding variant sequence. Using the reference match score and the variant match score, the sequencing reads are labeled in step 316 as having a variant, not having a variant, or null reads. In step 318, the frequency of gene variants is determined using the number of sequencing reads labeled as having a variant and the number of sequencing reads labeled as not having a variant. The determined variant frequencies can be used to determine the disease status of a subject, which indicates the severity of the disease at the time when a second sample is obtained from the subject.

[0137] In some embodiments, the disease being monitored is cancer. For example, in some embodiments, the disease is B-cell cancer, e.g., multiple myeloma, melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or adnexal cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer Cancer, osteosarcoma, chondrosarcoma, hematological cancers, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NH) L) Soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endosplenic synovoma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonic carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharynx These include cranioma, ependymoma, pineal glandoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenous myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, the familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or cancerous tumors.

[0138] In some embodiments, cancers include B-cell carcinoma, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendiceal cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, hematological cancer, adenocarcinoma, inflammatory myofibroblastoma, and digestive cancer. Tubostromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteosarcoma, Chordoma, angiosarcoma, endosarcoma, lymphangiosarcoma, lymphangioendosarcoma, synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, liver cancer, cholangiocarcinoma, choriocarcinoma, seminocarcinoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal gland These include cell tumors, glioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, aplastic myelogenesis, eosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumors.

[0139] In some embodiments, the methods described herein are used to identify viral or bacterial strains. Bacteria and viruses can mutate, and clearly distinguishing specific strain types can be particularly important for treating infected subjects. For example, it is important to know whether a Staphylococcus aureus strain infecting a subject is resistant to methicillin and / or vancomycin. Antibiotic or other drug-resistant bacteria and viruses have genomic signatures, and the methods described herein can be used to rapidly characterize different strains.

[0140] In some examples, disclosed methods for detecting gene variants or determining mutant allele frequencies in subject-derived test samples can be implemented as part of a genomic profiling process that includes identifying the presence of mutant sequences at one or more loci in subject-derived samples, as part of the detection, monitoring, risk factor prediction, or treatment selection for a particular disease, e.g., cancer. In some examples, a variant panel selected for genomic profiling may include the detection of mutant sequences at a selected set of loci. In some examples, a variant panel selected for genomic profiling may include the detection of mutant sequences at several loci via comprehensive genomic profiling (CGP), a next-generation sequencing (NGS) approach used to evaluate hundreds of genes (including relevant cancer biomarkers) in a single assay. Including disclosed methods for detecting gene variants or determining mutant allele frequencies as part of a genomic profiling process can improve the effectiveness of, for example, disease detection calls made based on genomic profiles, for example, by independently confirming the presence of disease or cancer driver mechanisms (e.g., impaired DNA mismatch repair (MMR) mechanisms) in a given patient sample.

[0141] In some cases, a genomic profile may include information about the presence of genes (or their variant sequences), copy number variations, epigenetic traits, proteins (or their modifications), and / or other biomarkers in an individual's genome and / or proteome, as well as the individual's corresponding phenotypic traits, and information about the interactions between genetic or genomic traits, phenotypic traits, and environmental factors.

[0142] In some cases, a subject's genomic profile may include results from comprehensive genomic profiling (CGP) tests, nucleic acid sequencing-based tests, gene expression profiling tests, cancer hotspot panel tests, DNA methylation tests, DNA fragmentation tests, RNA fragmentation tests, or any combination thereof.

[0143] Genomic profiles can be used to select anticancer drugs, administer anticancer drugs, or apply anticancer treatments to subjects (i.e., decisions regarding the selection, administration, or application of anticancer treatments can be based on the generated genomic profile). In some implementations of this method, genomic profiles are used as a basis for enrolling subjects in clinical trials for selected disease treatments (e.g., anticancer therapy). Disease treatment and detection assays

[0144] The methods described herein can be used when treating subjects with a disease. For example, the detection of gene variants or the determination of allele frequencies in a test sample can be used when making treatment decisions (e.g., cancer treatment) or when suggesting treatment decisions for a subject. In another example, the detection of gene variants or the determination of allele frequencies in a test sample can be used when tailoring disease (e.g., cancer) therapy. As described above, the methods may include monitoring disease progression, such as cancer progression, in subjects. Monitoring disease progression can enable clinicians to provide better treatment decisions and can be used to screen for disease (e.g., cancer) recurrence or metastasis.

[0145] A first test sample can be obtained from a subject with the disease, and nucleic acid molecules from the test sample can be sequenced to generate a first sequencing read, which is used to generate an individualized variant panel for the subject. The disease therapy is then administered to the subject, and after a certain period, a second test sample is obtained from the subject at a second time point. Nucleic acid molecules from the second test sample can be sequenced to generate a second sequencing read, which can be labeled using the method described herein. For example, the second sequencing read is generated by (a) selecting a gene variant at a mutant locus from an individualized mutant panel, (b) obtaining one or more sequencing reads related to a test sample that overlaps with the mutant locus, (c) aligning each sequencing read to a corresponding reference sequence, thereby generating a reference match score for each of the one or more sequencing reads, wherein the corresponding reference sequence does not contain the gene variant, (d) aligning each sequencing read to a corresponding mutant sequence, thereby generating a mutant match score for each of the one or more sequencing reads, wherein the corresponding mutant sequence contains the gene variant, and (e) the reference match score and Each of one or more sequencing reads can be labeled as either having a gene variant, not having a gene variant, or being a null read, based on the variant match score. A sequencing read is labeled as having a gene variant if the reference match score and variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence. A sequencing read is labeled as not having a gene variant if the reference match score and variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence. A sequencing read is labeled as a null read if the reference match score and variant match score are equal.A first disease state can be determined using a first sequencing read, and a second disease state can be determined using a labeled second sequencing read. Disease progression can be determined by comparing the first and second disease states. Disease therapy administered to a subject can be adjusted based on disease progression, and the adjusted disease therapy can then be administered to the subject.

[0146] The detected gene variants or determined variant allele frequencies can be used as a basis for adjusting the dosage of disease therapy (e.g., anti-cancer therapy) or for selecting different disease therapies depending on disease progression. The adjusted disease therapy can then be administered to the subject.

[0147] In some implementations of this method, the detected gene variants or determined variant allele frequencies are used as a basis for enrolling subjects in clinical trials of selected disease treatments (e.g., anti-cancer therapy). For example, a clinical trial may enroll patients who have (or do not have) one or more predetermined gene variants and who may be treated in the clinical trial with a selected disease treatment (e.g., anti-cancer therapy).

[0148] In an exemplary embodiment, a method for treating a subject having a disease (such as cancer) is to obtain a first test sample from the subject, sequence nucleic acid molecules in the first test sample to generate a first sequencing read, determine a first disease state using the first sequencing read, generate a personalized variant panel for the subject, administer a disease therapy to the subject, obtain a second test sample from the subject after the disease therapy has been administered to the subject, sequence nucleic acid molecules in the second test sample to generate a second sequencing read, (a) label the second sequencing read by selecting a gene variant at a variant locus from the variant panel, (b) obtain one or more sequencing reads associated with a test sample overlapping with the variant locus, (c) generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence such that the corresponding reference sequence does not contain a gene variant, and (d) label each sequencing read to a corresponding variant sequence (e) Labeling a second sequencing read by aligning it to generate a variant match score for each of one or more sequencing reads, wherein the corresponding variant sequence contains a gene variant; and labeling each of one or more sequencing reads based on the reference match score and the variant match score, wherein the sequencing read is labeled to contain a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence; the sequencing read is labeled to not contain a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence; and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal; and determining a second disease state using the labeled second sequencing read.This includes determining disease progression by comparing a first disease state with a second disease state, adjusting the disease therapy administered to the subject based on disease progression, and administering the adjusted disease therapy to the subject.

[0149] In some embodiments, disease therapy (e.g., anticancer therapy to treat cancer) includes surgery (e.g., resection surgery to remove one or more cancers). In some embodiments, disease therapy includes radiotherapy (e.g., external beam radiotherapy, stereotactic radiotherapy, intensity-modulated radiotherapy, volume-modulated arc therapy, particle therapy (e.g., proton therapy), auger therapy, close-range radiotherapy, or whole-body radioisotope therapy). In some embodiments, disease therapy includes the administration of one or more chemical agents (e.g., anticancer agents), such as one or more chemotherapeutic agents to treat cancer. Exemplary chemotherapeutic agents include, but are not limited to, anthracyclines (e.g., daunorubicin, epirubicin, idarubicin, mitoxantrone, barrubicin), alkylating agents or alkylating agent-like agents (e.g., carboplatin, carmustine, cisplatin, cyclophosphamide, melphalan, procarbazine, or thiotepa), or taxanes (e.g., paclitaxel, docetaxel, or taxotere). In some examples, the method may further include administering anticancer agents or applying anticancer treatments to subjects based on the generated genomic profile. Anticancer agents or anticancer treatments may refer to compounds effective in treating cancer cells. Examples of anticancer agents or anticancer therapies include, but are not limited to, alkylating agents, antimetabolites, natural products, hormones, chemotherapy, radiotherapy, immunotherapy, surgery, or treatments configured to target defects in specific cellular signaling pathways, such as defects in the DNA mismatch repair (MMR) pathway.

[0150] In some embodiments, the therapy is immunotherapy. In some embodiments, the therapy is an immune checkpoint inhibitor.

[0151] In some embodiments, disease therapy is targeted therapy. Exemplary targeted therapies include tyrosine kinase inhibitors (e.g., imatinib, gefitinib, erlotinib, sorafenib, sunitinib, dasatinib, lapatinib, nilotinib, bortezomib), JAK inhibitors (e.g., tofacitinib), ALK inhibitors (e.g., crizotinib), BCL-2 inhibitors (e.g., ovatoclax, navitoclax, gossypol), PARP inhibitors (e.g., iniparib, olaparib), PI3K inhibitors (e.g., perifosine), apatinib, BRAF inhibitors ( For example, these include vemurafenib, dabrafenib, LGX818), MEK inhibitors (e.g., trametinib, MEK162), CDK inhibitors, Hsp90 inhibitors, or salinomycin), serine / threonine kinase inhibitors (e.g., temsirolimus, everolimus, vemurafenib, trametinib, or dabrafenib), or monoclonal antibodies (e.g., pembrolizumab, rituximab, trastuzumab, alemtuzumab, cetuximab, panitumumab, or bevacizumab).

[0152] In some embodiments, the therapeutic or anti-cancer therapy administered to a subject is selected based on (e.g., in response to) the detection of gene variants in a sample using the methods described herein. The selected anti-cancer therapy can be administered to the subject. Exemplary selected cancer therapies may include chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. For example, detection of a specific biomarker using the methods described herein can be used as a basis for selecting a particular mode of treatment. The selected anti-cancer therapy can be administered to the subject. Exemplary selected cancer therapies may include chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. Exemplary personalized treatment selections for a given identified mutation are listed in Table 1. [Table 1]

[0153] In some embodiments, the disease being treated is cancer. For example, in some embodiments, the disease is B-cell cancer, e.g., multiple myeloma, melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or adnexal cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer Cancer, osteosarcoma, chondrosarcoma, hematological cancers, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NH) L) Soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endosplenic synovoma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonic carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharynx These include cranioma, ependymoma, pineal glandoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenous myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, the familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or cancerous tumors.

[0154] The detection of gene variants or determined variant allele frequencies can be used to diagnose or confirm a disease (such as cancer) in a subject. For example, one or more gene variants may be associated with a disease (e.g., cancer or a specific type of cancer), and a diagnosis may be made regarding such an association.

[0155] The detection of gene variants or determined variant allele frequencies can be used to identify patients as eligible for clinical trials of disease treatments (e.g., anti-cancer treatments for patients with cancer). Once identified, patients can be enrolled in the clinical trial. This method may further include administering the disease treatment to the patient. Computer systems and methods

[0156] The methods described herein can be implemented using one or more computer systems. Such computer systems may include one or more programs configured to run one or more processors in order for the computer systems to perform such methods. One or more steps of the computer implementation method may be performed automatically.

[0157] In some embodiments, a computer implementation method for detecting the presence of gene variants in a test sample derived from a subject and / or determining the frequency of a variant allele, or for labeling sequencing reads associated with a test sample derived from a subject, includes: (a) using one or more processors to select a gene variant at a variant locus from a variant panel stored in memory; (b) using one or more processors to receive one or more sequencing reads stored in memory, wherein the sequencing read is associated with a test sample that overlaps with the variant locus; (c) using one or more processors to generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence obtained from memory, wherein the corresponding reference sequence does not contain the gene variant; and (d) using one or more processors to align each sequencing read to a corresponding reference sequence obtained from memory (e) generating a mutant match score for each of one or more sequencing reads by aligning to a mutant sequence, wherein the corresponding mutant sequence contains a gene mutant; and (a) using one or more processors to label each of the one or more sequencing reads as having a gene mutant, not having a gene mutant, or being a null read, wherein the sequencing read is labeled to have a gene mutant if the reference match score and the mutant match score indicate that the sequencing read matches the corresponding mutant sequence more closely than the corresponding reference sequence; the sequencing read is labeled not to have a gene mutant if the reference match score and the mutant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding mutant sequence; and the sequencing read is labeled as a null read if the reference match score and the mutant match score are equal.

[0158] In some embodiments of the computer implementation method, the method further includes generating a corresponding reference sequence and / or a corresponding variant sequence. In some embodiments, the corresponding reference sequence and the corresponding variant sequence are identical except for the gene variant.

[0159] In some embodiments of the computer implementation method, one or more sequencing reads include multiple sequencing reads that overlap with the mutant locus, and the method further includes determining some sequencing reads from multiple sequencing reads that have the gene mutant or some sequencing reads from multiple sequencing reads that do not have the gene mutant. In some embodiments, the method further includes determining the mutant frequency for the gene mutant using the number of sequencing reads that have the gene mutant and the number of sequencing reads that do not have the gene mutant.

[0160] In some embodiments of the computer implementation method, the method includes labeling one or more sequencing reads related to a test sample for multiple gene variants at different variant loci selected from a variant panel.

[0161] In some embodiments of the computer implementation method, the method includes determining the disease status of a subject. For example, the disease status may be a value proportional to the percentage of circulating tumor DNA (ctDNA) compared to total cell-free DNA (cfDNA) in the test sample.

[0162] In some embodiments, the reference match score and variant match score are determined using a sequence alignment algorithm. In some embodiments, the reference match score and variant match score are determined using the Smith-Waterman alignment algorithm. In some embodiments, the reference match score and variant match score are determined using the Needleman-Wunsch alignment algorithm.

[0163] Figure 4 shows an exemplary computer implementation method for determining the frequency of variants in a test sample derived from a subject. Step 402 includes using one or more processors to select gene variants at variant loci from a variant panel stored in memory. In some embodiments, this step includes receiving information on gene variants and variant loci for one or more variants from the variant panel stored in memory. For example, a processor can access memory to obtain information on gene variants and variant loci that can be enumerated in a table or file stored in memory. The selection is made from the variant panel by any appropriate process (e.g., randomly, sequentially, using a prioritization rank). In some embodiments, the computer implementation method is repeated until a desired number (or all) of variants in the variant panel have been analyzed.

[0164] Step 404 includes receiving one or more sequencing reads stored in memory in one or more processors, the sequencing reads being associated with test samples that overlap with the mutant locus. For example, a processor can access memory to retrieve one or more sequencing reads that overlap with the mutant locus. Memory can store tables or files containing sequencing reads (e.g., BAM or SAM files) containing reads and read loci. Sequence reads in the table or file that overlap with the selected mutant locus can then be selected and received in one or more processors.

[0165] Step 406 includes generating a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence obtained from memory using one or more processors, wherein the corresponding reference sequence does not include any gene variants. In some embodiments, this step includes receiving the reference sequence corresponding to the selected variant (i.e., the corresponding reference sequence). For example, the corresponding reference sequence may be stored in a table or file in memory. In some embodiments, the table or file storing the corresponding reference sequence is the same table or file that stores information about the selected variant or variant panel. In some embodiments, the table or file storing the corresponding reference sequence is a different table or file from the table or file that stores information about the selected variant or variant panel. Each sequencing read received in one or more processors corresponding to the selected variant is aligned to the corresponding reference sequence using an alignment module. The alignment module implements an alignment algorithm (e.g., the Smith-Waterman alignment algorithm or the Needleman-Wunsch alignment algorithm) to generate a reference match score. In some embodiments, the reference match score is stored in memory, for example, by automatically updating a table or file that stores sequencing reads, or by automatically generating a new table or file that contains the reference match score and associated reads or read identifiers.

[0166] Step 408 includes generating a mutant match score for each of one or more sequencing reads by aligning each sequencing read to a corresponding mutant sequence obtained from memory using one or more processors, where the corresponding mutant sequence includes a gene variant. In some embodiments, this step includes receiving the mutant sequence (i.e., the corresponding mutant sequence) corresponding to the selected mutant. For example, the corresponding mutant sequence can be stored in a table or file in memory (this may be the same file or table as the table or file that stores the corresponding reference sequence, or it may be a different file). In some embodiments, the table or file that stores the corresponding mutant sequence is the same table or file that stores information about the selected mutant or mutant panel. In some embodiments, the table or file that stores the corresponding mutant sequence is a different table or file from the table or file that stores information about the selected mutant or mutant panel. Each sequencing read received in one or more processors corresponding to the selected mutant is aligned to the corresponding mutant sequence using an alignment module. The alignment module implements an alignment algorithm (generally the same alignment algorithm used to align sequencing reads with a reference alignment module) to generate variant match scores. In some embodiments, the variant match scores are stored in memory, for example, by automatically updating a table or file that stores sequencing reads, or by automatically generating a new table or file that contains the reference match scores and associated reads or read identifiers. In some embodiments, a table or file containing both the reference match scores and variant match scores is automatically generated.

[0167] Step 410 includes labeling each of one or more sequencing reads using one or more processors as having a gene variant, not having a gene variant, or being a null read, based on a reference match score and a variant match score, wherein the sequencing read is labeled as having a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled as not having a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal. In some embodiments, labeling each of one or more sequencing reads using one or more processors as having a gene variant, not having a gene variant, or being a null read is based on the reference match score, the variant match score is implemented by a labeling module. The labeling module can compare the variant match score with the reference match score. A sequencing read is labeled as having a gene variant if the reference match score and variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence. A sequencing read is labeled as not having a gene variant if the reference match score and variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence. Furthermore, in some embodiments, a sequencing read is labeled as a null read if the reference match score and variant match score are equal. In some embodiments, the labels associated with sequencing reads are automatically stored in memory. For example, in some embodiments, one or more processors automatically access a table or file stored in memory and update the file to include the labels for the sequencing reads.In some embodiments, one or more processors automatically generate a table or file and store it in memory containing markers for sequence determination reads.

[0168] Step 412 includes using one or more processors to determine the frequency of gene variants using some sequencing reads that have variants and some sequencing reads that do not have variants. In some embodiments, one or more processors automatically generate or update a table or file in memory to record the frequencies of gene variants.

[0169] A computer implementation method for detecting gene variants in a test sample derived from a subject or determining the allele frequencies of gene variants may include the use of an electronic system comprising one or more processors and a memory for storing reference sequences and variant sequence pairs. The reference sequences and variant sequence pairs correspond to gene variants queried by a method that can be selected from a panel of variants stored in memory using one or more processors. One or more processors may receive one or more sequencing reads from the test sample, the sequencing reads overlapping with the loci of the queried gene variant. One or more processors may also receive reference sequences from memory and generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to the corresponding reference sequence. Furthermore, one or more processors may receive variant sequences from memory and generate a variant match score for each of the one or more sequencing reads by aligning each sequencing read to the corresponding variant sequence. Based on the reference match score and the variant match score, the sequencing reads may be labeled as having a gene variant, not having a gene variant, or being a null read. A sequencing read is labeled as having a gene variant if its reference match score and variant match score indicate that it matches the corresponding variant sequence more closely than the corresponding reference sequence. A sequencing read is labeled as not having a gene variant if its reference match score and variant match score indicate that it matches the corresponding reference sequence more closely than the corresponding variant sequence. Finally, a sequencing read is labeled as a null read if its reference match score and variant match score are equal. Labeled sequencing reads can be stored in memory, or some sequencing reads having gene variants and / or some sequencing reads not having gene variants (and possibly a number of null reads) can be stored in memory.In some embodiments, the computer implementation process can call a sample having a variant and / or determine the variant allele frequency for the sample using the number of sequencing reads labeled as having a gene variant and / or the number of sequencing reads labeled as not having a gene variant. This process can be repeated for any number of gene variants to be queried.

[0170] In some embodiments, a computer implementation method for detecting gene variants in a test sample derived from a subject or determining the allele frequency of a gene variant comprises an electronic device comprising one or more processors and a memory that stores a reference sequence that does not contain the gene variant and a variant sequence that contains the gene variant at the variant locus, wherein one or more processors receive one or more sequencing reads related to the test sample corresponding to the reference sequence and the variant sequence; one or more processors receive the reference sequence from memory; one or more processors generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to the corresponding reference sequence; one or more processors receive the variant sequence from memory; and one or more processors generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to the corresponding variant The method includes generating a variant match score for each of one or more sequencing reads by aligning them to a body sequence, and labeling each of the one or more sequencing reads in one or more processors as having a gene variant, not having a gene variant, or being a null read, based on the reference match score and the variant match score, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled to not have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal. In some embodiments, the method further includes storing the labels associated with each sequencing read in memory.

[0171] In some embodiments, the computer implementation method may further include using one or more processors to recall the presence of gene variants in the test sample based on one or more labeled sequence reads. The recall of gene variants can be stored in memory by one or more processors.

[0172] In some embodiments, the computer implementation method may further include using one or more processors to determine the mutant allele frequencies of gene variants in a test sample based on one or more labeled sequence reads. The mutant allele frequency recall can be stored in memory.

[0173] A computer implementation method can generate reference sequences and / or variant sequences used in this method by relying on the use of a variant panel stored in memory. This method may include using one or more processors to select gene variants from a variant panel, using one or more processors to generate reference sequences and / or variant sequences, and storing the reference sequences and / or variant sequences in memory. In other embodiments, the reference sequences and / or sequenced variants used in this method are pre-stored in memory and correspond to the queried gene variants.

[0174] In some embodiments, the computer implementation method includes the automatic generation or updating of reports (such as electronic medical records). The reports may include a call for the presence or absence of gene variants, a call for the frequency of variant alleles, and / or one or more disease states. The reports may also include subject identification information (e.g., name, identification number, etc.). The reports may be stored in memory and / or transmitted to a second electronic device (e.g., the subject's electronic device or the subject's healthcare provider).

[0175] Figure 5A shows an example of a computing device according to one embodiment. Device 500 can be a host computer connected to a network. Device 500 can be a client computer or a server. As shown in Figure 5A, device 500 can be any preferred type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device, e.g., telephone or tablet). The device may include, for example, one or more of a processor 510, an input device 520, an output device 530, storage 540, and a communication device 560. The input device 520 and the output device 530 can generally correspond to those described above and can be connected to or integrated with the computer.

[0176] The input device 520 can be any suitable device that provides input, such as a touchscreen, keyboard or keypad, mouse, or voice recognition device. The output device 530 can be any suitable device that provides output, such as a touchscreen, haptic device, or speaker. In some embodiments, the input device 520 and the output device 530 can be the same or different devices.

[0177] Storage 540 can be any suitable device that provides storage, such as electrical, magnetic, or optical memory, including RAM (volatile and non-volatile), cache, hard drive, or removable storage disk. Communication device 560 can include any suitable device that can send and receive signals over a network, such as a network interface chip or device. The computer components can be connected in any suitable way, such as via the physical bus 580 or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).

[0178] The software 550, which is stored in the storage 540 and can be executed by the processor 510, may include, for example, programming to implement the functions of this disclosure (for example, as implemented in the device described above).

[0179] The software 550 may also be stored and / or transferred in any non-temporary computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device such as those described above, and may fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, the computer-readable storage medium may be any medium such as storage 540, which may contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.

[0180] The software 550 can also be propagated by an instruction execution system, apparatus, or device such as those described above, or in any transmission medium for use in connection with them, to fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, the transmission medium can be any medium that can communicate, propagate, or transmit transmission programming by or for use in connection with the instruction execution system, apparatus, or device. The transmission-readable medium may include, but is not limited to, wired or wireless transmission media of electronic, magnetic, optical, electromagnetic, or infrared.

[0181] Device 500 can be connected to a network which can be any preferred type of interconnected communication system. The network can implement any preferred communication protocol and can be protected by any preferred security protocol. The network may include network links of any preferred configuration that can implement the transmission and reception of network signals, such as wireless network connections (T1 or T3 lines), cable networks, DSL, or telephone lines.

[0182] Device 500 can implement any operating system suitable for running over a network. Software 550 can be written in any suitable programming language such as C, C++, Java®, or Python. In various embodiments, application software embodying the functionality of this disclosure can be deployed in different configurations (e.g., in a client / server configuration, or via a web browser as a web-based application or web service). In some embodiments, the operating system is run by one or more processors, for example, processor 510.

[0183] Device 500 may further include a sequencer 570, which can be any suitable nucleic acid sequencing instrument.

[0184] Figure 5B shows an example of a computing system according to one embodiment. In the computing system 590, device 500 (for example, as described above and shown in Figure 5A) is connected to a network 592, which is also connected to device 594. In some embodiments, device 594 is a sequencer (e.g., a next-generation sequencer). Exemplary sequencers, but not limited to, the Roche / 454 Genome Sequencer (GS) FLX System, the Illumina / Solexa Genome Analyzer (GA), the Illumina HiSeq 2500, HiSeq 3000, HiSeq 4000, and NovaSeq 6000 sequencing systems, and the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope This includes the Gene sequencing system, or the Pacific BioSciences PacBio RS system.

[0185] Devices 500 and 594 can communicate using an appropriate communication interface over a network 592, such as a local area network (LAN), a virtual private network (VPN), or the internet. In some embodiments, network 592 can be, for example, the internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 500 and 594 can communicate partially or entirely over wireless or wired communication, such as Ethernet® or IEEE 802.11b wireless. Furthermore, devices 500 and 594 can communicate over a second network, such as a mobile / cellular network, using a suitable communication interface. Communication between devices 500 and 594 can further include, or communicate with, various servers, such as mail servers, mobile servers, media servers, and telephone servers. In some embodiments, devices 500 and 594 can communicate directly (instead of, or in addition to, communication over network 592) over wireless or wired communication, such as Ethernet® or IEEE 802.11b wireless. In some embodiments, devices 500 and 594 communicate via communications 596, which can be a direct connection or occur over a network (e.g., network 592).

[0186] One or all of devices 500 and 594 are generally programmed to include logic (e.g., HTTP web server logic) accessed from local or remote databases or other sources of data and content, or to format data, in order to provide and / or receive information over the network 592 in accordance with the various examples described herein.

[0187] In an exemplary embodiment, the system comprises one or more processors, memory, and one or more programs, wherein one or more programs are stored in memory and configured to be executed by one or more processors, and one or more programs perform the following functions: (a) select gene variants at mutant loci from a mutant panel; (b) obtain one or more sequencing reads related to a test sample that overlaps with the mutant loci; (c) generate a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene variant; and (d) generate a mutant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding mutant sequence. There is an electronic device that includes instructions for generating a corresponding variant sequence containing a gene variant, and (e) labeling each of one or more sequencing reads, based on a reference match score and a variant match score, as either having a gene variant, not having a gene variant, or being a null read, wherein the sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled to not have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and the sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0188] In another exemplary embodiment, a non-temporary computer-readable storage medium storing one or more programs, wherein when one or more programs are executed by one or more processors of an electronic device having a display, the electronic device performs the following actions: (a) selecting gene variants at mutant loci from a mutant panel; (b) obtaining one or more sequencing reads related to a test sample overlapping with the mutant loci; (c) generating a reference match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding reference sequence, wherein the corresponding reference sequence does not contain the gene variant; and (d) generating a mutant match score for each of the one or more sequencing reads by aligning each sequencing read to a corresponding mutant sequence, A non-temporary computer-readable storage medium includes instructions to generate a corresponding variant sequence containing a gene variant, and to label each of one or more sequencing reads as either having a gene variant, not having a gene variant, or being a null read, based on a reference match score and a variant match score, wherein a sequencing read is labeled to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, a sequencing read is labeled not to have a gene variant if the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, and a sequencing read is labeled as a null read if the reference match score and the variant match score are equal.

[0189] While the present disclosure and examples are fully described with reference to the accompanying drawings, it should be noted that various modifications and changes will be apparent to those skilled in the art. Such modifications and changes should be understood to fall within the scope of the present disclosure and examples as defined by the claims.

[0190] The above description is provided for illustrative purposes with reference to specific embodiments. However, the above exemplary description is not intended to be exhaustive or to limit the invention to the exact form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been selected and described to best illustrate the principles of the art and their practical applications. Therefore, other persons skilled in the art will be able to best utilize the art and various embodiments with various modifications suitable for the specific use intended. [Examples]

[0191] The examples provided herein are for illustrative purposes only and are not intended to limit the scope of the invention. [Example 1]

[0192] Sequence reads from Sample 1 and Sample 2 were initially obtained using a targeted sequencing method invoked with a standard variant invocation protocol to generate a curated set of variants from baseline samples, as well as the depths of the variants and alleles. For Sample 1 and Sample 2, variant panels and allele depths were selected. The variants in the variant panel of Sample 1 ranged from 1 to 22 nucleotide lengths (Figure 6A), while the variant panel of Sample 2 contained only single-nucleotide-length variants (Figure 6B).

[0193] A reference sequence (i.e., corresponding reference sequence) and a mutant sequence (i.e., mutant reference sequence) corresponding to each mutant in the mutant panel were generated. The mutant or reference base had 200 base pairs adjacent to both sides of the mutant locus to generate the corresponding mutant sequence and corresponding reference sequence.

[0194] Each sequence reading from Sample 1 and Sample 2, which overlapped with the mutant loci of the mutants in the mutant panel, was aligned with the corresponding reference sequence and corresponding mutant sequence using the Striped Smith-Waterman alignment algorithm to generate reference match scores and mutant match scores, respectively. The match scores were used to label reads as either containing a mutant, not containing a mutant, or null reads. 199 mutants were detected from Sample 1 and 374 mutants were detected from Sample 2. Figures 7A and 8A show plots of the number of mutant reads detected by comparing the match score (y-axis) with the number of mutant reads detected using a standard mutant calling protocol (x-axis), for Sample 1 (Figure 7A) and Sample 2 (Figure 8A), on a logarithmic scale (left) and normalized scale (right). Figures 7B and 8B show plots (y-axis) of the depth of the mutant allele at each mutant locus, for the total number of sequencing reads labeled as having or not having a mutant (i.e., excluding null reads), against the depth of the mutant locus at each mutant locus, for the total number of sequencing reads from the initial pool of overlapping sequencing reads (x-axis), for Sample 1 (Figure 7B) and Sample 2 (Figure 8B), on a logarithmic scale (left) and normalized scale (right). [Example 2]

[0195] Sequence reads from Sample 1 and Sample 2 were initially obtained using a targeted sequencing method invoked with a standard variant invocation protocol to generate a curated set of variants from baseline samples, as well as the depths of the variants and alleles. For Sample 1 and Sample 2, variant panels and allele depths were selected. The variants in the variant panel of Sample 1 ranged from 1 to 22 nucleotide lengths (Figure 6A), while the variant panel of Sample 2 contained only single-nucleotide-length variants (Figure 6B).

[0196] A reference sequence (i.e., corresponding reference sequence) and a mutant sequence (i.e., mutant reference sequence) corresponding to each mutant in the mutant panel were generated. The mutant or reference base had 500 base pairs adjacent to both sides of the mutant locus to generate the corresponding mutant sequence and corresponding reference sequence.

[0197] Each sequence read from Sample 1 and Sample 2, which overlapped with a single nucleotide of the mutant locus of the mutants in the mutant panel, was aligned with the corresponding reference sequence and corresponding mutant sequence using the Striped Smith-Waterman alignment algorithm to generate reference match scores and mutant match scores, respectively. The match scores were used to label reads as either containing a mutant, not containing a mutant, or null reads. 202 mutants were detected from Sample 1 and 375 mutants were detected from Sample 2. Figures 9A and 10A show plots of the number of mutant reads detected by comparing the match score (y-axis) with the number of mutant reads detected using a standard mutant calling protocol (x-axis), on a logarithmic scale (left) and normalized scale (right) for Sample 1 (Figure 9A) and Sample 2 (Figure 10A), respectively. Figures 9B and 10B show plots (y-axis) of the depth of the mutant locus at each mutant locus, for the total number of sequencing reads labeled as having or not having a mutant (i.e., excluding null reads), against the depth of the mutant locus at each mutant locus, for the total number of sequencing reads from the initial pool of overlapping sequencing reads (x-axis), for Sample 1 (Figure 9B) and Sample 2 (Figure 10B), on a logarithmic scale (left) and normalized scale (right). The present invention provides, for example, the following items: (Item 1) A method for detecting gene variants in a test sample derived from a subject or for determining the frequency of a variant allele, To provide multiple nucleic acid molecules obtained from test samples from subjects, Ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules, The amplification of one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules, The process involves capturing the amplified nucleic acid molecules from the aforementioned amplified nucleic acid molecules, The process involves sequencing the captured nucleic acid molecule using a sequencer to obtain a plurality of sequence reads representing the captured nucleic acid molecule, wherein one or more of the plurality of sequence reads overlap with a variant locus within a subgenome section in the sample. One or more processors receive one or more sequencing reads corresponding to a reference sequence and a variant sequence, In the aforementioned one or more processors, receiving the reference array from memory, In the one or more processors, a reference match score is generated for each of the one or more sequence determination reads by aligning each sequence determination read to the corresponding reference sequence. In the one or more processors, receiving the variant sequence from the memory, In the one or more processors, a variant match score is generated for each of the one or more sequence reads by aligning each sequence read to the corresponding variant sequence. The one or more processors include labeling each of the one or more sequencing reads as having the gene variant, not having the gene variant, or being a null read, based on the reference match score and the variant match score. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled as having the gene variant. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, the sequencing read is labeled as not having a gene variant. A method in which, if the reference match score and the variant match score are equal, the sequencing read is labeled as a null read. (Item 2) The method according to item 1, wherein the one or more adapters include an amplification primer, a flow cell adapter sequence, a substrate adapter sequence, or a sample index sequence. (Item 3) The captured nucleic acid molecule undergoes hybridization into one or more bait molecules. The method according to item 2, wherein the nucleic acid molecules amplified by are captured by the aforementioned method. (Item 4) The method according to item 3, wherein the one or more bait molecules comprise one or more nucleic acid molecules, and each nucleic acid molecule comprises a region complementary to the region of the captured nucleic acid molecule. (Item 5) The method according to any one of items 1 to 4, wherein the amplification of nucleic acid molecules includes performing polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques. (Item 6) The method according to any one of items 1 to 5, wherein the sequencing includes the use of massively parallel sequencing (MPS) technology, whole-genome sequencing (WGS), whole-exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. (Item 7) The method according to item 6, wherein the sequencing includes massively parallel sequencing, and the massively parallel sequencing technique includes next-generation sequencing (NGS). (Item 8) The method according to any one of items 1 to 7, wherein the sequencer includes a next-generation sequencer. (Item 9) It is a method, One or more processors receive one or more sequencing reads related to a test sample corresponding to a reference sequence and a variant sequence, In the aforementioned one or more processors, receiving the reference array, In the one or more processors, a reference match score is generated for each of the one or more sequence determination reads by aligning each sequence determination read to the corresponding reference sequence. In the one or more processors, receiving the variant sequence, In the one or more processors, a variant match score is generated for each of the one or more sequence reads by aligning each sequence read to the corresponding variant sequence. The one or more processors include labeling each of the one or more sequencing reads as having the gene variant, not having the gene variant, or being a null read, based on the reference match score and the variant match score. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled as having the gene variant. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, the sequencing read is labeled as not having a gene variant. A method in which, if the reference match score and the variant match score are equal, the sequencing read is labeled as a null read. (Item 10) The method according to any one of items 1 to 9, comprising storing in the memory a label associated with each sequencing read labeled as having the gene variant and / or each sequencing read labeled as not having the variant. (Item 11) The further includes using one or more processors to invoke the presence or absence of the gene variant in the test sample based on one or more labeled sequence determination reads, and storing the invoked gene variant in memory. The method described in any one of items 1 through 10. (Item 12) The method according to any one of items 1 to 11, further comprising using one or more processors to determine the mutant allele frequencies of the gene variant in the test sample based on one or more labeled sequencing reads, and storing the mutant allele frequencies in the memory. (Item 13) Using the one or more processors, Selecting the gene variant from the variant panel stored in the memory using one or more of the aforementioned processors, Using the one or more processors, generate the reference sequence or the variant sequence, The method according to any one of items 1 to 12, comprising storing the reference sequence or the variant sequence in the memory. (Item 14) The method according to any one of items 1 to 13, wherein the one or more sequencing reads include a plurality of sequencing reads that overlap with the mutant locus, and the method further comprises using the one or more processors to determine some sequencing reads from the plurality of sequencing reads having the gene mutant or some sequencing reads from the plurality of sequencing reads that do not have the gene mutant. (Item 15) The method according to any one of items 1 to 14, comprising using one or more processors to label one or more sequence readings related to the test sample for multiple gene variants at different variant loci selected from a variant panel. (Item 16) The method according to any one of items 1 to 15, comprising using one or more of the processors to determine the disease state of the subject. (Item 17) The method according to any one of items 1 to 16, comprising using one or more processors to generate a report including (1) identification information of the subject, and (2) a call for the presence or absence of the gene variant or a call for the variant allele frequency. (Item 18) The method of item 17, which includes transmitting the aforementioned report to a second electronic device. (Item 19) The method described in item 18, wherein the report is transmitted via a computer network or peer-to-peer connection. (Item 20) A method for detecting gene variants in a test sample derived from a subject or for determining the frequency of a variant allele, Selecting the gene variant at the mutant locus from the mutant panel, To obtain one or more sequencing reads related to the test sample that overlap with the aforementioned mutant gene locus, The process involves aligning each sequencing read to a corresponding reference sequence to generate a reference match score for each of the one or more sequencing reads, wherein the corresponding reference sequence does not contain the gene variant. The process involves aligning each sequencing read to a corresponding mutant sequence to generate a mutant match score for each of the one or more sequencing reads, wherein the corresponding mutant sequence includes the gene mutant. Based on the reference match score and the variant match score, labeled sequencing To generate reads, each of the one or more sequencing reads is labeled as either having the gene variant, not having the gene variant, or being a null read. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled as having the gene variant. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, the sequencing read is labeled as not having a gene variant. A method in which, if the reference match score and the variant match score are equal, the sequencing read is labeled as a null read. (Item 21) The method according to item 20, further comprising calling for the presence of the gene variant in the test sample based on one or more labeled sequencing reads. (Item 22) The method according to item 20, further comprising calling for the presence or absence of the gene variant in the test sample based on one or more labeled sequencing reads. (Item 23) The method according to any one of items 20 to 22, comprising generating the corresponding reference sequence or the corresponding variant sequence. (Item 24) The method according to any one of items 20 to 23, wherein the one or more sequencing reads include a plurality of sequencing reads that overlap with the mutant locus, and the method further comprises determining the number of sequencing reads from the plurality of sequencing reads having the gene mutant or the number of sequencing reads from the plurality of sequencing reads that do not have the gene mutant. (Item 25) The method according to item 24, comprising determining the mutant allele frequency for the gene variant using the number of sequencing reads having the gene variant and the number of sequencing reads not having the gene variant. (Item 26) The method according to any one of items 20 to 25, comprising labeling one or more sequencing reads related to the test sample for multiple gene variants at different variant loci selected from the variant panel. (Item 27) The method according to any one of items 20 to 26, comprising generating or updating a report that includes (1) identifying information about the subject, and (2) a call for the presence or absence of the gene variant, or a call for the variant allele frequency for the gene variant. (Item 28) The method of item 27, which includes sending the report to the subject or the healthcare provider for the subject. (Item 29) The method described in item 27 or 28, wherein the report is transmitted via a computer network or peer-to-peer connection. (Item 30) The method according to any one of items 20 to 29, which includes determining the disease state of the subject. (Item 31) The aforementioned disease state was compared to the circulation of total cell-free DNA (cfDNA) in the test sample. The method described in item 16 or 30, which is a value proportional to the percentage of tumor DNA (ctDNA). (Item 32) The method according to item 16 or 30, wherein the disease state is the largest somatic allele fraction of cfDNA. (Item 33) The method according to item 16 or 30, wherein the disease state includes qualitative factors indicating cancer recurrence in the subject, the presence of cancer resistant to the treatment mode in the subject, or the presence of cancer that can be treated by a specific treatment mode. (Item 34) The method according to any one of items 1 to 33, wherein the reference match score and the variant match score are determined using a sequence alignment algorithm. (Item 35) The method according to item 34, wherein the sequence alignment algorithm is the Smith-Waterman alignment algorithm, the Striped Smith-Waterman alignment algorithm, or the Needleman-Wunsch alignment algorithm. (Item 36) The method according to any one of items 1 to 35, wherein the gene variant includes a single nucleotide variant (SNV), a multiple nucleotide variant (MNV), an indel, or a rearrangement conjugate. (Item 37) The method according to any one of items 1 to 36, wherein the variant panel is determined by sequencing nucleic acid molecules in a previous test sample obtained from the subject and calling one or more gene variants. (Item 38) The method according to item 37, wherein the subject receives an intervention for a disease between the previously acquired test sample and the acquired test sample. (Item 39) The method described in item 38, wherein the disease is cancer. (Item 40) The method according to item 38 or 39, further comprising adjusting the treatment based on the difference between the subject's disease state determined using the test sample and the subject's previous disease state based on the previous test sample. (Item 41) The method according to any one of items 9 to 40, comprising generating one or more sequencing reads by sequencing nucleic acid molecules in the test sample. (Item 42) The method according to any one of items 1 to 41, wherein the corresponding reference sequence and the corresponding mutant sequence include a mutant locus, a 5' flanking region, and a 3' flanking region. (Item 43) The method according to item 42, wherein the 5' flanking region and the 3' flanking region are each approximately 5 nucleotides long to approximately 5000 nucleotides long. (Item 44) The method according to any one of items 1 to 43, wherein the corresponding reference sequence and the corresponding variant sequence are identical except for the gene variant. (Item 45) The method described in any one of items 1 to 44, which includes generating a genomic profile of the subject using the detected gene variant or determined variant allele frequency. The method. (Item 46) The method according to item 45, wherein the subject's genomic profile includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. (Item 47) The method according to item 45 or 46, further comprising selecting an anticancer agent, administering an anticancer agent, or applying an anticancer treatment to the subject based on the generated genomic profile. (Item 48) The method according to any one of items 1 to 47, wherein the detection of the gene variant or the determined variant allele frequency is used to diagnose or confirm a disease in the subject. (Item 49) The method described in any one of items 1 to 48, wherein the subject has cancer, is at risk of having cancer, is routinely screened for cancer, or is suspected of having cancer. (Item 50) The method according to item 49, wherein the cancer is a solid tumor. (Item 51) The method described in item 49, wherein the cancer is a blood cancer. (Item 52) The method according to any one of items 1 to 51, further comprising selecting an anticancer therapy to be administered to the subject based on the detection of the gene variant or the determined variant allele frequency. (Item 53) The method according to item 52, further comprising administering the selected anticancer therapy to the subject. (Item 54) The method according to item 52 or 53, wherein the selected anticancer therapy includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. (Item 55) A method for diagnosing a disease, comprising diagnosing a subject having the disease based on the detection of a gene variant detected or a determined variant allele frequency in accordance with the method described in any one of items 1 to 54. (Item 56) A method for identifying patients as eligible for clinical trials of disease treatment based on the detection of gene variants or determined variant allele frequencies detected in accordance with the method described in any one of items 1 to 54. (Item 57) The method of item 56, further comprising enrolling the patient in the clinical trial. (Item 58) The method according to item 56 or 57, further comprising administering the disease treatment to the patient. (Item 59) A method for monitoring disease progression or recurrence, To generate a first sequencing read, nucleic acid molecules in a first test sample obtained from a subject with the disease are sequenced, To generate an individualized variant panel for the aforementioned subjects, To generate a second sequencing read, subjects were analyzed at a later time than the first test sample. The sequencing of nucleic acid molecules in the second test sample obtained from, A method comprising detecting the gene variant using the second sequencing read, or determining the variant allele frequency using the second sequencing read according to the method described in any one of items 1 to 54. (Item 60) The method according to item 59, comprising administering a disease therapy to the subject after the first test sample has been obtained from the subject and before the second test sample has been obtained from the subject. (Item 61) A first disease state is generated based on several first sequencing reads having variants input to the aforementioned variant panel, The method according to item 59 or 60, comprising generating a second disease state based on several second sequencing reads having variants from the aforementioned variant panel. (Item 62) The method according to item 61, further comprising determining disease progression by comparing the first disease state with the second disease state. (Item 63) After the first test sample is obtained from the subject, and before the second test sample is obtained from the subject, the disease therapy is administered to the subject. The method according to item 62, comprising adjusting the disease therapy based on the determined disease progression. (Item 64) The method according to item 63, wherein adjusting the disease therapy includes adjusting the dosage of the disease therapy or selecting a different disease therapy in response to the progression of the disease. (Item 65) The method according to item 63 or 64, further comprising administering the adjusted disease therapy to the subject. (Item 66) The method according to any one of items 59 to 65, wherein the first sample is obtained from the subject before the subject is administered the disease therapy, and the second sample is obtained from the subject after the subject has been administered the disease therapy. (Item 67) The method according to any one of items 60 to 66, wherein the disease therapy includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. (Item 68) A method for treating a subject with a disease, Obtaining a first test sample from the aforementioned subject, To generate a first sequencing read, the nucleic acid molecules in the first test sample are sequenced, Determining a first disease state using the first sequencing reads, To generate an individualized variant panel for the aforementioned subjects, Administering disease therapy to the aforementioned subjects, Obtaining a second test sample from the subject after the disease therapy has been administered to the subject, To generate a second sequencing read, the nucleic acid molecules in the aforementioned second test sample are sequenced, Using the second sequencing read described above, detect the gene variant, or use the second sequencing read described above to determine the variant allele frequency according to the method described in any one of items 1 to 54. Determining a second disease state based on the second sequencing reads, The progression of the disease is determined by comparing the first disease state and the second disease state, Adjusting the disease therapy administered to the subject based on the progression of the disease, A method comprising administering the aforementioned modified disease therapy to the subject. (Item 69) A method for selecting an anticancer therapy, comprising selecting an anticancer therapy for a subject in response to detecting a gene variant or determining the frequency of a mutant allele in a test sample derived from the subject, wherein the gene variant is detected or the frequency of the mutant allele is determined in accordance with the method described in any one of items 1 to 54. (Item 70) A method for treating cancer in a subject, comprising administering an effective dose of anticancer therapy to the subject in response to the detection of a gene variant or the determination of the frequency of a variant allele in a test sample from the subject, wherein the gene variant is detected or the frequency of the variant allele is determined in accordance with the method described in any one of items 1 to 54. (Item 71) The method according to any one of items 1 to 54, wherein the detection of the gene variant or the determination of the allele frequency in the test sample is used in making or suggesting a treatment decision for the subject. (Item 72) The method according to any one of items 1 to 54, wherein the detection of the gene variant or the determination of the allele frequency in the test sample used when administering treatment to the subject is performed. (Item 73) The method described in any one of items 16, 30-33, or 48-72, wherein the disease is cancer. (Item 74) The method according to any one of items 1 to 73, wherein the test sample is derived from a liquid biopsy sample from the subject. (Item 75) The method according to item 74, wherein the liquid biopsy sample includes blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. (Item 76) The method according to item 74 or 75, wherein the liquid biopsy sample contains circulating tumor cells (CTCs). (Item 77) The method according to any one of items 1 to 76, wherein the test sample contains cfDNA. (Item 78) The method according to any one of items 1-30 and 33-77, wherein the test sample is derived from a solid tissue biopsy sample from the subject. (Item 79) The method according to any one of items 1 to 78, wherein the aforementioned mutant is a somatic mutation. (Item 80) The method according to any one of items 1 to 78, wherein the mutant is a germline mutation. (Item 81) The method according to any one of items 1 to 80, wherein the subject is suspected of having cancer or is determined to have cancer. (Item 82) The method according to any one of items 1 to 81, further comprising obtaining the test sample from the subject. (Item 83) The method according to any one of Items 1 to 82, wherein the test sample contains a mixture of a tumor nucleic acid molecule and a non-tumor nucleic acid molecule. (Item 84) The method according to Item 83, wherein the tumor nucleic acid molecule is derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecule is derived from a normal portion of the heterogeneous tissue biopsy sample. (Item 85) The method according to Item 84, wherein the sample contains a liquid biopsy sample, the tumor nucleic acid molecule is derived from a circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecule is derived from a non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample. (Item 86) The cancer is B-cell cancer, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendiceal cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, cancer of the blood tissue, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myelogenous leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-H-H Hodg (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteosarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchioloalveolar carcinoma, renal cell carcinoma, liver cancer, cholangiocarcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineocytoma, glioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, myelofibrosis with myeloid metaplasia, eosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine cancer, or carcinoid tumor, and is the method according to any one of items 33, 39, 49 to 51, 73, and 81. (Item 87) An electronic device, One or more processors, A memory, One or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs being Selecting a gene variant at a variant locus from a variant panel, Obtain one or more sequencing reads related to the test sample that overlaps with the aforementioned mutant gene locus, The process involves aligning each sequencing read to a corresponding reference sequence to generate a reference match score for each of the one or more sequencing reads, wherein the corresponding reference sequence does not contain the gene variant. The process involves aligning each sequencing read to a corresponding mutant sequence to generate a mutant match score for each of the one or more sequencing reads, wherein the corresponding mutant sequence includes the gene mutant. The system comprises one or more programs, each of which includes instructions for labeling each of the one or more sequencing reads as having the gene variant, not having the gene variant, or being a null read, based on the reference match score and the variant match score. The reference match score and the variant match score are such that the sequencing read corresponds to the If the sequence determination read shows a closer match to the corresponding mutant sequence than to the reference sequence, the sequence determination read is labeled as having the gene mutant. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, the sequencing read is labeled as not having a gene variant. An electronic device in which, if the reference match score and the variant match score are equal, the sequencing read is labeled as a null read. (Item 88) The electronic device according to item 87, wherein the one or more programs further include instructions for calling for the presence of the gene variant in the test sample based on the one or more labeled sequencing reads. (Item 89) The electronic device according to item 87, wherein the one or more programs further include instructions for calling for the presence or absence of the gene variant in the test sample based on the one or more labeled sequencing reads. (Item 90) The electronic device according to any one of items 87 to 89, wherein the one or more programs further include instructions for generating the corresponding reference sequence or the corresponding variant sequence. (Item 91) The electronic device according to any one of items 87 to 90, wherein the one or more sequencing reads include a plurality of sequencing reads that overlap with the variant locus, and the one or more programs further include instructions for determining some sequencing reads from the plurality of sequencing reads having the gene variant or some sequencing reads from the plurality of sequencing reads that do not have the gene variant. (Item 92) The electronic device according to item 91, wherein the one or more programs further include instructions for determining the mutant allele frequency for the gene variant using the number of sequencing reads having the gene variant and the number of sequencing reads not having the gene variant. (Item 93) The electronic device according to any one of items 87 to 92, wherein the one or more programs further include instructions for labeling one or more sequencing reads related to the test sample for a plurality of gene variants at different variant loci selected from the variant panel. (Item 94) The electronic device according to any one of items 87 to 93, wherein the one or more programs further include instructions for generating or updating a report that includes (1) identification information about the subject, and (2) a call for the presence or absence of the gene variant, or a call for the variant allele frequency for the gene variant. (Item 95) The electronic device according to item 94, wherein the one or more programs further include instructions for transmitting the report to the subject or a healthcare provider for the subject. (Item 96) The electronic device described in item 94 or 95, which transmits the aforementioned report via a computer network or peer-to-peer connection. (Item 97) The electronic device according to any one of items 87 to 96, wherein the one or more programs further include instructions for determining the disease state of the subject. (Item 98) The electronic device according to item 97, wherein the disease state is a value proportional to the percentage of circulating tumor DNA (ctDNA) compared to total cell-free DNA (cfDNA) in the test sample. (Item 99) The electronic device described in item 97, wherein the disease state is the largest somatic allele fraction of cfDNA. (Item 100) The electronic device according to item 97, wherein the disease state includes qualitative factors indicating cancer recurrence in the subject, the presence of cancer resistant to the treatment mode in the subject, or the presence of cancer that can be treated by a specific treatment mode. (Item 101) An electronic device according to any one of items 87 to 100, wherein the reference match score and the variant match score are determined using a sequence alignment algorithm. (Item 102) The electronic device according to item 101, wherein the array alignment algorithm is the Smith-Waterman alignment algorithm, the Striped Smith-Waterman alignment algorithm, or the Needleman-Wunsch alignment algorithm. (Item 103) The electronic device according to item 102, wherein the gene variant includes a single nucleotide variant (SNV), a multiple nucleotide variant (MNV), an indel, or a rearrangement conjugate. (Item 104) The electronic device according to any one of items 87 to 103, wherein the variant panel is determined by sequencing nucleic acid molecules in a previous test sample obtained from the subject, and the one or more programs further include instructions for invoking one or more gene variants. (Item 105) The electronic device according to item 104, wherein the subject has received an intervention for a disease between the previously acquired test sample and the acquired test sample. (Item 106) The electronic device described in item 105, wherein the disease is cancer. (Item 107) The electronic device according to any one of items 87 to 106, wherein the one or more programs further include instructions for operating a sequencer to generate the one or more sequencing reads by sequencing nucleic acid molecules in the test sample. (Item 108) The electronic device according to any one of items 87 to 107, wherein the corresponding reference sequence and the corresponding variant sequence include the variant locus, the 5' flanking region, and the 3' flanking region. (Item 109) The electronic device according to item 108, wherein the 5' flanking region and the 3' flanking region are each approximately 5 nucleotides long to approximately 5000 nucleotides long. (Item 110) An electronic device according to any one of items 87 to 109, wherein the corresponding reference sequence and the corresponding variant sequence are identical except for the gene variant. (Item 111) The electronic device according to any one of items 87 to 106, wherein the one or more programs further include instructions for generating a genomic profile of the subject using the detected gene variant or determined variant allele frequencies. (Item 112) The electronic device according to item 111, wherein the genomic profile of the subject includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hot spot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. (Item 113) The electronic device according to item 111 or 112, wherein the one or more programs further include instructions for selecting an anticancer agent based on the generated genomic profile. (Item 114) The electronic device according to any one of items 87 to 113, wherein the subject has cancer, has a risk of having cancer, is routinely examined for cancer, or is suspected of having cancer. (Item 115) The electronic device according to item 114, wherein the cancer is a solid tumor. (Item 116) The electronic device according to item 114, wherein the cancer is a hematological cancer. (Item 117) The electronic device according to any one of items 87 to 116, wherein the one or more programs further include instructions for selecting an anticancer therapy to administer to the subject based on the detection of the gene variant or the determined variant allele frequency. (Item 118) The electronic device according to item 117, wherein the selected anticancer therapy includes chemotherapy, radiation therapy, immunotherapy, targeted therapy, or surgery. (Item 119) A non - transient computer - readable storage medium storing one or more programs, which when executed by one or more processors of an electronic device, cause the electronic device to select gene variants at variant loci from a variant panel, obtain one or more sequencing reads related to test samples overlapping the variant loci, The process involves aligning each sequencing read to a corresponding reference sequence to generate a reference match score for each of the one or more sequencing reads, wherein the corresponding reference sequence does not contain the gene variant. The process involves aligning each sequencing read to a corresponding mutant sequence to generate a mutant match score for each of the one or more sequencing reads, wherein the corresponding mutant sequence includes the gene mutant. The instruction includes labeling each of the one or more sequencing reads as having the gene variant, not having the gene variant, or being a null read, based on the reference match score and the variant match score. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding variant sequence more closely than the corresponding reference sequence, the sequencing read is labeled as having the gene variant. If the reference match score and the variant match score indicate that the sequencing read matches the corresponding reference sequence more closely than the corresponding variant sequence, the sequencing read is labeled as not having a gene variant. A non-temporary computer-readable storage medium in which, if the reference match score and the variant match score are equal, the sequencing read is labeled as a null read. (Item 120) The one or more programs are executed by the one or more processors. The non-temporary computer-readable storage medium according to item 119, further comprising an instruction to the electronic device to recall the presence of the gene variant in the test sample based on one or more labeled sequencing reads. (Item 121) The non-temporary computer-readable storage medium according to item 119, further comprising, when the one or more programs are executed by the one or more processors, instructions causing the electronic device to retrieve whether or not the gene variant is present in the test sample based on the labeled one or more sequencing reads. (Item 122) A non-temporary computer-readable storage medium according to any one of items 119 to 121, wherein the one or more programs, when executed by the one or more processors, further include instructions causing the electronic device to generate the corresponding reference array or the corresponding variant array. (Item 123) A non-temporary computer-readable storage medium according to items 119-122, wherein the one or more sequencing reads include a plurality of sequencing reads that overlap with the variant gene locus, and the one or more programs, when executed by the one or more processors, further include instructions causing the electronic device to determine the number of sequencing reads from the plurality of sequencing reads having the gene variant or the number of sequencing reads from the plurality of sequencing reads not having the gene variant. (Item 124) A non-temporary computer-readable storage medium according to item 123, further comprising instructions to cause the electronic device to determine the mutant allele frequency for the gene variant using the number of sequencing reads having the gene variant and the number of sequencing reads not having the gene variant, when the one or more programs are executed by the one or more processors. (Item 125) A non-temporary computer-readable storage medium according to any one of items 119 to 124, further comprising instructions to cause the electronic device to label one or more sequencing reads related to the test sample for a plurality of gene variants at different variant loci selected from the variant panel, when the one or more programs are executed by the one or more processors. (Item 126) A non-temporary computer-readable storage medium according to any one of items 119 to 125, further comprising instructions causing the electronic device to generate or update a report including (1) identification information about the subject, and (2) a call for the presence or absence of the gene variant, or a call for the variant allele frequency for the gene variant. (Item 127) A non-temporary computer-readable storage medium according to item 126, wherein the one or more programs, when executed by the one or more processors, further include instructions causing the electronic device to transmit the report to the patient or a healthcare provider for the patient. (Item 128) The report is transmitted via a computer network or peer-to-peer connection on a non-temporary computer-readable storage medium as described in item 126 or 127. (Item 129) A non-temporary computer-readable storage medium according to any one of items 119 to 128, wherein the one or more programs, when executed by the one or more processors, further include instructions causing the electronic device to determine the disease state of the subject. (Item 130) A non-temporary computer-readable storage medium as described in item 129, wherein the disease state is proportional to the percentage of circulating tumor DNA (ctDNA) compared to total cell-free DNA (cfDNA) in the test sample. (Item 131) The non-temporary computer-readable storage medium described in item 129, wherein the disease state is the largest somatic allele fraction of cfDNA. (Item 132) A non-temporary computer-readable storage medium according to item 129, wherein the disease state includes qualitative factors indicating cancer recurrence in the subject, the presence of cancer resistant to the treatment method in the subject, or the presence of cancer that can be treated by a specific treatment method. (Item 133) A non-temporary computer-readable storage medium according to any one of items 119 to 132, wherein the reference match score and the variant match score are determined using a sequence alignment algorithm. (Item 134) The non-temporary computer-readable storage medium described in item 133, wherein the sequence alignment algorithm is the Smith-Waterman alignment algorithm, the Striped Smith-Waterman alignment algorithm, or the Needleman-Wunsch alignment algorithm. (Item 135) The non-temporary computer-readable storage medium described in item 134, wherein the gene variant includes a single nucleotide variant (SNV), a multiple nucleotide variant (MNV), an indel, or a rearrangement conjugate. (Item 136) A non-temporary computer-readable storage medium according to any one of items 119 to 135, wherein the variant panel is determined by sequencing nucleic acid molecules in a previous test sample obtained from the subject, and the one or more programs, when executed by the one or more processors, further include instructions causing the electronic device to invoke one or more gene variants. (Item 137) A non-temporary computer-readable storage medium as described in item 136, in which the subject has received an intervention for a disease between the previously acquired test sample and the acquired test sample. (Item 138) A non-temporary computer-readable storage medium as described in item 137, wherein the disease is cancer. (Item 139) A non-temporary computer-readable storage medium according to any one of items 119 to 138, wherein when the one or more programs are executed by the one or more processors, the electronic device further includes instructions to operate a sequencer to generate the one or more sequencing reads by sequencing nucleic acid molecules in the test sample. (Item 140) A non-temporary computer-readable storage medium according to any one of items 119 to 139, wherein the corresponding reference sequence and the corresponding variant sequence include the variant locus, the 5' flanking region, and the 3' flanking region. (Item 141) The non-temporary computer-readable storage medium described in item 140, wherein the 5' flanking region and the 3' flanking region are each approximately 5 nucleotides long to approximately 5000 nucleotides long. (Item 142) A non-temporary computer-readable storage medium according to any one of items 119 to 141, wherein the corresponding reference sequence and the corresponding variant sequence are identical except for the gene variant. (Item 143) A non-temporary computer-readable storage medium according to any one of items 119 to 142, further comprising, when the one or more programs are executed by the one or more processors, instructions to cause the electronic device to generate a genomic profile of the subject using the detected gene variants or determined variant allele frequencies. (Item 144) A non-temporary computer-readable storage medium as described in item 143, wherein the subject's genomic profile includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. (Item 145) A non-temporary computer-readable storage medium according to item 143 or 144, further comprising, when the one or more programs are executed by the one or more processors, instructions to cause the electronic device to select an anticancer drug based on the generated genome profile. (Item 146) A non-temporary computer-readable storage medium as described in any one of items 119-145, wherein the subject has cancer, is at risk of having cancer, is routinely screened for cancer, or is suspected of having cancer. (Item 147) A non-temporary computer-readable storage medium as described in item 146, wherein the cancer is a solid tumor. (Item 148) A non-temporary computer-readable storage medium as described in item 146, wherein the cancer is a blood cancer. (Item 149) A non-temporary computer-readable storage medium according to any one of items 119 to 148, further comprising instructions to the electronic device to select an anti-cancer therapy to administer to the subject based on the detection of the gene variant or the determined variant allele frequency, when the one or more programs are executed by the one or more processors. (Item 150) A non-temporary computer-readable storage medium as described in item 145, wherein the selected anti-cancer therapy includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery.

Claims

[Claim 1] The invention described in the specification.