Methods and systems for detecting residual disease

Through individualized nucleic acid sequencing data analysis, utilizing the genomic differences between diseased and healthy tissues, combined with false positive error rate and sampling variance, the problem of difficult measurement of the proportion of nucleic acid molecules in diseased tissues was solved, and accurate monitoring and early detection of disease levels were achieved.

CN114127308BActive Publication Date: 2025-09-23ULTIMA GENOMICS INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202080051437.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-07
Filing Date
2020-05-15
Publication Date
2025-09-23
Estimated Expiration
2040-05-15

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately measure the proportion of nucleic acid molecules from diseased tissues in individual samples, especially when the ratio of diseased tissue to healthy tissue is extremely asymmetric, making it difficult to detect and monitor the severity of the disease.

Method used

By using personalized nucleic acid sequencing data, the signal of small nucleotide variant loci associated with diseased tissue is compared with the background factor of sequencing false positive error rate, combined with sampling variance, to determine the level of disease, recurrence, progression or regression.

Benefits of technology

It improves the accuracy and efficiency of disease detection, reduces the sequencing depth requirement, reduces time and cost, and enables early detection of disease recurrence or regression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

Described herein are methods, devices, and systems for measuring disease (e.g., cancer) levels, such as fractions of nucleic acid molecules (e.g., cell-free DNA) in samples associated with diseased tissue (e.g., cancer tissue) from individuals. Methods, devices, and systems for measuring the presence, recurrence, progression, or regression of a disease in an individual are also described. Some methods include using nucleic acid sequencing data associated with an individual, where an indication is selected from a combination of individualized disease-related small nucleotide variations (SNV) loci, and a signal indicating a ratio of the sequenced loci derived from the diseased tissue is compared to a background factor indicating a false positive error rate for sequencing across the selected loci, or a noise factor indicating a sampling variance across the selected loci.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 849,414, filed May 17, 2019, and U.S. Provisional Patent Application Serial No. 62 / 971,530, filed February 7, 2020, the contents of which are hereby incorporated by reference in their entirety.

[0003] Submission of Sequence Listings as ASCII Text Files

[0004] The contents of the following ASCII text file submission are hereby incorporated by reference in their entirety: Computer Readable Form (CRF) of the Sequence Listing (file name: 165272000140SEQLIST.TXT, recording date: May 14, 2020, size: 1 KB). Technical Field

[0005] Described herein are methods, systems, and apparatus for measuring the fraction of disease (e.g., cancer)-associated nucleic acid molecules in a sample using nucleic acid sequencing data. Also described are methods, systems, and apparatus for measuring the level, presence, recurrence, progression, or regression of a disease (e.g., cancer). Background Art

[0006] The detection and quantification of residual disease before, during, and after cancer treatment can be used to monitor the effectiveness of cancer treatment or cancer remission in patients. Targeted nucleic acid sequencing has previously been used to determine differences (i.e., variants) between disease-free and cancerous tissue. Targeted sequencing methods typically look for mutations in known driver genes or known mutation hotspots in the cancer genome or exome, or use deep sequencing methods to ensure accurate variant calls at specific targeted loci.

[0007] The amount of cell-free DNA ("cfDNA") derived from tumors (also called "circulating tumor DNA" or "ctDNA") in an individual can correlate with the severity of the disease. Except for the most rapidly progressive disease states, only a small fraction of the DNA in a sample originates from diseased tissue, with the vast majority of the DNA coming from non-diseased tissue in the individual. This makes accurately measuring the amount of cfDNA from diseased tissue particularly challenging. Current methods typically involve very high-sensitivity protocols, such as custom qPCR or custom enrichment targeting a relatively small number of cancer-specific variants. SUMMARY OF THE INVENTION

[0009] Described herein are methods, systems, and devices for measuring the level of a disease (eg, cancer) in an individual, as well as methods for measuring the presence, recurrence, progression, or regression of a disease in an individual.

[0010] In some embodiments, a method of measuring the level of a disease in an individual comprises: using nucleic acid sequencing data associated with the individual, comparing a signal indicating the ratio at which sequenced loci selected from a personalized disease-associated small nucleotide variation (SNV) locus panel are derived from diseased tissue with a background factor indicating the sequencing false-positive error rate across the selected loci; and determining the level of the disease in the individual based on the comparison of the signal to the background factor.

[0011] In some embodiments, a method of measuring disease recurrence in an individual comprises: using nucleic acid sequencing data associated with the individual, comparing a signal indicating the ratio at which sequenced loci selected from a personalized combination of disease-associated small nucleotide variation (SNV) loci are derived from diseased tissue to a background factor indicating the sequencing false-positive error rate across the selected loci; and determining the level of disease in the individual based on the comparison of the signal to the background factor.

[0012] In some embodiments, a method for measuring disease progression or regression in an individual comprises: using nucleic acid sequencing data associated with the individual, comparing a signal indicating the ratio of sequenced loci selected from a personalized disease-associated small nucleotide variation (SNV) locus combination derived from diseased tissue with a background factor indicating the sequencing false positive error rate across the selected loci; and determining the level of disease in the individual based on the comparison of the signal with the background factor; and comparing the measured disease level with a previously measured disease level in the individual. In some embodiments, the progression or regression of the disease is based on a statistically significant change in the measured disease level.

[0013] In some embodiments of any of the above methods, the disease level is the fraction of nucleic acid molecules associated with the disease in a sample from the individual.

[0014] In some embodiments of any of the above methods, comparing comprises subtracting a background factor from the signal.

[0015] In some embodiments of any of the above methods, the method further comprises determining an error in the measurement of the disease level. In some embodiments, the error is a confidence interval for the disease level. In some embodiments, the error is proportional to the total number of individual small nucleotide variant reads detected at the selected locus. In some embodiments, the disease level is the fraction of nucleic acid molecules associated with the disease in a sample from an individual, and wherein the fraction and the error are defined as:

[0016]

[0017] Where: F is the score; N total is the total number of individual small nucleotide variant reads detected at the selected locus; N varis the number of selected loci; D is the average sequencing depth.

[0018] In some embodiments, the method for detecting a disease in an individual includes: using nucleic acid sequencing data associated with the individual, comparing a signal indicating the ratio of the sequenced loci derived from the diseased tissue selected from a combination of individualized disease-associated small nucleotide variations (SNV) loci with a noise factor indicating the sampling variance across the selected loci; and determining whether the individual has a disease based on a comparison of the signal with the background factor. In some embodiments, if the signal exceeds the noise factor by more than a predetermined threshold, the individual is determined to have a recurrence of the disease or a residual level of the disease. In some embodiments, if the signal exceeds the noise factor by k times or more, the individual is determined to have a recurrence of the disease or a residual level of the disease, wherein k is about 1.5. In some embodiments, k is about 3.0. In some embodiments, k is about 5.0. In some embodiments, k is about 10. In some embodiments, the method includes detecting the recurrence of the disease.

[0019] In some embodiments, a method of detecting recurrence, progression, or regression of a disease in an individual comprises measuring at least one of: (a) a likelihood that a value indicating a fraction F of nucleic acid molecules in a sample derived from diseased tissue of the individual is greater than zero, wherein F greater than zero indicates the presence of disease in the individual, and (b) a statistically significant change in a value indicating a fraction F of nucleic acid molecules in a sample derived from diseased tissue of the individual, wherein the statistically significant change is relative to a previously measured fraction F. prior , and wherein a statistically significant change in F indicates progression or regression of the disease in the individual; wherein the fraction F is the total number N of small nucleotide variations (SNVs) detected in the cell-free nucleic acid sequencing data total (wherein SNVs are selected from individualized disease-associated SNV loci combinations) and the number of SNVs selected from the SNV combination N var The dfs were determined by comparison and adjusted by the average sequencing depth D and further adjusted by the sequencing false-positive error rate E across the selected SNVs.

[0020] In some embodiments of the above method, the method further comprises generating a personalized disease-associated SNV locus combination. In some embodiments, generating a personalized disease-associated SNV locus combination comprises: sequencing nucleic acid molecules derived from a diseased tissue sample to determine a disease-associated SNV set; and filtering the disease-associated SNV set to remove germline variation and non-cancer-associated somatic variation. In some embodiments, the sample of the diseased tissue is a tumor biopsy sample obtained from an individual. In some embodiments, germline variation or somatic variation or both are determined by sequencing nucleic acid molecules derived from a non-diseased tissue sample obtained from an individual. In some embodiments, the sample of the non-diseased tissue comprises white blood cells. In some embodiments, the sample of the non-diseased tissue is a buffy coat. In some embodiments, the method further comprises filtering the disease-associated SNV set to remove SNVs supported by only one sequencing read. In some embodiments, the method further comprises filtering the disease-associated SNV set to remove SNVs that are not supported by complementary sequencing reads. In some embodiments, the method further comprises filtering the disease-associated SNV set to remove SNVs present in a general population of individuals whose allele frequency is greater than a predetermined threshold. In some embodiments, the predetermined threshold is about 0.01. In some embodiments, the method further comprises filtering SNVs within low complexity genomic regions (i.e., homopolymer regions or short tandem repeats (STRs)). In some embodiments, the nucleic acid sequencing data is obtained by sequencing nucleic acid molecules from a fluid sample obtained from an individual using non-terminating nucleotides provided in separate nucleotide flows according to a flow cycle order comprising a plurality of flow positions, wherein the flow positions correspond to nucleotide flows; and generating the personalized disease-associated SNV locus combination further comprises filtering the disease-associated SNV set to include only those SNVs that, when the nucleic acid sequencing data and the reference sequencing data are sequenced using non-terminating nucleotides provided in separate nucleotide flows according to the flow cycle order, result in nucleic acid sequencing data that is different from the reference sequencing data associated with the reference sequence at two or more flow positions.

[0021] In some embodiments of the above method, the nucleic acid sequencing data is obtained by sequencing nucleic acid molecules from a fluid sample obtained from an individual using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order comprising a plurality of flow positions, wherein the flow positions correspond to the nucleotide streams; and the method further comprises generating a personalized disease-associated SNV locus combination, which comprises sequencing nucleic acid molecules derived from a diseased tissue sample to determine a set of disease-associated SNVs; and generating the personalized disease-associated SNV locus combination further comprises filtering the set of disease-associated SNVs to include only those SNVs that, when the nucleic acid sequencing data and the reference sequencing data are sequenced using non-termination nucleotides provided in separate nucleotide streams according to the flow cycle order, result in nucleic acid sequencing data that is different from the reference sequencing data associated with the reference sequence at two or more flow positions.

[0022] In some embodiments of any of the above methods, the nucleic acid molecule is a cell-free nucleic acid molecule. In some embodiments, the nucleic acid molecule is a DNA molecule. In some embodiments, the nucleic acid molecule is an RNA molecule.

[0023] In some embodiments of any of the above methods, the nucleic acid sequencing data is derived from nucleic acid molecules in a fluid sample obtained from an individual. In some embodiments, the fluid sample is a blood sample, a plasma sample, a saliva sample, a urine sample, or a feces sample.

[0024] In some embodiments of any of the above methods, the disease is cancer. In some embodiments, the cancer is metastatic cancer.

[0025] In some embodiments of any of the above methods, the method further comprises sequencing the nucleic acid molecule to obtain sequencing data.

[0026] In some embodiments of any of the above methods, the nucleic acid sequencing data is obtained by sequencing the nucleic acid molecule according to a predetermined sequence of nucleotide sequencing cycles. In some embodiments, the nucleic acid sequencing data is further obtained by resequencing the nucleic acid molecule according to different predetermined sequence of nucleotide sequencing cycles, wherein the different predetermined sequence of nucleotide sequencing cycles results in a different false positive variant rate at the subset of the sequenced loci compared to the first predetermined sequence of nucleotide sequencing cycles.

[0027] In some embodiments of any of the above methods, the sequencing data is non-targeted sequencing data. In some embodiments, the sequencing data is obtained from a non-targeted whole genome.

[0028] In some embodiments of any of the above methods, the average sequencing depth of the sequencing data is at least 0.01. In some embodiments, the average sequencing depth of the sequencing data is less than about 100. In some embodiments, the average sequencing depth of the sequencing data is less than about 10. In some embodiments, the average sequencing depth of the sequencing data is less than about 1.

[0029] In some embodiments of any of the above methods, the combination of disease-associated SNV loci comprises a passenger mutation and / or a driver mutation.

[0030] In some embodiments of any of the above methods, the combination of disease-associated SNV loci comprises a small nucleotide polymorphism (SNP) locus. In some embodiments of the method, the combination of disease-associated SNV loci comprises an indel mutation locus.

[0031] In some embodiments of any of the above methods, the loci selected from the combination of disease-associated SNV loci comprise about 300 or more loci.

[0032] In some embodiments of any of the above methods, the loci selected from the combination of disease-associated SNVs are selected based on false positive rates for individual loci.

[0033] In some embodiments of any of the above methods, loci are selected from the combination of disease-associated SNVs based on unique SNVs associated with a selected subclone of the disease.

[0034] In some embodiments of any of the above methods, the disease-associated SNV combination is determined by comparing sequencing data associated with the diseased tissue to sequencing data associated with the non-diseased tissue. In some embodiments, the method further comprises sequencing nucleic acid molecules derived from the diseased tissue to obtain sequencing data associated with the diseased tissue. In some embodiments, the method further comprises sequencing nucleic acid molecules derived from non-diseased tissue to obtain sequencing data associated with the non-diseased tissue.

[0035] In some embodiments of any of the above methods, the nucleic acid sequencing data is obtained using surface-based nucleic acid molecule sequencing, and wherein the nucleic acid molecules are not amplified prior to attaching the nucleic acid molecules to the surface.

[0036] In some embodiments of any of the above methods, the nucleic acid sequencing data is obtained without the use of unique molecular identifiers (UMIs).

[0037] In some embodiments of any of the above methods, the nucleic acid sequencing data is obtained without the use of a sample identification barcode.

[0038] In some embodiments of any of the above methods, the sequencing false positive error rate is measured using a control locus combination.

[0039] In some embodiments of any of the above methods, sequencing data is obtained by sequencing the nucleic acid molecules in the combined sample obtained from a plurality of individuals. In some embodiments, selected locus is unique for each individual in a plurality of individuals. In some embodiments, at least one locus in the selected locus is common between at least two individualities in a plurality of individuals. In some embodiments, the sequencing depth of each individual is determined, and the signal of this individual is adjusted based on the sequencing depth relevant to each individual. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Exemplary methods for measuring the fraction of nucleic acid molecules associated with a disease in a sample from an individual are shown.

[0042] Figure 2 Another exemplary method of measuring the fraction of nucleic acid molecules associated with a disease in a sample from an individual is shown.

[0043] Figure 3 Exemplary methods of measuring disease levels in an individual are shown.

[0044] Figure 4 Exemplary methods of measuring disease levels in an individual are shown.

[0045] Figure 5 Exemplary methods of monitoring disease recurrence, progression, or regression in an individual are shown.

[0046] Figure 6 Another exemplary method of monitoring disease recurrence, progression, or regression in an individual is shown.

[0047] Figure 7 An example of a computing device according to one embodiment is shown, which may be used to perform the methods described herein.

[0048] Figure 8A Shown are sequencing data obtained by extending a primer with the sequence TATGGTCGTCGA (SEQ ID NO: 1) using a repeated flow cycle sequence of TACG. The sequencing data represent the extended primer strand, and it can be readily determined that the sequencing information of the complementary template strand is effectively identical.

[0049] Figure 8B Shows Figure 8ASequencing data shown. Given the sequencing data, the most likely sequence was selected based on the highest likelihood for each flow position (as indicated by an asterisk).

[0050] Figure 8C Shows Figure 8A Sequencing data shown, where the traces represent two different candidate sequences: TATGGTCATCGA (SEQ ID NO: 2) (filled circles) and TATGGTCGTCGA (SEQ ID NO: 1) (open circles). The likelihood of the sequencing data matching a given sequence can be determined by multiplying the likelihood of each flow position matching the candidate sequence. In some embodiments, the first candidate sequence (SEQ ID NO: 2) can also be considered the reverse complement of an exemplary reference sequence, and the second candidate sequence (SEQ ID NO: 1) can be considered the sequence comprising the SNV.

[0051] Figure 8D Shown is sequencing data for a nucleic acid molecule containing a SNV (SEQ ID NO: 1) using AGCT sequencing cycles and compared to a reference sequence (SEQ ID NO: 2). Detailed Description of the Invention

[0053] Methods, equipment and systems described herein relate to detection and / or measurement of disease level in individual.Disease level can be relevant to the score of the nucleic acid molecules (such as cell-free DNA) in the sample derived from diseased tissue (such as cancerous tissue).For example, the signal of the ratio of small nucleotide variation (SNV) reading section detected in the selected locus derived from diseased tissue can be detected by measuring indication, and the background factor of the false positive error rate of order-checking or the noise factor of the sampling variance of indication across locus can be compared, to detect disease or measurement level. The nucleic acid molecules score relevant to diseased tissue detected in the sample can inform individual disease level.By detecting individual disease level, the recurrence of pre-existing disease (or previously thought to be disappearing disease) can be determined, and the progress of disease state or disappearance.

[0054] Compared with the normal healthy genome of an individual, some diseased tissues, especially cancers, can include thousands (or tens of thousands, hundreds of thousands or more) mutations in the entire diseased genome. These mutations can be driver mutations, which give cancer growth advantages (for example, proliferation or survival), or can be passenger mutations, which can be found in the entire coding or non-coding regions of the genome, but are not considered to give any growth advantages. In some cases, passenger mutations accumulate in cells that are cancerous before becoming cancerous, because even healthy tissues have a certain mutation rate. The broad spectrum mutations of any particular disease in the patient are unique for the patient, even for specific diseased tissue clones or subclones, thereby giving the diseased tissue a unique genetic signature. By comparing the genome (or part thereof) of the diseased tissue with the genome (or corresponding genome) of the non-diseased tissue of the same patient, individualized disease-related small nucleotide variations (SNV) locus combinations can be established for the diseased tissue. Optionally, a subset of the loci from this combination can be selected for analysis, and selection can be based on, for example, the false positive error rate at a given locus, for example, lower than the false positive error rate at other loci. The SNV combination can include passenger mutations and / or driver mutations.

[0055] When measuring the diseased part or disease level of the nucleic acid molecule of the patient, by considering false positive error rate and / or sampling variance, the overall sequencing depth can be reduced, thereby significantly saving time and cost. In the sequencing process, due to chemical damage, wrong base incorporation or fluorescence reading error, false positive error can occur, and SNV can be mistakenly indicated to be present in a given locus. Sampling variance is related to the number of SNV reads detected, which includes false positive error and true positive judgment. In order to prevent potential errors at a specific locus, other disease detection methods generally require multiple independent SNV judgments at a given locus, which can only be obtained by sequencing the locus at a depth that is inversely proportional to the sick nucleic acid score in the sample. In some cases, other methods relate to the consensus sequence determined at the locus from multiple sequencing reads. The depth sequencing used by other methods generally requires a narrow subset (for example, mutation hotspots or full exon group sequencing) of targeted specific locus or genome. Additionally, other sequencing methods generally require amplified nucleic acid molecules in the library preparation process, so that multiple copies of the same nucleic acid molecule are independently sequenced. This amplification process has the risk of introducing additional false errors.

[0056] Described method does not relate to the false positive error at any specific locus place, but uses false positive error rate and / or the sampling variance across the locus that is selected for analyzing to measure mark or the disease level of sick nucleic acid molecule.In case locus has been selected, the false positive of any specific locus does not significantly affect measurement.Therefore, although the locus that is selected for analyzing can use the false positive error rate at each specific locus place to select, do not consider the influence of any specific mistake that may produce at given locus place order-checking.

[0057] definition

[0058] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0059] Reference herein to "about" a value or parameter includes (and describes) variations with respect to that value or parameter itself. For example, a description referring to "about X" includes a description of "X."

[0060] As used herein, the term "average" refers to the mean or median, or any value used to approximate the mean or median.

[0061] As used herein, "variance" or "variance" refers to any statistical measure that defines the width of a distribution, which may be, but is not limited to, standard deviation, variance, or interquartile range.

[0062] The terms "individual," "patient," and "subject" are used synonymously and refer to animals, including humans.

[0063] As used herein, the term "tissue" refers to any cellular material and may include circulating or non-circulating cells.

[0064] It should be understood that the aspects and variations of the present invention described herein include "consisting of" and / or "consisting essentially of" the aspects and variations.

[0065] When a range of values ​​is provided, it is understood that each intervening value between the upper and lower limits of that range, as well as any other stated or intervening values ​​in that stated range, is included within the scope of the present disclosure. Where the stated range includes an upper or lower limit, ranges excluding one of those included limits are also included in the present disclosure.

[0066] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. This description is presented to enable one of ordinary skill in the art to make and use the present invention, and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the embodiments shown, but rather to the widest scope consistent with the principles and features described herein.

[0067] Figure 1-8D The method of showing multiple examples. For example, one or more electronic devices that execute the software platform can be used to perform these exemplary methods. In some instances, one or more exemplary methods are performed using a client-server system, and the modules of the methods shown can be divided in any way between the server and the client device. In other instances, the components of the exemplary methods are divided between the server and multiple client devices. Therefore, although part of the exemplary method is described herein as being performed by a specific device of the client-server system, it should be understood that these methods are not limited thereto. In other instances, only a client device (e.g., a user device) or only one or more client devices are used to perform one or more exemplary methods. In the exemplary methods, some modules are optionally combined, the order of some modules is optionally changed, and some modules are optionally omitted. In some instances, additional steps can be performed in combination with the exemplary methods. Therefore, the operations shown (and described in more detail below) are exemplary in nature and should not be considered as limiting.

[0068] The disclosures of all publications, patents, and patent applications mentioned herein are each incorporated by reference in their entirety. In the event of a conflict between any reference incorporated by reference and the present disclosure, the present disclosure shall control.

[0069] Personalized locus combination

[0070] Some diseases in individual, such as cancer, can produce mutant nucleic acid sequence, provide feature for disease.The nucleic acid molecule sequence relevant to diseased tissue (that is, diseased genome) can be compared with the nucleic acid molecule sequence relevant to non-diseased tissue (that is, healthy or non-diseased genome) from the same individual.The difference between diseased genome (or its part) and non-diseased genome (or its part) determines the variation of diseased tissue.Some or all of small nucleotide variations (for example, small nucleotide polymorphism (SNP) or small insertion and deletion mutation (length is generally 1-5 base)) between genome (or genome part) can be used for setting up the individual disease-specific individualized disease-related SNV genome combination of disease.SNV locus combination can be computer simulation, for example, is not included in one group of oligonucleotide primers.Therefore, individualized disease-related SNV locus combination is based on the difference between the nucleotide sequence relevant to diseased tissue and the nucleotide sequence relevant to health (that is, non-diseased) tissue to build.In some embodiments, the sequencing data relevant to diseased tissue and / or healthy tissue is targeted sequencing data. In some embodiments, the sequencing data associated with diseased tissue and / or healthy tissue is non-targeted (eg, whole genome or whole genome) sequencing data.

[0071] In some embodiments, the SNV locus combination is generated by filtering germline variation and / or non-disease (e.g., non-cancer) related somatic variation from SNVs associated with diseased (e.g., cancer) tissue. For example, diseased tissue can be sequenced to determine multiple variations associated with diseased tissue. For example, the resulting sequencing reads can be compared with a reference genome, and variation can be selected based on the difference between the sequencing reads and the reference genome. The variation identified can include not only variations unique to diseased tissue, but also variations found in healthy tissue (e.g., variations found in white blood cells or other healthy tissues). For example, variations found in white blood cells can be obtained by sequencing matching buffy coat samples from the same subject and comparing sequencing data with a reference genome. Although these variations can include cancer variations, a large number of variations can be caused by age-related clonal hematopoiesis. In some embodiments, variations identified by buffy coat / leukocyte sequencing are considered to be approximate representative sets of non-cancer related somatic variations. Therefore, germline variation and / or non-disease related somatic variation (relative to a reference genome) can be determined by sequencing healthy tissue and comparing sequencing reads with a reference genome. Then, when generating disease-associated SNV loci panels, SNVs associated with diseased tissue can be filtered to remove germline variants and / or somatic variants.

[0072] In some embodiments, sequence data associated with diseased tissue and / or sequence data associated with healthy tissue are determined in advance (i.e., before sequencing and / or analyzing nucleic acid molecules in a fluid sample). For example, any healthy tissue obtained from an individual can be used to determine the sequence of a healthy genome (or portion thereof). Healthy tissue can be obtained, for example, from a fluid sample (e.g., from cell-free nucleic acid molecules (e.g., cfDNA) or healthy blood cells in a fluid sample), a cheek swab, a biopsy of healthy tissue, or any other suitable method. In some embodiments, healthy tissue includes white blood cells, such as white blood cells obtained from a buffy coat. In some embodiments, healthy tissue includes non-diseased tissue. For example, a tumor biopsy sample (e.g., a solid tumor biopsy sample, such as an FFPE tissue sample) may include healthy (i.e., non-diseased) tissue and diseased tissue. In some embodiments, healthy tissue includes a healthy cfDNA sample; for example, an individual can undergo a routine health checkup, including whole genome sequencing (WGS) analysis of a blood sample (e.g., plasma and / or a sample containing white blood cells). These data can be stored in an individual's health record. When an individual subsequently develops a disease such as cancer, the sequencing data previously obtained can be used to establish an individual's healthy baseline. On the contrary, for individuals with a known disease condition (e.g., liver cancer or breast cancer) and who have received treatment (e.g., surgical treatment), healthy tissue may include one or more samples collected immediately after treatment when the disease condition is no longer detected. This healthy tissue can be used as a baseline sample to compare with subsequent samples to assess whether the disease recurs in the individual. Nucleic acid sequencing libraries can be prepared from healthy tissue and sequenced to obtain sequencing data attributed to the genome (or part thereof) of healthy tissue. Although a small amount of disease tissue can be extracted together with healthy tissue, for obtaining sequencing data of healthy tissue, disease tissue is typically a negligible minor component.

[0073] The sequence data of nucleic acid molecules (e.g., genomes or portions thereof) associated with diseased tissues can be determined by obtaining tissue samples of diseased tissues, such as primary or secondary cancers that can be resected, biopsied, or otherwise sampled, and sequencing the nucleic acid molecules in the obtained tissues. In some embodiments, multiple samples are obtained from diseased tissues that can capture mosaics within the diseased tissues (e.g., different clones or subclones of diseased tissues). In some embodiments, sequence data associated with diseased tissues are obtained by sequencing nucleic acid molecules obtained from fluid samples (e.g., cell-free nucleic acid molecules (e.g., cfDNA) or healthy blood cells in fluid samples). Fluid samples may also include nucleic acid molecules associated with healthy tissues, but sequencing data associated with healthy tissues typically have substantially higher depth counts and can be ignored for the purpose of determining sequencing data associated with diseased tissues. For example, diseased tissues may be sampled before or after disease treatment (e.g., chemotherapy for cancer treatment) begins.

[0074] The individualized disease-related SNV locus combination includes the variation (including the locus of variation and mutation change) of the nucleic acid molecules from the diseased tissue relative to the nucleic acid molecules from non-diseased tissue. The combination can include less than all nucleic acid differences between healthy tissue and the diseased tissue, because some variation can not be detected due to the limitation of healthy tissue and / or diseased tissue sequencing data, or appear in the genomic regions that are technically difficult to order-checking, for example, low complexity regions or regions with mapping degeneracy. In some embodiments, the individualized combination includes driving mutations, passenger mutations or driving mutations and passenger mutations. In some embodiments, the locus combination includes the mutation in the coding region of the genome, the non-coding region of the genome or both. The variation number in the individualized combination depends on the diseased tissue, including the type of diseased tissue or the severity of the disease. In some embodiments, this individualized combination includes 2 or more, 5 or more, 10 or more, 25 or more, 50 or more, 100 or more, 200 or more, 300 or more, 500 or more, 1000 or more, 2500 or more, 5000 or more, 10000 or more, 25000 or more, 50000 or more, 100000 or more, 250000 or more, 500000 or more, 1000000 or more, 5000000 or more loci.In some embodiments, only when two or more (for example, 3 or more, 4 or more or 5 or more) redundant variation are carried out at any given locus place, variant loci are just included in the individualized locus combination.The screening site of redundant variation judgment limits the quantity of the false positive variant loci introducing this combination.In some cases, this combination only includes the variation that is verified as having difference between diseased tissue and non-diseased tissue by the common nucleic acid sequencing determined with high confidence.

[0075] For the methods described herein, it is not necessary to analyze all loci in the individualized disease-related SNV locus combination. In some embodiments, a portion of the loci in the individualized disease-related SNV locus combination is selected for analysis. Certain loci or variations can be more prone to false positive errors than other loci or variations. Additionaly, certain sequencing methods can be more prone to false positive errors than other methods. In some embodiments, loci are selected from the individualized locus combination based on the false positive error rate at the locus. For example, if the false positive error rate at the locus is about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, about 0.01% or less, about 0.005% or less, about 0.0025% or less, or about 0.0001% or less, then the locus can be selected. As only an example, a specific sequencing method can have a lower false positive error rate for sequencing, for detecting a specific mutation (such as G→A), being different from the mutation (such as G→C) of other mutation types, and a variation with a lower false positive error rate can be selected. In some embodiments, selected loci include 2 or more, 5 or more, 10 or more, 25 or more, 50 or more, 100 or more, 200 or more, 300 or more, 500 or more, 1000 or more, 25000 or more, 50000 or more, 100000 or more, 250000 or more, 50000 or more, 100000 or more, 250000 or more or 500000 or more loci. In some embodiments, all loci in the individualized locus combination are selected.

[0076] Filtering germline and non-disease-related somatic variations from SNVs associated with diseased tissues is a technique that can be used to select loci from disease-related SNV loci combinations (or to produce disease-related SNV loci combinations). cfDNA in blood can be derived from several cell sources, including cancer cells and non-cancerous cells. Hematopoietic stem cells may include clonal hematopoiesis-related somatic variations, which can lead to the expansion of clonal blood cell populations. These clonal hematopoiesis-related somatic variations are usually non-malignant, and the clonal expansion driven by these somatic variations can be referred to as clonal hematopoiesis of indeterminate potential (CHIP). See Steensma et al., Clonal hematopoiesis of indeterminate potential and its distinction from myelodysplastic syndromes, Blood, Vol. 126, pp. 9-16 (2015). Some studies have shown that at least 10% of elderly people over 70 years old carry CHIP due to oligoclonal expansion of mutant hematopoietic stem cells. See Jaiswal et al., Age-Related Clonal Hematopoiesis Associated with Adverse Outcomes, N.Engl.J.Med., Volume 371, Issue 26, Pages 2488-2498 (2014). Therefore, these non-disease-related somatic variations can be significantly expressed in cfDNA, even if they are not related to the disease. See also US2019 / 0385700 A1, US2019 / 0355438 A1, US2020 / 0013484 A1, the contents of which are hereby incorporated by reference for all purposes. Removing these non-disease-related somatic variations from the SNV locus combination can significantly reduce the background error rate. Non-disease-related somatic variations, such as clonal hematopoiesis-related somatic variations, can be identified by, for example, sequencing nucleic acid molecules derived from leukocytes (e.g., leukocytes in buffy coats).

[0077] In some embodiments, the SNV locus combination includes SNVs associated with diseased tissues, which have been filtered to remove germline and non-disease-related somatic variations (i.e., somatic variations unrelated to the disease). For example, these non-disease-related somatic variations can be determined by sequencing nucleic acid molecules derived from healthy tissues (e.g., samples containing white blood cells, such as buffy coats). When measuring disease levels by sequencing cfDNA, it is particularly useful to remove germline and non-disease-related somatic variations detected by sequencing nucleic acid molecules obtained from white blood cells (e.g., from buffy coats). When cfDNA is sequenced for analysis, disease-related variations and non-disease-related somatic variations and germline variations caused by tumors are detected. Removing germline and non-disease-related somatic variations from the analysis can reduce errors attributed to ctDNA. Therefore, by removing non-disease-related somatic variations, the false positive error rate (i.e., SNVs mistakenly attributed to diseased tissues) can be reduced.

[0078] Other techniques or alternative methods can also be used to select loci from disease-associated SNV loci combinations or generate disease-associated SNV loci combinations. For example, in some embodiments, loci can be selected from disease-associated SNV loci combinations (or disease-associated SNV loci combinations can be generated to include SNVs) only when disease-associated variation is supported by two or more (e.g., 3, 4, 5 or more) sequencing reads obtained when sequencing nucleic acid molecules from diseased tissues. By requiring two or more sequencing reads to support variation associated with diseased tissues, the possibility of false positives can be reduced (e.g., by limiting the number of variations determined by sequencing or other errors when analyzing diseased tissues). Therefore, the false positive error rate (i.e., SNVs mistakenly attributed to diseased tissues) can be reduced by removing SNVs that are unreliably supported by sequencing data obtained by sequencing nucleic acid molecules derived from diseased tissues.

[0079] In some embodiments, loci in a disease-associated SNV locus combination can be selected (or disease-associated SNV locus combinations can be generated) by excluding common variant alleles, for example, excluding variants with a frequency greater than a predetermined frequency threshold from the general population. Common variants may be germline mutations that are not specific to diseased tissue and can therefore be excluded to reduce errors. In some embodiments, the predetermined frequency threshold is about 0.005 (or higher), about 0.01 or higher, about 0.02 or higher, or about 0.05 or higher. Thus, the false positive error rate (i.e., SNVs that are incorrectly attributed to diseased tissue) can be reduced by removing SNVs that are common in the general population and therefore attributable to germline variants.

[0080] In some embodiments, the locus in the disease-associated SNV locus combination can be selected (or disease-associated SNV locus combination can be produced by excluding the variation of the allele frequency detected in the nucleic acid sequencing data being greater than a predetermined threshold or greater than a statistical threshold). The cfDNA derived from the diseased tissue is usually a small portion of the cfDNA, and the variation with high allele frequency may be attributed to germline and / or somatic variation (e.g., non-disease-related somatic variation or somatic variation associated with different disorders or diseases) that is unrelated to the disease, and can be excluded from the analysis of measuring disease levels. Drawing an allele frequency histogram generally provides lower allele frequency clusters (usually attributed to diseased tissue or sequencing noise), and higher allele frequency clusters (usually attributed to germline and / or somatic variation). In some embodiments, statistical parameters are determined to distinguish between lower allele frequency clusters and higher allele frequency clusters, and variations associated with higher allele frequency clusters can be excluded. In some embodiments, the predetermined threshold is used to exclude variations in higher allele frequency clusters. The predetermined threshold can be, for example, about 0.2 or higher, about 0.25 or higher, or about 0.3 or higher.

[0081] In some embodiments, disease-associated SNV combinations can be selected (or disease-associated SNV locus combinations can be generated by) by excluding variations in homopolymer regions (a stretch of consecutive nucleotides with the same base type). In some embodiments, the homopolymer region comprises 3, 4, 5, 6, 7, 8, 9, 10 or more consecutive nucleotides with the same base type. Variations in homopolymer regions are prone to becoming false positive variations and may not accurately reflect diseased tissues. Therefore, by removing SNVs that fall within homopolymer regions, the false positive error rate (i.e., SNVs that are incorrectly attributed to diseased tissues) can be reduced.

[0082] In some embodiments, loci in a disease-associated SNV locus combination can be selected (or a disease-associated SNV locus combination can be generated) by excluding variants that are not supported by the complementary strand in a nucleic acid molecule derived from disease tissue. For example, if a variant is called in a sequencing read associated with a first strand, but a complementary variant is not called in a second strand that is complementary to the first strand, then this can be considered a sequencing error or other artifact and the variant can be excluded from further analysis. Thus, the false positive error rate (i.e., SNVs that are incorrectly attributed to diseased tissue) can be reduced by removing SNVs that are not reliably supported by sequencing data obtained by sequencing nucleic acid molecules derived from diseased tissue.

[0083] In some embodiments, the locus in the disease-associated SNV locus combination can be selected (or disease-associated SNV locus combination can be produced by it) by including only those variations that induce cyclic shifts (e.g., based on the flow cycle order, the flow graph signal is offset from a reference to one or more flow cycles) and / or generate new zeros or new non-zero signals in the sequencing data. For example, see U.S. Patent Application No. 16 / 864981 and International Patent Application No. PCT / US2020 / 031147, the contents of which are hereby incorporated by reference in their entirety for all purposes. Since cyclic shift events are unlikely to occur in the absence of true positive events (as further explained herein), in some embodiments, if the variation at the locus results in a cyclic shift event, the locus from the disease-associated SNV locus combination can be selected. Therefore, by including only SNVs that provide strong signals, the false positive error rate (i.e., SNVs mistakenly attributed to diseased tissues) can be reduced.

[0084] Methods described herein can be used for analyzing the different clones or different subclones of the diseased tissue in the same individual simultaneously.The different clones (for example, independent cancer clones) of diseased tissue generally have unique or almost unique variation characteristics.The subclones of diseased tissue can have some overlapping variations, although there are usually enough unique variations to select unique or almost unique variation subsets.In some embodiments, the locus to be sequenced is selected from the logical union of the variation loci relevant to several disease subclones, and the analysis detects the sample scores comprising all disease subclones, and also detects the disease score from each subclone.In some embodiments, the sequenced locus selected for analyzing a given clone or subclone is to avoid variation overlapping (that is, not selecting any variation shared by two or more clones or subclones).Therefore, the same sample from individual can be used to determine the disease level of a separate clone or subclone, or the score of the nucleic acid molecule related to a separate clone or subclone.In some embodiments, one or more of the clones or subclones are refractory to one or more cancer treatments, and the method can be used for monitoring the progress or disappearance of this refractory clone or subclone.

[0085] Patient samples and sequencing

[0086] Fluid sampling is a relatively non-invasive method for obtaining samples from individuals. Such fluid samples may include, for example, blood, plasma, saliva, feces or urine samples. In addition, for residual, malignant or other diseases without (or without obvious) primary or entity diseased tissue, fluid sampling allows obtaining nucleic acid molecules associated with the diseased tissue without tumor biopsy. Therefore, these methods are particularly useful when the location of the diseased tissue is unknown or the entity diseased tissue is too small to be sampled.

[0087] Fluid samples collected from individuals with diseases such as cancer typically have cell-free DNA (or "cfDNA"), which includes nucleic acid molecules derived from cancer tissue and nucleic acid molecules derived from non-diseased tissue. The nucleic acid sample from which sequencing data is obtained may be, but is not necessarily, cfDNA. For example, a fluid sample may provide other nucleic acids from which sequencing data can be obtained. For example, if the disease is a blood disease (e.g., a blood cancer), blood cells may be obtained from the blood sample, and the nucleic acid molecules from the blood cells may be sequenced to obtain sequencing data. In some embodiments, the nucleic acid molecule is a cell-free RNA molecule obtained from the fluid sample.

[0088] Any suitable sequencing method can be used to sequence nucleic acid molecules to obtain sequencing data from nucleic acid molecules. Exemplary sequencing methods may include but are not limited to high-throughput sequencing, next-generation sequencing, synthetic sequencing, flow sequencing, large-scale parallel sequencing, shotgun sequencing, single-molecule sequencing, nanopore sequencing, pyrophosphate sequencing, semiconductor sequencing, connection sequencing, hybridization sequencing, RNA-Seq, digital gene expression, synthetic single-molecule sequencing (SMSS), cloned single-molecule arrays, connection sequencing and Maxim-Gilbert sequencing. In some embodiments, nucleic acid molecules can be sequenced using a high-throughput sequencer, such as the Illumina HiSeq2500, Illumina HiSeq3000, Illumina HiSeq4000, Illumina HiSeqX, Roche 454, Life Technologies ion proton or open sequencing platforms described in U.S. Patent 10267790, which are hereby incorporated by reference in their entirety. Other sequencing methods and sequencing systems are known in the art. In some embodiments, nucleic acid molecules are sequenced using synthetic sequencing (SBS) methods. In some embodiments, nucleic acid molecules are sequenced using "natural sequencing by synthesis" or "non-terminated sequencing by synthesis" methods (see US Pat. No. 8,772,473, which is incorporated herein by reference in its entirety).

[0089] In some embodiments, the sequence of nucleic acid molecules is sequenced using two or more different sequencing methods. For example, the sequence of nucleic acid molecules is sequenced using two or more different sequencing methods with different false positive error rates. For example, the sequence of nucleic acid molecules is sequenced using two or more different sequencing methods with different false positive error rates. For example, the sequence of nucleic acid molecules is sequenced using two or more different sequencing methods with different false positive error rates. For example, some sequencing methods depends on predetermined nucleotide sequencing cycles (for example, CTAG, ATCG, TCAG, etc.), and the sequence of nucleic acid molecules is sequenced using different predetermined nucleotide sequencing cycles. For example, the sequence of nucleic acid molecules is sequenced using two, three, four or more different ...

[0090] In some embodiments, sequencing data is non-targeted.Some sequencing methods rely on the specific region of the targeted genome or locus to limit the breadth and / or enrichment specific region of sequencing. Conventional targeting methods include hybridization targeting (for example, using the nucleic acid probe attached on the label or the bead for selectively targeting the nucleic acid molecule region in the sample to carry out targeted sequencing), primer-based targeting (for example, using nucleic acid primers to amplify the targeted nucleic acid region by amplification (for example PCR), array-based capture and solution capture methods.For example, the targeted region can be a mutation hotspot in the gene or genome of a known cancer proliferation driver factor in the previously identified variation, genome. However, targeted sequencing ignores the important part information of the whole diseased tissue genome that can be used by methods described herein.

[0091] The method is alternatively performed using the sequencing data obtained by whole genome sequencing (WGS). By utilizing whole genome sequencing, more variant loci can be detected and used for analysis. Along with the increase of the number of loci analyzed, the detected signal increases with a ratio larger than noise, and by utilizing the whole genome, more data can be analyzed with less complicated preparation. Therefore, in some embodiments, the region of non-targeted genome. In some embodiments, the sequencing data is obtained from non-targeted whole genome sequencing.

[0092] Since the methods described herein can be used for a wide range of sequencing data (e.g., non-targeted or whole genome sequencing data), the average sequencing depth does not have to be as high as the targeted enrichment method. For example, in some embodiments, the average sequencing depth of sequencing data is about 100 or less, about 50 or less, about 25 or less, about 10 or less, about 5 or less, about 1 or less, about 0.5 or less, about 0.25 or less, about 0.1 or less, about 0.05 or less, about 0.025 or less, or about 0.01 or less. In some embodiments, the average sequencing depth is about 0.01 to about 1000, or any depth therebetween.

[0093] In some embodiments, sequencing data is obtained without amplifying nucleic acid molecules before establishing sequencing colonies (also referred to as sequencing clusters). Methods for generating sequencing colonies include bridge amplification or emulsion PCR. Methods that rely on shotgun sequencing and determine consensus sequences typically use unique molecular identifiers (UMIs) to label nucleic acid molecules and amplify nucleic acid molecules to produce many copies of the same independently sequenced nucleic acid molecules. The amplified nucleic acid molecules can then be attached to a surface and bridge amplified to produce independently sequenced sequencing clusters. The UMIs can then be used to associate independently sequenced nucleic acid molecules. However, the amplification process may introduce errors into the nucleic acid molecules, for example due to the limited fidelity of DNA polymerases. As discussed above, the currently provided method can be performed without determining a consensus sequence, so the initial amplification process is not required and can be avoided to reduce the false positive error rate. In some embodiments, nucleic acid molecules are not amplified before amplification to generate colonies for obtaining sequencing data. In some embodiments, nucleic acid sequencing data is obtained without using unique molecular identifiers (UMIs).

[0094] The proportion of individual samples in a sample pool can be determined using pooled sequencing data and sequencing data associated with an individual. The genome of an individual has a unique variation signature that can be used to determine the proportion of nucleic acid molecules attributable to that individual. Therefore, samples from multiple individuals can be pooled, and the fraction of nucleic acid molecules associated with an individual in the pooled sample can be determined without the use of sample identification barcodes.

[0095] In some embodiments, the individual has a disease or has previously had a disease. In some embodiments, the disease is cancer. Exemplary cancers encompassed by the methods described herein include, but are not limited to, acute lymphoblastic leukemia, acute myeloid leukemia, adenocarcinoma (e.g., prostate, small intestine, endometrium, cervix, large intestine, lung, pancreas, esophagus, colorectum, uterus, stomach, breast, and ovary), B-cell lymphoma, breast cancer, cancer, cervical cancer, chronic myeloid leukemia, colon cancer, esophageal cancer, glioblastoma, glioma, blood cancer, Hodgkin lymphoma, leukemia, lymphoma, lung cancer (e.g., non-small cell lung cancer), liver cancer, melanoma (e.g., metastatic malignant melanoma), multiple myeloma, neoplastic malignancies, neuroblastoma, non-Hodgkin lymphoma, ovarian cancer, pancreatic cancer, prostate cancer (e.g., hormone-refractory prostate cancer), kidney cancer (e.g., clear cell carcinoma), squamous carcinoma (e.g., cervix, eyelid, conjunctiva, vagina, lung, oral cavity, skin, bladder, tongue, larynx, and esophagus), head and neck squamous cell carcinoma, T-cell lymphoma, and thyroid cancer. In some embodiments, the cancer is refractory to one or more therapies.In some embodiments, the cancer is in remission or suspected of being in remission.

[0096] Flow sequencing and circular shift detection

[0097] An exemplary method for sequencing a nucleic acid molecule may include sequencing a nucleic acid molecule using a flow sequencing method to generate sequencing data. For example, by selecting loci or variations with a low error rate, a flow sequencing method may allow selection of variant loci with high confidence in a disease-associated SNV combination. For example, in some embodiments, the loci in a disease-associated SNV locus combination may be selected (or a disease-associated SNV locus combination may be generated by it) by including only induced cyclic shifts (i.e., based on the flow cycle order, the flow graph signal is offset by a complete cycle (e.g., 4 flow positions) relative to a reference) and / or generating new zeros or new non-zero signals in the sequencing data, as further described herein.

[0098] Flow sequencing can include extending primers bound to template polynucleotide molecules according to predetermined flow cycles, wherein a single type of nucleotide is accessible to the primer being extended at any given flow position. In some embodiments, at least some specific types of nucleotides include a label, and when the labeled nucleotide is incorporated into the primer being extended, the label produces a detectable signal. The resulting sequence incorporated into the extension primer by such nucleotides should be the reverse complementary sequence of the template polynucleotide molecule sequence. In some embodiments, for example, flow sequencing is used to generate sequencing data, which includes extending the primer using labeled nucleotides and detecting the presence or absence of the labeled nucleotides incorporated into the primer being extended. Flow sequencing can also be referred to as "natural sequencing by synthesis" or "non-terminated sequencing by synthesis" methods. Exemplary methods are described in U.S. Patent No. 8,772,473, which is incorporated herein by reference in its entirety. Although the following description is provided with reference to flow sequencing, it should be understood that other sequencing methods can be used to sequence all or part of the sequencing region. For example, the sequencing data discussed herein can be generated using pyrophosphate sequencing.

[0099] Flow sequencing includes using nucleotides to extend primers hybridized with polynucleotides. If there are complementary bases in the template strand, then the nucleotides of a given base type (e.g., A, C, G, T, U, etc.) can be mixed with the template of hybridization to extend the primer. Nucleotide can be, for example, non-terminal nucleotides. When nucleotides are non-terminal, if there are more than one continuous complementary bases in the template strand, then more than one continuous base can be incorporated into the primer strand being extended. In contrast to non-terminal nucleotides, there are nucleotides with 3' reversible terminators, wherein blocking groups are generally removed before attaching continuous nucleotides. If there are no complementary bases in the template strand, primer extension stops until the nucleotides complementary to the next base in the template strand are introduced. At least a portion of the nucleotides can be labeled so that their incorporation can be detected. Most commonly, only a single nucleotide type (i.e., discrete addition) is introduced once, although two or three different types of nucleotides can be introduced simultaneously in certain embodiments. This methodology can be contrasted with sequencing methods that use reversible terminators, in which primer extension stops after each single base extension, after which the terminator reverses to allow incorporation of the next subsequent base.

[0100] Nucleotide can be introduced in the primer extension process with flow order, which can be further divided into flow cycle.Flow cycle is the repetitive order of nucleotide flow, and can have any length. Nucleotide is added stepwise, and this allows the nucleotide of addition to be incorporated into the end of the sequencing primer of the complementary base in the template chain. As an example only, the flow order of flow cycle can be ATGC, or the flow cycle order can be ATCG. Those skilled in the art can easily envision alternative sequences. The flow cycle order can be any length, although the flow cycle containing four unique base types (A, T, C and G in any order) is the most common. In some embodiments, the flow cycle includes 5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20 or more nucleotide flows respectively in the flow cycle order. For example only, the flow cycle order can be TCACGATGCATGCTAG, and these 16 nucleotides provided separately are provided to several cycles with this flow cycle order. Between introducing different nucleotides, unincorporated nucleotides can be removed, for example, by washing the sequencing platform with washing solution.

[0101] Polymerase can be used to extend the sequencing primer by incorporating one or more nucleotides at the primer end in a template-dependent manner. In some embodiments, the polymerase is a DNA polymerase. The polymerase can be a naturally occurring polymerase or a synthetic (e.g., mutant) polymerase. Polymerase can be added in the initial step of primer extension, but a supplementary polymerase can be optionally added during sequencing, such as with the gradual addition of nucleotides or after multiple flow cycles. Exemplary polymerases include DNA polymerases, RNA polymerases, thermostable polymerases, wild-type polymerases, modified polymerases, Bst DNA polymerases, Bst 2.0 DNA polymerases, Bst 3.0 DNA polymerases, Bsu DNA polymerases, Escherichia coli DNA polymerase I, T7 DNA polymerases, bacteriophage T4 DNA polymerases, Φ29 (phi29) DNA polymerases, Taq polymerases, Tth polymerases, Tli polymerases, Pfu polymerases, and SeqAmp DNA polymerases.

[0102] When determining the sequence of the template strand, the nucleotides introduced may include labeled nucleotides, and the presence or absence of the labeled nucleic acid incorporated can be detected to determine the sequence. The label can be, for example, an optically active label (e.g., a fluorescent label) or a radioactive label, and a detector can be used to detect the signal emitted or changed by the label. The presence or absence of labeled nucleotides incorporated into the primer hybridized with the template polynucleotide can be detected, which allows determination of the sequence (e.g., by generating a flow graph). In some embodiments, the labeled nucleotides are labeled with fluorescence, luminescence, or other luminescent moieties. In some embodiments, the label is connected to the nucleotide via a joint. In some embodiments, the joint is cleavable, for example, by photochemical or chemical cleavage reactions. For example, the label can be cut after detection and before continuous nucleotides are incorporated. In some embodiments, the label (or joint) is connected to the nucleotide base, or is connected to another site on the nucleotide, without interfering with the extension of the nascent chain of the DNA. In some embodiments, the joint comprises a disulfide or a PEG-containing part.

[0103] In some embodiments, the introduced nucleotides include only unlabeled nucleotides, and in some embodiments, the nucleotides include a mixture of labeled and unlabeled nucleotides. For example, in some embodiments, the portion of labeled nucleotides compared to the total nucleotides is about 90% or less, about 80% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 20% or less, about 10% or less, about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2% or less, about 1.5% or less, about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, or about 0.01% or less. In some embodiments, the portion of labeled nucleotides compared to the total nucleotides is about 100%, about 95% or more, about 90% or more, about 80% or more, about 70% or more, about 60% or more, about 50% or more, about 40% or more, about 30% or more, about 20% or more, about 10% or more, about 5% or more, about 4% or more, about 3% or more, about 2.5% or more, about 2% or more, about 1.5% or more, about 1% or more, about 0.5% or more, about 0.25% or more, about 0.1% or more, about 0.05% or more, about 0.025% or more, or about 0.01% or more. In some embodiments, the portion of labeled nucleotides compared to the total nucleotides is from about 0.01% to about 100%, such as from about 0.01% to about 0.025%, from about 0.025% to about 0.05%, from about 0.05% to about 0.1%, from about 0.1% to about 0.25%, from about 0.25% to about 0.5%, from about 0.5% to about 1%, from about 1% to about 1.5%, from about 1.5% to about 2%, from about 2% to about 1%. About 2.5%, about 2.5% to about 3%, about 3% to about 4%, about 4% to about 5%, about 5% to about 10%, about 10% to about 20%, about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 90% to less than 100%, or about 90% to about 100%.

[0104] Before generating sequencing data, polynucleotides are hybridized with sequencing primers to generate hybridization templates. Polynucleotides can be connected to connectors during sequencing library preparation. Connectors can include hybridization sequences that hybridize with sequencing primers. For example, the hybridization sequence of the connector can be a uniform sequence across multiple different polynucleotides, and the sequencing primer can be a uniform sequencing primer. This allows multiple sequencing of different polynucleotides in the sequencing library.

[0105] Polynucleotide can be connected to surface (such as solid support) for sequencing.Polynucleotide can be amplified (for example, by bridge amplification or other amplification techniques) to produce polynucleotide sequencing colony.The polynucleotide amplified in the cluster is substantially identical or complementary (some errors may be introduced during the amplification process so that a part of the polynucleotide may not necessarily be identical with the original polynucleotide).Colony formation allows signal amplification, so that the detector can accurately detect the incorporation of the labeled nucleotides of each colony.In some cases, emulsion PCR is used to form colonies on beads, and beads are distributed on the sequencing surface.The example of the system and method for sequencing can be found in U.S. Patent No. 10,344,328, which is hereby incorporated by reference in its entirety.

[0106] According to the flow sequence (which can be cycled according to the flow cycle sequence), primers hybridized to the polynucleotide are extended across the nucleic acid molecule using separate nucleotide streams, and the incorporation of the nucleotides can be detected as described above, thereby generating a sequencing data set for the nucleic acid molecule.

[0107] Use the primer extension of flow order-checking to allow length to be hundreds or even thousands of base magnitude long range order-checking.Can increase or reduce the quantity of flow step or circulation to obtain required order-checking length.For using the primer with one or more different base types, the extension of primer progressively extends and can comprise one or more flow steps.In some embodiments, the extension of primer comprises 1 to approximately 1000 flow steps, as 1 to approximately 10 flow steps, approximately 10 to approximately 20 flow steps, approximately 20 to approximately 50 flow steps, approximately 50 to approximately 100 flow steps, approximately 100 to approximately 250 flow steps, approximately 250 to approximately 500 flow steps or approximately 500 to approximately 1000 flow steps.Flow step can be divided into identical or different flow circulations.The base number that is incorporated in the primer depends on the sequence in the order-checking zone, and the flow order that is used to extend the primer. In some embodiments, the sequencing region is about 1 base to about 4000 bases long, such as about 1 base to about 10 bases long, about 10 bases to about 20 bases long, about 20 bases to about 50 bases long, about 50 bases to about 100 bases long, about 100 bases to about 250 bases long, about 250 bases to about 500 bases long, about 500 bases to about 1000 bases long, about 1000 bases to about 2000 bases long, or about 2000 bases to about 4000 bases long.

[0108] Sequencing data can be produced based on the detection of the nucleotides incorporated and the order in which the nucleotides are introduced. For example, the extended sequence (i.e., each reverse complementary sequence of the corresponding template sequence) is taken: CTG, CAG, CCG, CGT, and CAT (assuming that there is no sequencing method in the preceding sequence or in the following sequence), and the repeated flow cycle of TACG (i.e., sequentially adding T, A, C, and G nucleotides in the repeated cycle). Only when there is a complementary base in the template polynucleotide, the nucleotides of the specific type at the given flow position will be incorporated into the primer. An exemplary resulting flow diagram is shown in Table 1, wherein 1 represents the incorporation of the nucleotides introduced and 0 represents the nucleotides not incorporated into the introduced nucleotides. The flow diagram can be used to derive the sequence of the template chain. For example, the sequencing data discussed herein (e.g., flow diagram) represent the sequence of the primer chain extended, and its reverse complementary sequence can be easily determined as the sequence representing the template chain. The asterisk (*) in Table 1 represents that if additional nucleotides are incorporated into the sequencing chain extended (e.g., a longer template chain), then there may be a signal in the sequencing data.

[0109] Table 1

[0110]

[0111] Flow plots can be binary or non-binary. Binary flow plots detect the presence (1) or absence (0) of incorporated nucleotides. Non-binary flow plots can more quantitatively determine the number of nucleotides incorporated from each step-wise introduction. For example, an extension sequence of CCG includes the incorporation of two C bases into the extension primer within the same C flow (e.g., at flow position 3), and the signal emitted by the labeled base will have an intensity greater than the intensity level corresponding to a single base incorporation. This is shown in Table 1. Non-binary flow plots also indicate the presence or absence of bases and can provide additional information, including the number of bases that may be incorporated into each extension primer at a given flow position. These values ​​do not need to be integers. In some cases, these values ​​can reflect the uncertainty and / or probability of the number of bases incorporated at a given flow position.

[0112] In some embodiments, the sequencing data set includes a flow signal representing a base count, which indicates the number of bases incorporated into the sequenced nucleic acid molecule at each flow position. For example, as shown in Table 1, a primer extended with a CTG sequence using a TACG flow cycle sequence has a value of 1 at position 3, indicating that the base count at that position is 1 (the 1 base is C, which is complementary to the G in the template strand being sequenced). Also in Table 1, a primer extended with a CCG sequence using a TACG flow cycle sequence has a value of 2 at position 3, indicating that the base count of the extended primer at that position during that flow position is 2. Here, the 2 bases refer to the CC sequence at the start of the CCG sequence in the extension primer sequence, and it is complementary to the GG sequence in the template strand.

[0113] The flow signal in the sequencing data set can include one or more statistical parameters, which indicate the possibility or confidence interval of one or more base counts at each flow position. In some embodiments, the flow signal is determined by an analog signal detected during the sequencing process, such as a fluorescent signal of one or more bases incorporated into the sequencing primer during sequencing. In some cases, the analog signal can be processed to generate statistical parameters. For example, a machine learning algorithm can be used to correct the contextual effect of the analog sequencing signal, as described in the disclosed international patent application WO2019084158A1, which is incorporated herein by reference in its entirety. Although the integer incorporated at any given flow position is zero or more bases, a given analog signal may not fully match the analog signal. Therefore, considering the detected signal, a statistical parameter indicating the possibility of the number of bases incorporated at the flow position can be determined. For example only, for the CCG sequence in Table 1, the flow signal indicates that the possibility of incorporating 2 bases at flow position 3 can be 0.999, and the flow signal indicates that the possibility of incorporating 1 base at flow position 3 can be 0.001. Sequencing data sets can be formatted as sparse matrices, where the flow signal includes statistical parameters indicating the likelihood of multiple base counts at each flow position. By way of example only, a sequence of repeated flow cycles using TACG with a primer extended with the sequence TATGGTCGTCGA (SEQ ID NO: 1) (i.e., the reverse complement of the sequencing read) can generate Figure 8A . The statistical parameter or likelihood value can vary, for example, based on noise or other artifacts present during detection of the analog signal during sequencing. In some embodiments, if the statistical parameter or likelihood is below a predetermined threshold, the parameter can be set to a predetermined non-zero value (i.e., some very small or negligible value) that is substantially zero to assist in the statistical analysis discussed further herein, where a true zero value may cause computational errors or insufficiently distinguish between levels of improbability, e.g., very unlikely (0.0001) and incredible (0).

[0114] A value indicating the likelihood of a given sequence in a sequencing data set can be determined from a sequencing data set without sequence alignment. For example, given this data, the most likely sequence can be determined by selecting the base count with the highest likelihood at each flow position, such as Figure 8B As shown by the star in (using Figure 8A ). Thus, the sequence for primer extension can be determined based on the most likely base count at each flow position: TATGGTCGTCGA (SEQ ID NO: 1). From this, the reverse complement sequence (i.e., the template strand) can be easily determined. Furthermore, given the TATGGTCGTCGA (SEQ ID NO: 1) sequence (or its reverse complement), the likelihood of this sequencing dataset can be determined as the product of the selected likelihoods at each flow position.

[0115] In some embodiments, the sequencing data set associated with the nucleic acid molecule is compared with one or more (e.g., 2, 3, 4, 5, 6 or more) possible candidate sequences. The close match between the sequencing data set and the candidate sequence (based on the matching score discussed below) indicates that the sequencing data set may come from a nucleic acid molecule with the same sequence as the candidate sequence of the close match. In some embodiments, the sequence of the nucleic acid molecule sequenced can be mapped to a reference sequence (e.g., using Burrows-Wheeler alignment (BWA) algorithm or other suitable alignment algorithms) to determine the locus (or one or more loci) of the sequence. The sequencing data set in the flow space can be easily converted to base space (or vice versa if the flow order is known), and can be mapped in flow space or base space. The locus (or multiple loci) corresponding to the mapped sequence can be associated with one or more variant sequences, which can be operated as the candidate sequence (or haplotype sequence) of the analytical method described herein. An advantage of the methods described herein is that in some cases, the sequence of the nucleic acid molecule sequenced does not need to be aligned with each candidate sequence using an alignment algorithm, which is typically computationally expensive. Instead, the matching score for each candidate sequence can be determined using sequencing data in flow space, which is a more computationally efficient operation.

[0116] The match score indicates the degree to which the sequencing data set supports the candidate sequence. For example, given the expected sequencing data for the candidate sequence, a match score indicating the likelihood that the sequencing data set matches the candidate sequence can be determined by selecting a statistical parameter (e.g., likelihood) at each flow position that corresponds to the base count at that flow position. The product of the selected statistical parameters can provide the match score. For example, assuming Figure 8A Sequencing datasets of extended primers shown in , and TATGGTCA Candidate primer extension sequence of TCGA (SEQ ID NO: 2). Figure 8C (show Figure 8A The same sequencing dataset in (shown in Figure 2) shows the traces of candidate sequences (solid circles). For comparison, TATGGTC G TCGA (SEQ ID NO: 1) sequence (see Figure 8B ) trace in Figure 8C The match score indicating the likelihood that the sequencing data matches the first candidate sequence TATGGTCATCGA (SEQ ID NO: 2) is substantially different from the match score indicating the likelihood that the sequencing data matches the second candidate sequence TATGGTCGTCGA (SEQ ID NO: 1), even though the sequences vary by only a single base change. Figure 8C As shown in , the difference between the traces is observed at flow position 12 and propagates for at least 9 flow positions (and possibly longer if the sequencing data extends across additional flow positions). This persistent spread across one or more flow cycles can be referred to as a "cyclic shift" and is generally a very unlikely event if the sequencing data set matches the candidate sequence.

[0117] When sequencing nucleic acid sequencing data and reference sequencing data using non-terminal nucleotides provided in separate nucleotide streams according to a flow cycle order, SNV-induced cyclic shifts occur when the sequencing data associated with a nucleic acid molecule having an SNV is offset by one or more flow cycles relative to the reference sequencing data associated with the reference sequence (i.e., a sequence having the same sequence as the nucleic acid molecule except that it does not have the SNV). That is, the sequencing data and the reference sequencing are different across one or more flow cycles. The reference sequencing data need not be obtained by sequencing a reference nucleic acid molecule, but can be generated in a computer based on the reference sequence.

[0118] Figure 8C Showing exemplary circular shift-induced SNVs. Figure 8C The second candidate sequence shown in FIG is the reverse complement of the sequence read associated with the nucleic acid molecule containing the SNV (and associated with the sequencing data shown in the flow chart at the top of the figure) TATGGTC G TCGA (SEQ ID NO: 1), and the first candidate sequence is the reverse complement of the sequence read of the reference sequence TATGGTC ATCGA (SEQ ID NO: 2). A→GSNP (at base position 8 of both sequences) induces a cyclic shift, which can be observed by the sequencing data associated with the nucleic acid molecule containing the SNV being shifted to the left by one cycle compared to the reference sequencing data. For example, based on the sequencing data associated with the nucleic acid molecule containing the SNV, the T base at base position 9 is sequenced at flow position 13, and is sequenced at position 17 according to the reference sequencing data. Similarly, based on the sequencing data associated with the nucleic acid molecule containing the SNV, the CG bases at base positions 10 and 11 are sequenced at flow positions 15 and 16, and the CG bases at positions 19 and 20 are sequenced according to the reference sequencing data.

[0119] Because recurrent translocation events are unlikely to occur in the absence of true positive events, in some embodiments, loci from the combination of disease-associated SNV loci can be selected only if the variation at the locus produces a recurrent translocation event.

[0120] The sensitivity of short genetic variations to induce circular shifts can depend on the flow cycle order used to sequence nucleic acid molecules harboring SNVs. Figure 8C In the example shown in , the TACG flow cycle sequence is included, but other flow cycle sequences can be used to induce cyclic shifts in other variations. By generating a new zero signal or a new non-zero signal in the sequencing data, any flow sequence can be used to observe the possibility of SNV inducing a cyclic shift event. Therefore, even if the selected flow sequence does not induce a cyclic shift event, SNV can also use different flow sequences to induce a cyclic shift event. In some embodiments, when sequencing nucleic acid sequencing data and reference sequencing data are sequenced using the non-terminal nucleotides provided in the respective nucleotide streams according to the flow cycle sequence, only when the variation at the locus causes sequencing data and reference sequencing data to be different from the sequencing data with a new zero signal or a new non-zero signal, the locus from the disease-related SNV locus combination is selected. In some embodiments, the signal variation can be continuous. In some embodiments, when sequencing nucleic acid sequencing data and reference sequencing data are sequenced using the non-terminal nucleotides provided in the respective nucleotide streams according to the flow cycle sequence, only when the variation at the locus causes sequencing data and reference sequencing data to be different at two or more flow positions (which can be continuous), the locus from the disease-related SNV locus combination is selected.

[0121] Since the nucleic acid molecules are sequenced using different flow cycle orders, the sequencing data sets are different. Figure 8DAn exemplary sequencing dataset of a nucleic acid molecule containing a SNV having the reverse complement sequence of TATGGTCGTCGA (SEQ ID NO: 1) determined using a different flow cycle order (AGCT) is shown. Figure 8C The reference sequencing data is mapped onto the sequencing data of the nucleic acid molecule containing the SNV. The SNV generates a new zero signal at position 17 and a new non-zero signal at position 18. Therefore, even if the SNV is the same, the TACG flow cycle induces a circular shift (see Figure 8C ), and AGCT flow cycling also did not induce cyclic shift. However, the new zero and new non-zero signals suggest that SNV may use different cycling orders to induce cyclic shift.

[0122] Variant signals, false positive errors, and noise

[0123] The nucleic acid molecules in the fluid sample obtained from the individual are sequenced to obtain sequencing data associated with the individual. The sequencing data includes sequencing data associated with non-diseased tissue and sequencing data associated with diseased tissue. However, due to the existence of false positive errors that occur during sequencing, the differences between the sequencing data associated with non-diseased tissue and the sequencing data associated with diseased tissue are not all attributed to mutations in the genome of the diseased tissue. In other words, the total number N of individual small nucleotide variant (SNV) reads detected at the loci selected from the individualized locus combination in the sequencing data is total , is the number of SNV reads N detected at the position selected from the individualized locus combination that can be attributed to the diseased tissue det and the number of SNV reads N detected in positions selected from the individualized locus combination that are attributable to false positive errors (ie, background) bkg The sum of . That is:

[0124] N total =N det +N bkg The number of SNV reads detected in the selected locus that can be attributed to the diseased tissue is N det The number of loci N selected from the individualized loci combination var , the average sequencing depth D, and the fraction F of nucleic acid molecules in the fluid sample that originate from diseased tissue. In some embodiments, N det has a first-order relationship with the fraction F. In some embodiments:

[0125] N det =N var DF.

[0126] Similarly, the number of SNV reads N detected in the selected loci that can be attributed to false positive errors isbkg The number of loci N selected from the individualized loci combination var , the average sequencing depth D, and the error rate E across the selected loci, e.g., in some embodiments, Nbkg has a first-order relationship with the error rate E. That is, in some embodiments:

[0127] N bkg =N var DE.

[0128] Therefore, in some embodiments, N total It can be schematically determined as:

[0129] N total =N var D(F+E).

[0130] The number of SNV reads N detected in the selected loci due to false positive errors bkg The error rate E is proportional to the error rate, and thus can be reduced by excluding those loci that are more likely to cause false positive errors. Exemplary methods for selecting loci with lower false positive errors are further described herein.

[0131] The fraction of nucleic acid molecules in a sample that are associated with an individual's disease can be calculated using N det In some embodiments:

[0132]

[0133] When N det When not directly measured, for example due to false positive errors, the fraction of nucleic acid molecules in a sample that are associated with an individual disease can be determined by adding a signal (e.g., ) is determined by comparing it with a background factor indicating the false positive error rate of sequencing across the selected loci. In some embodiments, F is determined by comparing it with N total For example, the first-order relationship In some embodiments, the score is determined as:

[0134]

[0135] By assuming the number of false positive errors and the Poisson sampling noise of the true detections, the signal-to-noise ratio (SNR) of the number of SNVs detected in the SNVs selected from the personalized locus combination that can be attributed to diseased tissue can be determined. total The sampling noise (i.e. ) can be set to Thus, in some embodiments, the signal-to-noise ratio (SNR) of a SNV detected in a selected locus attributable to diseased tissue can be determined as:

[0136]

[0137] In some embodiments, the false positive error rate E is determined independently of the loci selected, eg, the balance of the genome outside of the individualized locus combination or the loci selected from the individualized locus combination.

[0138] The error in the determined fraction F can also be determined based on sampling noise. For example, in some embodiments, the error in F is

[0139]

[0140] Alternatively, in some embodiments:

[0141]

[0142] Therefore, in some embodiments, the score is treated as a nominal value with an error, which can be defined as a confidence interval for the score.

[0143] The disease level of an individual can be related to the score F of the nucleic acid molecules in the sample derived from the diseased tissue. Therefore, the presence or level of the disease can be measured by determining, for example, this score. Disease recurrence, progression, or regression can be determined by measuring the disease level of an individual at multiple time points. In some embodiments, the confidence intervals of two or more measurement scores are compared, which can be used to determine the statistically significant differences (e.g., for measuring the progression or regression of a disease) between the measurement scores.

[0144] In some embodiments, the signal-to-noise ratio is used to detect the presence or recurrence of a disease. A higher signal-to-noise ratio indicates an increased likelihood of the presence or recurrence of a disease.

[0145] In some embodiments, multiple samples from different individuals are merged together to obtain the nucleic acid sequencing data merged, which includes the nucleic acid sequencing data relevant to the individual subject. The nucleic acid molecules relevant to the diseased tissue of a given individual have unique or almost unique variation characteristics, which allows many detected variation reads to be assigned to an individual. In some embodiments, the sequencing locus selected for analysis is used to avoid variation overlap (that is, any variation shared by two or more individuals is not selected). In other embodiments, the variation reads of the variation shared by two or more individuals are included in the analysis, for example, by counting the variation reads of the individual sharing the variation, or by weighting the variation read counts of the individual sharing the variation (for example, based on the relative amount of the nucleic acid molecules derived from an individual), or by performing maximum likelihood analysis of sample and disease scores on the entire sequence library. The measurement score of the nucleic acid molecules relevant to the disease in the individual in the individual pool (that is, using the nucleic acid sequencing data merged) will first be determined as the score of the nucleic acid molecules in the sample pool, and can be adjusted based on the ratio of the sample in the sample pool. By way of example only, if the measured fraction of nucleic acid molecules in a sample pool that originate from an individual's diseased tissue is 0.5%, and the sample from that individual represents 5% of the nucleic acid molecules in the sample pool, then the fraction of nucleic acid molecules in the sample from that individual that originate from the diseased tissue is 10%.

[0146] Accurately determining the false positive error rate E provides a more accurate determination of the score F and the signal-to-noise ratio SNR. In some embodiments, the false positive error rate is determined empirically. In some embodiments, sequencing data from one or more other individuals is used to determine the false positive error rate. In some embodiments, the false positive error rate is determined using sequencing data from the same individual, such as in regions outside the individualized locus combination. In some embodiments, the false positive error rate is essentially determined from sequencing data related to the individual used to determine the score, signal-to-noise ratio, or disease level. For example, in some embodiments, a group of control loci can be selected to determine the false positive error rate. The control loci can select loci that are extremely unlikely to mutate, such as highly conserved regions of the genome. For example, the control loci can be located in the coding region of an essential gene, for which true variation will result in cell death. Therefore, true variation at the control loci is extremely unlikely, and any variation detected can be attributed to false positive errors. The total number of SNV base reads N detected at the control loci total,con , the total number of control loci N con and the average sequencing depth D can be used to determine the false positive error rate. That is, in some embodiments:

[0147]

[0148] Figure 1Shown is an exemplary method 100 for measuring the level of a disease (such as cancer) in an individual, for example, the fraction of nucleic acid molecules (such as cfDNA molecules) associated with a disease in a sample from an individual. The sample can be a fluid sample, such as a blood sample, a plasma sample, a saliva sample, a urine sample, or a fecal sample. In step 105, the signal is compared with a background factor using nucleic acid sequencing data associated with the individual. Optionally, the nucleic acid sequencing data is non-targeted and / or non-enriched nucleic acid sequencing data (such as whole genome sequencing data). In some embodiments, the sequencing depth of the sequencing data is less than about 100, less than about 10, or less than about 1. In some embodiments, the sequencing depth of the sequencing data is at least 0.01. The signal indicates the ratio of the sequenced loci selected from the individualized disease-associated SNV locus combination to the diseased tissue. Optionally, the loci selected from the disease-associated SNV combination are selected based on the false positive rate of the individual loci. In some embodiments, the signal is: or N det In some embodiments, the amplitude of the signal depends at least on the average sequencing depth associated with the multiple selected loci and the nucleic acid sequencing data. The background factor indicates the false positive error rate of sequencing across the selected loci. In step 110, the disease level in the individual (e.g., the fraction of nucleic acid molecules in the sample associated with the disease) is determined based on a comparison of the signal with the background factor. For example, the fraction can be determined based on:

[0149]

[0150] Figure 2Another exemplary method 200 for measuring disease (such as cancer) levels in an individual is shown, for example, the score of nucleic acid molecules (such as cfDNA molecules) associated with the disease in a sample from an individual. The sample can be a fluid sample, such as a blood sample, a plasma sample, a saliva sample, a urine sample, or a fecal sample. In step 205, individualized disease-related small nucleotide variations (SNV) locus combinations are constructed using the sequencing data associated with the diseased tissue and the sequencing data associated with non-diseased tissue. The individualized locus combination is based on the difference between the sequencing data associated with the diseased tissue and the sequencing data associated with non-diseased tissue. In step 210, loci are selected from the individualized locus combination. In some embodiments, all loci in the individualized locus combination are selected, and in some embodiments, a subset of loci in the individualized locus combination is selected. For example, loci can be selected from the individualized locus combination based on the false positive rate of the individual locus. In step 215, the sequencing data associated with the sample from an individual is obtained. For example, sequencing data can be obtained by sequencing the nucleic acid molecules in the sample or by receiving sequencing data from a record. Optionally, the nucleic acid sequencing data is non-targeted and / or non-enriched nucleic acid sequencing data (such as whole genome sequencing data). In some embodiments, the sequencing depth of the sequencing data is less than about 100, less than about 10, or less than about 1. In some embodiments, the sequencing depth of the sequencing data is at least 0.01. In step 220, the nucleic acid sequencing data associated with the individual is used to compare the signal with the background factor. The signal indicates the ratio of the sequenced loci selected from the individualized disease-associated SNV locus combination to the diseased tissue. In some embodiments, the signal is: or N det In some embodiments, the amplitude of the signal depends on at least a plurality of selected loci and an average sequencing depth associated with the nucleic acid sequencing data. The background factor indicates the sequencing false positive error rate of the selected loci. In step 225, the disease level in the individual (e.g., the fraction of nucleic acid molecules associated with the disease in a sample from the individual) is determined based on a comparison of the signal with the background factor. For example, the fraction can be determined based on:

[0151]

[0152] Methods for detecting the presence, level, recurrence, progression or regression of a disease

[0153] The methods described herein can be used to detect the presence (e.g., recurrence) of a disease, measure the level of a disease, or measure or detect the progression or regression of a disease. In some embodiments of the methods described herein, the individual has previously been treated for the disease. In some embodiments, the suspected disease is in remission, such as complete remission or partial remission. After treatment of the disease, such as by chemotherapy or cancer resection, the disease may recur, such as due to incomplete removal or killing of all diseased tissues. For example, cancer may metastasize and be relocated to a different location in the individual, or may be too small to be detected by known imaging modalities (e.g., MRI, PET scans, etc.). Individuals may be regularly monitored for disease recurrence, regression, or progression so that when the disease recurs or progresses, the individual is retreated.

[0154] For example, by using nucleic acid sequencing data associated with an individual, a signal indicating the ratio of the sequenced loci selected from a combination of individualized disease-associated small nucleotide variations (SNV) loci derived from diseased tissue is compared with a noise factor indicating the sampling variance across the selected loci; and determining whether the individual has a disease based on a comparison of the signal with the background factor, the presence or residual level of a disease such as cancer can be detected. In some embodiments, the signal-to-noise ratio is determined, for example, as described herein.

[0155] The statistical significance of the detected signal can be determined by comparing the signal to the statistical noise (e.g., sampling variance, which can be based on at least the number of true detections and the number of false positive errors). If the signal is greater than the statistical noise, for example, the signal-to-noise ratio (SNR) is greater than about 1.5, about 2, about 3, about 5, about 8, about 10, or more, the disease can be detected with certainty. Conversely, in some embodiments, a lower SNR indicates that the disease was not detected, for example, less than about 1.5, less than about 1.4, less than about 1.3, less than about 1.2, or less than about 1.1.

[0156] Figure 3The exemplary method 300 of detecting disease or disease (such as cancer) recurrence in an individual is presented. In step 305, the nucleic acid sequencing data related to the individual is used to compare signal with noise factor. Nucleic acid sequencing data can be derived from the nucleic acid molecules in the fluid sample obtained from the individual. For example, in some embodiments, nucleic acid sequencing data is derived from the cell-free DNA in the fluid sample (such as, blood sample, plasma sample, saliva sample, urine sample or fecal sample) from the individual. Optionally, nucleic acid sequencing data is non-targeted and / or non-enriched nucleic acid sequencing data (such as whole genome sequencing data). In some embodiments, the sequencing depth of sequencing data is less than about 100, less than about 10 or less than about 1. In some embodiments, the sequencing depth of sequencing data is at least 0.01. The signal indicates the ratio of the sequenced locus derived from the diseased tissue of the individualized disease-related small nucleotide variation (SNV) genome combination. Optionally, the locus selected from the disease-related SNV combination is selected based on the false positive rate of the individual locus. The noise factor indicates the sequencing sampling noise across the selected locus. In step 310, determine whether there is disease in the individual based on the comparison of signal and noise factor. For example, in some embodiments, a statistically significant signal above a noise factor indicates that the individual has a disease.

[0157] Figure 4The exemplary method 400 of the existence or recurrence of disease (such as cancer) in display individuality.In step 405, use the sequencing data relevant to diseased tissue and the sequencing data relevant to non-diseased tissue to build individualized disease-related small nucleotide variation (SNV) locus combination.Individualized locus combination is based on the difference between the sequencing data relevant to diseased tissue and the sequencing data relevant to non-diseased tissue.In step 410, from individualized locus combination, select locus.In some embodiments, select all loci in individualized locus combination, and select the locus subset in individualized locus combination in some embodiments.For example, locus can be selected from individualized locus combination based on the false positive rate of individual locus.In step 415, obtain the nucleic acid sequencing data relevant to the sample from individuality.For example, sequencing data can be obtained by the nucleic acid molecule in sample being checked or by receiving the sequencing data of sample from record.Sample can be the fluid sample obtained from individuality.For example, in some embodiments, nucleic acid sequencing data is derived from the cell-free DNA in the fluid sample (for example, blood sample, plasma sample, saliva sample, urine sample or fecal sample) from individuality. Optionally, nucleic acid sequencing data is non-targeted and / or non-enriched nucleic acid sequencing data (such as whole genome sequencing data). In some embodiments, the sequencing depth of sequencing data is less than about 100, less than about 10 or less than about 1. In some embodiments, the sequencing depth of sequencing data is at least 0.01. In step 420, the nucleic acid sequencing data related to the individual are used to compare signal and noise factors. The signal indicates the ratio of the sequenced loci derived from the diseased tissue selected from the individualized disease-related small nucleotide variation (SNV) genome combination. The noise factor indicates the sampling noise across the selected loci. In step 425, based on the comparison of signal and noise factor, determine whether there is disease in the individual. For example, in some embodiments, the statistically significant signal indication individuality higher than the noise factor suffers from disease.

[0158] For example, by measuring the disease level of an individual, the presence or disease residue of a disease (such as cancer) can also be detected. Optionally, the disease level is indicated by the score of the nucleic acid molecules in the sample from an individual derived from diseased tissue. The score of nucleic acid molecules (such as cfDNA) in the fluid sample obtained from an individual derived from diseased tissue is related to the severity or level of the disease in the individual. Therefore, the nucleic acid molecule score attributable to diseased tissue can be used as a marker for residual disease level or recurrence. By using the nucleic acid sequencing data associated with an individual, the signal indicating the ratio of the sequenced locus derived from the diseased tissue of the individualized disease-related small nucleotide variation (SNV) locus combination is compared with the background factor indicating the false positive error rate of sequencing across the selected locus, and the level of the disease in the individual is determined based on the comparison of the signal with the background factor, the level can be determined.

[0159] Optionally determine the error in the disease measurement level (e.g., the error in the measurement score), such as the confidence interval of the level. In some embodiments, the error is proportional to the total number of individual small nucleotide variant reads detected at the selected locus. For example, the error in the measurement level can be used to determine whether the measurement level is statistically significant. For example, in some embodiments, if the lower limit of the confidence interval of the score is higher than zero, the measured level indicates the presence or recurrence of the disease. The error can also be used to measure the possibility that the measured score is greater than a predetermined value. In some embodiments, the likelihood that the measured fraction of nucleic acid molecules attributable to diseased tissue is greater than a predetermined threshold (e.g., 0 or above, about 0.1% or above, about 0.2% or above, about 0.5% or above, about 1% or above, about 1.5% or above, about 2.5% or above, about 3% or above, about 4% or above, about 5% or above, about 6% or above, about 7% or above, about 8% or above, about 9% or above, or about 10% or above) is measured as compared to nucleic acid molecules attributable to non-diseased tissue, wherein a score above the predetermined threshold indicates the presence or recurrence of disease in the individual.

[0160] The progression or regression of a disease can be determined and / or monitored by measuring the disease level at two or more time points (e.g., the fraction of nucleic acid molecules in an individual sample that are attributable to diseased tissue, or a signal indicating the ratio of sequenced loci selected from a personalized set of disease-associated small nucleotide variant (SNV) loci that are derived from diseased tissue, compared to a background factor indicating the sequencing false positive error rate across the selected loci). Thus, the measured score can be compared to a previous score F. prior In some embodiments, the time point of the present invention can be used to compare the time points of the disease.Time point can include the first time point before the disease is treated and the second time point after the disease is treated.In some embodiments, the increase of score or signal (compared with background factor) indicates the progress of disease, and the reduction of score or signal (compared with background factor) indicates the disappearance of disease.In some embodiments, the statistical significance of score or signal increases (compared with background factor) indicates the progress of disease, and the statistical significance of score or signal reduces (compared with background factor) indicates the disappearance of disease.The level of two or more time points determines that error (such as confidence interval) can be used to determine whether the change of measurement level has statistical significance.

[0161] Figure 5The exemplary method 500 of displaying the recurrence, progress or disappearance of disease (such as cancer) in monitoring individual is compared.In step 505, the nucleic acid sequencing data relevant to individual is used to compare signal with background factor.Nucleic acid sequencing data can be derived from the nucleic acid molecules in the fluid sample obtained from individual.For example, in some embodiments, nucleic acid sequencing data is derived from the cell-free DNA in the fluid sample (such as, blood sample, plasma sample, saliva sample, urine sample or fecal sample) from individual.Optionally, nucleic acid sequencing data is non-targeted and / or non-enriched nucleic acid sequencing data (such as whole genome sequencing data).In some embodiments, the sequencing depth of sequencing data is less than about 100, less than about 10 or less than about 1.In some embodiments, the sequencing depth of sequencing data is at least 0.01.This signal indicates the ratio of the sequencing locus derived from the diseased tissue of the small nucleotide variation (SNV) locus combination of individualized disease.Optionally, the locus selected from the disease-related SNV combination is selected based on the false positive rate of individual locus.Background factor indicates the sequencing false positive error rate variance across selected locus. In step 510, the disease level in the individual is determined based on a comparison of the signal with the background factor. For example, in some embodiments, a statistically significant signal that is higher than the background factor indicates that the individual has a disease. In step 515, the disease level in the individual is compared with the previous disease level in the individual. Compared to the previously measured disease level, a statistically significant change in the measured disease level indicates that the disease has recurred, progressed, or subsided. For example, compared to the previously measured disease level, the measured disease level has increased statistically significantly, indicating that the disease has progressed. Compared to the previously measured disease level, the measured disease level has decreased statistically significantly, indicating that the disease has subsided.

[0162] Figure 6Another exemplary method 600 showing the recurrence, progress or regression of a disease (such as cancer) in a monitoring individual is shown. In step 605, individualized disease-related small nucleotide variation (SNV) locus combination is constructed using the sequencing data relevant to the diseased tissue and the sequencing data relevant to non-diseased tissue. The individualized locus combination is based on the difference between the sequencing data relevant to the diseased tissue and the sequencing data relevant to non-diseased tissue. In step 610, locus is selected from the individualized locus combination. In some embodiments, all loci in the individualized locus combination are selected, and in some embodiments, a subset of the loci in the individualized locus combination is selected. For example, locus can be selected from the individualized locus combination based on the false positive rate of the individual locus. In step 615, the nucleic acid sequencing data relevant to the sample from the individual is obtained. For example, sequencing data can be obtained by sequencing the nucleic acid molecules in the sample or by receiving the sequencing data of the sample from the record. The sample can be a fluid sample obtained from the individual. For example, in some embodiments, the nucleic acid sequencing data is derived from the cell-free DNA in the fluid sample (for example, blood sample, plasma sample, saliva sample, urine sample or fecal sample) from the individual. Optionally, nucleic acid sequencing data is non-targeted and / or non-enriched nucleic acid sequencing data (such as whole genome sequencing data). In some embodiments, the sequencing depth of sequencing data is less than about 100, less than about 10 or less than about 1. In some embodiments, the sequencing depth of sequencing data is at least 0.01. In step 620, the nucleic acid sequencing data related to the individual are used to compare the signal with the background factor. The signal indicates the ratio of the sequenced locus derived from the diseased tissue of the individualized disease-related small nucleotide variation (SNV) genome combination. The background factor indicates the sequencing false positive error rate variance of the selected locus. In step 625, the disease level in the individual is determined based on the comparison of the signal with the background factor. For example, in some embodiments, the statistically significant signal indication individuality higher than the background factor suffers from disease. In step 630, the disease level in the individual is compared with the disease level before in the individual. Compared with the disease level previously measured, the statistically significant change indication disease of the disease level measured has recurred, progressed or disappeared. For example, compared with the disease level previously measured, the disease level measured is statistically significantly increased, indicating that the disease has progressed. A statistically significant decrease in the measured disease level compared to the previously measured disease level indicates that the disease has resolved.

[0163] Optionally, the measured scores, measured levels, progression, regression and / or recurrence of the disease are recorded in a record, such as an electronic medical record (EMR) or a patient profile. In some embodiments of any of the methods described herein, the measured scores, measured levels, progression, regression and / or recurrence of the disease are informed to the individual. In some embodiments of any of the methods described herein, the individual is diagnosed as having a disease, recurrence of disease or progression of disease. In some embodiments of any of the methods described herein, the disease is treated individually.

[0164] Systems and equipment

[0165] The above operations (including reference to Figure 1-6 those described) are optionally Figure 7 It is clear to those skilled in the art how to perform the Figure 7 The components depicted in the figure may be used to perform other processes, for example, combinations or sub-combinations of all or part of the operations described above. It will also be clear to those skilled in the art how the methods, techniques, systems, and devices described herein may be combined with each other in whole or in part, and whether those methods, techniques, systems, and / or devices are combined with each other. Figure 7 The components depicted in are performed and / or provided.

[0166] Figure 7 1 illustrates an example of a computing device according to one embodiment. Device 700 may be a host computer connected to a network. Device 400 may be a client computer or a server. Figure 7 As shown, device 700 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device), such as a phone or tablet. The device can include, for example, one or more of a processor 710, an input device 720, an output device 730, a memory 740, and a communication device 760. The input device 720 and the output device 730 can generally correspond to those described above and can be connected to or integrated with the computer.

[0167] Input device 720 may be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice recognition device. Output device 730 may be any suitable device that provides output, such as a touch screen, tactile device, or speaker.

[0168] Memory 740 can be any suitable device that provides storage, such as electrical, magnetic, or optical memory, including RAM, cache, hard drive, or removable storage disk. Communication device 760 can include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a physical bus or wireless connection.

[0169] Software 750 , which may be stored in memory 740 and executed by processor 710 , may include, for example, programming embodying the functionality of the present disclosure (eg, as embodied in the devices described above).

[0170] The software 750 may also be stored and / or transmitted in any non-transitory computer-readable storage medium for use by or in conjunction with an instruction execution system, apparatus, or device, such as those described above, which may retrieve instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, a computer-readable storage medium may be any medium, such as memory 740, that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0171] The software 750 may also be transmitted over any transmission medium for use by or in conjunction with an instruction execution system, apparatus, or device, such as those described above, which may retrieve instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transmission medium may be any medium that can convey, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. Transmission-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless communication media.

[0172] Device 700 can be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links that can perform transmission and reception of network signals, such as a wireless network connection, a T1 or T3 line, a cable network, DSL, or a telephone line.

[0173] Device 700 can execute any operating system suitable for operating on a network. Software 750 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or as a web-based application or web service, for example, through a web browser.

[0174] The methods described herein optionally further include reporting the information determined using the analytical method and / or generating a report including the information determined using the analytical method. For example, in some embodiments, the method further includes reporting or generating a report including information related to individual disease levels. The reported information or the information within the report can be related to the score of cfDNA attributable to a disease (such as cancer) or the presence or absence of a detectable amount of disease (such as cancer) in a sample obtained from an individual. The report can be distributed to a recipient, or the information can be reported to a recipient, such as a clinician, subject, or researcher. Example

[0175] By reference to the following non-limiting examples provided as the exemplary embodiments of the application, the application can be better understood. The following examples are provided to more fully illustrate the embodiments, but should never be construed as limiting the broad scope of the application. Although some embodiments of the application have been shown and described herein, it is apparent that these embodiments are only provided by way of example. Without departing from the spirit and scope of the present invention, it will be appreciated by those skilled in the art that many variations, changes and replacements can be expected. It should be understood that the various alternatives of the embodiments described herein can be used for implementing the methods described herein.

[0176] Example 1

[0177] DNA obtained from a cancer tissue biopsy obtained from an individual is sequenced by whole genome sequencing to obtain sequencing data associated with the cancer tissue. A blood sample is obtained from the individual, and DNA from the whole blood is sequenced to obtain sequencing data associated with healthy tissue. The sequencing data associated with the cancer tissue is compared with the sequencing data associated with the healthy tissue, and the differences are listed in the personalized disease-associated SNV locus combination. The variants in the personalized locus combination are filtered based on the false positive error rate of the variant, and the variants with the lowest false positive error rate are selected for analysis. A total of Nvar loci were selected.

[0178] Cell-free DNA is obtained from a fluid sample from an individual and sequenced using non-targeted and non-enriched whole-genome sequencing to obtain sequencing data with an average sequencing depth of D. The sequencing method generates a sequencing false positive error rate E. The number of sequencing reads with variant calls from the personalized locus combination is measured, and the fraction of nucleic acid molecules associated with the disease in the fluid sample (F) is determined. prior ) and the error of the score.

[0179] An individual receives cancer treatment. After treatment, cell-free DNA is obtained from a subsequent fluid sample of the individual and the cfDNA is sequenced using non-targeted and non-enriched whole-genome sequencing to obtain sequencing data with an average sequencing depth of D (same or different depth than the previous sample). The sequencing method generates a sequencing false positive error rate E (same or different than the previous sample). The number of sequencing reads with variant calls from the personalized locus combination is measured. total , and determine the fraction of disease-associated nucleic acid molecules in the fluid sample (F present ) and the error of the score.

[0180] The fraction associated with the latter sample (F present ) and the score associated with the previous sample (F prior) to monitor the progression or regression of the cancer. A statistically significant increase in the score indicates that the disease has progressed, and a statistically significant decrease in the score indicates that the disease has regressed.

[0181] Example 2

[0182] DNA obtained from a cancer tissue biopsy obtained from an individual is sequenced by whole genome sequencing to obtain sequencing data associated with the cancer tissue. A blood sample is obtained from the individual, and DNA from the whole blood is sequenced to obtain sequencing data associated with healthy tissue. The sequencing data associated with the cancer tissue is compared with the sequencing data associated with the healthy tissue, and the differences are listed in the personalized disease-associated SNV locus combination. The variants in the personalized locus combination are filtered based on the false positive error rate of the variant, and the variants with the lowest false positive error rate are selected for analysis. A total of N var A gene locus.

[0183] An individual receives cancer treatment. After treatment, cell-free DNA is obtained from a subsequent fluid sample of the individual and the cfDNA is sequenced using non-targeted and non-enriched whole-genome sequencing to obtain sequencing data with an average sequencing depth of D (same or different depth than the previous sample). The sequencing method generates a sequencing false positive error rate E (same or different than the previous sample). The number of sequencing reads with variant calls from the personalized locus combination is measured. total , and determining a signal-to-noise ratio (SNR) of nucleic acid molecules associated with the disease in the fluid sample. A signal-to-noise ratio above a set threshold (k) indicates that the individual has residual disease.

[0184] Example 3

[0185] Cancer samples were purchased from the Analytical Biological Services (ABS) biobank. The biobank's collection of normal and diseased human tissue biospecimens is based on strict legal requirements and includes appropriate informed consent for commercial research. The biospecimens included archival FFPE tumor biopsies matched with buffy coat and plasma (cfDNA) from cancer donors. This study evaluated the genetic characteristics of these samples.

[0186] Samples. FFPE, buffy coat, and plasma samples were obtained from patient 1, a 40-year-old woman with metastatic colon adenocarcinoma. The FFPE sample consisted of ~80% cancer cells, ~10-20% fibroblasts, infiltrating mononuclear cells, and necrotic tissue (dead tissue).

[0187] A plasma sample was obtained from Patient 2, a 69-year-old male with metastatic melanoma. This sample served as a control to determine the sequencing error rate. The plasma sample appears red, indicating the presence of red and white blood cells during the blood draw. Lysed blood cells can result in higher-than-expected background non-tumor cfDNA compared to cancer cfDNA (i.e., ctDNA).

[0188] Nucleic acid extraction and library preparation. Use DNeasy blood and tissue kit or Nucleic acid molecules were extracted from 100 μL of buffy coat (patient 1) using a DNA / RNA kit. gDNA extracted from both kits was combined, and 1000 ng of the extracted gDNA was used for library construction using the Roche KAPA HyperPrep kit.

[0189] Use the DNeasy Blood and Tissue Kit or RecoverAll TM The total nucleic acid isolation kit was used to extract nucleic acid molecules from 30 μm FFPE tissue sections (patient 1). 173 ng of gDNA extracted from FFPE samples using the DNeasy blood and tissue kit containing xylene on the slide was used for the first FFPE-based library construction. TM 446 ng of gDNA extracted from FFPE samples (xylene-free on glass slides) using the Total Nucleic Acid Isolation Kit was used for library construction of the second FFPE-based library. The library was constructed using the Roche KAPA HyperPrep Kit, followed by 7 cycles of PCR using the KAPA HiFi HotStart ReadyMix Kit.

[0190] Using MagMAX TM The Cell-Free Total Nucleic Acid Isolation Kit was used to extract nucleic acid molecules from 4 mL of plasma (patient 1 or patient 2). 100 ng of cfDNA was extracted from the plasma sample of patient 1, and 25 ng of cfDNA was extracted from the plasma sample of patient 2 using the Roche KAPA HyperPrep Kit. Then, 7 cycles of PCR were performed using the KAPA HiFi HotStart ReadyMix Kit.

[0191] Accurate quantification of adapter-ligated libraries was performed using the KAPA Library Quantification Kit.

[0192] Whole-genome sequencing: Emulsion PCR and sequencing were performed on each sample using the Ultima Genomics instrument and pipeline (TACG flow cycling) at coverages of x30-150.

[0193] Bioinformatics analysis. 917,319,868 raw reads were obtained for the buffy coat (patient 1) sample library (library 1, average length 228 bases, median coverage). 2,136,822,000 raw reads were obtained for the cfDNA (plasma, patient 1) sample library (library 2, average length 183 bases). For two different FFPE-based sequencing libraries, 553,298,760 raw reads were obtained (library 3) and 1,768,786,851 raw reads were obtained (library 4) (average length 186 bases).

[0194] 211,8786,000 raw reads (average length 187 bases) were obtained for the cfDNA (plasma, patient 2) sample library (library 5).

[0195] Raw reads were aligned to the reference genome (hg38) using BWA (version 0.7.15-r1140), and duplicates were marked using Picard Tools (version 2.15.0, Broad Institute) for buffy coat and FFPE reads or the SAM Tools rmdup program for cfDNA reads. After alignment and duplicate removal, the median genome coverage for libraries 1–5 was 45x, 84x, 8x, 18x, and 56x, respectively.

[0196] Using the HaplotypeCaller program in the GATK4 software package (modified to process sequencing data generated by Ultima Genomics instruments and processes), variations in FFPE reads were determined for the hg38 reference genome. 4,694,198 variations were determined from the first FFPE-based library (library 3) and 6,702,421 variations were determined from the second FFPE-based library (library 4). The baseline variations from the two FFPE samples were combined into a list of 7,682,808 unique variations (i.e., "baseline variations") to account for differences in sample processing, and for each baseline variation, the number of reads supporting the baseline variation in each sample was tabulated. The baseline variation was then filtered to remove germline variation, variation due to DNA damage caused by sample preparation, and variation due to sequencing errors. First, the baseline variation was filtered to include only SNP variations supported by 2 or more sequencing reads, resulting in 4,179,203 unique variations. These variants were then filtered to remove variants with an allele frequency greater than 0.01 (thought to be likely germline mutations) from a population database (gnomAD v3, available from the Broad Institute), resulting in 1,292,135 unique variants. These variants were then filtered to remove variants within homopolymer regions of 8 bases or longer, resulting in 1,176,179 unique variants. These variants were then filtered to remove unsupported variants in the complementary strand (suspected to be sequencing errors), resulting in 505,500 unique variants. These variants were then filtered to remove variants detected by reads from buffy coat samples (putative germline and / or non-cancerous somatic mutations), resulting in 67,660 unique variants. From the 67,660 unique variant combinations, 17,073 variants that were present in both FFPE sample libraries and were expected to induce circular shifts (i.e., the flow profile signal was shifted by one full cycle (e.g., 4 flow positions) or more relative to the reference based on the flow cycle order) were selected for further analysis. For comparison, 17,509 variants that were present in both FFPE sample libraries and were expected to induce circular shifts (i.e., contained new zero or new non-zero flow profile signals) under different flow order conditions were analyzed, as well as 5,748 variants that were not expected to induce circular shifts (i.e., did not contain new zero or new non-zero flow profile signals).

[0197] Bioinformatics analysis was performed using patient 1 data, and cfDNA from patient 2 was used to estimate the sequencing error rate for selected variants in the same cohort. The estimated fraction of cfDNA in patient 1 that was associated with cancer was determined when analyzing circulating shift-induced variants. =4.65%, and the background level was determined to be ~0.35%. See Table 2. Therefore, the error correction fraction F' = F - E is about 4.3%.

[0198] Table 2

[0199]

[0200] In the analysis of potential circulating shift variants, the estimated fraction of cfDNA associated with cancer in Patient 1 was determined to be 4.34%, and the background level was determined to be ∼0.44%, providing an error-corrected fraction of 3.9%. See Table 3.

[0201] Table 3

[0202]

[0203]

[0204] When analyzing variants that did not induce or had the potential for circulating shift, the estimated fraction of cfDNA associated with cancer in Patient 1 was determined to be 3.92%, and the background level was determined to be ∼0.55%, providing an error-corrected fraction of 3.37%. See Table 4.

[0205] Table 4

[0206]

[0207] Example 4

[0208] The genome of DNA sample NA12878 (sample available from the Coriell Institute for Medical Research) was sequenced using non-terminated fluorescently labeled nucleotides in four flow cycles (TACG). The sequencing run generated 415,900,002 reads with an average length of 176 bases. 399,804,925 reads were aligned (using BWA, version 0.7.17-r1188) to the hg38 reference genome.

[0209] After alignment, reads that were perfectly aligned to the reference genome (178,634,625 reads) or reads that contained a single mismatch with the reference genome and aligned with a mapping quality score of 20 or higher were selected (27,265,661 reads). That is, 193,904,639 were excluded for further analysis, for example, due to insertion / deletion mutations (indels), multiple mismatches, or potential incorrect (artificial) alignments with the reference genome. Therefore, it was assumed that 27,265,661 reads included true positive NA12878 SNPs, as well as any false positive SNPs caused by sequencing errors. From this pool of 27,265,661 reads, sequencing reads that spanned the mismatch locus more than once were removed to reduce the impact of true positive NA12878 SNP variations, resulting in a total of 3,413,700 reads containing mismatches of depth 1).

[0210] The remaining 3,413,700 reads each included a mismatch that: (1) is expected to induce a cyclic shift if the flow map flow signal is offset by one full cycle (e.g., 4 flow positions) relative to the reference based on the flow cycle order, (2) potentially could induce a cyclic shift if a different flow cycle was used (e.g., it produces a new zero or a new non-zero signal in the flow map), or (3) could not induce a cyclic shift regardless of the flow cycle order. Of the 3,413,700 mismatches, 1,184,954 (34%) induced a cyclic shift, while 1,546,588 (43%) could induce a cyclic shift with a different flow order (i.e., "potential cyclic shifts"). In comparison, the theoretical expectation of random mismatches would nominally indicate 42% cyclic shifts and 46% potential cyclic shift mismatches. Overall, the mismatch rate for induced cyclic shifts was 3.7×10 -5 events / base, and the mismatch rate that induced potential circular shifts was 4.8×10 -5 Table 5 shows the 10 most common single mismatches that induce circular shifts and their relative incidence percentages.

[0211] Table 5

[0212] refer to Read segment % cases TTT TCT 7.18 AAA AGA 7.18 GAG GGG 4.63 CTC CCC 4.62 CAG CGG 4.12 CTG CCG 4.09 AAC AGC 3.86 GTT GCT 3.83 CAT CGT 3.63 GAT GGT 3.62

[0213] Then based on the mismatch in each category of three different categories (i.e., induced circular shift, potential induced circular shift, or not induced and cannot induce circular shift), the performance of variation determination was evaluated. Reads were aligned with the reference genome using BWA, and variation determination was performed using the HaplotypeCaller tool of GATK (version 4). Mismatch determinations were filtered by discarding variations within homopolymers longer than 10 bases or within 10 bases adjacent to homopolymers of 10 bases or longer.

[0214] Mismatch calls were compared with calls generated by the Genome in a Bottle (GIAB) project for the same NA12878 to determine the accuracy of each class of mismatch #TP / (#FP+#FN+#TP). Sequencing data were randomly downsampled to a specified average genome depth. Mismatches that induced cyclic shifts and mismatches that potentially induced cyclic shifts had higher accuracy than mismatches that did not induce cyclic shifts, as shown in Table 6.

[0215] Table 6

[0216] Mismatch type 30x 22x 15x 8x Circular shift 0.9834 0.981 0.981 0.9772 No cyclic shift 0.9799 0.9759 0.9775 0.9696 Potential cyclic shift 0.9826 0.9808 0.9795 0.9767

Claims

1. A system for determining a fraction of nucleic acid molecules associated with cancer, comprising: one or more processors; and A non-transitory computer-readable medium storing one or more programs comprising instructions for measuring the fraction of nucleic acid molecules associated with cancer in a sample from an individual, the instructions being: using nucleic acid sequencing data associated with the individual, comparing a signal indicating a rate at which selected sequenced loci are derived from cancer tissue to a background factor indicating a sequencing false-positive error rate across the selected loci, wherein the sequenced loci are selected from a personalized cancer-associated small nucleotide variation (SNV) locus combination, the comparing comprising subtracting the background factor from the signal, wherein the personalized cancer-associated SNV locus combination comprises one or more loci at which one or more SNVs are identified in cancer tissue from the individual and not identified in non-cancerous tissue, and wherein the average sequencing depth of the nucleic acid sequencing data is less than 100; and determining a fraction of nucleic acid molecules associated with cancer in the sample from the individual based on a comparison of the signal to the background factor, Where the score is defined as: in: F is the score; N total is the total number of individual small nucleotide variant reads detected at the selected locus; N var is the number of selected loci; D is the average sequencing depth; and E is the false positive error rate across the selected loci, where the false positive error rate is defined as: where Ntotal,con is the total number of SNV reads detected at the control locus, Ncon is the number of control loci, and D is the average sequencing depth.

2. The system of claim 1, wherein the one or more programs further comprise instructions for determining an error for measuring the score. The system of claim 2 , wherein the error is a confidence interval for the score.

4. The system of claim 2, wherein the error is proportional to the total number of individual small nucleotide variation reads detected at the selected locus.

5. The system of claim 4, wherein the score and the error are defined as: in: F is the score; N total is the total number of individual small nucleotide variant reads detected at the selected locus; N var is the number of the selected loci; D is the average sequencing depth; and E is the false positive error rate across the selected loci.

6. The system of claim 1, wherein the one or more programs further comprise instructions for measuring cancer recurrence.

7. The system of claim 1, wherein the one or more programs further comprise instructions for measuring cancer progression or regression by comparing the measured score to a previously measured score.

8. The system of claim 7, wherein progression or regression of the cancer is based on a statistically significant change in the measured score.

9. The system of claim 1, wherein the amplitude of the signal depends on at least the number of selected loci and an average sequencing depth associated with the nucleic acid sequencing data.

10. A system for determining whether an individual has cancer, comprising: one or more processors; and A non-transitory computer-readable medium storing one or more programs comprising instructions for detecting cancer in an individual, the instructions comprising: (a) determining a signal to noise factor ratio using nucleic acid sequencing data associated with nucleic acid molecules from an individual, wherein the signal is a ratio indicating that selected sequenced loci are derived from cancer tissue, wherein the selected sequenced loci are selected from a personalized cancer-associated small nucleotide variation (SNV) locus set, and the noise factor indicates sampling variance across the selected loci, wherein the personalized cancer-associated SNV locus set includes one or more loci at which one or more SNVs are identified in cancer tissue from the individual and not identified in non-cancerous tissue, and wherein the nucleic acid sequencing data has an average sequencing depth of less than 100, The signal-to-noise ratio SNR det Defined as: in, N total is the total number of individual small nucleotide variant reads detected at the selected locus; N var is the number of selected loci; D is the average sequencing depth; and E is the false positive error rate across the selected loci, where the false positive error rate is defined as: where Ntotal,con is the total number of SNV reads detected at the control locus, Ncon is the number of control loci, and D is the average sequencing depth; and (b) determining whether the individual has cancer based on a ratio of the signal to the noise factor.

11. The system of claim 10, wherein if the signal exceeds the noise factor by more than a predetermined threshold, then the individual is determined to have cancer recurrence or residual cancer.

12. The system of claim 10, wherein the individual is determined to have cancer recurrence or residual cancer if the signal exceeds the noise factor by a factor of k or more, where k is 1.

5.

13. The system of claim 10, wherein the individual is determined to have cancer recurrence or residual cancer if the signal exceeds the noise factor by a factor of k or more, where k is 3.

0.

14. The system of claim 10, wherein the individual is determined to have cancer recurrence or residual cancer if the signal exceeds the noise factor by a factor of k or more, where k is 5.

0.

15. The system of claim 10, wherein the individual is determined to have cancer recurrence or residual cancer if the signal exceeds the noise factor by a factor of k or more, where k is 10.

16. The system of claim 10, wherein the one or more programs further comprise instructions for detecting cancer recurrence.

17. The system of claim 10, wherein the amplitude of the signal depends on at least the number of selected loci and an average sequencing depth associated with the nucleic acid sequencing data.

18. A system for determining the presence, progression, or regression of cancer in an individual, comprising: one or more processors; and A non-transitory computer-readable medium storing one or more programs comprising instructions for detecting the likelihood of the presence of cancer in an individual, or detecting the progression or regression of cancer in an individual, the instructions comprising: Measure at least one of the following: (a) indicating a likelihood that the value of the fraction F of nucleic acid molecules in the sample originating from cancerous tissue of the individual is greater than zero, wherein F greater than zero indicates the presence of cancer in the individual, and (b) indicating a statistically significant change in the value of a fraction F of nucleic acid molecules in a sample derived from cancer tissue of the individual, wherein the statistically significant change is relative to a previously measured fraction Fprior, and wherein a statistically significant change in F indicates progression or regression of cancer in the individual; Wherein the fraction F is determined by comparing the total number Ntotal of individual small nucleotide variation (SNV) reads detected in the cell-free nucleic acid sequencing data to the number Nvar of SNVs selected from the SNV panel, adjusted by the average sequencing depth D, and further adjusted by the sequencing false positive error rate E across the selected loci, wherein the SNV reads are detected at loci selected from the personalized cancer-associated SNV locus panel, wherein the personalized cancer-associated SNV locus panel comprises one or more loci at which one or more SNVs identified in cancer tissue from the individual and not identified in non-cancerous tissue are located, wherein F is defined as: wherein the average sequencing depth of the nucleic acid sequencing data is less than 100, and The false positive error rate is defined as: where Ntotal,con is the total number of SNV reads detected at the control locus, Ncon is the number of control loci, and D is the average sequencing depth.

19. The system of claim 18, wherein the one or more programs further comprise instructions for generating personalized cancer-associated SNV locus combinations.

20. The system of claim 19, wherein generating a personalized cancer-associated SNV locus combination comprises: sequencing nucleic acid molecules derived from a cancer tissue sample of an individual to identify a set of cancer-associated SNVs specific to the individual; and This cancer-associated SNV set was filtered to remove germline variants and non-cancer-associated somatic variants.

21. The system of claim 20, wherein the sample of cancerous tissue is a tumor biopsy sample obtained from the individual.

22. The system of claim 20, wherein the germline variation or the non-cancer associated somatic variation, or both, is determined by sequencing nucleic acid molecules derived from a non-cancerous tissue sample obtained from the individual.

23. The system of claim 22, wherein the sample of non-cancerous tissue comprises white blood cells.

24. The system of claim 23, wherein the non-cancerous tissue sample is a buffy coat.

25. The system of claim 20, wherein the one or more programs further comprise instructions for filtering the set of cancer-associated SNVs to remove SNVs supported by only one sequencing read.

26. The system of claim 20, wherein the one or more programs further comprise instructions for filtering the set of cancer-associated SNVs to remove SNVs that are not supported by complementary sequencing reads.

27. The system of claim 20, wherein the one or more programs further comprise instructions for filtering the set of cancer-associated SNVs to remove SNVs present in a general population of individuals having an allele frequency greater than a predetermined threshold.

28. The system according to claim 27, wherein the predetermined threshold is 0.

01.

29. The system of claim 20, wherein the one or more programs further comprise instructions for filtering SNVs within homopolymer regions or filtering SNVs within short tandem repeats.

30. The system of claim 20, wherein the nucleic acid sequencing data is obtained by sequencing nucleic acid molecules from a fluid sample obtained from an individual using non-terminal nucleotides provided in separate nucleotide streams according to a flow cycle sequence comprising a plurality of flow positions, wherein The flow position corresponds to a nucleotide flow; and Generating a personalized cancer-associated SNV locus combination also includes filtering the cancer-associated SNV set to include only SNVs that, when the nucleic acid sequencing data and the reference sequencing data are sequenced according to the flow cycle order using non-terminal nucleotides provided in separate nucleotide streams, result in nucleic acid sequencing data that is different from the reference sequencing data associated with the reference sequence at two or more flow positions.

31. The system of claim 18, wherein the nucleic acid sequencing data is obtained by sequencing nucleic acid molecules from a fluid sample obtained from an individual using non-terminal nucleotides provided in respective nucleotide streams according to a flow cycle sequence comprising a plurality of flow positions, wherein the flow positions correspond to nucleotide streams; and The one or more programs further include instructions for generating a personalized cancer-associated SNV locus combination comprising: Nucleic acid molecules derived from a cancer tissue sample are sequenced to determine a set of cancer-associated SNVs; and generating an individualized cancer-associated SNV locus combination further includes filtering the set of cancer-associated SNVs to include only SNVs that, when the nucleic acid sequencing data and the reference sequencing data are sequenced according to a flow cycle order using non-terminal nucleotides provided in separate nucleotide flows, result in nucleic acid sequencing data that is different from reference sequencing data associated with a reference sequence at two or more flow positions.

32. The system of claim 30, wherein generating the personalized cancer-associated SNV locus combination comprises filtering the cancer-associated SNV set to include only SNVs that, when the nucleic acid sequencing data and the reference sequencing data are sequenced according to a flow cycle order using non-terminal nucleotides provided in separate nucleotide streams, result in nucleic acid sequencing data that differs from the reference sequencing data associated with the reference sequence across one or more flow cycles.

33. The system of any one of claims 1-17, wherein the nucleic acid molecule is a cell-free nucleic acid molecule.

34. The system according to any one of claims 1-32, wherein the nucleic acid molecule is a DNA molecule.

35. The system of any one of claims 1-32, wherein the nucleic acid molecule is an RNA molecule.

36. The system of any one of claims 1-32, wherein the nucleic acid sequencing data is derived from nucleic acid molecules in a fluid sample obtained from the individual.

37. The system of claim 36, wherein the fluid sample is a blood sample, a plasma sample, a saliva sample, a urine sample, or a stool sample.

38. The system of any one of claims 1-32, wherein the cancer is a metastatic cancer.

39. The system of any one of claims 1-32, wherein the one or more programs further comprise instructions for sequencing nucleic acid molecules to obtain sequencing data.

40. The system according to any one of claims 1-32, wherein the nucleic acid sequencing data is obtained by sequencing nucleic acid molecules according to a predetermined nucleotide sequencing cycle order.

41. The system of claim 40, wherein the nucleic acid sequencing data is further obtained by resequencing the nucleic acid molecule according to different predetermined nucleotide sequencing cycles, wherein The different predetermined nucleotide sequencing cycles result in a different false positive variation rate at the subset of sequenced loci compared to the predetermined nucleotide sequencing cycle order.

42. The system of any one of claims 1-32, wherein the sequencing data is non-targeted sequencing data.

43. The system of claim 42, wherein the sequencing data is obtained from a non-targeted whole genome.

44. The system of any one of claims 1-32, wherein the sequencing data has an average sequencing depth of at least 0.

01.

45. The system of any one of claims 1-32, wherein the average sequencing depth of the sequencing data is less than 10.

46. ​​The system of any one of claims 1-32, wherein the average sequencing depth of the sequencing data is less than 1.

47. The system of any one of claims 1-32, wherein the combination of cancer-associated SNV loci comprises a passenger mutation.

48. The system of any one of claims 1-32, wherein the combination of cancer-associated SNV loci comprises a driver mutation.

49. The system of any one of claims 1-32, wherein the combination of cancer-associated SNV loci comprises single nucleotide polymorphism (SNP) loci.

50. The system of any one of claims 1-32, wherein the combination of cancer-associated SNV loci comprises indel mutation loci.

51. The system of any one of claims 1-32, wherein the loci selected from the combination of cancer-associated SNV loci comprise 300 or more loci.

52. The system of any one of claims 1-32, wherein the loci selected from the combination of cancer-associated SNV loci are selected based on false positive rates of individual loci.

53. The system of any one of claims 1-32, wherein the loci selected from the combination of cancer-associated SNV loci are based on unique SNVs associated with a selected subclone of the cancer.

54. The system of any one of claims 1-32, wherein the cancer-associated SNV locus combination is determined by comparing sequencing data associated with cancerous tissue and sequencing data associated with non-cancerous tissue.

55. The system of claim 54, wherein the one or more programs further comprise instructions for sequencing nucleic acid molecules derived from cancer tissue to obtain sequencing data associated with the cancer tissue.

56. The system of claim 54, wherein the one or more programs further comprise instructions for sequencing nucleic acid molecules derived from non-cancerous tissue to obtain sequencing data associated with the non-cancerous tissue.

57. The system of any one of claims 1-32, wherein the nucleic acid sequencing data is obtained using surface-based nucleic acid molecule sequencing, and wherein the nucleic acid molecules are not amplified prior to attaching the nucleic acid molecules to the surface.

58. The system of any one of claims 1-32, wherein the nucleic acid sequencing data is obtained without the use of unique molecular identifiers (UMIs).

59. The system of any one of claims 1-32, wherein the nucleic acid sequencing data is obtained without the use of a sample identification barcode.

60. The system of any one of claims 1-32, wherein the sequencing data is obtained by sequencing nucleic acid molecules in a pooled sample obtained from a plurality of individuals.

61. The system of claim 60, wherein the selected loci are unique to each individual in the plurality of individuals.

62. The system of claim 61, wherein at least one locus within the selected loci is common between at least two individuals in the plurality of individuals.

63. The system of claim 60, wherein the sequencing depth of each individual is determined, and wherein the signal for each individual is adjusted based on the sequencing depth associated with the individual.

64. The system of any one of claims 1-32, wherein the one or more programs further comprise instructions for generating a report indicating the presence, absence, or fraction of cancer in the individual.

65. The one or more programs or systems of claim 64, wherein the one or more programs further comprise instructions for providing a report to the patient or the patient's medical representative.

Citation Information

Patent Citations

  • Systems for biological sample processing and analysis

    US10267790B1

  • Methods for biological sample processing and analysis

    US10344328B2

  • Inferring selection in white blood cell matched cell-free DNA variants and / or in RNA variants

    US20190355438A1

  • METHODS AND SYSTEMS FOR DETERMINING The CELLULAR ORIGIN OF CELL-FREE NUCLEIC ACIDS

    US20190385700A1

  • Machine learning variant source assignment

    US20200013484A1