Methods for classifying samples into clinically relevant categories
A novel multi-parameter strategy for analyzing ctDNA sequence coordinates and motifs enhances the sensitivity and specificity of liquid biopsy methods, enabling accurate classification of ctDNA levels in cancer diagnosis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2026-03-27
AI Technical Summary
Current liquid biopsy methods for cancer diagnosis have limited sensitivity and specificity, failing to accurately classify samples into clinically relevant categories due to the complexity and inaccuracies in analyzing cell-free tumor DNA (ctDNA).
A novel multi-parameter strategy is employed to extract and analyze sequence coordinates and nucleic acid motifs from ctDNA, calculating diagnostic scores to classify samples by comparing them to reference values, enhancing sensitivity and specificity.
The method achieves significantly improved sensitivity and specificity in distinguishing normal from abnormal samples, with sensitivity ranging from 96.8% to 99.9997% and outperforming existing methods, allowing for accurate classification of ctDNA levels as low, medium, or high.
Smart Images

Figure 0007836823000013 
Figure 0007836823000014 
Figure 0007836823000015
Abstract
Description
[Technical Field]
[0001] This invention relates to the fields of biology, medicine, and chemistry, particularly to the field of molecular biology, and more particularly to the field of molecular diagnostics. [Background technology]
[0002] Eukaryotic genomes are organized within chromatin, enabling not only DNA compactification but also the regulation of DNA metabolism (replication, transcription, repair, and recombination). The signature of eukaryotic chromatin structure, particularly nucleosome arrangement, has been shown to be useful for identifying rare nucleic acid fragments in complex mixtures present in eukaryotes (Heitzer E. et al., Nat. Rev. Genet., 2019, 20(2):71-88).
[0003] It has been hypothesized that nucleosome-mediated DNA protection is related to the existence of non-random fragmentation hotspots (HSNRFs), which are defined as regions within the genome where the ends of nucleic acid fragments with a specific size distribution occur more frequently than expected compared to nearby genomic locations.
[0004] Cancer is often found in locations within the human body that are not easily accessible. Invasive surgical biopsy, the "gold standard" for cancer diagnosis, carries significant clinical risks, including bleeding and infection. The drawbacks of such invasive procedures include the fact that the sample taken from tumor tissue represents only a spatially limited representation from the time the procedure was performed. However, cancer does not remain static; it undergoes continuous change, resulting in genetic heterogeneity within the tumor and between primary and metastatic tumors. Much effort has been dedicated to developing non-invasive / minimally invasive methods for cancer diagnosis, monitoring, and therapeutic guidance. Successful developments in non-invasive prenatal testing for numerical abnormalities using cell-free DNA from maternal plasma have also been applicable to the discovery of biomarkers for cancer diagnosis. The discovery of circulating tumor DNA in plasma has offered the possibility of employing liquid-based biopsy testing to detect, predict, and track responses to cancer treatment, without having to address the risks associated with invasive surgical procedures. This technology benefits cancer patients by detecting cancer at an early stage, increasing the chances of successful recovery, assisting in the selection of the most appropriate therapy, and further facilitating the detection of microresidual disease after treatment, thereby assisting clinicians in making necessary medical interventions. Unlike current invasive testing methods that carry the risk of complications, liquid biopsy is inherently safe for patients because it uses samples such as blood, urine, and sputum.
[0005] To date, only a very limited number of methods have been described that attempt to provide estimates of the tumor-derived contribution to the total amount of cell-free DNA (cfDNA) found in plasma, for use as a prognostic biomarker, an indicator of response and / or resistance to therapy, and a disease recurrence (Smith CGet al., Genome Med., 2020, 12(1):23; Peiyong Jiang et al., PNAS, 2018, 115(46):E10925-E10933; Cristiano S. et al., Nature, 2019, 570:385-389; Mouliere et al., Sci.Transl.Med., 2018, 10(466):eaat4921; Newman A. et al., Nat. Med., 2014, 20(5):548-554).
[0006] Current liquid-based biopsy tests are complex and have limited sensitivity and specificity, failing to meet the needs of accurate oncology (De Rubis G. et al., Trends Pharmacol Sci., 2019, 40(3):172-186; Peiyong Jiang et al., Cancer Discov., 2020, CD-19-0622). Therefore, the accuracy of such methods is not sufficiently high and may lead to misleading results. [Overview of the project] [Means for solving the problem]
[0007] This invention provides a solution to the limitations faced by conventional liquid biopsy approaches by expanding the range of information that can be extracted from circulating tumor DNA (ctDNA) sequencing, realizing a novel multi-parameter strategy, and establishing a robust, sensitive, and specific liquid biopsy assay for classifying samples into clinically relevant categories.
[0008] The present invention provides solutions to the accuracy limitations currently faced by other liquid biopsy approaches. The present invention expands the scope of information extractable from the sequencing of cell-free tumor DNA or ctDNA to implement a novel multi-parameter strategy and establish a robust, sensitive, and specific liquid biopsy assay for the classification of samples into clinically relevant categories, thereby overcoming the said accuracy limitations.
[0009] In one embodiment, the present invention relates to a method of classifying a sample as containing cell-free tumor DNA, the method comprising: (i) determining, by alignment to a reference sequence, the start and / or stop sequence coordinates of at least 100,000 cfDNA fragments in a sample comprising a plurality of cell-free DNA (cfDNA) fragments; (ii) a) within, but within a range of 1 to 5 base pairs adjacent to, each start and / or stop sequence coordinate determined in (i), and / or b) outside, but within a range of 1 to 5 base pairs adjacent to, each start and / or stop sequence coordinate determined in (i) determining all nucleic acid motifs composed of trinucleotides, tetranucleotides, and pentanucleotides in the reference sequence; (iii) a) each + and / or - 1 base pair of the sequence coordinates determined in (i) in a plurality of cfDNA fragments contained in the sample, b) each of the nucleic acid motifs determined in (ii) a) and b) in a plurality of cfDNA fragments contained in the sample determining the frequency; (iv) calculating the ratio of each of the frequencies determined in (iii) a) and b) to the corresponding reference frequency; (v) calculating a diagnostic score separately for each ratio determined in step (iv), the score being the weighted sum of each of the respective frequency ratios of step (iv). (vi) A step of calculating a combined diagnostic score from at least two of the diagnostic scores determined in (v), wherein the score is a weighted sum of the two or more diagnostic scores determined in (v), (vii) A step to determine the classification of a sample by comparing the combined diagnostic score with the reference score. A sample is classified as containing tumor cfDNA if its combined diagnostic score is at least one standard deviation higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0010] In one embodiment, the combined diagnostic score is calculated from all the diagnostic scores calculated for each ratio calculated in step (v) of the method described above.
[0011] In one embodiment, the present invention relates to a method for classifying a sample as containing cell-free tumor DNA, wherein the method is (i) In a sample containing multiple cell-free DNA (cfDNA) fragments, the steps of determining the sequence coordinates of the start and / or stop and start and / or stop + and / or -1 base pairs of at least 100,000 cfDNA fragments by alignment to a reference sequence, (ii) A step of determining the frequency of each coordinate determined in (i) in multiple cfDNA fragments contained in the sample, (iii) The step of calculating the ratio of the frequency of each coordinate determined in (ii) to the corresponding reference frequency, (iv) A step of calculating a diagnostic score from all the ratios determined in (iii), wherein the score is a weighted sum of all the frequency ratios determined in (iii), (v) A step in which the classification of the sample is determined by comparing the diagnostic score with the reference score. A sample is classified as containing tumor cfDNA if its diagnostic score is at least one standard deviation higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0012] In one embodiment, the present invention relates to a method for classifying a sample as containing cell-free tumor DNA, wherein the method is (i) In a sample containing multiple cell-free DNA (cfDNA) fragments, the steps of determining the sequence coordinates of the start and / or stop of at least 100,000 cfDNA fragments by alignment to a reference sequence, (ii) A step of determining in the reference sequence all nucleic acid motifs consisting of trinucleotides, tetranucleotides, and pentanucleotides within the range of 1 to 5 base pairs inside, but adjacent to, each start and / or stop sequence coordinate determined in (i), (iii) A step of determining the frequency of each of the nucleic acid motifs determined in (ii) in multiple cfDNA fragments contained in the sample, (iv) The step of calculating the ratio of each frequency determined in (iii) to the corresponding reference frequency, A step of calculating a diagnostic score from all the ratios determined in (v)(iv), wherein the score is a weighted sum of all the frequency ratios determined in (iv), (vi) A step in which the classification of the sample is determined by comparing the diagnostic score with the reference score. A sample is classified as containing tumor cfDNA if its diagnostic score is at least one standard deviation higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0013] In another embodiment, the present invention relates to a method for classifying a sample as containing cell-free tumor DNA, wherein the method is (i) In a sample containing multiple cell-free DNA (cfDNA) fragments, the steps of determining the sequence coordinates of the start and / or stop of at least 100,000 cfDNA fragments by alignment to a reference sequence, (ii) A step of determining in the reference sequence all nucleic acid motifs consisting of trinucleotides, tetranucleotides, and pentanucleotides within the range of 1 to 5 base pairs outside of, but adjacent to, each start and / or stop sequence coordinate determined in (i), (iii) A step of determining the frequency of each of the nucleic acid motifs determined in (ii) in multiple cfDNA fragments contained in the sample, (iv) The step of calculating the ratio of each frequency determined in (iii) to the corresponding reference frequency, A step of calculating a diagnostic score from all the ratios determined in (v)(iv), wherein the score is a weighted sum of all the frequency ratios determined in (iv), (vi) A step in which the classification of the sample is determined by comparing the diagnostic score with the reference score. A sample is classified as containing tumor cfDNA if its diagnostic score is at least one standard deviation higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0014] In one embodiment, the range of base pairs inside, but adjacent to, each start and / or stop sequence coordinate may be 2 bp to 6 bp, or 3 bp to 7 bp, or 4 bp to 8 bp, or 5 bp to 9 bp, or 6 bp to 10 bp from each start and / or stop coordinate.
[0015] In one embodiment, the minimum amount of cfDNA fragments contained in the sample to be analyzed is 100,000 to 500,000, 500,000 to 1,000,000, 1,000,000 to 2,000,000, 2,000,000 to 5,000,000, or 5,000,000 to 10,000,000, or 10,000,000 to 20,000,000, or 20,000,000 to 50,000,000, or 50,000,000 to 500,000,000.
[0016] In one embodiment, the amount of tumor cfDNA in the sample can be classified as low if the combined diagnostic score is 2 to 4 standard deviations of the reference score, as medium if the combined score is 4 to 6.5 standard deviations of the reference score, and as high if the combined score is greater than 6.5 standard deviations of the reference score.
[0017] In one embodiment, the reference sample may be a sample from a patient without cancer, or a patient without recurrence, or a patient with cancer who has been successfully treated.
[0018] In one embodiment, in a sample containing a plurality of cell-free DNA (cfDNA) fragments, step (i) of any of the above methods, in which the start and / or stop sequence coordinates of at least 100,000 cfDNA fragments are determined by alignment to a reference sequence, includes determining the nucleic acid sequences of at least a portion of the plurality of cfDNA fragments in the sample before alignment to a reference sequence.
[0019] In one embodiment, in a sample containing a plurality of cell-free DNA (cfDNA) fragments, step (i) of any of the above methods, in which the sequence coordinates of the start and / or stop of at least 100,000 cfDNA fragments are determined by alignment to a reference sequence, further comprises enriching the cfDNA fragments before determining the nucleic acid sequence of the cfDNA fragments.
[0020] In one embodiment, the sample is classified as containing tumor cfDNA originating from a tumor selected from the group of hematological cancers, liver cancers, lung cancers, pancreatic cancers, prostate cancers, breast cancers, gastric cancers, glioblastomas, colorectal cancers, head and neck cancers, solid tumors, benign tumors, malignant tumors, advanced-stage cancers, metastatic or precancerous tissues.
[0021] In another embodiment, the present invention is (i) A component for carrying out any of the above methods, a) One or more components for isolating cell-free DNA from a biological sample, b) One or more components for preparing and enriching a sequencing library, and / or c) One or more components for amplifying and / or sequencing the enriched library Ingredients containing, (ii) Software for performing statistical analysis Regarding the kit that includes this.
[0022] Twenty normal samples from cancer-free patients and 27 abnormal samples from patients diagnosed with advanced non-small cell lung cancer (NSCLC) or colon cancer were analyzed. In Examples 1-4, 10 randomly selected normal samples and 10 randomly selected abnormal samples were used in the training step to estimate unknown parameters. [Brief explanation of the drawing]
[0023] [Figure 1] The distribution of scores obtained in Examples 1-4 is shown for a “normal” sample (a control sample of healthy, cancer-free individuals not included in the training step) compared to scores obtained by a prior art method (referred to herein as “other” methods) (Peiyong Jiang et al., Cancer Discov., 2020, CD-19-0622). The other methods for measuring the amount of sequence-end motifs of cfDNA fragments contained in the sample being analyzed include the start and / or stop coordinates of the fragments, which differ from the present disclosure that excludes the start and / or stop. Non-significant Kruskal-Wallis rank-sum tests (p-value = 0.9966) suggest that none of the methods are probabilistically superior to the other approaches for a normal sample. The mean of the calculated scores is set to zero for each example. [Figure 2]Examples 1-4 illustrate the score values and their respective distributions obtained by the method of the present invention and by the prior art method (referred to herein as "other" methods) for samples containing acellular tumor ("abnormal") DNA (these samples are not included in the training step). When these scores are compared with scores obtained from normal samples (Figure 1), the maximum distinction is achieved by the method of the present invention in Examples 1-4, clearly illustrating the improved (increased) sensitivity of the method of the present invention (Examples 1-4) compared to the prior art method in distinguishing abnormal samples from normal samples. [Figure 3] Examples 1-4 illustrate the comparison of sensitivity performance between the methods described and prior art methods (referred to herein as "other" methods). Estimated sensitivity for all methods in Examples 1-4 and the prior art methods ("other") was calculated from the empirical distributions of scores for normal and abnormal samples. The specificity (i.e., significance level in statistical hypothesis testing) for all methods was set to 99.9%, resulting in estimated sensitivities for this dataset of 96.8%, 99.94%, 99.48%, and 99.9997% for each of the methods in Examples 1-4. All of the methods of the present invention are significantly superior to prior art methods, which achieve only 84.3% sensitivity, and to other currently available methods in the literature (Mouliere et al. 2018 and Adalsteinsson et al. 2017) (data not shown), which classify samples into clinically notifiable categories using fragment size and copy number variation information, achieving sensitivity in the range of 60%-90%. [Figure 4] Table 1: The table illustrates the scores obtained by the method of the present invention in Example 4 for four additional normal samples and three additional abnormal samples. The abnormal samples were from cancer patients diagnosed with NSCLC (Stage I). The table highlights the classification of ctDNA levels into low, medium, and high. The amount of ctDNA in a sample is classified as low if the combined diagnostic score is between 2 and 4.5, as medium if the combined diagnostic score is between 4.5 and 6, and as high if the combined diagnostic score is greater than 6. [Modes for carrying out the invention]
[0024] This invention describes a liquid biopsy method that establishes a robust, sensitive, and specific liquid biopsy assay for classifying samples into clinically relevant categories by realizing a novel multi-parameter strategy using novel bioinformatics based on an expanded range of information that can be extracted from ctDNA sequencing.
[0025] One embodiment of the present invention relates to a method for classifying a sample as containing cell-free tumor DNA, the method comprising determining the sequence coordinates of the ends or “start and / or stop” of a plurality of cfDNA fragments contained in the sample, and optionally the start and / or stop+ and / or -1 base pairs. The “start and / or stop” of a cfDNA fragment, as herein it refers to the ends, boundaries, or outermost base pairs or nucleotides of a cfDNA fragment. Determination of the sequence coordinates of a cfDNA fragment can be achieved by alignment to a reference sequence, the reference sequence may be a DNA sequence of an organism, preferably a human DNA sequence, such as an hg19 or hg38 human genome sequence, or a genome sequence of a human subject (in one embodiment it may be a healthy or cancer-free human subject).
[0026] In one embodiment of the present invention, the determination of sequence coordinates may include the analysis and / or determination of nucleic acid sequences of a plurality of cfDNA fragments by sequencing analysis or the like. In one embodiment, the determination of sequence coordinates may further include the extraction or purification of nucleic acids and / or specifically cfDNA fragments from a sample and / or the enrichment of cfDNA fragments from a sample and / or the preparation of a sequencing library from isolated DNA, RNA or cfDNA before sequencing analysis.
[0027] The analysis of sequencing data may include the alignment of the obtained cfDNA nucleic acid sequence information to a reference genome sequence. This alignment enables the mapping of the "start and / or stop" or terminal sequence coordinates of the analyzed cfDNA fragment to the reference genome sequence. In a preferred embodiment of the present invention, in addition to the start and / or stop coordinates of the sequenced cfDNA fragment, the sequence coordinates at +1 bp and 1 bp positions from the start and / or stop are also determined from the reference genome sequence.
[0028] Next, the frequencies of each determined start and / or stop sequence coordinate can be determined from multiple cfDNA fragments contained in the sample. All coordinates detected for the same cfDNA fragment (technical duplicate) or for two different cfDNA fragments (biological duplicate) are taken into consideration in the calculation of the frequency (abundance) of each start and / or stop sequence coordinate detected in the multiple cfDNA fragments. In a preferred embodiment of the present invention, in addition to the frequency of each start and / or stop coordinate, the frequencies of each sequence coordinate +1 bp and 1 bp from the start and / or stop coordinate are also determined within the multiple cfDNA fragments in the sample.
[0029] In one embodiment of the present invention, the ratio of the frequency of each determined reference genome coordinate to the corresponding reference frequency is determined. In a preferred embodiment, this ratio of the frequency of the coordinate in the sample to the reference frequency is also calculated for the frequencies of the start and / or stop +1 bp and 1 bp sequence coordinates.
[0030] Next, a diagnostic score can be calculated from all frequency ratios according to the method of the present invention. The diagnostic score is defined as a weighted sum of all frequency ratios obtained as described in Example 1, and an analyzed sample is classified as containing tumor cfDNA if the diagnostic score value is at least one standard deviation of the reference score higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0031] In one embodiment of the present invention, after determining the start and / or stop coordinates of multiple cfDNA fragments contained in a sample, all nucleic acid motifs in a reference sequence, consisting of, for example, trinucleotides (3 consecutive nucleotides), tetranucleotides (4 consecutive nucleotides), and / or pentanucleotides (5 consecutive nucleotides), can be determined within a specific range of base pairs inward from each start and / or stop sequence coordinate, but adjacent to it by 1 bp or more. In one embodiment of the present invention, the specific range of base pairs inward from each start and / or stop sequence coordinate, but adjacent to it by 1 bp or more, may be 1 bp to 5 bp, 2 bp to 6 bp, 3 bp to 7 bp, 4 bp to 8 bp, 5 bp to 9 bp, or 6 bp to 10 bp. In a preferred embodiment, the range inward from each start and / or stop sequence coordinate determined by multiple cfDNA fragments in the sample may be 1 bp to 5 bp. Motifs are extracted from a reference genome sequence to avoid inter-individual variability (i.e., single nucleotide polymorphisms).
[0032] Nucleic acid motifs can be determined based on each detected start and / or stop position in the reference sequence to which the cfDNA fragments are aligned, which is not the actual sequence of the fragments.
[0033] Next, the frequency (abundance) of each detected nucleic acid motif in multiple cfDNA fragments in the sample can be determined. All motifs detected in the same cfDNA fragment or in two different cfDNA fragments are taken into consideration in the calculation of the frequency (abundance) of each motif detected in the multiple cfDNA fragments. Subsequently, the ratio of each nucleic acid motif frequency in the multiple cfDNA fragments to the corresponding reference frequency is calculated. Then, a diagnostic score is calculated from all frequency ratios according to the method of the present invention. The diagnostic score is defined as the weighted sum of all frequency ratios described in Example 2, and the analyzed sample is classified as containing tumor cfDNA if the diagnostic score value is at least one standard deviation of the reference score higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0034] In one embodiment of the present invention, after determining the start and / or stop coordinates of multiple cfDNA fragments contained in a sample, all nucleic acid motifs in a reference sequence, consisting of, for example, trinucleotides (3 consecutive nucleotides), tetranucleotides (4 consecutive nucleotides), and / or pentanucleotides (5 consecutive nucleotides), can be determined within a specific range of base pairs inward from each start and / or stop sequence coordinate, but adjacent to it by 1 bp or more.
[0035] In one embodiment of the present invention, the specific range of base pairs outside each start and / or stop sequence coordinate, but adjacent to it by 1 bp or more, may be 1 bp to 5 bp, 2 bp to 6 bp, 3 bp to 7 bp, 4 bp to 8 bp, 5 bp to 9 bp, or 6 bp to 10 bp. In a preferred embodiment, the range outside each start and / or stop sequence coordinate determined by a plurality of cfDNA fragments in the sample may be 1 bp to 5 bp. A nucleic acid motif may be determined based on each detected start and / or stop position in the reference sequence to which the cfDNA fragment is aligned. Such a nucleic acid motif may consist only of nucleic acid sequences of the reference sequence adjacent to the position to which the cfDNA fragment is aligned by 1 bp or more. Such a motif does not include the nucleic acid sequence of the cfDNA fragment, but includes sequences in the reference sequence that start directly from outside the start or stop coordinate, for example, the start coordinate, and are 1 bp to 5 bp outside the start and / or stop, but adjacent to it.
[0036] Next, the frequencies of each detected nucleic acid motif in multiple cfDNA fragments in the sample can be determined. All motifs detected in the same cfDNA fragment or in two different cfDNA fragments are taken into consideration in the calculation of the frequency (abundance) of each motif detected in the multiple cfDNA fragments. Subsequently, the ratio of each nucleic acid motif frequency in the multiple cfDNA fragments to the corresponding reference frequency is calculated. Subsequently, a diagnostic score can be calculated from all frequency ratios according to the method of the present invention. The diagnostic score is defined as a weighted sum of all frequency ratios as described in Example 3, and an analyzed sample is classified as containing tumor cfDNA if the diagnostic score is at least one standard deviation of the reference score higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0037] In one embodiment of the present invention, a score is calculated from the ratio of the frequencies of (a) the frequencies of the start and / or stop sequence coordinates (arbitrarily -1 bp and / or +1 bp), (b) the frequencies of all nucleic acid motifs located inward with respect to the start and / or stop coordinates of the cfDNA fragment, but adjacent to them by at least 1 bp, and (c) the frequencies of all nucleic acid motifs located outward with respect to the start and / or stop coordinates of the cfDNA fragment, but adjacent to them by at least 1 bp, without containing the cfDNA sequence. All steps of the previously described method may be performed in parallel or in a specific order, and then two or all of the diagnostic score values from steps (a), (b), and (c) may be used to calculate a combined diagnostic score value according to the method of the present invention, as described in Example 4. According to this combined diagnostic score value, an analyzed sample is classified as containing tumor cfDNA or circulating tumor DNA (ctDNA) if the combined diagnostic score value is at least one standard deviation of the reference score higher than the mean of the reference score, and the reference score is calculated from one or more reference values.
[0038] In one embodiment, by comparing the combined diagnostic score obtained for each abnormal sample with the reference score, the amount of tumor cfDNA or ctDNA in the sample can be classified as follows: (a) low if the combined diagnostic score is 2 to 4 standard deviations of the reference score, (b) medium if the combined score is 4 to 6.5 standard deviations of the reference score, and (c) high if the combined score is greater than 6.5 standard deviations of the reference score (Table 1).
[0039] cell-free nucleic acids In this specification, the mixture of nucleic acid fragments is preferably isolated from a sample taken from a eukaryote, preferably a primate, and more preferably a human. The sample may contain cells or nucleic acids from different tissue types. For this reason, the sample may intrinsically contain a mixture of nucleic acid fragments.
[0040] In this specification, “nucleic acid” or “nucleic acid sequence” may be used interchangeably with, but are not limited to, DNA, RNA, genomic DNA, cell-free DNA and / or RNA, as well as tRNA, messenger RNA (mRNA), synthetic DNA, or RNA.
[0041] In relation to the present invention, the terms "nucleic acid fragment" and "fragmented nucleic acid" can be used interchangeably. In preferred embodiments of the method according to the present invention, the nucleic acid fragment is circulating cell-free DNA or RNA.
[0042] In one embodiment, a minimum of 100,000 cfDNA fragments contained in the sample can be analyzed. In another embodiment, the number of cfDNA fragments contained in the sample to be analyzed may be in the range of 100,000 to 500,000, 500,000 to 1,000,000, 1,000,000 to 2,000,000, 2,000,000 to 5,000,000, 5,000,000 to 10,000,000, 10,000,000 to 20,000,000, 20,000,000 to 50,000,000, or 50,000,000 to 500,000,000.
[0043] In one embodiment of the present invention, “sample” is a blood sample, serum sample, plasma sample, liquid biopsy sample, or DNA sample (e.g., a mixture of nucleic acid fragments) containing cell-free DNA (cfDNA), cell-free tumor DNA (cftDNA), circulating tumor DNA (ctDNA), or circulating cftDNA. In connection with the present invention, the terms “cfDNA,” “cftDNA,” “ctDNA,” or “circulating cftDNA” may be used interchangeably.
[0044] In one embodiment, the sample is selected from the group consisting of plasma samples, blood samples, urine samples, sputum samples, cerebrospinal fluid samples, ascites samples, and pleural fluid samples from a subject having or suspected of having a tumor. In one embodiment, the sample or DNA sample is derived from a tissue sample from a subject having or suspected of having a tumor or a group of malignant cells.
[0045] In connection with the present invention, the terms “tumor,” “cancer,” or “abnormal” may be used interchangeably. In this specification, the terms “cancer” or “tumor” may also include tissue or cells of early-stage cancer or advanced cancer, metastasis, or precancerous conditions. In this specification, a tumor sample or abnormal sample may refer to a sample containing (cell-free) DNA or RNA originating from a primary tumor or metastatic tumor. A normal sample or reference sample may refer to a sample containing only (cell-free) DNA or RNA originating from non-cancerous, healthy, or “normal” tissue or cells. In connection with the present invention, the terms “normal,” “control,” or “reference” may be used interchangeably.
[0046] The method of the present invention can be used with a variety of biological samples. Essentially, any biological sample containing genetic material, such as RNA or DNA, particularly cell-free DNA (cfDNA) or cell-free RNA, can be used as a sample by this method, which enables genetic analysis of the RNA or DNA contained therein. For example, in one embodiment, the DNA sample is a plasma sample or blood sample containing cell-free DNA (cfDNA).
[0047] In another embodiment, the sample is a biological sample obtained from a subject who has or is suspected of having a tumor or cancer. In one embodiment, the sample includes circulating acellular tumor DNA (cftDNA). In another embodiment, the sample is urine, sputum, ascites, cerebrospinal fluid, or pleural exudate of the subject. In yet another embodiment, the oncological sample is a subject plasma sample prepared from the subject's peripheral blood. Therefore, since the sample may be a liquid biopsy sample obtained non-invasively from the subject's blood sample, it may potentially enable early detection of cancer before the development of a detectable or palpable tumor, or enable monitoring of disease progression, disease treatment, or disease recurrence.
[0048] In this specification, cell-free DNA (cfDNA) means DNA that is not contained within a cell. Samples may contain cfDNA from normal or healthy cells and / or cancer cells. Cell-free DNA may be released into the blood or serum via secretion, apoptosis, or necrosis. When cfDNA is released from tumor or cancer cells, it may be called cell-free tumor DNA (cftDNA).
[0049] In relation to the present invention, the term “subject” means an animal, preferably a mammal, more preferably a human or a human patient. As used herein, the term “subject” may mean a subject that is suffering from or suspected to have a tumor.
[0050] In this specification, "tumor" means cancer in general, including, but not limited to, solid tumors, adenomas, hematological malignancies, liver cancers, lung cancers, pancreatic cancers, prostate cancers, breast cancers, gastric cancers, glioblastomas, colorectal cancers, head and neck cancers, advanced-stage cancerous tumors, benign or malignant tumors, metastatic or precancerous tissues.
[0051] In this specification, the “ends” of a cfDNA fragment define the outermost nucleotides at the 3' and 5' ends of the nucleic acid fragment, and may also be referred to as the “start and / or stop (position),” “break point,” or “boundary” of the cfDNA fragment. When aligned to a reference sequence, the “(start and / or stop) coordinates” or “sequence coordinates” of a cfDNA fragment are defined by the outermost nucleic acid sequence position in the reference sequence to which the ends of the cfDNA fragment are aligned. For example, if a cfDNA fragment is complementary to or aligned to a reference nucleic acid sequence spanning sequence positions 1500 bp to 1700 bp, the sequence coordinates would be 1500 and 1700 bp, defining the 200 bp length of the cfDNA fragment.
[0052] The size profile of cfDNA, exhibiting a major peak at 166 bp and smaller peaks with 10 bp intervals, suggested that the biological characteristics of cfDNA may be related to nucleosomal organization. Similar patterns were observed in plasma DNA from cancer patients. The non-random fragmentation pattern of cfDNA, related to the tissue of origin, may also be related to the patient's health status. Therefore, the coordinates and frequency of terminal or start and / or stop points of cell-free DNA fragments may serve as indicators of disease progression. They differ depending on the tumor mass, reflecting the origin of the tumor, the extent of the disease, and its response to a given therapy.
[0053] As used herein, the term “inside” from the “start and / or stop” coordinates means the direction from the “start and / or stop” coordinates of the nucleic acid fragment in the reference sequence to which the sequence or motif extends. “Inside” may refer to a nucleic acid fragment sequence or a nucleic acid sequence or motif contained in the reference sequence to which it is aligned. “Inside” may refer to base pairs such as +1, +2, +3, +4, +5 from the start coordinate of the nucleic acid fragment and / or -1, -2, -3, -4, -5 from the stop coordinate. In one embodiment, the range of base pairs inside, but adjacent to, each start and / or stop sequence coordinate may be 1 bp to 5 bp, 2 bp to 6 bp, or 3 bp to 7 bp, or 4 bp to 8 bp, or 5 bp to 9 bp, or 6 bp to 10 bp from each start and / or stop coordinate.
[0054] As used herein, the term “outside” from the “start and / or stop” coordinates means the direction from the “start and / or stop” coordinates of the nucleic acid fragment in the reference sequence to which the sequence extends. “Outside” may refer to the nucleic acid fragment sequence or the nucleic acid sequence or motif contained in the reference sequence to which it is aligned. “Outside” may refer to base pairs such as +1, +2, +3, +4, +5 from the stop coordinate of the nucleic acid fragment and / or base pairs such as -1, -2, -3, -4, -5 from the start coordinate. In one embodiment, the range of base pairs outside, but adjacent to, each start and / or stop sequence coordinate may be 1 bp to 5 bp, 2 bp to 6 bp, or 3 bp to 7 bp, or 4 bp to 8 bp, or 5 bp to 9 bp, or 6 bp to 10 bp from each start and / or stop coordinate.
[0055] Because the observed end site of a fragment may not necessarily be the true cleavage / digestion site, this method analyzes the frequency and / or sequence motifs within ±1 bp of the start and / or stop coordinates (Peiyong Jiang et al., Genome Res., 2020, doi:10.1101 / gr.261396.120). Therefore, by taking into account the likelihood that nearby genomic bases are the true digestion sites, the present invention provides a significant improvement in accuracy compared to prior art in classifying biological samples into clinically relevant categories.
[0056] In this specification, “nucleic acid motif,” “sequence motif,” or “motif” means an array of consecutive nucleotides in a nucleic acid sequence, consisting of consecutive nucleotides such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, etc. This array of consecutive nucleotides may also be called “trinucleotide,” “tetranucleotide,” “pentanucleotide,” “hexanucleotide,” etc. The motif is a subset of human genomic sites that are preferentially cleaved by specific nucleases, etc., when cell-free and / or circulating DNA molecules are generated and released into plasma. Such plasma DNA-end motifs, arising from nucleases that cleave nucleic acids such as DNA during apoptosis, may contain or be specific to HSNRFs and present an identifiable signature. In preferred embodiments, “motif” means an array of 3, 4, or 5 consecutive nucleotides from a reference genome sequence.
[0057] In one embodiment, the nucleic acid motif may be located at the end or break point of the cfDNA fragment, and the motif may be contained within the nucleic acid sequence of the cfDNA fragment, or it may be located outside the boundary of the cfDNA fragment sequence and within the reference nucleic acid sequence (for example, adjacent to the aligned position of the cfDNA fragment).
[0058] cfDNA analysis In this specification, “reference sequence” may be any nucleic acid sequence, genomic sequence, genomic sequence of an organism or subject, preferably a human genome (e.g., HG19 or HG38), or a sequence of a healthy individual or subject.
[0059] In this specification, “reference frequency” for the frequency of start and / or stop sequence coordinates may be the frequency of the corresponding start and / or stop sequence coordinates in one or more reference genomes, reference sequences, or one or more healthy or “normal” control samples, subjects, or patients. In this specification, “reference frequency” for nucleic acid motifs may be the frequency of the corresponding nucleic acid motifs in one or more reference genomes, reference sequences, or one or more healthy or “normal” control samples, subjects, or patients.
[0060] In this specification, “frequency” may be used interchangeably with “abundance” and “incidence.” In one embodiment of the present invention, “frequency” describes, for example, the abundance and incidence or number of nucleic acid sequence motifs, nucleic acid (cfDNA) fragments, or start and / or stop sequence coordinates detected or counted in a plurality of nucleic acids or cfDNA fragments contained in a sample.
[0061] In this specification, “ratio” may mean, for example, a mathematical relationship or proportion of the frequencies of nucleic acid sequence motifs detected in multiple nucleic acid fragments in a sample to the frequency of the same nucleic acid sequence motif in a reference sample. In this specification, a ratio may be calculated by dividing the frequency of each coordinate or motif by the corresponding reference frequency of the corresponding coordinate or motif.
[0062] For sample preparation, nucleic acids such as DNA and / or RNA are extracted from the sample using standard techniques known in the art (non-limited examples include the QIAsymphony (QIAGEN) protocol, QIAamp Circulating Nucleic Acid (QIAGEN), KingFisher (Thermofisher) protocol, MagMAX® Cell-Free DNA (Thermofisher), or any other manual or automated extraction method suitable for cell-free DNA isolation).
[0063] After isolation, the cell-free DNA of the sample can be used for sequencing library preparation to make the sample compatible with downstream sequencing technologies such as next-generation sequencing (NGS). Typically, this involves ligating adapters to the ends of the cell-free DNA fragments. Sequencing library preparation kits are commercially available or can be developed.
[0064] Targeted enrichment of cfDNA is performed using targeted capture sequences (TACS) that bind to target regions on the human genome. Each sequence in the pool is 125–260 base pairs long and / or 125–300 bp long and / or 125–350 bp long, each sequence has a 5' end and a 3' end, and each sequence in the pool binds to a target region at least 10 base pairs away from regions containing copy number variation, segmental duplication, or repeating DNA elements, with both the 5' and 3' ends binding to the target region. The GC content of the TACS is 20%–50% and / or 20%–60% and / or 20%–70% and / or 20%–80%.
[0065] In this specification, the terms “target capture sequence” or “TACS” mean a DNA sequence complementary to a target region on a target genomic sequence, which is used as a “bait” to capture and enrich the target region from a larger sequence library, such as a whole genomic sequencing library prepared from a biological sample. In the context of this invention, the terms “target capture sequence” or “TACS” or “probe” may be used interchangeably.
[0066] In another embodiment, the pool of TACS may include, but is not limited to, AKT1, ALK, APC, AR, ARAF, ATM, BAP1, BARD1, BMPR1A, BRAF, BRCA1, BRCA2, BRIP1, CDH1, CDK4, CDKN2A(pl4ARF), CDKN2A(pl6INK4a), CHEK2, CTNNB1, DDB2, DDR2, DICERl, eGFR, EPCAM, ERBB2, ERBB3, ERBB4, ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ESR1, FANCA, FANCB, FANCC, FANCD2, FANCE, FANCF, FANCG, FANCI, FANCL, FANCM, FBXW7, FGFR1, FGFR2, FLT3, FOXA1, FOXL2, GATA3, G NA11, GNAQ, GNAS, GREM1, HOXB13, IDH1, IDH2, JAK2, KEAP1, KIT, KRAS, MAP2K1, MAP3K1, MEN1, MET, MLH 1, MPL, MRE11A, MSH2, MSH6, MTOR, MUTYH, MYC, MYCN, NBN, NPM1, NRAS, NTRK1, PALB2, PDGFRA, PIK3CA, PI K3CB, PMS2, POLD1, POLE, POLH, PTEN, RAD50, RAD51C, RAD51D, RAF1, RBI, RET, ROS1, RUNX1, SDHA, SDHAF 2, SDHB, SDHC, SDHD, SLX4, SMAD4, SMARCA4, SPOP, STAT, STK11, TMPRSS2, TP53, VHL, XPA, XPC and combinations thereof It binds to multiple target tumor biomarker sequences selected from a group including EGFR_6240, KRAS_521, EGFR_6225, NRAS_578, NRAS_580, PIK3CA_763, EGFR_13553, EGFR_18430, BRAF_476, KIT_1314, NRAS_584, EGFR_12378 and combinations thereof.
[0067] In another embodiment, the TACS pool binds to multiple target tumor biomarker sequences selected from the group including, but not limited to, COSM6240 (EGFR_6240), COSM521 (KRAS_521), COSM6225 (EGFR_6225), COSM578 (NRAS_578), COSM580 (NRAS_580), COSM763 (PIK3CA_763), COSM13553 (EGFR_13553), COSM18430 (EGFR_18430), COSM476 (BRAF_476), COSM1314 (KIT_1314), COSM584 (NRAS_584), COSM12378 (EGFR_12378), and combinations thereof. Here, the identifier means the COSMIC database ID number of the biomarker. Generally, probe hybridization or enrichment steps can be performed before or after creating the sequencing library.
[0068] In one embodiment of the present invention, a sequencing library can be enriched with respect to target sequence regions by hybridizing the library to one or more probes that cover non-random fragmentation hotspots (HSNRFs), etc. Such HSNFR regions are regions that are highly likely to contain within a short distance numerous nucleic acid sequence variations that facilitate the identification of different tissue origin types (e.g., cancer and normal) present in a cfDNA mixture.
[0069] The target region on the target chromosome where the HSNRF is located is enriched by hybridizing a pool of HSNRF capture probes into a sequencing library, followed by isolation of the sequences in the sequencing library that bind to the probes. In one embodiment, the probe straddles the HSNRF site such that only the 5' end of the non-fragmented cell nucleic acid is captured by the probe. In another embodiment, the probe straddles the HSNRF site such that only the 3' end of the non-fragmented cell nucleic acid resulting from the HSNRF is bindable to the probe. In another preferred embodiment, the probe straddles both HSNRF sites associated with the fragmented nucleic acid such that both the 5' and 3' ends of the cell-free nucleic acid associated with a given HSNRF site are captured by the probe.
[0070] To facilitate the isolation of desired enriched sequences (HSNRFs), probe sequences are typically modified to separate sequences that hybridize to the probe from sequences that do not. Typically, this is achieved by immobilizing the probe on a carrier. This allows for the physical separation of probe-binding sequences from sequences that do not bind to the probe. For example, each sequence in a pool of probes can be labeled with biotin, and the pool can then be bound to beads coated with a biotin-binding substance such as streptavidin or avidin. In a preferred embodiment, if the probe is labeled with biotin and bound to streptavidin-coated magnetic beads, separation can be achieved by utilizing the magnetic properties of the beads. However, as will be apparent to those skilled in the art, other affinity binding systems are known in the art and can be used instead of biotin-streptavidin / avidin. For example, an antibody-based system can be used in which the probe is labeled with an antigen and then bound to antibody-coated beads. Furthermore, the probe can incorporate an array tag at one end and can be bound to a carrier via a complementary array on the carrier that hybridizes to the array tag. In addition to magnetic beads, other types of carriers, such as polymer beads or glass, can also be used.
[0071] In certain embodiments, members of the sequencing library that bind to the probe pool are fully complementary to the probes. In other embodiments, members of the sequencing library that bind to the probe pool are partially complementary to the probes. For example, in certain situations, it may be desirable to use and analyze data from DNA fragments (i.e., such DNA fragments are probe-binding due to partial homology) that are products of enrichment processes but do not necessarily belong to the target genomic region, and which, when sequenced, may result in very low coverage across non-probe coordinates throughout the genome.
[0072] After enriching the target sequence with the probe to form an enriched library of DNA containing HSNRF sites, members of the enriched HSNRF library are eluted, amplified, and sequenced using standard methods known in the art. In another embodiment, the probe is supplied with a carrier, such as a biotinylated probe supplied with streptavidin-coated magnetic beads.
[0073] For the detection of tumor biomarkers, probes are designed based on the design criteria described herein and known sequences of tumor biomarker genes and the genetic mutations contained therein that are associated with cancer. In one embodiment, multiple probes used in this method bind to multiple target tumor biomarker sequences. In this case, the probes may be located at non-random fragmentation hotspots adjacent to the mutation sites.
[0074] While next-generation sequencing (NGS) may be used for nucleic acid sequence analysis in this specification, other sequencing techniques that provide highly accurate counting in addition to sequence information are also acceptable. Therefore, other accurate counting methods such as digital PCR, single-molecule sequencing, nanopore sequencing, DNA nanoball sequencing, ligation sequencing, ion semiconductor sequencing, synthetic sequencing, and microarrays can be used instead of NGS, although these are not limited to NGS.
[0075] In one embodiment, the present invention relates to a method for when nucleic acid fragments that are detected or whose origin is determined are present in a mixture at a lower concentration than nucleic acid fragments from the same genetic locus but of different origins.
[0076] This method is particularly suitable for analyzing such low concentrations of target cfDNA. In the method according to the present invention, nucleic acid fragments to be detected or whose origin is determined, and nucleic acid fragments from the same genetic locus but of different origins, are present in a mixture in ratios selected from the group 1:2, 1:4, 1:10, 1:20, 1:50, 1:100, 1:200, 1:500, 1:1000, 1:2000, and 1:5000. The ratios should be understood as approximate ratios meaning ±30%, 20%, or 10%. It will be known to those skilled in the art that such ratios do not occur strictly at the values cited above. The ratios represent the number of rare types of locus-specific molecules relative to the number of abundant types of locus-specific molecules.
[0077] Data Analysis Information obtained from sequencing of enriched libraries is analyzed using an innovative biomathematical / biostatistical data analysis pipeline. Because this method uses a reference genome sequence and may not represent the true digestion site, it utilizes features of cfDNA fragments that include all possible motif combinations adjacent to the terminal coordinates by more than 1 bp, excluding the observed cfDNA terminal region. Furthermore, by combining the analysis of different cfDNA features, including location and motifs, the present invention achieves the unexpected technical effect of improved accuracy, i.e., increased sensitivity at the same specificity level.
[0078] According to a preferred embodiment of the present invention, targeted paired-end next-generation sequencing is performed. Multiplex data for all samples are demultiplexed using the Illumina bcltofastq tool. The sequencing data for the samples are processed using cutadapt software to remove adapter sequences and poor-quality reads (Q score < 25) (Martin, M. et al. 2011 EMB.net Journal 17.1).
[0079] At least 25 nucleotides in length were processed and aligned to the human reference genome build GRCh37(hg19) (UCSC Genome Bioinformatics) using the Burrows-Wheel alignment algorithm (Li, H. and Durbin, R. (2009) Bioinformatics 25:1754-1760). Paired reads with insert sizes exceeding a threshold were removed. The threshold ranged from 100 to 600. Where applicable, duplicate reads were identified after alignment, grouped by unique molecular identifier (UMI) family, and used to generate consensus reads for each UMI family.
[0080] Where applicable, sequencing outputs from the same sample, but processed on separate sequencing lanes, were merged into a single sequencing output file. Duplicate and merging procedures were performed using fgbio, the picard tool software suite (Broad Institute), and the Sammbamba tool software suite (Sambamba reference, Tarasov, Artem, et al. Sammbamba: fast processing of NGS alignment formats. Bioinformatics 31.12(2015):2032-2034). Information regarding mapping locations (outermost and nearest coordinates), base-by-base read depth of target loci, and fragment size was obtained using the mpileup option of the SAMtools software suite (hereinafter referred to as the mpileup file) and processed using custom-built application programming interfaces (APIs) written in Python and the R programming language (Python Software Foundation (2015) Python, The R Foundation (2015) The R Project for Statistical Computing).
[0081] The terminal coordinates of a fragment are defined as the outermost coordinates of the reference genome that the fragment spans. That is, each align fragment has two terminal coordinates: the coordinates of the start / leftmost position (5' end) and the stop / rightmost position (3' end) relative to the reference genome.
[0082] In various embodiments of the present invention, the target panel consisted of a minimum of 500 target genomic bases. The minimum number of fragments required per sample is 100,000.
[0083] In this specification, the “diagnostic score value” is calculated as the weighted sum of all frequency ratios described in Examples 1, 2, and 3 of the “Examples Section.”
[0084] In this specification, the “combined diagnostic score value” is calculated as a weighted sum of at least two frequency ratios from all the steps described in the present invention, as described in Example 4.
[0085] In one embodiment of the present invention, the "reference score" can be calculated from one or more "reference values".
[0086] In one embodiment, a reference value or reference score may be calculated from data obtained from one or more normal or reference samples. In one embodiment, the reference value or reference score and the values of the analysis sample on which it is compared (e.g., the frequency of nucleic acid motifs, the frequency of start and / or stop coordinates) or the diagnostic score of the analysis sample are calculated according to the same calculation method as disclosed herein.
[0087] Sample Classification In this specification, the classification of samples includes binary classification (i.e., cancer, no cancer, good prognosis, poor / poor prognosis, recurrence, non-recurrence) and classification of cftDNA levels into low, medium, and high.
[0088] Clinically relevant categories for sample classification may include the presence or absence of cancer, remission of the disease or cancer, recurrence of the disease or cancer, early cancer stage, and prognosis.
[0089] Inquiry regarding missing translations is underway.
[0090] Oncology use The present invention may be used in the treatment of cancer or for the assessment of tumor burden, detection of microresidual disease, monitoring of treatment outcomes, and long-term monitoring of patient outcomes. The present invention may be further used for the identification of mutations suitable for targeted therapy and for the detection of cancer somatic and germline mutations. This method facilitates the early detection of small tumors that are undetectable by other methods and enables more targeted and customized treatment approaches.
[0091] kit In another embodiment, the present invention provides a kit for carrying out the method of the present invention. In one embodiment, the kit includes a container comprising a pool of probes, as well as software and instructions for carrying out the method.
[0092] In addition to a pool of probes, the kit may include (i) one or more components for isolating cell-free DNA from a biological sample, (ii) one or more components for preparing and enriching a sequencing library (e.g., primers, adapters, buffers, linkers, DNA-modifying enzymes, ligation enzymes, polymerase enzymes, probes, etc.), (iii) one or more components for amplifying and / or sequencing the enriched library, and / or (iv) software for performing statistical analysis. Suitable components for performing the steps referenced in (i), (ii), and (iii) are well known to those skilled in the art.
[0093] In one embodiment, the probe is provided in a form that can be bound to a solid carrier, such as a biotinylated probe. In another embodiment, the probe is provided with a solid carrier, such as a biotinylated probe provided with streptavidin-coated magnetic beads.
[0094] In various other embodiments, the kit may include additional components for carrying out other aspects of the method. For example, in addition to the probe pool, the kit may include (i) one or more components for isolating cell-free DNA from a maternal plasma sample, (ii) one or more components for preparing a sequencing library (e.g., primers, adapters, linkers, restriction enzymes, ligation enzymes, polymerase enzymes), (iii) one or more components for amplifying and / or sequencing the enriched library, and / or (iv) software for performing statistical analysis. Suitable components for carrying out the steps referenced in (i), (ii), and (iii) are well known to those skilled in the art. [Examples]
[0095] Example 1 The start and / or stop (+ and / or -1 base pair) of multiple cfDNA fragments contained in the sample was determined by alignment to a reference sequence. Subsequently, the frequency of each determined start and / or stop sequence coordinate of the multiple cfDNA fragments contained in the sample was determined. The ratio of the frequency of each determined reference genomic coordinate to the corresponding reference frequency was determined, and a weighted sum of all obtained frequency ratios (referred to herein as the “diagnostic score”) was calculated.
[0096] According to one embodiment of the present invention, for each base i (where i = 1, ..., B, where B is equal to the total number of target bases in the panel), the following conditions apply: (A1) Base i has starting position coordinates, or (A2) Base i has a stopping position coordinate, or (A3) Base i has a start--1 base position coordinate, or (A4) Base i has a start + 1 base position coordinate, or (A5) Base i has a stop -1 base position coordinate, or (A6) Base i has a stop + 1 base position coordinate. The total number of map reads that satisfy at least one of the following conditions, is the random variable X. i This was defined.
[0097] Under the null hypothesis (i.e., background model), it is expected that a different, but stationary, number of reads satisfying at least one of conditions A1-A6 will be observed for different bases in the genome. The background probability distribution model for each base is estimated from a group of normal samples. i From the definition, X i ~Bin(x i ;n i ,p i ) is obtained. Here, n i This is equal to the total number of reads that span base i, and p i This is estimated for all i, for example,
number
[0098] After training the background model for each base, the following steps were taken. For each sample k, in one embodiment of the present invention, the following was performed. That is, for each X i , the observed value, for example x i , was compared with the estimated background model for each base. If the p - value, that is, P(X i >x i ) = 1 - P(X i ≦x i ) was less than 0.001, the observed value of X i was divided by the total number of reads spanning base i . That is, Y i =X i / n i , otherwise Y i =0. Subsequently, the sample - specific score was [Number] It is calculated as follows. Here, n2 is Y i This is the total number of bases that have >0. Next, S 0,k The formula is as follows:
number
[0099] Example 2 After determining the start and / or stop (+ and / or -1 base pair) sequence coordinates of the cfDNA fragments, all nucleic acid motifs in the reference sequence of the reference genome were determined. These motifs consisted of trinucleotides, tetranucleotides, and / or pentanucleotides and were within a specific range of base pairs adjacent to the start and / or stop coordinates, but at least one base pair adjacent to them. The frequency ratio of each nucleic acid motif frequency in multiple cfDNA fragments to the corresponding reference frequency was determined, and a weighted sum of all the resulting frequency ratios (referred to herein as the “diagnostic score”) was calculated.
[0100] According to one embodiment of the present invention, for each sample, e.g., k, two sequences are determined for each cfDNA fragment aligned to the hg19 reference genome, the sequences containing the hg19 genome sequence within a range of 1 to 5 base pairs inward from the two ends of the aligned cfDNA fragment (excluding nucleic acid sequences that the fragment spans), and the absolute frequency of all trinucleotides (e.g., ACC, GGT, etc.), tetranucleotides, and pentanucleotide sequence motifs within the sequence, e.g., T ij (Here, i=1, ..., n j And j=3, 4, 5 are the number of nucleotides, and n j The number of all possible J-nucleotide motifs was calculated (n3=64, n4=256, n5=1024). Sample-specific score S 2,k teeth,
number
[0101] In the above equation, D k is the total number of consensus fragments for sample k, and r ij This is calculated from the training dataset of samples that do not contain ctDNA. ij It is a reference value of m ij and s ij This was calculated from a training dataset of samples that did not contain ctDNA.
number
number
[0102] Example 3 After determining the start and / or stop (+ and / or -1 base pair) sequence coordinates of the cfDNA fragments, all nucleic acid motifs in the reference sequence of the reference genome were determined. These motifs consisted of trinucleotides, tetranucleotides, and / or pentanucleotides and were located outside the start and / or stop coordinates, but within a specific range of base pairs adjacent to them by at least one base pair. The frequency ratio of each nucleic acid motif frequency in multiple cfDNA fragments to the corresponding reference frequency was determined, and a weighted sum of all obtained frequency ratios (referred to herein as the “diagnostic score”) was calculated.
[0103] In one embodiment of this method, for each sample, e.g., k, two sequences are determined for each cfDNA fragment aligned to the hg19 reference genome, wherein the sequences include the hg19 genome sequence within a range of 1 to 5 base pairs outward from the two ends of the aligned cfDNA fragment (excluding nucleic acid sequences that the fragment spans), and the absolute frequency of all trinucleotides (e.g., ACC, GGT, etc.), tetranucleotides, and pentanucleotide sequence motifs within the sequence, e.g., T ij (Here, i=1, ..., n j And j=3, 4, 5 are the number of nucleotides, and n j The number of all possible J-nucleotide motifs was calculated (n3=64, n4=256, n5=1024). Sample-specific score S 3,k teeth,
number
[0104] In the above equation, D k is the total number of consensus fragments for sample k, and r ij This is calculated from the training dataset of samples that do not contain ctDNA. ij It is a reference value of m ij and s ij This was calculated from a training dataset of samples that did not contain ctDNA.
number
number
[0105] Example 4 In one embodiment of this method, a weighted sum of at least two of the scores calculated in Examples 1, 2, and 3 was calculated for each sample. This weighted sum will hereafter be referred to as the "combined diagnostic score." The diagnostic score for sample k, for example, DS k This is defined as the weighted average of at least two of the scores described in Examples 1, 2, and 3 above. That is,
number
Claims
1. A method for classifying samples as containing cell-free tumor DNA, (i) In a sample containing multiple cell-free DNA (cfDNA) fragments, the steps of determining the sequence coordinates of the start and / or stop of at least 100,000 cfDNA fragments by alignment to a reference sequence, (ii) a) Within the range of 1 to 5 base pairs adjacent to each start and / or stop sequence coordinate determined in (i), and b) Within the range of 1 to 5 base pairs adjacent to each start and / or stop sequence coordinate determined in (i) The steps include determining all nucleic acid motifs composed of trinucleotides, tetranucleotides, and pentanucleotides in the reference sequence, (iii) a) Each sequence coordinate + and / or -1 base pair determined in (i) in the plurality of cfDNA fragments contained in the sample, b) Each of the nucleic acid motifs determined in (ii) a) and b) in the plurality of cfDNA fragments contained in the sample The step of determining the frequency, (iv) A step of calculating the ratio of each of the frequencies determined in (iii) a) and b) to the corresponding reference frequency, (v) A step of calculating a diagnostic score separately for each ratio determined in step (iv), wherein the score is the respective weighted sum of all the respective frequency ratios in step (iv), (vi) A step of calculating a combined diagnostic score from at least two of the diagnostic scores determined in (v), wherein the score is a weighted sum of the two or more diagnostic scores determined in (v), (vii) The step of determining the classification of the sample by comparing the combined diagnostic score with the reference score. A method comprising: a sample being classified as containing tumor cfDNA if the combined diagnostic score value is at least one standard deviation of the reference score higher than the mean of the reference score, the reference score being calculated from one or more reference values.
2. The method according to claim 1, wherein the combined diagnostic score is calculated from all of the diagnostic scores calculated in step (v) of claim 1.
3. The method according to claim 1 or 2, wherein the range of base pairs inside, but adjacent to, each start and / or stop sequence coordinate is in the range of 2 bp to 6 bp, or 3 bp to 7 bp, or 4 bp to 8 bp, or 5 bp to 9 bp, or 6 bp to 10 bp from each start and / or stop coordinate.
4. The method according to any one of claims 1 to 3, wherein the minimum number of cfDNA fragments contained in the sample to be analyzed is 100,000 to 500,000, 500,000 to 1,000,000, 1,000,000 to 2,000,000, 2,000,000 to 5,000,000, or 5,000,000 to 10,000,000, or 10,000,000 to 20,000,000, or 20,000,000 to 50,000,000, or 50,000,000 to 500,000,000.
5. The method according to any one of claims 1 to 4, wherein the amount of tumor cfDNA in the sample is classified as low when the combined diagnostic score is 2 to 4 standard deviations of the reference score, as medium when the combined score is more than 4 to 6.5 standard deviations of the reference score, and as high when the combined score is more than 6.5 standard deviations of the reference score.
6. The method according to any one of claims 1 to 5, wherein the reference sample is a sample from a patient without cancer, or a patient without recurrence, or a patient with cancer who has been successfully treated.
7. The method according to any one of claims 1 to 6, wherein step (i) is to determine the nucleic acid sequences of at least a portion of the plurality of cfDNA fragments in the sample before alignment to a reference sequence.
8. The method according to any one of claims 1 to 7, further comprising step (i) enriching the cfDNA fragment before determining the nucleic acid sequence of the cfDNA fragment.
9. The method according to any one of claims 1 to 8, wherein the sample is classified as containing tumor cfDNA originating from a tumor selected from the group consisting of hematological cancer, liver cancer, lung cancer, pancreatic cancer, prostate cancer, breast cancer, gastric cancer, glioblastoma, colorectal cancer, head and neck cancer, solid tumor, benign tumor, malignant tumor, advanced stage cancer, metastatic or precancerous tissue.
10. (i) A component for performing the method described in any one of claims 1 to 9, a) One or more components for isolating cell-free DNA from a biological sample, b) One or more components for preparing and enriching a sequencing library, wherein the components for enriching the sequencing library include a pool of non-random fragmentation hotspot (HSNRF) capture probes for hybridization, the HSNRF capture probes binding to multiple tumor biomarkers, and c) One or more components for amplifying and / or sequencing the enriched library. Ingredients containing, (ii) Software for performing statistical analysis A kit that includes this.
Citation Information
Patent Citations
Power transmission gearing
CA306811A
Identification and Use of Circulating Nucleic Acid Tumor Markers
US20160032396A1
Systems and methods for determining tumor fraction in cell-free nucleic acid
WO2019204360A1