Method for detecting cancer using fragment end sequence frequency and size by position of cell-free nucleic acid
By aligning and analyzing the relative sequence frequency and size of nucleic acid fragments with a trained AI model, the method achieves high sensitivity and accuracy in cancer diagnosis, addressing the limitations of invasive and inaccurate current techniques.
Patent Information
- Application Number
- JP2024526514
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-03
- Filing Date
- 2022-11-01
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-11-01
AI Technical Summary
Current cancer diagnostic methods, such as tissue biopsies and tumor markers, are invasive and have low accuracy, while liquid biopsies using cell-free DNA lack effective methods for high sensitivity and specificity in cancer diagnosis based on sequence analysis.
A method involving the extraction of nucleic acids from a biological sample, alignment of sequence information to a reference genome, derivation of relative sequence frequency and size of nucleic acid fragments by position, and inputting this information into a trained artificial intelligence model for cancer diagnosis.
Enables highly sensitive and accurate cancer diagnosis by distinguishing between normal and cancerous samples using the relative sequence frequency and size of nucleic acid fragments, overcoming the limitations of existing methods.
Smart Images

Figure 0007805453000018 
Figure 0007805453000019 
Figure 0007805453000020
Abstract
Description
[Technical Field]
[0001] The present invention relates to a cancer diagnostic method using the relative sequence frequency and size of cell-free nucleic acid fragments by position, more specifically, a cancer diagnostic method using a method in which nucleic acids are extracted from a biological sample, sequence information is obtained, and the relative sequence frequency and size of nucleic acid fragments by position are derived based on aligned reads, and the results are input into a trained artificial intelligence model and the calculated values are analyzed. [Background technology]
[0002] Clinical cancer diagnosis is typically confirmed by a tissue biopsy after a medical history, physical examination, and clinical evaluation. Clinical cancer diagnosis is only possible when the cancer cell count is greater than 1 billion and the tumor is greater than 1 cm in diameter. In this case, the cancer cells already have the ability to metastasize, and at least half of these have already metastasized. Furthermore, tissue biopsies are invasive, causing significant discomfort to patients, and are often not possible during cancer treatment. Additionally, tumor markers are used in cancer screening to monitor substances produced directly or indirectly by cancer. However, their accuracy is limited because more than half of tumor marker screening results are normal even when cancer is present, and they frequently show positive results even when cancer is absent.
[0003] Due to the demand for a relatively simple, non-invasive cancer diagnostic method with high sensitivity and specificity that can overcome the problems of conventional cancer diagnostic methods, liquid biopsy, which utilizes a patient's bodily fluids, has recently been widely used for cancer diagnosis and follow-up testing. Liquid biopsy is a non-invasive method and is a diagnostic technology that is attracting attention as an alternative to conventional invasive diagnostic and testing methods.
[0004] Recently, methods have been developed for diagnosing cancer and differentiating cancer types using cell-free DNA obtained by liquid biopsy (US 10975431, Zhou, Xionghui et al., bioRxiv, 2020.07.16.201350). In particular, methods are known for analyzing motif frequency information in the terminal sequences of cell-free nucleic acids and using it for cancer diagnosis, prenatal diagnosis, or organ transplant monitoring (WO 2020-125709, Peiyong Jiang et al., Cancer Discovery, Vol. 10, 2020, pp. 664-673).
[0005] Meanwhile, the gradient boosting algorithm (GBM) is a predictive model that can perform regression analysis or classification analysis, and is an algorithm that belongs to the boosting family of predictive model ensemble methodologies. The gradient boosting algorithm has shown incredible performance in predicting tabular data (data in an XY grid format, such as Excel), and is known to have the highest predictive performance among machine learning algorithms.
[0006] There are various papers using this gradient boosting algorithm in the biotechnology field (Daping Yu et al., Thoracic Cancer Vol. 11, pp. 95-102. 2020, KR 10-2061800, KR 10-2108050, KR 10-2021-0081547), but there is currently a lack of research on methods for diagnosing cancer through GBM based on sequence analysis information from cell-free DNA (cfDNA) in the blood.
[0007] Therefore, the present inventors have made extensive efforts to solve the above problems and develop an AI-based cancer diagnosis method with high sensitivity and accuracy. As a result, they have confirmed that cancer diagnosis can be performed with high sensitivity and accuracy by selecting an optimal combination of relative sequence frequency and size based on the relative sequence frequency by position of cell-free nucleic acid fragments and size information of nucleic acid fragments, and analyzing this with a trained AI model. This has led to the completion of the present invention. Summary of the Invention [Problem to be solved by the invention]
[0008] An object of the present invention is to provide a method for diagnosing cancer using the relative frequency and size of sequences of cell-free nucleic acid fragments by position. Another object of the present invention is to provide a cancer diagnostic device that uses the relative frequency and size of sequences of cell-free nucleic acid fragments by position.
[0009] It is yet another object of the present invention to provide a computer-readable storage medium containing instructions configured to be executed by a processor to perform cancer diagnosis in the above-described manner.
[0010] To achieve the above object, the present invention provides a method for providing information for cancer diagnosis using cell-free nucleic acids, comprising: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) deriving the relative sequence frequency and size of nucleic acid fragments by position using the aligned sequence information (reads); and (d) inputting the derived relative sequence frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine the presence or absence of cancer, wherein the artificial intelligence model is trained to distinguish between normal samples and cancer samples based on the relative sequence frequency and size information of nucleic acid fragments by position.
[0011] The present invention also provides a cancer diagnostic device including: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment analysis unit that derives the relative sequence frequency by position of nucleic acid fragments and the size of the nucleic acid fragments based on the aligned sequences; and a cancer diagnosis unit that inputs the derived relative sequence frequency by position of nucleic acid fragments and the size information of nucleic acid fragments into a trained artificial intelligence model for analysis and compares the information with a reference value to determine the presence or absence of cancer.
[0012] The present invention also provides a computer-readable storage medium containing instructions configured to be executed by a processor that provides information for cancer diagnosis, the instructions being configured to be executed by a processor, the computer-readable storage medium providing information for cancer diagnosis through the steps of: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) deriving the relative sequence frequency and size of nucleic acid fragments by position using the aligned sequence information (reads); and (d) inputting the derived relative sequence frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine the presence or absence of cancer, wherein the artificial intelligence model in step (d) is trained to distinguish between normal samples and cancer samples based on the relative sequence frequency and size information of nucleic acid fragments by position.
[0013] The present invention also provides a method for diagnosing cancer using cell-free nucleic acids, comprising: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) deriving the relative sequence frequencies and sizes of nucleic acid fragments by position using the aligned sequence information (reads); and (d) inputting the derived relative sequence frequencies and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output value, and comparing it with a cut-off value to determine the presence or absence of cancer, wherein the artificial intelligence model is trained to distinguish between normal samples and cancer samples based on the relative sequence frequencies of nucleic acid fragments by position and the size information of the nucleic acid fragments. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a flowchart showing the overall procedure for carrying out the method for diagnosing cancer using the relative sequence frequency and size of cell-free nucleic acid fragments according to position in the present invention. [Figure 2] 1 illustrates an example of a process for selecting nucleic acid fragment sizes that have statistically significant differences in relative frequency between healthy individuals and cancer patients in accordance with one embodiment of the present invention. [Figure 3] 1 is a graph showing the statistical values of the relative frequency of nucleic acid fragments by size as determined in one embodiment of the present invention and the size distribution of the selected nucleic acid fragments. [Figure 4] FIG. 1 is a diagram visualizing the FESS table created in one embodiment of the present invention in the form of a heat map. [Figure 5] The left panel is an enlarged view of the area indicated by the dotted line in Figure 4, and the two right panels show the results of a statistical analysis of the relative frequency of base sequences at each position. [Figure 6] The results are obtained by calculating the relative frequency of each of the base sequences A, T, G, and C at the position of the nucleic acid fragment selected in one embodiment of the present invention, and statistically confirming the similarity between each of the base sequences. [Figure 7](A) shows the results of checking the performance of the machine learning model constructed in one embodiment of the present invention using accuracy and AUC, and (B) shows the confusion matrix. [Figure 8] This is the result of checking the degree to which the probability values of healthy individuals and neuroblastoma patients predicted by the machine learning model constructed in one embodiment of the present invention match with the actual patients, through the distribution of XPI values output by the machine learning model. [Figure 9] 1 is a graph showing the statistical values of the relative frequency of nucleic acid fragments by size as determined in one embodiment of the present invention and the size distribution of selected nucleic acid fragments at different positions and bases. [Figure 10] These are the results of checking the performance of a machine learning model built with a small number of features according to the importance of the features selected in one embodiment of the present invention. The upper panel shows accuracy, and the lower panel shows AUC (Area Under Curve). DETAILED DESCRIPTION OF THE INVENTION
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs. Generally, the nomenclature used herein and the laboratory procedures described below are those well known and commonly employed in the art.
[0016] Terms such as "first," "second," "A," and "B" may be used to describe various components, but the components are not limited by these terms and are used solely to distinguish one component from another. For example, a first component may be designated a "second component," and similarly, a second component may be designated a "first component" without departing from the scope of the technology described below. The term "and / or" includes a combination of multiple related listed items or any of multiple related listed items.
[0017] In the terms used in this specification, singular expressions should be understood to include plural expressions unless the context clearly indicates otherwise, and terms such as "comprise" should be understood to mean the presence of a stated feature, number, step, operation, component, part, or combination thereof, but not to exclude the possible presence or addition of one or more other features, numbers, step operations, components, parts, or combinations thereof.
[0018] Before describing the drawings in detail, it should be clarified that the division of components in this specification is merely a division according to the main function of each component. That is, two or more components described below may be combined into one component, or one component may be divided into two or more components according to more specific functions. Furthermore, each component described below may additionally perform some or all of the functions performed by other components in addition to its own main function, and some of the main functions performed by each component may be exclusively performed by other components.
[0019] Furthermore, in performing a method or method of operation, the steps constituting the method may be performed out of the order specified unless the context clearly dictates a particular order, i.e., the steps may be performed in the same order as specified, substantially simultaneously, or in the reverse order.
[0020] In the present invention, it was confirmed that cancer diagnosis can be performed with high sensitivity and accuracy by aligning sequence analysis data obtained from a sample to a reference gene, deriving the relative sequence frequency by position of nucleic acid fragments and the size of nucleic acid fragments based on the aligned sequence information, inputting the derived relative sequence frequency by position of nucleic acid fragments and the size information of nucleic acid fragments into a trained artificial intelligence model, and then calculating and analyzing XPI values.
[0021] That is, in one embodiment of the present invention, DNA extracted from blood is sequenced, aligned to a reference chromosome, and then used to derive the relative sequence frequency by position of nucleic acid fragments and the size of nucleic acid fragments. The optimal combination of relative sequence frequency by position of nucleic acid fragments and the size of nucleic acid fragments is then derived, and this is trained into a deep learning model to calculate the XPI value, which is then compared with a reference value to diagnose cancer (Figure 1). Thus, in one aspect, the present invention provides a method for manufacturing a semiconductor device comprising: A method for providing information for cancer diagnosis using cell-free nucleic acids, comprising the steps of: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) deriving the relative sequence frequencies and sizes of nucleic acid fragments by position using the aligned sequence reads; and (d) inputting the derived sequence relative frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine whether or not cancer is present; The artificial intelligence model is characterized by being trained to distinguish between normal samples and cancer samples based on the relative frequency of sequences at each position of nucleic acid fragments and size information of nucleic acid fragments.
[0022] In the present invention, the nucleic acid fragment can be any fragment of nucleic acid extracted from a biological sample, and preferably may be a fragment of cell-free nucleic acid or intracellular nucleic acid, but is not limited thereto.
[0023] In the present invention, the nucleic acid fragment can be obtained by any method known to those skilled in the art, preferably, but not limited to, direct sequencing, next-generation sequencing, non-specific whole genome amplification, or probe-based sequencing.
[0024] In the present invention, the cancer may be a solid cancer or a blood cancer, and may be preferably selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphocytic leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colorectal / rectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, liver cancer, thyroid cancer, gastric cancer, gallbladder cancer, biliary tract cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary, kidney cancer, esophageal cancer, neuroblastoma, and mesothelioma, and more preferably, but not limited to, neuroblastoma.
[0025] In the present invention, The step (a) comprises: (ai) obtaining nucleic acids from a biological sample; (a-ii) removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acids; (a-iii) preparing a single-end sequencing or paired-end sequencing library from the purified nucleic acids or the nucleic acids randomly fragmented by enzymatic cleavage, crushing, or hydroshear method; (a-iv) reacting the prepared library with a next-generation sequencer; and (av) obtaining nucleic acid sequence information (reads) using a next-generation gene sequencer; The method can be characterized by including:
[0026] In the present invention, the step (a) of obtaining sequence information may be characterized by, but is not limited to, obtaining sequence information by whole genome sequencing of the isolated cell-free DNA at a depth of 1 million to 100 million reads.
[0027] In the present invention, the biological sample refers to any substance, biological fluid, tissue, or cell obtained from or derived from an individual, such as whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic juice, etc. The fluid may include, but is not limited to, blood, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, hair, buccal cells, placental cells, cerebrospinal fluid, and mixtures thereof.
[0028] In the present invention, the term "reference population" refers to a reference population that can be compared with a standard base sequence database, and refers to a group of people who are currently free of a specific disease or condition. In the present invention, the standard base sequence in the standard chromosome sequence database of the reference population may be a reference chromosome registered with a public health organization such as NCBI.
[0029] In the present invention, the nucleic acid in step (a) may be cell-free DNA, more preferably circulating tumor cell DNA, but is not limited thereto.
[0030] In the present invention, the next-generation sequencer can be used with any sequencing method known in the art. Sequencing of nucleic acids isolated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of individual nucleic acid molecules or clonally extended proxies for individual nucleic acid molecules in a highly similar manner (e.g., 10 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species in a library can be estimated by measuring the relative occurrence of their cognate sequences in data generated by a sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, incorporated herein by reference.
[0031] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequence of individual nucleic acid molecules (e.g., Helicos BioSciences' HeliScope Gene Sequencing system and Pacific Biosciences' PacBio RS system). In other embodiments, sequencing, such as massively parallel short-read sequencing (e.g., Illumina Inc. Solexa sequencer, San Diego, CA), which generates more bases of sequence per sequence unit than other sequencing methods that generate fewer but longer reads, determines the nucleotide sequence of clonally extended proxies for individual nucleic acid molecules (e.g., Illumina Inc. Solexa sequencer, San Diego, CA; 454 Life Sciences (Branford, CT) and Ion Torrent). Other methods or machines for next-generation sequencing include, but are not limited to, those provided by 454 Life Sciences (Branford, Connecticut), Applied Biosystems (Foster City, California; SOLiD sequencer), Helicos Biosciences Corporation (Cambridge, Massachusetts), and emulsion and microfluidic sequencing technology nano-drip (e.g., GnuBio drip).
[0032] Platforms for next-generation sequencing include, but are not limited to, the Roche / 454 Genome Sequencer (GS) FLX system, the Illumina / Solexa Genome Analyzer (GA), the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, the Polonator G.007 system, Helicos BioSciences' HeliScope Gene Sequencing system, Oxford Nanopore Technologies' PromethION, GriION, MinION systems, and the Pacific Biosciences PacBio RS system.
[0033] In the present invention, the sequence alignment in step (b) includes a computational method or approach used as a computer algorithm to determine identity from possible read sequences in a genome (e.g., short-read sequences from next-generation sequencing) by evaluating the similarity between the most read sequences and a reference sequence. Various algorithms can be applied to the sequence alignment problem. Some algorithms are relatively slow but allow for relatively high specificity. These include, for example, dynamic programming-based algorithms. Dynamic programming is a method of solving complex problems by breaking them down into simpler steps. Other approaches are relatively more efficient but typically not thorough. These include, for example, heuristic algorithms and probabilistic methods designed for large database searches.
[0034] Typically, the alignment process can have two stages: candidate testing and sequence alignment. Candidate testing reduces the search space for sequence alignment from the entire genome to a shorter list of possible alignment positions. As the term suggests, sequence alignment involves aligning sequences with the sequences provided in the candidate testing stage. This can be done using global alignments (e.g., Needleman-Wunsch alignments) or local alignments (e.g., Smith-Waterman alignments).
[0035] Most attribute sorting algorithms can be characterized as one of three types based on indexing methods: hash table (e.g., BLAST, ELAND, SOAP), suffix tree (e.g., Bowtie, BWA), and merge sort (e.g., Slider) based algorithms. Short read sequences are typically used for alignment.
[0036] In the present invention, the alignment step (b) may be performed using, but is not limited to, the BWA algorithm and the Hg19 sequence. In the present invention, the BWA algorithm can include, but is not limited to, BWA-ALN, BWA-SW, or Bowtie2, etc.
[0037] In the present invention, the length of the sequence information (reads) in step (b) is 5 to 5,000 bp, and the number of sequence information used can be, but is not limited to, 5,000 to 5,000,000.
[0038] In the present invention, the method may further include a step of selecting reads having a mapping quality score of the aligned nucleic acid fragments equal to or greater than a reference value before performing step (c). The reference value may be any value that can confirm the quality of the aligned nucleic acid fragments, and may be, but is not limited to, 50 to 70 points, and more preferably 60 points. In the present invention, the size of the nucleic acid fragment in step (c) is the number of bases from the 5' end to the 3' end of the nucleic acid fragment.
[0039] In the present invention, the size of the nucleic acid fragment in step (c) can be used without limitation as long as it is a size that can distinguish between healthy subjects and cancer patients, and may be preferably 90 to 250 bp, and more preferably can be selected from the group consisting of 127 to 129 bp, 137 to 139 bp, 148 to 150 bp, 156 to 158 bp, and 181 to 183 bp, but is not limited to these.
[0040] For example, given a nucleic acid fragment sequenced by paired-end sequencing as follows: Forward strand: 5'-TACAGACTTTGGAAT-3' (SEQ ID NO: 1) Reverse strand: 3'-ATGACTGAAACCTTA-5' (SEQ ID NO: 2) The number of bases from the 5' end to the 3' end of the forward strand, 15, is the size value of the nucleic acid fragment.
[0041] In the present invention, the relative sequence frequency by position of the nucleic acid fragment in step (c) may be characterized as a value obtained by normalizing the number of nucleic acid fragments having A, T, G, and C bases detected at each position in nucleic acid fragments of the same size by the total number of nucleic acid fragments. In the present invention, the position of the nucleic acid fragment in step (c) may be characterized as being 1 to 10 bases from the 5' end of the nucleic acid fragment.
[0042] In the present invention, the relative sequence frequencies by position of the nucleic acid fragment in step (c) may be characterized in that the frequencies of A, T, G, and C bases are at positions 1 to 5 from the 5' end of the nucleic acid fragment, and the frequency of A bases is at positions 6 to 10 from the 5' end of the nucleic acid fragment.
[0043] In the present invention, the relative sequence frequencies and nucleic acid fragment sizes of the nucleic acid fragments in step (c) may be one or more selected from those listed in Table 3, and may preferably be the relative sequence frequencies and nucleic acid fragment sizes of the nucleic acid fragments from Top 1 to Top 5 listed in Table 7, more preferably be the relative sequence frequencies and nucleic acid fragment sizes of the nucleic acid fragments from Top 50 listed in Table 7, and most preferably be the relative sequence frequencies and nucleic acid fragment sizes of the nucleic acid fragments from Top 375 to Top 75. [Table 3] JPEG0007805453000002.jpg193121 JPEG0007805453000003.jpg251166
[0044] In the present invention, the position of a nucleic acid fragment is defined relative to the 5' end of the nucleic acid fragment. For example, the positions of the nucleic acid fragment from the 5' end of the forward strand of SEQ ID NO: 1 can have values of For1, For2, ..., For15, and similarly for the reverse strand. The For1 value of SEQ ID NO: 1 is T, and the Rev1 value of the reverse strand is A. In the present invention, the frequency of a base sequence at each position of a nucleic acid fragment can be calculated by the following process. a) dividing the total nucleic acid fragments into populations of nucleic acid fragments having the same size; b) counting the number of A, T, G, and C bases at each position of the nucleic acid fragment within each group; and c) normalizing the number of bases at each position of the nucleic acid fragment using Equation 2.
number
[0045] In the present invention, it is obvious to those skilled in the art that the size, position, and base in Equation 2 vary depending on the size, position, and base to be normalized.
[0046] In the present invention, the artificial intelligence model in step (d) can be any model that can be trained to distinguish between healthy individuals and cancer patients, and is preferably a machine learning model.
[0047] In the present invention, the artificial intelligence model may be selected from the group consisting of AdaBoost, Random forest, CatBoost, Light Gradient Boosting Model, and XGBoost, but is not limited thereto.
[0048] In the present invention, when the artificial intelligence model is XGBoost and binary classification is learned, the loss function can be characterized as shown in the following Equation 1.
number
[0049] In the present invention, the binary classification means that an artificial intelligence model learns to distinguish between the presence and absence of cancer. In the present invention, when the artificial intelligence model is XGBoost, the learning may be characterized by including the following steps:
[0050] i) classifying relative sequence frequency and size information of nucleic acid fragments by position into training, validation, and test data; In this case, the training data is used to train the XGBoost model, the validation data is used to validate hyper-parameter tuning, and the performance evaluation data is used to evaluate the performance after producing the optimal model. ii) Building an optimal XGBoost model through hyper-parameter tuning and learning process; iii) comparing the performance of multiple models obtained through hyper-parameter tuning using validation data, and determining the model with the best performance on the validation data as the optimal model; In the present invention, the hyper-parameter tuning process is a process of optimizing the values of a plurality of parameters (maximum depth of learner trees, number of learner trees, learning rate, etc.) that constitute an XGBoost model, and the hyper-parameter tuning process can be characterized by using Bayesian optimization and grid search techniques.
[0051] In the present invention, the learning process may be characterized by optimizing the internal parameters (weights) of the XGBoost model using predetermined hyper-parameters, determining that the model is overfitting when the validation loss begins to increase relative to the training loss, and interrupting the model learning before that occurs.
[0052] In the present invention, in step d), the result value obtained by analyzing the relative sequence frequency and size information of the input nucleic acid fragments by position in the artificial intelligence model may be any specific score or real number, and may preferably be, but is not limited to, an XPI (XGBoost Probability Index) value.
[0053] In the present invention, the XGBoost probability index means a value obtained by adjusting the output of an artificial intelligence model on a scale of 0 to 1 and expressing it as a probability value. In the case of binary classification, a sigmoid function is used to train the model so that the XPI value for cancer is 1. For example, when a neuroblastoma sample and a normal sample are input, the model is trained so that the XPI value for the neuroblastoma sample approaches 1 and that for the normal sample approaches 0.
[0054] In the present invention, the artificial intelligence model was trained so that the output result would be close to 1 if cancer was present and close to 0 if cancer was not present. Performance was measured based on the standard of 0.5, with a value of 0.5 or higher, determining that cancer was present, and a value of 0.5 or lower determining that cancer was not present (accuracy of learning, verification, and performance evaluation).
[0055] It is obvious to any ordinary engineer that the reference value of 0.5 is a value that can change at any time. For example, if you want to reduce false positives, you can set a reference value higher than 0.5 and set stricter standards for determining whether cancer is present, and if you want to reduce false negatives, you can measure a lower reference value and set a slightly weaker standard for determining whether cancer is present.
[0056] Most preferably, a trained artificial intelligence model can be used to apply unknown data (data for which the answers are known and not trained) to check the probability of the XPI value and determine a reference value. In another aspect, the present invention includes a decoding unit that extracts nucleic acid from a biological sample and decodes sequence information; an alignment section that aligns the decoded sequences to a standard chromosome sequence database; a nucleic acid fragment analysis unit that derives relative sequence frequencies and nucleic acid fragment sizes by position of the aligned sequences; and a cancer diagnosis unit that inputs the derived relative frequency of sequences by position of nucleic acid fragments and size information of nucleic acid fragments into a trained artificial intelligence model for analysis and compares the results with reference values to determine the presence or absence of cancer; The present invention relates to a cancer diagnostic device including:
[0057] In the present invention, the decoding unit may include a nucleic acid injection unit that injects nucleic acids extracted from an independent device; and a sequence information analysis unit that analyzes the sequence information of the injected nucleic acids, and may preferably be, but is not limited to, an NGS analysis device. In the present invention, the decoding unit may receive and decode sequence information data generated by an independent device.
[0058] In yet another aspect, the present invention provides a computer-readable storage medium including instructions configured to be executed by a processor to provide information for diagnosing cancer, (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive relative sequence frequencies by position of nucleic acid fragments and sizes of nucleic acid fragments; and (d) inputting the derived sequence relative frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine whether or not cancer is present; The artificial intelligence model in step (d) is trained to distinguish between normal samples and cancer samples based on the relative frequency of sequences at each position of nucleic acid fragments and size information of the nucleic acid fragments, and provides information for cancer diagnosis through this step. The present invention relates to a computer-readable storage medium containing instructions configured to be executed by a processor to provide information for cancer diagnosis through this step.
[0059] In still another aspect, the present invention relates to a method for diagnosing cancer using cell-free nucleic acids, comprising the following steps: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) deriving the relative sequence frequencies and sizes of nucleic acid fragments by position using the aligned sequence reads; and (d) inputting the derived sequence relative frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine whether or not cancer is present; The artificial intelligence model is characterized by being trained to distinguish between normal samples and cancer samples based on the relative frequency of sequences at each position of nucleic acid fragments and size information of nucleic acid fragments.
[0060] In another aspect, the method of the present application can be implemented using a computer. In one embodiment, the computer includes one or more processors coupled to a chipset. The chipset is also coupled to a memory, a storage device, a keyboard, a graphics adapter, a pointing device, and a network adapter. In one embodiment, the chipset's functionality is enabled by a memory controller hub and an I / O controller hub. In another embodiment, the memory can be directly coupled to the processor instead of the chipset. The storage device can be any device capable of maintaining data, including a hard drive, a CD-ROM (Compact Disk Read-Only Memory), a DVD, or other memory device. The memory contains data and instructions used by the processor. The pointing device can be a mouse, trackball, or other type of pointing device, and is used in combination with a keyboard to transfer input data to the computer system. The graphics adapter displays images and other information on a display. The network adapter is coupled to the computer system via a local or long-distance communication network. The computer used in the present application is not, however, limited to the above configuration, and may include none or additional configurations, may be part of a storage area network (SAN), and the computer of the present application may be configured to execute modules in a program for implementing the method of the present application.
[0061] In this application, the term "module" may refer to a functional and structural combination of hardware for implementing the technical idea of the present application and software for driving the hardware. For example, the term "module" may refer to a logical unit of predetermined code and hardware resources for executing the predetermined code, and it is obvious to those skilled in the art that the term "module" does not necessarily refer to physically connected code or to a single type of hardware.
[0062] The methods of the present application may be implemented in hardware, firmware, or software, or a combination thereof. When implemented in software, the storage medium includes any medium that stores or transmits information in a form readable by a device such as a computer. For example, computer-readable media include read-only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices, and other electrical, optical, or acoustic signal transmission media. [Example]
[0063] The present invention will be described in more detail below through examples. It will be obvious to those skilled in the art that these examples are merely for the purpose of illustrating the present invention and should not be construed as limiting the scope of the present invention.
[0064] Example 1. Extracting DNA from blood and conducting next-generation base sequence analysis Ten milliliters of blood was collected from 202 healthy volunteers and 61 neuroblastoma patients and stored in EDTA tubes. Within two hours of collection, the plasma fraction was centrifuged at 1200 g for 15 minutes at 4°C. The resulting plasma was then centrifuged at 16000 g for 10 minutes at 4°C to separate the precipitate. Cell-free DNA was extracted from the isolated plasma using the Chemagic ccfNA 2K Kit (Chemagen). Library preparation was performed using the MGI Easy cell-free DNA library prep set kit, followed by sequencing on a DNBseq G400 instrument (MGI) in 100-base paired-end mode. Approximately 170 million reads were generated per sample.
[0065] Example 2. Selection of optimal nucleic acid fragment position-specific sequence relative frequency and nucleic acid fragment size 2-1. Definition and measurement of the relative frequency of nucleic acid fragment positions and base sequences The position of the nucleic acid fragment was defined based on the 5' end of the nucleic acid fragment. The reads obtained in Example 1 were paired-end sequence reads and 100 bp in length, so the positions of For1, For2, ..., For100 from the 5' end were set for the forward strand, and the positions of Rev1, Rev2, ..., Rev100 from the 5' end were also set for the reverse strand. The nucleic acid fragments were assembled using the bamtobed-bedpe option of the bedtools program.
[0066] To briefly explain the process of determining the relative frequency of base sequences at each position in nucleic acid fragments, first, 17M reads were arbitrarily selected and downsampled from the sequence data of approximately 170M reads produced in Example 1, and then QC filtering was performed to count the number of nucleic acid fragments that satisfied the combination of size, position, and base (e.g., Size120_For1_A).Then, the number was divided by the total number of sequence reads remaining after the above 3 QC filtering and normalized.
[0067] More specifically, the method was as follows. 1. The total nucleic acid fragments were divided into groups of nucleic acid fragments having the same size, for example, a group of 101 nucleic acid fragments, a group of 150, ..., a group of 200, etc. 2. The number of A, T, G, and C bases at each nucleic acid fragment position within each group was counted. For example, if the number of bases at each nucleic acid fragment position in a group of 120 nucleic acid fragments is counted, the results can be organized as shown in Table 1 below. [Table 1]
[0068] Interpreting the above table, this means that there were a total of 23,135 nucleic acid fragments with a size of 120, of which 5,683, 4,680, 4,194, and 8,566 nucleic acid fragments had A, T, G, and C bases at the For1 position, respectively. 3. After counting the number of bases at each nucleic acid fragment position in the above process, the total number of sequenced reads (the total number of reads produced for the analysis target, regardless of nucleic acid fragment size; in Example 1, 15,063,130) was divided and normalized using Equation 2 to create a FESS (Fragment End Sequence frequency and Size) table, which calculates the relative frequency as shown in Table 2 (FESS_Table_120) below.
number
[0069] 2-2. Selection of optimal nucleic acid fragment size The position and base sequence of the nucleic acid fragment to be analyzed were fixed to (For1_A), and the following analysis was carried out. The Kruskal-Wallis test was used to statistically confirm whether there was a difference in the relative frequency distribution of (For1_A) between healthy subjects and neuroblastoma patients, while changing the nucleic acid fragment size in increments of 1. Specifically, as shown in Figure 2, in the nucleic acid fragment population with a size of 118, it was confirmed that the relative frequency of (For1_A) was distributed at a statistically significant level higher in the neuroblastoma patient group than in the healthy subjects. Using the same method, it was confirmed that the relative frequency of (For1_A) was distributed with no significant difference between the two groups in the nucleic acid fragment population with a size of 168, and that the relative frequency of (For1_A) was distributed at a statistically significant level lower in the healthy subjects than in the neuroblastoma patient group in the nucleic acid fragment population with a size of 185. Using this method, the relative frequency difference of (For1_A) between healthy individuals and neuroblastoma patients was statistically confirmed (p-value) while changing the nucleic acid fragment size from 101 to 200.
[0070] As a result, the X-axis of Figure 3 represents the size of nucleic acid fragments, and the Y-axis represents -log 10 The graph shows the value of (p), where a larger value on the Y axis means a larger difference between healthy subjects and neuroblastoma patients. As shown in Figure 3, the difference in (For1_A) frequency between healthy subjects and neuroblastoma patients widens significantly with a period of about 10 nucleic acid fragment sizes (-log 10 (p) value peaks and then declines.
[0071] Furthermore, these patterns are repeated not only in the training dataset but also in an independent validation dataset, confirming that the nucleic acid fragment sizes shown in green are not coincidental patterns that have been overfitted to the training dataset.
[0072] The two datasets have a common -log 10 Nucleic acid fragment sizes showing peaks in (p) value were selected (127-129, 137-139, 148-150, 156-158, 181-183), and a total of 15 nucleic acid fragment sizes were selected. Furthermore, it was confirmed that similar patterns were observed for other bases at other positions (Figure 9).
[0073] 2-3. Selection of optimal nucleic acid fragment position The data obtained in Example 1 is 100 PE data, so there are a total of 200 types of nucleic acid fragment positions that can be used for analysis, from For1 to 100 and Rev1 to 100.
[0074] Figure 4 visualizes FESS_Table_120 in Table 2 in heatmap format. It shows that differences in the relative frequencies of A, T, G, and C base sequences are observed only in the part of both ends indicated by the dotted lines (For1-10, Rev1-10), and that the relative frequencies of nearly similar A, T, G, and C base sequences are repeated toward the end of the read (up to 100).
[0075] For example, the relative frequency of A, T, G, C base sequences in For1 shows a significant difference from the relative frequency of A, T, G, C in For2, but it can be confirmed that the relative frequency of A, T, G, C base sequences in For11, the relative frequency of A, T, G, C base sequences in For99, and the relative frequency of A, T, G, C base sequences in For100 are almost similar with no significant difference.
[0076] Therefore, to improve the performance of the learning model, only the positions of For1 to 10 and Rev1 to 10, excluding the positions of the latter part of the lead, were selected as features to be trained on the model.
[0077] Furthermore, if we enlarge the area indicated by the dotted line in Figure 4, which is the same as Figure 5 (Rev1-10 were aligned in reverse order from Rev10-1), we can see in the leftmost panel that the relative frequencies of the same sequence at the same position in the forward and reverse directions are quite similar to each other.
[0078] For example, (For1_A and Rev1_A), (For1_T and Rev1_T), (For1_G and Rev1_G), and (For1_C and Rev1_C) have similar relative frequency values to each other, and in the same way (For2_A and Rev2_A), (For2_T and Rev2_T), (For2_G and Rev2_G), and (For2_C and Rev2_C) have similar relative frequency values to each other.
[0079] When this similarity was measured using Pearson's correlation coefficient in a healthy population, the results were the same as those shown in the two right panels of Figure 5. We confirmed that the similarities measured in the healthy population between the relative frequency values of For1_A and Rev1_A, between the relative frequency values of For1_T and Rev1_T, between the relative frequency values of For1_G and Rev1_G, and between the relative frequency values of For1_C and Rev1_C were all 1.
[0080] Through this analysis, it was confirmed that the relative frequency of the 5'-terminal base sequence on the forward strand of the nucleic acid fragment was similar to the relative frequency of the 5'-terminal base sequence on the reverse strand. Therefore, only the For1 to 10 positions, excluding the Rev1 to 10 positions, were selected as features to be used for model training.
[0081] 2-4. Selection of optimal base sequences for nucleic acid fragments At the 10 positions selected in Example 2-3, the relative frequencies of four types of base sequences, A, T, G, and C, can be calculated. For example, at the For1 position, the relative frequencies of For1_A, For1_T, For1_G, and For1_C can be calculated. To reduce the number of variables to be trained on the model, the similarity between base sequences at the same position was confirmed and additional selection was performed. Selection of base sequences by position was performed in a healthy population using the following method. Calculate the relative frequency of A, T, G, and C base sequences at each position from For1 to For10. · Similarity between (For1_A and For1_T), (For1_A and For1_G), (For1_A and For1_C), (For1_T and For1_G), (For1_T and For1_C), and (For1_G and For1_C) was measured by Pearson's correlation coefficient.
[0082] As a result, as shown in Figure 6, it was confirmed that at positions For1 to 5, there was low similarity between the relative frequencies of the four types of base sequences, A, T, G, and C, while at positions For6 to 10, there was fairly high similarity between the relative frequencies of the four types of base sequences, A, T, G, and C.
[0083] Therefore, at positions For1 to 5, all four types of base sequences, A, T, G, and C, were selected, and at positions For6 to 10, only the A base sequence was selected as a representative value among A, T, C, and G.
[0084] In conclusion, the optimal nucleic acid fragment sizes and relative sequence frequencies by position are as follows: 1) Nucleic acid fragment sizes: 127, 128, 129, 137, 138, 139, 148, 149, 150, 156, 157, 158, 181, 182, 183. 15 in total. 2) Position of nucleic acid fragment: For1 to 10. Total 10. 3) Combination of base sequences by position of nucleic acid fragments For1~5:A, T, G,C For6~10:A 15 size * 25 position_base sequence = 375 features The 375 feature combinations are listed in Table 3.
[0085] Example 3. Machine learning model construction and learning process A machine learning model for classifying healthy subjects and neuroblastoma patients was trained using the relative frequency values of the 375 features selected in Example 2 as input. XGBoost was used as the machine learning algorithm.
[0086] The entire sample was divided into training, validation, and performance evaluation datasets. The training dataset was used for model training, the validation dataset for hyper-parameter tuning, and the performance evaluation dataset for evaluating the performance of the final model. The number of samples in each set is as follows: [Table 4]
[0087] The hyper-parameter tuning process is a process of optimizing the values of various parameters that make up the XGBoost model (maximum depth of the learner tree, number of learner trees, learning rate, etc.).
[0088] The hyper-parameter tuning process used Bayesian optimization and grid search techniques, and when the validation loss began to increase relative to the training loss, the model was determined to be overfitting and model training was discontinued.
[0089] The performance of multiple models obtained through hyper-parameter tuning was compared using a validation dataset, and the model with the best performance on the validation dataset was determined to be the optimal model, and final performance evaluation was performed on the performance evaluation dataset.
[0090] When a vector of relative frequency values of 375 features calculated from an arbitrary sample was input into the XGBoost model created through the above process, the probability that the sample was a healthy individual or a neuroblastoma patient was calculated, and this probability value was defined as the XGBoost Probability Index (XPI). If the calculated XPI value for any sample exceeded 0.5, it was determined to be a neuroblastoma patient, and if it was 0.5 or less, it was determined to be a healthy individual.
[0091] Example 4. Performance verification of the constructed model 4-1. Performance check The performance of the XPI values output by the machine learning model constructed in Example 3 was tested. All samples were divided into training, validation, and performance evaluation groups, and after constructing a model using the training samples, the performance of the model created using the training samples was confirmed using samples from the validation and performance evaluation groups.
[0092] [Table 5]
[0093] As a result, as shown in Table 5 and Figure 7, the accuracy was confirmed to be 1.000, 0.945, and 0.937 for the learning, validation, and performance evaluation groups, respectively, and the AUC values, which are the results of the ROC analysis, were confirmed to be 1.000, 0.952, and 0.987 for the learning, validation, and performance evaluation groups, respectively.
[0094] 4-2. Checking XPI distribution We confirmed how closely the XPI values, which are output values of the machine learning model constructed in Example 3, matched those of actual patients. The X axis of Fig. 8 shows the group (True label) information of the actual sample, and the Y axis shows, from left to right, the XPI values calculated by the machine learning model for healthy subjects (Normal) and neuroblastoma patients (NBT).
[0095] As a result, as shown in Figure 8, the XPI distribution confirmed that for all training, validation, and performance evaluation datasets, healthy subject samples had the highest probability of being healthy subjects, and neuroblastoma patient samples had the highest probability of being liver cancer patients.
[0096] Example 5. Checking model performance by feature 5-1. Deriving feature importance A learning model was constructed in Example 3 using the features selected in Example 2, and when the XGB model was learned using each feature, the importance value of each feature was as shown in Table 6 below.
[0097] [Table 6] JPEG0007805453000012.jpg253161 JPEG0007805453000013.jpg253161 JPEG0007805453000014.jpg253161 JPEG0007805453000015.jpg139166
[0098] 5-2. Checking TopN feature performance The performance of the XGB model constructed using only the top 1 feature using the method of Example 3, the model using features up to 2, and the XGB model constructed using features 3, 4, 5, 6, 7, 8, 9, 15, 20, 25, 30, 35, 40, 45 and 50 was confirmed using the method of Example 4.As a result, it was confirmed that sufficient performance was achieved even when using the five top features, as shown in Table 7 and Figure 10. [Table 7]
[0099] That is, the top three rows of Table 7 are the results of measuring performance using the accuracy (ACC) method, and the bottom three rows are the results of measuring performance using the AUC method. The training, validation, and performance evaluation sets used to measure ACC and AUC performance are composed of the same components. Accuracy (ACC) is a performance indicator that measures whether the probability values predicted by the model are higher or lower than a set cutoff value (cutoff = 0.5). Unlike ACC, AUC does not set a specific cutoff, but rather measures how clearly the distribution of predicted probability values differs between normal and cancer patient populations.
[0100] In the case of ACC, the results may vary depending on how the cutoff value is set, so it is correct to interpret the AUC value as the standard. Using the AUC value of the performance evaluation set as the standard, the results in Table 7 are as follows: i) When all 375 features are used, the AUC is 0.987, showing the best performance when compared to using a small subset of features. ii) When we search for the smallest number of features that can ensure similar performance evaluation AUC performance as when 375 features are used, we find that TopN=5.
[0101] Although certain parts of the present invention have been described in detail above, it will be apparent to those skilled in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the true scope of the present invention is to be defined by the appended claims and their equivalents. [Industrial Applicability]
[0102] The cancer diagnosis method using the relative sequence frequency and size of cell-free nucleic acid fragments by position according to the present invention is useful because it obtains optimal relative sequence frequency and size information by position of nucleic acid fragments and analyzes it using an AI algorithm, thereby showing high sensitivity and accuracy even with low read coverage.
Claims
1. (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) deriving the relative sequence frequencies and sizes of nucleic acid fragments by position using the aligned sequence reads; and (d) inputting the derived sequence relative frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine whether or not cancer is present; the artificial intelligence model is trained to distinguish between normal samples and cancer samples based on the relative frequency of sequences at each position of nucleic acid fragments and size information of nucleic acid fragments; Including, The method for providing information for cancer diagnosis using cell-free nucleic acids, wherein the relative sequence frequency by position of the nucleic acid fragment in step (c) is a value obtained by normalizing the number of nucleic acid fragments having A, T, G, and C bases detected at each position in nucleic acid fragments of the same size by the total number of nucleic acid fragments.
2. The method according to claim 1, wherein step (a) is carried out by a method comprising the steps of: (ai) obtaining nucleic acids from the biological sample; (a-ii) removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acids; (a-iii) preparing a single-end sequencing or paired-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, crushing, or hydroshear method; (a-iv) subjecting the prepared library to a next-generation sequencer; and (av) A step of obtaining nucleic acid sequence information (reads) using a next-generation gene sequence testing machine.
3. 2. The method of claim 1, wherein the size of the nucleic acid fragment in step (c) is selected from the group consisting of 127-129 bp, 137-139 bp, 148-150 bp, 156-158 bp, and 181-183 bp.
4. The method of claim 1, wherein the position of the nucleic acid fragment in step (c) is 1 to 10 bases at the 5' end of the nucleic acid fragment.
5. The method of claim 1, wherein the relative sequence frequencies by position of the nucleic acid fragment in step (c) are the frequencies of A, T, G, and C bases at positions 1 to 5 from the 5' end of the nucleic acid fragment, and the frequency of A bases at positions 6 to 10 from the 5' end of the nucleic acid fragment.
6. The method of claim 1, wherein the relative sequence frequency by position of the nucleic acid fragments and the size of the nucleic acid fragments in step (c) are at least one selected from those listed in Table 3.
7. The method of claim 1, wherein the artificial intelligence model in step (d) is selected from the group consisting of AdaBoost, Random forest, CatBoost, Light Gradient Boosting Model, and XGBoost.
8. The method according to claim 7, wherein the artificial intelligence model is XGBoost and when learning binary classification, the loss function is expressed by the following Equation 1: [Equation 1]
9. 2. The method of claim 1, wherein the result value output by the artificial intelligence model in step (d) after analyzing the input sequence relative frequency and size information is an XPI (XGBoost Probability Index) value.
10. The method according to claim 1, wherein the reference value for step (d) is 0.5, and a value of 0.5 or more is determined to be cancer.
11. a decoding unit that extracts nucleic acids from the biological sample and decodes the sequence information; an alignment section for aligning the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment analysis unit that derives relative positional sequence frequencies and nucleic acid fragment sizes of the aligned sequences based on the nucleic acid fragments; and a cancer diagnosis unit that inputs the derived relative sequence frequency of each position of the nucleic acid fragment and the size information of the nucleic acid fragment into a trained artificial intelligence model, analyzes the data, and compares the data with a reference value to determine whether or not cancer is present; Including, A cancer diagnostic device characterized in that the relative sequence frequency by position of the nucleic acid fragment is a value obtained by normalizing the number of nucleic acid fragments having A, T, G, and C bases detected at each position in nucleic acid fragments of the same size by the total number of nucleic acid fragments.
12. A computer-readable storage medium comprising instructions configured to be executed by a processor for cancer diagnosis, the instructions comprising: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive relative sequence frequencies by position of nucleic acid fragments and sizes of nucleic acid fragments; and (d) inputting the derived sequence relative frequency and size information into an artificial intelligence model trained to diagnose cancer, analyzing the output result, and comparing it with a cut-off value to determine whether or not cancer is present; The artificial intelligence model of step (d) is trained to distinguish between normal samples and cancer samples based on the relative frequency of sequences at each position of nucleic acid fragments and size information of nucleic acid fragments, A computer-readable storage medium comprising instructions configured to be executed by a processor for cancer diagnosis, wherein the relative sequence frequency by position of the nucleic acid fragment in step (c) is a value obtained by normalizing the number of nucleic acid fragments having A, T, G, and C bases detected at each position in nucleic acid fragments of the same size by the total number of nucleic acid fragments.
Citation Information
Patent Citations
Models for targeted sequencing
JP2021503922A
Methods and systems for analyzing nucleic acid molecules
JP2023501376A
Significance modeling of clonal absence of targeted variants
JP2023512239A
Method for classifying genetic mutations detected in cell-free nucleic acids as of tumor or non-tumor origin
JP2023517029A
Circulating Tumor DNA Detection Method Using Sample comprising Cell free DNA and Uses thereof
KR1020190085667A