Method for diagnosing and predicting cancer type based on single nucleotide variant in cell-free DNA
A method for cancer diagnosis and type prediction using single nucleotide variants in cfDNA involves extracting and filtering mutations, then using AI models to achieve high sensitivity and accuracy, addressing the limitations of existing cfDNA WGS and AI-based methods.
Patent Information
- Application Number
- JP2024573381
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-15
- Filing Date
- 2023-06-15
- Publication Date
- 2025-07-23
AI Technical Summary
Current methods for cancer diagnosis using cell-free DNA (cfDNA) WGS lack accuracy in mutation discovery and are not effectively applied for cancer diagnosis due to the absence of a reliable filtering method, and existing AI-based methods for predicting cancer types from cfDNA sequencing data are inaccurate.
Extract cancer-specific single nucleotide mutations from a biological sample, calculate their distribution and frequency, and input these values into trained artificial intelligence models for precise cancer diagnosis and type prediction.
The method achieves high sensitivity and accuracy in cancer diagnosis and type prediction, comparable to tissue-based methods, using AI models trained on single nucleotide variant data from cfDNA.
Smart Images

Figure 2025523429000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cancer diagnosis and cancer type prediction method using single nucleotide variants of cell-free nucleic acids. More specifically, nucleic acids are extracted from a biological sample, sequence information is obtained and aligned reads are used to extract cancer-specific single nucleotide variants through filtering. Then, the distribution of single nucleotide variants and the frequency by type of single nucleotide variant are calculated and input into a learned artificial intelligence model, and the output value is analyzed. The present invention relates to a cancer diagnosis and cancer type prediction method using single nucleotide variants of cell-free nucleic acids, including the above method.
Background Art
[0002] In clinical cancer diagnosis, tissue biopsy is usually performed and confirmed after medical history taking, physical examination, and clinical evaluation. Cancer diagnosis by clinical experiment is possible only when the number of cancer cells is 1 billion or more and the diameter of the cancer is 1 cm or more. In this case, the cancer cells already have the ability to metastasize, and at least half of them are already in a metastasized state. In addition, since tissue biopsy is invasive, it causes considerable discomfort to the patient, and there is often a problem that tissue biopsy cannot be performed when treating cancer patients. In addition, in cancer screening, tumor markers for monitoring substances produced directly or indirectly from cancer are used. However, even when cancer is present, more than half of the results of tumor marker screening are shown to be normal, and even when there is no cancer, it is frequently shown to be positive, so its accuracy has limitations.
[0003] Studies on diagnosing cancer through single nucleotide variant analysis of cell-free DNA have been actively conducted. Methods that target recurrent mutations frequently found in cancer by increasing the sequencing depth for targeted sequencing have been widely used (Chabon JJ et al., nature, Vol. 580, pp. 245-251, 2020). However, recently, it has been revealed that even with a lower sequencing depth than targeted sequencing, it is more sensitive to examine more types of mutations using whole-genome sequencing (WGS) data of cell-free DNA (Zviran A et al., Nat Med, Vol. 26, pp. 1114-1124, 2020).
[0004] However, with current technologies, there are problems with the accuracy of mutation discovery in cell-free DNA WGS, and cell-free DNA WGS cannot be used for cancer diagnosis. If the patient has mutation information through tumor tissue WGS of cancer, cell-free DNA WGS is only used for recurrence monitoring of cancer by filtering and tracking only those mutations (Zviran A et al., Nat Med, Vol. 26, pp. 1114-1124, 2020). That is, although it is effective to use cell-free DNA WGS for cancer diagnosis, due to the absence of an effective filtering method, cell-free DNA WGS could not be used for cancer diagnosis.
[0005] On the one hand, the mutation rate in cancer varies by region on the genome, and furthermore, the mechanisms by which mutations occur and the patterns by which mutations accumulate also differ among cancer types. Using such characteristics, it has been reported that cancer types can be distinguished using the distribution of mutations (regional mutation density) and the types of mutations (mutation signature) in cancer tissues (Jia Wei et al., Nat. Communications, Vol. 11, no. 728, 2020). However, in this case, the theoretical possibilities were explored after cancer diagnosis and cancer type differentiation had already been completed through surgery, and it has not been applied to cancer diagnosis technology through cell-free DNA WGS.
[0006] On the other hand, although there are various patents (KR 10-2017-0185041, KR 10-2017-0144237, KR 10-2018-0124550) that utilize artificial neural networks in the bio field, the current situation is that methods for predicting cancer types by analyzing mutations based on the sequence analysis information of cell-free DNA (cell-free DNA, cfDNA) WGS in blood are lacking due to the problem of inaccurate discovery of cancer-specific mutations.
[0007] Therefore, the inventors of the present invention made diligent efforts to solve the above problems and develop a method for cancer diagnosis and cancer type prediction with high sensitivity and accuracy for single nucleotide mutations of cell-free nucleic acids. As a result, nucleic acids were extracted from a biological sample, sequence information was obtained and aligned reads were used to extract cancer-specific single nucleotide mutations through filtering, calculate the distribution of single nucleotide mutations and the frequency by type of single nucleotide mutation, and when analyzing the value output by inputting this into a learned artificial intelligence model, it was confirmed that cancer diagnosis and cancer types can be predicted with high sensitivity and accuracy, and the present invention was completed.
Summary of the Invention
Problems to be Solved by the Invention
[0008] An object of the present invention is to provide a method for cancer diagnosis and cancer type prediction using single nucleotide mutations of cell-free nucleic acids.
[0009] Another object of the present invention is to provide an apparatus for cancer diagnosis and cancer type prediction using single nucleotide mutations of cell-free nucleic acids.
[0010] Still another object of the present invention is to provide a computer-readable recording medium including instructions configured to be executed by a processor that predicts cancer diagnosis and cancer type by the above method.
Means for Solving the Problems
[0011] To achieve the above object, the present invention includes: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) for the aligned sequence information (reads), detecting single nucleotide variants and performing filtering to extract cancer-specific single nucleotide mutations; (d) dividing the reference chromosome into certain intervals and calculating the regional mutation density of the single nucleotide mutations extracted for each interval; (e) calculating the frequency of each type of mutation signature of the extracted mutations; and (f) inputting the distribution of the single nucleotide mutations calculated in step (d) and the frequency value of each type of mutation signature calculated in step (e) into an artificial intelligence model trained to perform cancer diagnosis, and comparing the output value with a reference value. The present invention provides a method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations.
[0012] The present invention also includes: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequences with a standard chromosomal sequence database; a mutation discovery unit that discovers single nucleotide variants from the aligned sequences and performs filtering to extract cancer-specific single nucleotide variants; a single nucleotide variant distribution calculation unit that divides a standard chromosome into certain intervals and calculates the distribution (regional mutation density) of single nucleotide variants extracted for each interval; a mutation frequency calculation unit that calculates the frequency of each type (mutation signature) of single nucleotide variant of the extracted mutations; a cancer diagnosis unit that inputs the calculated single nucleotide variant distribution value and mutation frequency into an artificial intelligence model learned to perform cancer diagnosis, compares the output value with a reference value, and determines the presence or absence of cancer; and a cancer type prediction unit that inputs the single nucleotide variant distribution value and mutation frequency of a sample determined to have cancer into a second artificial intelligence model learned to classify cancer types, compares the resulting output values, and predicts the cancer type, thereby providing an artificial intelligence-based cancer diagnosis and cancer type prediction device.
[0013] The present invention also relates to a computer-readable recording medium comprising instructions configured to be executed by a processor for diagnosing cancer and predicting cancer types, the instructions including: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) detecting single nucleotide variants in the aligned sequence information (reads) and performing filtering to extract cancer-specific single nucleotide variants; (d) dividing the reference chromosome into certain intervals and calculating the distribution of single nucleotide variants (regional mutation density) extracted for each interval; (e) calculating the frequency of each type of single nucleotide variant (mutation signature) of the extracted mutations; (f) inputting the distribution of single nucleotide variants calculated in step (d) and the frequency value of each type of single nucleotide variant calculated in step (e) into an artificial intelligence model trained to perform cancer diagnosis, comparing the output value with a reference value to determine the presence or absence of cancer; and (g) inputting the distribution of single nucleotide variants and the frequency value of each type of single nucleotide variant of the sample determined to be cancer in step (f) into a second artificial intelligence model trained to classify cancer types, comparing the output values to predict the cancer type. The present invention provides a computer-readable recording medium comprising instructions configured to be executed by a processor for diagnosing cancer and predicting cancer types through the above steps.
Advantages of the Invention
[0014] The method for diagnosing cancer and predicting cancer types using single nucleotide variants of cell-free nucleic acids according to the present invention is not only highly sensitive and accurate compared to other methods for diagnosing cancer and predicting cancer types using the genetic information of cell-free nucleic acids, but also can ensure the same level of sensitivity and accuracy as cancer tissue cell-based methods, and can be utilized in other analyses using single nucleotide variants of cell-free nucleic acids, so it is useful.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. In general, the nomenclature used herein and the experimental methods described below are well known and commonly used in the technical field.
[0017] Terms such as first, second, A, B, etc. can be used to describe various components, but these components are not limited by the above terms and are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the technology described below, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes any combination of a plurality of related listed items or any one of a plurality of related listed items.
[0018] In the terms used in this specification, a singular expression should be understood to include a plural expression unless the context clearly indicates a different interpretation. Terms such as "including" mean that the indicated features, numbers, steps, operations, components, parts, or combinations thereof exist, and should not be understood to exclude the existence or addition possibility of one or more other features, numbers, step operations, components, parts, or combinations thereof.
[0019] Prior to the detailed description of the drawings, it is clarified that the classification of the components in this specification is merely based on the main functions assumed by each component. That is, two or more of the components described below may be combined into one component, or one component may be further divided into two or more components according to more refined functions. And each of the components described below may additionally perform some or all of the functions assumed by other components in addition to its own main function. Of course, it is also possible that some functions of the main function assumed by each component are exclusively performed by other components.
[0020] Also, when performing a method or an operation method, each process constituting the method may be performed in an order different from the specified order unless a specific order is clearly described in the context. That is, each process may be performed in the same order as the specified order, may be performed substantially simultaneously, or may be performed in the reverse order.
[0021] In the present invention, after aligning the array analysis data obtained from a sample with a reference genetic material, nucleic acid is extracted from a biological sample, and based on the reads obtained by acquiring and aligning sequence information, cancer-specific single nucleotide mutations are extracted through filtering, the distribution of single nucleotide mutations and the frequency by type of single nucleotide mutation are calculated, and when analyzing the value calculated by inputting this into a learned artificial intelligence model, it was confirmed that cancer diagnosis and the type of cancer can be predicted with high sensitivity and accuracy.
[0022] That is, in one embodiment of the present invention, after sequencing DNA extracted from blood and aligning it with a reference chromosome, cancer-specific single nucleotide mutations are extracted from the aligned reads through filtering, the reference chromosome is divided into certain intervals, the distribution of single nucleotide mutations for each interval is calculated, the frequency by type of each single nucleotide mutation is calculated, the distribution of single nucleotide mutations and the frequency by type of single nucleotide mutation are input into an artificial intelligence model learned to perform cancer diagnosis, after performing cancer diagnosis by comparing the output value with a reference value, the distribution of single nucleotide mutations and the frequency by type of single nucleotide mutation of the sample determined to be cancer are input into a second artificial intelligence model learned to distinguish cancer types, and the cancer type showing the highest value among the output values is determined as the cancer type of the sample (Figure 1).
[0023] Therefore, the present invention, in one aspect, (a) a step of extracting nucleic acid from a biological sample to obtain sequence information; (b) a step of aligning the obtained sequence information (reads) with a reference genome database; (c) a step of detecting single nucleotide variants in the aligned sequence information (reads) and performing filtering to extract cancer-specific single nucleotide mutations; (d) a step of dividing the reference chromosome into certain intervals and calculating the distribution (regional mutation density) of single nucleotide mutations extracted for each interval; (e) calculating the frequency of each type of single nucleotide mutation (mutation signature) of the extracted mutation; and (f) inputting the distribution of single nucleotide mutations calculated in step (d) and the frequency values for each type of single nucleotide mutation calculated in step (e) into an artificial intelligence model trained to perform cancer diagnosis, and comparing the output value with a reference value; It relates to a method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations including the above.
[0024] In the present invention, the cancer may be a solid cancer or a blood cancer, preferably selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphoid leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, thyroid cancer, liver cancer, gastric cancer, gallbladder cancer, bile duct cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary origin, kidney cancer, and mesothelioma, and most preferably may be liver cancer or ovarian cancer, but is not limited thereto.
[0025] In the present invention, The step (a) is (a-i) obtaining nucleic acids from a biological sample; (a-ii) using a salting-out method, column chromatography method, or beads method from the collected nucleic acids to remove proteins, fats, and other residues, and obtaining purified nucleic acids; (a-iii) For the purified nucleic acid or the nucleic acid randomly fragmented by enzymatic cleavage, grinding, or the hydroshear method, the step of preparing a single-end sequencing or pair-end sequencing library; (a-iv) The step of reacting the prepared library with a next-generation sequencer; and (a-v) The step of obtaining sequence information (reads) of the nucleic acid with a next-generation sequencer; It can be characterized by including the above.
[0026] In the present invention, the step of obtaining the sequence information in the step (a) can be characterized by obtaining the separated cell-free DNA by whole-genome sequencing at a read depth of 1 million to 100 million.
[0027] In the present invention, the biological sample means any substance, biological fluid, tissue or cell obtained from or derived from an individual, for example, whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, hair, oral cells, placental cells, cerebrospinal fluid, and mixtures thereof, but is not limited thereto.
[0028] In the present invention, the term "reference population" is a reference population that can be compared, such as a standard base sequence database, and currently means a population of humans without a specific disease or medical condition. In the present invention, the standard base sequence in the standard chromosomal sequence database of the reference population may be a reference chromosome registered with a public health institution such as NCBI.
[0029] In the present invention, the nucleic acid in the step (a) may be cell-free DNA, and more preferably, may be circulating tumor DNA, but is not limited thereto.
[0030] In the present invention, the next-generation gene sequencer can be used with any sequencing method known in the art. Sequencing of the nucleic acid separated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of one of the proxies that are clonally amplified for individual nucleic acid molecules, either individually or in a highly similar manner (e.g., 105 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species in the library can be estimated by measuring the relative occurrence of its homologous sequences in the data generated by the sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in the literature incorporated herein by reference (Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46).
[0031] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequence of individual nucleic acid molecules (e.g., the HeliScope Gene Sequencing system of Helicos BioSciences and the PacBio RS system of Pacific Biosciences). In other embodiments, a massively parallel short-read sequencing method that generates more bases of sequence per sequencing unit than other sequencing methods, e.g., other methods that generate fewer but longer reads, (e.g., the Solexa sequencer of Illumina Inc. located in San Diego, Calif.) determines the nucleotide sequence of proxies that are clonally amplified for individual nucleic acid molecules (e.g., the Solexa sequencer of Illumina Inc. located in San Diego, Calif.; 454 Life Sciences (located in Branford, Conn.) and Ion Torrent). Other methods or machines for next-generation sequencing include, but are not limited to, 454 Life Sciences (located in Branford, Conn.), Applied Biosystems (located in Foster City, Calif.; SOLiD sequencer), Helicos BioSciences Corporation (located in Cambridge, Mass.) and those provided by emulsion and microfluidic sequencing techniques such as nanoporation (e.g., nanoporation by GnuBio).
[0032] Platforms for next-generation sequencing include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX system, Illumina / Solexa's Genome Analyzer (GA), Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene Sequencing system, and Pacific Biosciences' PacBio RS system.
[0033] In the present invention, the alignment step in step (b) may be performed using, but not limited to, the BWA algorithm and the Hg19 sequence.
[0034] In the present invention, the BWA algorithm may include, but is not limited to, BWA-ALN, BWA-SW, or Bowtie2, etc.
[0035] In the present invention, the length of the sequence information (reads) in step (b) is 5 to 5000 bp, and the number of sequence information used can be 5000 to 5 million, but is not limited thereto.
[0036] In the present invention, the filtering in step (c) can be used without limitation as long as it can distinguish single nucleotide mutations generated from healthy individuals from single nucleotide mutations specifically generated in cancer. Preferably, it can be characterized by extracting single nucleotide mutations where the read depth of the mutation region with the discovered single nucleotide mutation is 3 or more and the average sequencing quality is 30 or more, but is not limited thereto.
[0037] In the present invention, the mutation region means the exact position where there is a single nucleotide mutation, and the meaning that the read depth of the mutation region is 3 or more means that the number of reads aligned at that position is 3 or more.
[0038] In the present invention, the filtering in the step (c) can be further characterized by performing a process of removing artifacts and germline mutations generated during the sequence analysis process, and the process is i) Mutations detected only in one of the read pairs; ii) Mutations detected two or more types at one location; iii) Mutations where normal bases are not detected at each position; and iv) Mutations detected from a healthy individual database; It can be characterized by removing any one or more mutations selected from the group consisting of, but not limited to, these.
[0039] In the present invention, the healthy individual database can be used without limitation as long as it is a database containing base sequence mutation information of healthy individuals. Preferably, it may be a database containing cfDNA WGS data of healthy individuals, WGS data of tissue samples, etc. More preferably, it may be a publicly available database such as dbSNP, 1000 Genome, Hapmap, ExAC, Gnomad, etc., but not limited to these.
[0040] In the present invention, the section in the step (d) can be arbitrarily set to any length as long as it can calculate the distribution of single nucleotide mutations. Preferably, it may be 100 kb to 10 Mb, more preferably 500 kb to 5 Mb, and most preferably 1 Mb, but not limited to these.
[0041] In the present invention, the step of calculating the distribution of the single nucleotide mutations (regional mutation density, RMD) extracted in the step (d) can be characterized by being performed by a method including the following steps: (d-i) Calculating the number of single nucleotide mutations extracted for each interval excluding intervals where no mutations are detected above the reference value of the entire sample; and (d-ii) Normalizing by dividing the calculated number by the total number of mutations for each interval.
[0042] In the present invention, the reference value can be used without limitation as long as it can significantly distinguish the extracted single nucleotide mutations, and may preferably be 40 to 60%, more preferably 45 to 55%, and most preferably 50%, but is not limited thereto.
[0043] In the present invention, the interval excluding the intervals where no mutations are detected above the reference value of the entire sample means excluding intervals where no single nucleotide mutations extracted from 50% or more of the entire sample exist when the reference value is 50%.
[0044] In the present invention, the interval can be characterized by being one or more selected from the intervals described in Table 1.
[0045] In the present invention, the distribution of single gene mutations (regional mutation density, RMD) is used in the same sense as the background mutation rate, and means calculating the mutation frequency by dividing the entire genome into certain intervals.
[0046] In the present invention, the distribution of single gene mutations for each cancer type is a quantitative value indicating whether the region has a high or low frequency of mutations in that cancer. Single gene mutations in cancer are not uniformly distributed across the human genome. The amount of single gene mutations accumulated varies across the entire genomic region, and the patterns of accumulation also differ significantly among cancer types. In addition, epigenomic features (histone modification, replication timing) are the main causes of the distribution of single gene mutations in cancer types, and the distribution of single gene mutations encompasses the epigenomic features of the corresponding cancer type.
[0047] Since the distribution of single gene mutations varies by entire genomic region and by cancer type, it can be a useful indicator for cancer diagnosis and cancer type discrimination. Using the distribution of single gene mutations, it is possible to determine whether the discovered mutations are located in regions with a high probability of occurrence in that cancer.
[0048] In the present invention, the step of calculating the frequency of mutation signatures for each type of single nucleotide mutation in step (e) can be characterized by being performed by a method including the following steps: (e-i) Calculating the number of mutations for each of the following types of mutations; and (1) Mutations in which cytosine (C) is substituted with thymine (T), adenine (A), or guanine (G); (2) Mutations in which thymine is substituted with cytosine, adenine, or guanine; (3) Mutations in (1) or (2) that further contain one base in the 5' direction; (4) Mutations in (1) or (2) that further contain one base in the 3' direction; and (5) Mutations that further contain one base in the 5' direction and one base in the 3' direction, where adenine, guanine, cytosine, and thymine are each substituted with a different base; (e-ii) Normalizing by dividing the total of the calculated number of mutations by the grand total. In the present invention, it can be characterized in that the type of the mutation is one or more selected from the mutations described in Table 2.
[0049] In the present invention, the type of single nucleotide mutation (mutation signature) can be used without limitation as long as the normal base is mutated to another base and gene dysfunction occurs. Preferably, it can be characterized in that it is one or more selected from the group consisting of C->A, C->G, C->T, T->A, T->C, and T->G, but is not limited thereto.
[0050] In the present invention, C->A means confirming whether the detected mutation is one in which the normal base C is mutated to the mutant base A, C->G means confirming whether the detected mutation is one in which the normal base C is mutated to the mutant base G, and the others have the same meaning.
[0051] In the present invention, the reference value in the step (f) can be used without limitation as long as it is a value capable of diagnosing cancer. Preferably, it may be 0.5, but is not limited thereto. In the event that the reference value is 0.5, it can be characterized in that it is determined to be cancer when it is 0.5 or more.
[0052] In the present invention, when the artificial intelligence model is learning, it is learned so that the output result is close to 1 when there is cancer, and it is learned so that the output result is close to 0 when there is no cancer. Based on 0.5, if it is 0.5 or more, it is determined that there is cancer, and if it is 0.5 or less, it is determined that there is no cancer, and performance measurement was performed (accuracy of training, validation, and test accuracy).
[0053] Here, it is obvious to an ordinary technician that the reference value of 0.5 can be changed at any time. For example, if one wants to reduce false positives, a reference value higher than 0.5 can be set to make the criteria for determining the presence of cancer stricter. If one wants to reduce false negatives, the reference value can be measured lower to slightly weaken the criteria for determining the presence of cancer.
[0054] In the present invention, (g) Inputting the distribution of single nucleotide variations and the frequency values by type of single nucleotide variation of the sample determined to be cancer into a second artificial intelligence model learned to distinguish cancer types, and comparing the output values to predict the cancer type; can further be characterized by including the above.
[0055] In the present invention, the comparison of the output values in step (g) can be characterized by being performed by a method including a step of determining, as the cancer of the sample, the cancer type showing the highest value among the output values.
[0056] In the present invention, the artificial intelligence model can be used without limitation as long as it is a model capable of diagnosing cancer or discriminating cancer types. Preferably, it may be an artificial neural network model. More preferably, it may be selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and an autoencoder. Most preferably, it may be a deep neural network, but is not limited thereto.
[0057] In the present invention, when the artificial intelligence model is a DNN and learns binary classification, the loss function can be characterized by being binary crossentropy represented by the following mathematical formula 1:
Number
[0058] In the present invention, when the second artificial intelligence model is a DNN and learns multi-class classification, the loss function can be characterized in that it is the categorical cross-entropy shown in the following mathematical formula 2:
Number
[0059] In the present invention, when the artificial intelligence model is a DNN, the learning can be characterized in that it includes the following steps: i) Classifying the produced data into training, validation, and performance evaluation data; At this time, the training data is used when learning the DNN model, the validation data is used for hyper-parameter tuning verification, and the performance evaluation data is used for performance evaluation after the optimal model is produced. ii) constructing an optimal DNN model through hyperparameter tuning and the learning process; iii) comparing the performance of various models obtained through hyperparameter tuning using validation data, and determining the model with the best performance on the validation data as the optimal model;
[0060] In the present invention, the hyperparameter tuning process is a process of optimizing various parameter values (number of layers, number of filters, etc.) that make up the DNN model. As the hyperparameter tuning process, it is possible to use hyperband optimization, Bayesian optimization, and grid search techniques.
[0061] In the present invention, the learning process optimizes the internal parameters (weights) of the DNN model using the determined hyperparameters. When the validation loss begins to increase compared to the learning loss, it is determined that the model is overfitting, and the model learning is interrupted before that.
[0062] In another aspect of the present invention, a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; An alignment unit that aligns the decoded sequences with a standard chromosomal sequence database; A mutation discovery unit that discovers single nucleotide variants from the aligned sequences, performs filtering, and extracts cancer-specific single nucleotide variants; A single nucleotide variant distribution calculation unit that divides the standard chromosome into certain intervals and calculates the distribution (regional mutation density) of single nucleotide variants extracted for each interval; A mutation frequency calculation unit that calculates the frequency of single nucleotide variant types (mutation signature) of the extracted mutations; The cancer diagnosis unit inputs the calculated distribution values and mutation frequencies of single nucleotide variants into an artificial intelligence model learned to perform cancer diagnosis, compares the output value with a reference value, and determines the presence or absence of cancer; and An artificial intelligence-based cancer diagnosis and cancer type prediction apparatus including a cancer type prediction unit that inputs the distribution values and mutation frequencies of single nucleotide variants of a sample determined to have cancer into a second artificial intelligence model learned to classify cancer types, and predicts the cancer type by comparing the output result values.
[0063] In the present invention, the decoding unit may include a nucleic acid injection unit that injects nucleic acids extracted from an independent device; and a sequence information analysis unit that analyzes the sequence information of the injected nucleic acids. Preferably, it may be an NGS analysis device, but is not limited thereto.
[0064] In the present invention, the decoding unit may be characterized by receiving and decoding sequence information data generated by an independent device.
[0065] In another aspect, the present invention is a computer-readable recording medium including instructions configured to be executed by a processor for predicting cancer diagnosis and cancer types, while (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) detecting single nucleotide variants in the aligned sequence information (reads), performing filtering, and extracting cancer-specific single nucleotide variants; (d) dividing the reference chromosome into certain intervals and calculating the distribution (regional mutation density) of single nucleotide variants extracted for each interval; (e) calculating the frequencies of different types (mutation signatures) of single nucleotide variants of the extracted mutations; (f) Inputting the distribution of single nucleotide variations calculated in step (d) and the frequency values by type of single nucleotide variations calculated in step (e) into an artificial intelligence model trained to perform cancer diagnosis, and determining the presence or absence of cancer by comparing the output value with a reference value; and (g) Inputting the distribution of single nucleotide variations and the frequency values by type of single nucleotide variations of the sample determined to be cancer in step (f) into a second artificial intelligence model trained to classify cancer types, and predicting the cancer type by comparing the output values; Relates to a computer-readable recording medium including instructions configured to be executed by a processor that predicts the presence or absence of cancer and the cancer type through the above steps.
[0066] In other aspects, the method according to the present application can be implemented using a computer. In one embodiment, the computer includes one or more processors connected to a chipset. Also, connected to the chipset are a memory, a recording device, a keyboard, a Graphics Adapter, a Pointing Device, a Network Adapter, and the like. In one embodiment, the performance of the chipset is enabled by a Memory Controller Hub and an I / O Controller Hub. In other embodiments, the memory can be directly connected to the processor instead of the chipset for use. The recording device is any device capable of maintaining data including a hard drive, a CD-ROM (Compact Disk Read-Only Memory), a DVD, or other memory devices. The memory is involved in the data and instructions used by the processor. The Pointing Device may be a mouse, a Track Ball, or other types of pointing devices and is used in combination with the keyboard to transfer input data to the computer system. The Graphics Adapter displays images and other information on the display. The Network Adapter connects the computer system to a short-range or long-range communication network. However, the computer used in the present application is not limited to the above configuration, can include no part of the configuration or additional configurations, and may also be part of a Storage Area Network (SAN). The computer of the present application can be configured to conform to the execution of modules in a program for performing the method according to the present application.
[0067] In this application, a module can refer to a functional and structural combination of hardware for implementing the technical idea according to this application and software for driving the hardware. For example, the module can refer to a logical unit of a predetermined code and the hardware resources for executing the predetermined code. It is obvious to those skilled in the art of this application that it does not necessarily mean physically connected code or a certain type of hardware.
[0068] Example Hereinafter, the present invention will be described in more detail through examples. It is obvious to those with ordinary knowledge in the industry that these examples are solely for illustrating the present invention and should not be construed as limiting the scope of the present invention.
[0069] Example 1. Extracting DNA from blood and performing next-generation sequencing analysis Blood samples of 471 healthy individuals, 151 ovarian cancer patients, and 131 liver cancer patients were each collected in an amount of 10 mL and stored in EDTA tubes. Within 2 hours after collection, only the plasma portion was centrifuged for the first time at 1200 g, 4 °C for 15 minutes. Then, the plasma obtained from the first centrifugation was centrifuged for the second time at 16000 g, 4 °C for 10 minutes, and the supernatant of the plasma excluding the precipitate was separated. For the separated plasma, cell-free DNA was extracted using the Tiangen micro DNA kit (Tiangen), and after performing the library preparation process using the MGIEasy cell-free DNA library prep set kit, sequencing was performed on the DNBseq G400 equipment (MGI) in 100 base pair end mode. As a result, it was confirmed that approximately 170 million sequence information (reads) was produced per sample.
[0070] Example 2. Extracting single nucleotide variations, calculating the distribution of single nucleotide variations, and extracting feature of frequency by type 2-1. Filtering for extracting cancer-specific mutations The NGS data obtained in Example 1 was aligned to the reference chromosome (hg 19), and the resulting bam file was processed using the GATK pipeline. Mutations were mined using varscan (mutation caller) to ensure a mutation profile for each sample.
[0071] The Varscan mutation mining criteria were applied very leniently. Variant calling was performed with lenient criteria through the presence of one or more variant reads, an overall depth of 3 or more in the mutation region, an average base quality of 30 or more, removal of the minimum variant allele frequency criteria, removal of the strand filter, and removal of the varscan variant P-value criteria (variant allele frequency means the ratio of mutations, which is the ratio of the number of reads in which mutations were mined among all reads at the mutation position).
[0072] After mining all mutations that could be cancer-derived mutations with lenient criteria, artifacts and germline mutations were removed using various criteria. Four methods were used to remove mutations at incorrect positions.
[0073] First, when a position where mutations exist in both the forward and reverse reads of the fragment was sequenced, if a mutation was found only in one of the reads, it was removed. Second, if there were two or more mutations at one position, it was removed. Third, when the variant allele frequency was 1, since it means that mutations exist in all DNA present in the blood, it was removed assuming there is no probability of being a tumor-derived mutation.
[0074] Fourthly, mutations in various healthy individual mutation databases and blacklist regions were removed. The blacklist regions are regions such as repeats and centromeres where the probability of misalignment is high during alignment. The blacklist regions used were those organized by Haley M amemiya et al., Scientific report Vol. 9, no. 9354, 2019. In addition, a public database collecting healthy individual mutations was used to remove mutations with a high probability of being healthy individual mutations. The dbSNP (https: / / data.amerigeoss.org / ko_KR / dataset / dbsnp), 1000 Genome (https: / / www.internationalgenome.org / ), Hapmap (https: / / ftp.ncbi.nlm.nih.gov / hapmap / ), ExAC (https: / / gnomad.broadinstitute.org / downloads#exac-variants), and Gnomad (https: / / gnomad.broadinstitute.org / ) databases were used.
[0075] In addition, mutations in the cfDNA WGS database of 20,000 healthy individuals produced by Green Cross Corporation (South Korea) were filtered because they are less likely to be tumor-derived mutations. And for the input values of the algorithm for classifying cancer types, the mutations discovered by the cell-free DNA WGS of 412 healthy individuals in Example 1 were also removed.
[0076] 2-2. Calculating the distribution of single nucleotide variations The entire genome was segmented into 1-Mb intervals, and the distribution of single nucleotide mutations (regional mutation density, RMD) for each interval was calculated. Excluding the intervals where mutations were not present in more than 50% of the total samples extracted in Example 2-1, the distribution of single nucleotide mutations in a total of 2,726 intervals was used as the input value for the algorithm. The number of mutations in each interval was calculated and divided by the total sum of the number of mutations in the 2,726 intervals for normalization. Finally, features of the distribution of 2,726 single gene mutations were generated, and the feature list is as shown in Table 1 below.
Table 1
[0077] 2-3. Calculating the frequency by type of single nucleotide variations The frequencies of single-gene mutations by type (mutation signature) were calculated across the entire genome. The criteria for classifying mutation types were defined in four categories.
[0078] First, mutation types were classified using the types of standard bases and mutated bases, defining a total of six basic mutation types (C>A, C>G, C>T, T>A, T>C, T>G). Second, in the basic mutation types, one base in the 5' direction was further considered, defining 24 (4×6) mutation types. Third, in the basic mutation types, one base in the 3' direction was further considered, defining 24 (6×4) mutation types. Finally, one base in the 5' direction and one base in the 3' direction were further considered in the basic mutation types, determining 96 (4×6×4) mutation types commonly used in mutation signature analysis.
[0079] The occurrence frequencies were calculated for each of the total 150 mutation types thus classified. Then, the total number of mutations was calculated for each of the four mutation classification methods and normalized by dividing by the total sum of all mutations that occurred in the entire bases.
[0080] The defined mutation types are as shown in Table 2 below.
Table 2
[0081] Finally, a total of 2,876 features, combining 2,726 features of the distribution of single-gene mutations and 150 features of the types of single-gene mutations, were used as input values for the algorithm.
[0082] Example 3. Constructing a DNN model and the learning process For the development of an algorithm to distinguish cancer diagnosis and cancer types from cfDNA, a total of 2,876 features for the distribution and types of single-gene mutations previously obtained through analysis were used. A total of two artificial intelligence algorithms were developed.
[0083] First, a binary classification model for diagnosing whether a person is healthy or has cancer was constructed. Second, a multiple classification model for classifying cancer types was constructed. As the loss function for algorithm learning, binary cross-entropy was used for the binary classification model, and categorical cross-entropy was used for the multiple classification model. A deep neural network artificial intelligence model was used for algorithm learning.
[0084] The entire dataset was divided into a learning, validation, and performance evaluation dataset, and the model was learned using hyperparameter tuning with the method of Bayesian optimization. The entire dataset was divided into five learning, validation, and performance evaluation sets, and learning was performed five times to create five algorithm models. Then, each of the five algorithm models was used to make predictions on the five performance evaluation datasets, such that the entire dataset could be used once as a performance evaluation dataset. Thus, the performance of the model was evaluated using the prediction probability when the entire sample was the performance evaluation dataset.
[0085] Example 4. Constructing a deep learning model for cancer diagnosis and cancer type classification and verifying the performance To test the performance of the deep learning model constructed using the leads obtained in Example 1, the method of an artificial intelligence model (Cristiano, S. et al., Nature, Vol. 570(7761), pp. 385-389. 2019) used for conventional known cancer diagnosis and cancer type discrimination was applied. Based on the dataset of Example 1, a cancer diagnosis and cancer type discrimination comparison model based on fragmentation pattern and copy number variation (CNV) was constructed so that it could be applied to cfDNA.
[0086] More specifically, in the method of the fragment pattern, after GC correction of the entire genome, it was divided into 5Mb intervals, and the ratio of the number of short fragments in each interval to the total number of fragments was subjected to z-score normalization and used as the input value. Here, a short fragment means a fragment with a length between 100bp and 150bp. For the CNV method, after dividing the entire genome into non-overlapping 50KB regions and performing GC correction, the depth was calculated for each region and then converted to a log2 value and used as the input value. Xgboost was used for the learning of the fragment pattern and CNV models.
[0087] For the performance comparison of the cancer diagnosis model, the sensitivity at the predict probability threshold with specificities of 95%, 98%, and 99% was confirmed.
[0088] As a result, as shown in Figure 2, it was confirmed that the performance of the cancer diagnosis model constructed in the present invention is superior to the conventional method. Also, as shown in Figure 3, at all accuracies, not only is the performance of the cancer diagnosis model constructed in the present invention excellent in cancer diagnosis, but as shown in (B) of Figure 3, while the performance of the conventional method is inhibited in the early diagnosis (stage I) of cancer, the cancer diagnosis model constructed in the present invention shows excellent performance even in the early diagnosis of cancer.
[0089] Also, as a result of comparing the performance of the cancer type discrimination model, as shown in Figure 4, it was confirmed that the cancer type discrimination model constructed in the present invention is superior in cancer type discrimination performance at all stages compared to the conventional method.
[0090] Example 5. Verifying the effect of filtering conditions 5-1. Verifying the effect of filtering criteria After the present inventors discovered all mutations that could be cancer-derived mutations based on lenient criteria, they removed artifacts and germline mutations using various criteria. The present inventors compared the performance when discovering mutations using a strict method, a less strict method, and a lenient method. The strict method discovers a mutation when the variant read exists in both the forward and reverse reads, and the less strict method also discovers a mutation when there are two or more mutations in the variant read. The lenient method is the same as the method described in Example 2-1. After discovering the mutations, the performance was compared after model learning using the same filtering and learning process.
[0091] As a result, as shown in FIG. 5, it was confirmed that the performance was the best when filtering after discovering all mutations that could be cancer-derived mutations based on lenient criteria.
[0092] 5-2. Verifying the effect of filtering databases In the present invention, a method of filtering mutations that appear in cfDNA and tissues of healthy individuals was used together with the discovery of lenient mutations. By using the mutations discovered by large-scale cfDNA / tissue WGS of healthy individuals for mutation filtering, it was expected that the detection of artifacts and germline mutations that could occur in cfDNA could be effectively removed.
[0093] Since there was no large-scale cfDNA mutation database in the public database, 20,000 healthy individual cfDNA WGS produced by Green Cross was used.
[0094] As a result, as shown in FIG. 6, it was confirmed that the performance was improved when using cfDNA and tissue mutations of healthy individuals compared to when not using them. Therefore, all public databases regarding cfDNA mutations and tissue mutations of healthy individuals were used in the prediction model of the present invention.
[0095] Example 6. RMD distribution of cfDNA mutations in specific mutation regions by cancer type When the cfDNA mutations were discovered using the cfDNA mutation discovery method developed in the present invention and the RMD values were calculated, it was confirmed that the actual characteristics and distribution of the cancer were well reflected.
[0096] Using whole-genome sequencing (WGS) of cancer tissues of ovarian cancer and liver cancer in the PCAWG, a large-scale cancer cohort, tumor mutations were discovered for each sample. After calculating the RMD values in 1Mbp bin units, edgeR was used to search for regions with a high number of mutations specific to each cancer type and regions with a low number of mutations specific to each cancer type for each cancer type. Then, it was confirmed whether the regions with a high number of mutations specific to the cancer tumor tissue actually had a high RMD value in cfDNA and whether the regions with a low number of mutations specific to the cancer tumor tissue also had a low cfDNA RMD value.
[0097] As a result, as shown in Figure 7, it was confirmed that for both ovarian cancer and liver cancer, the characteristics of the RMD regions of the cancer in cfDNA were reflected in the same way as in the tissue samples. The liver and ovary on the X-axis of the figure mean the cancer tumors of the actual cfDNA samples. Also, the region type means the regions with a high / low number of mutations specific to the cancer tumor defined using the PCAWG data.
[0098] As described above, specific parts of the content of the present invention have been described in detail. However, for those with ordinary knowledge in the art, such specific descriptions are merely preferred embodiments, and it will be clear that the scope of the present invention is not limited thereby. Therefore, it can be said that the substantial scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. (a) Extracting nucleic acids from a biological sample to obtain sequence information; (b) Aligning the obtained sequence information (reads) with a reference genome database; (c) Detecting single nucleotide variants in the aligned sequence information (reads) and performing filtering to extract cancer-specific single nucleotide variants; (d) Dividing the reference chromosome into certain intervals and calculating the regional mutation density (RMD) of the single nucleotide variants extracted for each interval; (e) Calculating the frequencies of different mutation signatures of the extracted mutations; and (f) Inputting the distribution of single nucleotide variants calculated in step (d) and the frequencies of different mutation signatures calculated in step (e) into an artificial intelligence model trained for cancer diagnosis, and comparing the output value with a reference value; A method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants.
2. The method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants according to Claim 1, wherein step (a) is performed by a method including the following steps: (a-i) Obtaining nucleic acids from a biological sample; (a-ii) Using a salting-out method, column chromatography method, or beads method to remove proteins, fats, and other residues from the collected nucleic acids and obtain purified nucleic acids; (a-iii) Preparing a single-end sequencing or pair-end sequencing library for the purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method; (a-iv) Reacting the prepared library with a next-generation sequencer; and (a-v) Obtaining nucleic acid sequence information (reads) using a next-generation gene sequencer.
3. The filtering in step (c) extracts single nucleotide variations where the read depth of the variant region with the discovered single nucleotide variation is 3 or more and the average sequencing quality is 30 or more. The method for providing information for cancer diagnosis and cancer type prediction using the single nucleotide variation according to claim 1 is characterized by this.
4. The filtering in step (c) further includes a process of removing artifacts and germline mutations that occurred during the sequence analysis process. The method for providing information for cancer diagnosis and cancer type prediction using the single nucleotide variation according to claim 3 is characterized by this.
5. The process of removing the artifacts and germline mutations i) Variations detected only in one of the read pairs; ii) Variations detected two or more types at one location; iii) Variations where normal bases are not detected at each position; and iv) Variations detected from a healthy individual database; The method for providing information for cancer diagnosis and cancer type prediction using the single nucleotide variation according to claim 4 is characterized by removing any one or more variations selected from the group consisting of these.
6. The interval in step (d) is 100 kb to 10 Mb. The method for providing information for cancer diagnosis and cancer type prediction using the single nucleotide variation according to claim 1 is characterized by this.
7. The step of calculating the distribution (regional mutation density) of the single nucleotide variations extracted in step (d) is performed by a method including the following steps. The method for providing information for cancer diagnosis and cancer type prediction using the single nucleotide variation according to claim 1 is characterized by this: (d-i) Calculating the number of single nucleotide variations extracted for each interval excluding intervals where no variation is detected above the reference value of the entire sample; and (d-ii) Normalizing by dividing the calculated number by the total number of variations for each interval.
8. The reference value is 40 - 60%. The method for providing information for cancer diagnosis and cancer type prediction using the single nucleotide variation according to claim 7 is characterized by this.
9. The method for providing information for cancer diagnosis and cancer type prediction using a single nucleotide mutation according to claim 7, wherein the interval is one or more selected from the intervals described in Table 1 below. 【Table 1】
10. The step of calculating the frequency of each type of single nucleotide mutation (mutation signature) in step (e) is performed by a method including the following steps, in the method for providing information for cancer diagnosis and cancer type prediction using a single nucleotide mutation according to claim 1: (e-i) calculating the number of mutations for each of the following types of mutations; and (1) mutations in which cytosine (C) is substituted with thymine (T), adenine (A), or guanine (G); (2) mutations in which thymine is substituted with cytosine, adenine, or guanine; (3) mutations in which one additional base is included in the 5' direction in the mutation of (1) or (2); (4) mutations in which one additional base is included in the 3' direction in the mutation of (1) or (2); and (5) mutations in which one base in the 5' direction and one base in the 3' direction are each further included, where adenine, guanine, cytosine, and thymine are substituted with different bases from each other; (e-ii) normalizing by dividing the total number of calculated mutations by the total of all mutations that occurred in the entire bases.
11. The method for providing information for cancer diagnosis and cancer type prediction using a single nucleotide mutation according to claim 10, wherein the type of mutation is one or more selected from the mutations described in Table 2 below. 【Table 2】
12. The method for providing information for cancer diagnosis and cancer type prediction using a single nucleotide mutation according to claim 1, wherein the reference value in step (f) is 0.5, and when it is 0.5 or more, it is determined to be cancer.
13. (g) Inputting the distribution of single nucleotide mutations and the frequency values for each type of single nucleotide mutation of the sample determined to be cancer into a second artificial intelligence model learned to distinguish cancer types, and comparing the output values to predict the cancer type; The method for providing information for cancer diagnosis and cancer type prediction using a single nucleotide mutation according to claim 1, further comprising the above.
14. The comparison of the output values in the step (g) is performed by a method including a step of determining the cancer type showing the highest value among the output values as the cancer of the sample. A method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants according to claim 13.
15. The artificial intelligence model in the step (f) is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and an autoencoder. A method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants according to claim 1.
16. When the artificial intelligence model is a DNN and learns binary classification, the loss function is binary crossentropy represented by the following formula 1. A method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants according to claim 15: 【Number 1】 Here, in binary classification, N is the total number of samples. 【Number】 is the probability value predicted by the model that the i-th input value is close to class 1, and y i is the actual class of the i-th input value.
17. When the second artificial intelligence model is a DNN and learns multi-class classification, the loss function is categorical crossentropy represented by the following formula 2. A method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants according to claim 13: 【Number 2】 Here, in the categorical cross-entropy, N is the total number of samples, J is the number of overall classes, and y j is a value indicating the actual class of the sample. If the actual class is j, it is indicated by 1, and if the actual class is not j, it is indicated by 0. 【Number】 is the probability value predicted that the sample is in class j, and the closer it is to 1, the higher the probability is predicted that the sample is in that class.
18. A decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; An alignment unit that aligns the decoded sequences with a standard chromosomal sequence database; A mutation discovery unit that discovers single nucleotide variants (single nucleotide variant) from the aligned sequences, performs filtering, and extracts cancer-specific single nucleotide variants; A single nucleotide variant distribution calculation unit that divides the standard chromosome into certain intervals and calculates the distribution (regional mutation density) of the single nucleotide variants extracted for each interval; A mutation frequency calculation unit that calculates the frequency of single nucleotide mutations of the extracted mutations by mutation signature; A cancer diagnosis unit that inputs the calculated distribution value and mutation frequency of single nucleotide mutations into an artificial intelligence model learned to perform cancer diagnosis, and determines the presence or absence of cancer by comparing the output value with a reference value; and An artificial intelligence-based cancer diagnosis and cancer type prediction device including a cancer type prediction unit that inputs the distribution value and mutation frequency of single nucleotide mutations of a sample determined to have cancer into a second artificial intelligence model learned to distinguish cancer types, and predicts the cancer type by comparing the output result values.
19. A computer-readable recording medium including instructions configured to be executed by a processor for predicting cancer diagnosis and cancer type, (a) A step of extracting nucleic acid from a biological sample to obtain sequence information; (b) A step of aligning the obtained sequence information (reads) with a reference genome database; (c) A step of detecting single nucleotide variants in the aligned sequence information (reads), performing filtering, and extracting cancer-specific single nucleotide mutations; (d) A step of dividing the standard chromosome into certain intervals and calculating the regional mutation density of the single nucleotide mutations extracted for each interval; (e) A step of calculating the frequency of single nucleotide mutations of the extracted mutations by mutation signature; (f) A step of inputting the distribution of single nucleotide mutations calculated in step (d) and the frequency value of single nucleotide mutations by mutation signature calculated in step (e) into an artificial intelligence model learned to perform cancer diagnosis, and determining the presence or absence of cancer by comparing the output value with a reference value; and (g) A step of inputting the distribution of single nucleotide mutations and the frequency value of single nucleotide mutations by mutation signature of a sample determined to have cancer in step (f) into a second artificial intelligence model learned to distinguish cancer types, and predicting the cancer type by comparing the output values; A computer-readable recording medium including instructions configured to be executed by a processor for predicting the presence or absence of cancer and cancer type through the above steps.
Citation Information
Patent Citations
Ultrasound-sensitive detection of circulating tumor DNA by genome-wide integration
JP2021519607A
Multi-Assay Prediction Model for Cancer Detection
US20190316209A1
Ultra-sensitive detection of circulating tumor DNA through genome-wide integration
US20210043275A1
Methods and systems for detecting microsatellite instability of a cancer in a liquid biopsy assay
US20210098078A1
Methods and Systems for Analyzing Nucleic Acid Molecules
US20210172022A1