A method for diagnosing and predicting cancer type based on single nucleotide variants in cell-free nucleic acids.
Patent Information
- Application Number
- JP2024573381
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-06-15
- Filing Date
- 2023-06-15
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2043-06-15
AI Technical Summary
【0014】 本発明による無細胞核酸の単一塩基変異を用いたがん診断及びがん種予測方法は、無細胞核酸の遺伝情報を用いたがん診断及びがん種を予測する他の方法に比べて感度と正確度が高いだけでなく、がん組織細胞ベースの方法と同じレベルの感度と正確度を確保することができ、無細胞核酸の単一塩基変異を用いた他の解析でも活用することができるため、有用である。
Smart Images

Figure 0007927096000055 
Figure 0007927096000056 
Figure 0007927096000057
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for cancer diagnosis and cancer type prediction using single nucleotide mutations of cell-free nucleic acids. More specifically, the present invention relates to a method for cancer diagnosis and cancer type prediction using single nucleotide mutations of cell-free nucleic acids, which comprises: extracting nucleic acids from a biological sample, acquiring sequence information, extracting cancer-specific single nucleotide mutations through filtering based on aligned reads, then calculating the distribution of single nucleotide mutations and the frequency of each type of single nucleotide mutation, inputting the calculation results into a trained artificial intelligence model, and analyzing the output values. [Background Art]
[0002] In clinical practice, cancer diagnosis is usually confirmed by performing tissue biopsy after medical history investigation, physical examination and clinical evaluation. Cancer diagnosis based on clinical experiments is only possible when the number of cancer cells is more than 1 billion and the diameter of the cancer is 1 cm or more. In this case, cancer cells already have metastatic capacity, and at least half of them have already metastasized. In addition, since tissue biopsy is invasive, it causes considerable discomfort to patients, and it is often impossible to perform tissue biopsy when treating cancer patients, which is a problematic issue. Besides, in cancer screening, tumor markers are used to monitor substances produced directly or indirectly by cancer. However, even when cancer is present, more than half of the results of tumor marker screening are shown as normal, and even when there is no cancer, the results are frequently shown as positive, so the accuracy of tumor markers is limited.
[0003] Research into diagnosing cancer through single-nucleotide variant analysis of cell-free DNA is actively being conducted, and targeted sequencing methods that increase sequencing depth to target recurrent mutations frequently found in cancer have been widely used (Chabon JJ et al., nature, Vol. 580, pp. 245-251, 2020). However, it has recently become clear that examining a wider variety of mutations using whole-genome sequencing (WGS) data of cell-free DNA, even at a lower sequencing depth, is more sensitive than targeted sequencing (Zviran A et al., Nat Med, Vol. 26, pp. 1114-1124, 2020).
[0004] However, with current technology, there are issues with the accuracy of mutation detection using cell-free DNA WGS, making it unsuitable for cancer diagnosis. Instead, cell-free DNA WGS has only been used for cancer recurrence monitoring, where only the mutations are filtered and followed up if the patient's mutation information is available through tumor tissue WGS (Zviran A et al., Nat Med, Vol. 26, pp. 1114-1124, 2020). In other words, while cell-free DNA WGS is effective for cancer diagnosis, the lack of an effective filtering method has prevented its use in cancer diagnosis.
[0005] On the other hand, the mutation rate in cancer varies depending on the region of the gene, and furthermore, the mechanisms by which mutations occur and the patterns in which mutations accumulate differ depending on the type of cancer. It has been reported that cancer types can be distinguished using these characteristics, specifically the distribution of mutations in cancer tissue (regional mutation density) and the type of mutation (mutation signature) (Jia Wei et al., Nat. Communications, Vol. 11, no. 728, 2020). However, in this case, the theoretical possibility was explored in a situation where cancer diagnosis and cancer type differentiation had already been completed through surgery, and it was not applied to cancer diagnostic technology via cell-free DNA WGS.
[0006] On the other hand, while there are various patents (KR 10-2017-0185041, KR 10-2017-0144237, KR 10-2018-0124550) that utilize artificial neural networks in the bio-field, there is currently a lack of methods for predicting cancer types by analyzing mutations based on sequence analysis information of cell-free DNA (cfDNA) WGS in the blood, due to the inaccuracy of cancer-specific mutation discovery.
[0007] Therefore, the inventors have made diligent efforts to solve the aforementioned problems and develop a method for diagnosing cancer and predicting cancer type using single nucleotide mutations in cell-free nucleic acids with high sensitivity and accuracy. As a result, they have confirmed that by extracting nucleic acids from a biological sample, obtaining sequence information, and using the aligned reads, extracting cancer-specific single nucleotide mutations through filtering, calculating the distribution of single nucleotide mutations and their respective frequencies, and then inputting this into a trained artificial intelligence model and analyzing the output values, it is possible to diagnose cancer and predict cancer type with high sensitivity and accuracy, thus completing the present invention. [Overview of the Initiative] [Problems that the invention aims to solve]
[0008] The objective of this invention is to provide a method for cancer diagnosis and cancer type prediction using single nucleotide mutations in cell-free nucleic acids.
[0009] Another object of the present invention is to provide a cancer diagnostic and cancer type prediction device using single nucleotide mutations in cell-free nucleic acids.
[0010] Another object of the present invention is to provide a computer-readable recording medium that includes instructions configured to be executed by a processor that predicts cancer diagnosis and cancer type in the manner described above. [Means for solving the problem]
[0011] To achieve the above objective, the present invention provides a method for providing information for cancer diagnosis and cancer type prediction using single nucleotide variants, comprising: (a) the step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) the step of aligning the obtained sequence information (reads) in a reference genome database; (c) the step of discovering single nucleotide variants from the aligned sequence information (reads) and filtering them to extract cancer-specific single nucleotide variants; (d) the step of dividing the reference chromosome into certain intervals and calculating the distribution of single nucleotide variants (regional mutation density) extracted for each interval; (e) the step of calculating the frequency of each type of single nucleotide variant (mutation signature) of the extracted mutations; and (f) the step of inputting the distribution of single nucleotide variants calculated in step (d) and the frequency values of each type of single nucleotide variant calculated in step (e) into an artificial intelligence model trained to perform cancer diagnosis and comparing the output values with reference values.
[0012] The present invention also provides an artificial intelligence-based cancer diagnosis and cancer type prediction device, which includes: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequences in a standard chromosome sequence database; a mutation discovery unit that discovers single nucleotide variants from the aligned sequences and filters them to extract cancer-specific single nucleotide variants; a single nucleotide variant distribution calculation unit that divides a standard chromosome into certain intervals and calculates the distribution of single nucleotide variants (regional mutation density) extracted for each interval; a mutation frequency calculation unit that calculates the frequency of each type of single nucleotide variant (mutation signature) of the extracted mutations; a cancer diagnosis unit that inputs the calculated single nucleotide variant distribution values and mutation frequencies into an artificial intelligence model trained to perform cancer diagnosis and compares the output values with reference values to determine the presence or absence of cancer; and a cancer type prediction unit that inputs the single nucleotide variant distribution values and mutation frequencies of a sample determined to be cancerous into a second artificial intelligence model trained to classify cancer types and predicts the cancer type by comparing the output result values.
[0013] The present invention also includes instructions configured to be executed by a computer-readable recording medium and a processor for cancer diagnosis and cancer type prediction, including: (a) the step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) the step of aligning the obtained sequence information (reads) in a reference genome database; (c) the step of discovering single nucleotide variants from the aligned sequence information (reads) and filtering to extract cancer-specific single nucleotide variants; (d) the step of dividing the reference chromosome into certain intervals and calculating the distribution of single nucleotide variants extracted for each interval (regional mutation density); and (e) the step of classifying the extracted single nucleotide variants by type (mutation) The present invention provides a computer-readable recording medium that includes instructions configured to be executed by a processor that predicts the presence or absence of cancer and the type of cancer, through the steps of: (f) calculating the signature frequency; (d) inputting the distribution of single nucleotide mutations calculated in step (d) and the frequency values of each type of single nucleotide mutation calculated in step (e) into an artificial intelligence model trained to perform cancer diagnosis and comparing the output values with reference values to determine the presence or absence of cancer; and (g) inputting the distribution of single nucleotide mutations and the frequency values of each type of single nucleotide mutation of a sample determined to be cancerous in step (f) into a second artificial intelligence model trained to classify cancer types and comparing the output values to predict the type of cancer. [Effects of the Invention]
[0014] The cancer diagnosis and cancer type prediction method using single nucleotide mutations in cell-free nucleic acids according to the present invention is useful because it not only has higher sensitivity and accuracy compared to other methods for diagnosing cancer and predicting cancer type using genetic information from cell-free nucleic acids, but also ensures the same level of sensitivity and accuracy as cancer tissue cell-based methods, and can be utilized in other analyses using single nucleotide mutations in cell-free nucleic acids. [Brief explanation of the drawing]
[0015] [Figure 1]This is an overall flowchart for determining chromosomal abnormalities using single nucleotide mutations in cell-free nucleic acids according to the present invention. [Figure 2] The results of comparing the cancer diagnostic performance of a DNN model constructed according to one embodiment of the present invention with other models are shown, where (A) is the accuracy of cancer diagnostic performance and (B) is the cancer type discrimination performance. [Figure 3] (A) shows the results of comparing the cancer diagnostic performance of the DNN model constructed in one embodiment of the present invention with that of a conventional method for different types of cancer, and (B) shows the results of comparing that performance with that of a different stage of cancer progression. [Figure 4] (A) shows the results of comparing the cancer type discrimination performance of the DNN model constructed in one embodiment of the present invention with that of a conventional method for each cancer type, and (B) shows the results of comparing it with that of a cancer progression stage. [Figure 5] (A) shows the results of confirming the performance of a cancer diagnostic model constructed by changing the mutation discovery criteria according to one embodiment of the present invention, and (B) shows the results of confirming the cancer type discrimination performance. [Figure 6] (A) shows the results of confirming the performance of cancer diagnostic models using a WGS database of healthy individuals for filtering and a method using the technical characteristics of cfDNA, according to one embodiment of the present invention, and (B) shows the results of confirming the cancer type discrimination performance. [Figure 7] This is the result of confirming whether the cancer-type specific RMD value of cfDNA calculated by the method constructed according to one embodiment of the present invention accurately reflects the cancer-type specific RMD value in the tissue sample. [Modes for carrying out the invention]
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by skilled experts in the art to which this invention pertains. Generally, the nomenclature used herein and the experimental methods described below are well known and commonly used in the art.
[0017] Terms such as first, second, A, and B may be used to describe various components, but said components are not limited by the above terms, and are merely used for the purpose of distinguishing one component from another. For example, a first component can be named a second component, and similarly, a second component can also be named a first component, without departing from the scope of the technology described below. The term "and / or" includes a combination of a plurality of associated recited items or any one of the plurality of associated recited items.
[0018] As for the terms used in this specification, singular expressions should be understood to include plural expressions unless there is a clearly different interpretation from the context. Terms such as "comprise" mean that the stated features, numbers, steps, operations, components, parts, or combinations thereof are present, and it should be understood that they do not exclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0019] Prior to the detailed description of the drawings, it is clarified that the division into components in this specification is only based on the main function each component is responsible for. That is, two or more components described below may be combined into one component, or one component may be divided into two or more components according to more subdivided functions. In addition, each of the components described below may additionally perform some or all of the functions that are responsible for by other components in addition to its own main function, and it goes without saying that some functions of the main function that each component is responsible for may be exclusively performed by other components.
[0020] In addition, when carrying out a method or an operation method, each step constituting the method may be performed in a different order from the stated order unless the context clearly specifies a specific order. That is, each step may be performed in the same order as stated, may be performed substantially simultaneously, or may be performed in the reverse order.
[0021] In the present invention, after aligning sequence analysis data obtained from a sample to a reference genome, nucleic acid is extracted from a biological sample, sequence information is obtained, and cancer-specific single nucleotide variants are extracted through filtering based on the aligned reads, the distribution of single nucleotide variants and the frequency of each type of single nucleotide variant are calculated, and when this is input into a trained artificial intelligence model to analyze the calculated value, it has been confirmed that cancer diagnosis and cancer type prediction can be performed with high sensitivity and accuracy.
[0022] That is, in one embodiment of the present invention, after sequencing DNA extracted from blood and aligning it to a reference chromosome, cancer-specific single nucleotide variants are extracted from the aligned reads through filtering, the reference chromosome is divided into fixed intervals, the distribution of single nucleotide variants in each interval is calculated, the frequency of each type of single nucleotide variant is calculated, the distribution of single nucleotide variants and the frequency of each type of single nucleotide variant are input into an artificial intelligence model trained to perform cancer diagnosis, cancer diagnosis is performed by comparing the output value with a reference value, then the distribution of single nucleotide variants and the frequency of each type of single nucleotide variant of the sample determined to be cancer are input into a second artificial intelligence model trained to classify cancer types, and a method for determining the cancer type showing the highest value among the output values as the cancer type of the sample was developed (Figure 1).
[0023] Therefore, in one aspect, the present invention provides: (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) mining single nucleotide variants from said aligned sequence information (reads), and performing filtering to extract cancer-specific single nucleotide variants; (d) dividing said reference chromosome into fixed intervals, and calculating the regional mutation density of single nucleotide variants extracted for each interval; (e) A step of calculating the frequency of each type of single nucleotide variant (mutation signature) of the extracted mutation; and (f) The distribution of single nucleotide mutations calculated in step (d) and the frequency values of each type of single nucleotide mutation calculated in step (e) are input into an artificial intelligence model trained to perform cancer diagnosis, and the output values are compared with reference values; This relates to a method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, including those mentioned above.
[0024] In the present invention, the cancer may be a solid tumor or a hematological cancer, and may be selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphoblastic leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, thyroid cancer, liver cancer, stomach cancer, gallbladder cancer, biliary tract cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary origin, kidney cancer, and mesothelioma, and most preferably liver cancer or ovarian cancer, but is not limited to these.
[0025] In the present invention, The aforementioned step (a) is, (ai) The stage of obtaining nucleic acids from biological samples; (a-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using the salting-out method, column chromatography method, or beads method to obtain purified nucleic acids; (a-iii) The step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method; (a-iv) The step of reacting the prepared library with a next-generation sequencer; and (av) The stage of obtaining nucleic acid sequence information (reads) using a next-generation gene sequencing machine; It can be characterized by including
[0026] In the present invention, the step of obtaining sequence information in step (a) above is characterized by obtaining the isolated cell-free DNA by whole-genome sequencing at a read depth of 1 million to 100 million.
[0027] In the present invention, the biological sample means any substance, biological fluid, tissue, or cell obtained from or derived from an individual, such as whole blood, leukocytes, peripheral blood mononuclear cells, leukocyte buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid This may include, but is not limited to, lymphatic fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, hair, oral cells, placental cells, cerebrospinal fluid, and mixtures thereof.
[0028] In this invention, the term "reference population" refers to a reference population that can be compared, such as a standard nucleotide sequence database, and currently means a population of humans that does not have a specific disease or condition. In this invention, the standard nucleotide sequence in the standard chromosome sequence database of the reference population may be a reference chromosome registered with a public health organization such as the NCBI.
[0029] In the present invention, the nucleic acid in step (a) may be cell-free DNA, and more preferably circulating tumor DNA, but is not limited to these.
[0030] In the present invention, the next-generation sequencer can be used with any sequencing method known in the art. Sequencing of nucleic acids isolated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the sequence of one nucleotide from a cloned, extended proxy for each nucleic acid molecule, either for individual nucleic acid molecules or in a highly similar manner (e.g., 10⁵ or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species in a library can be estimated by measuring the relative occurrence of its genealogous sequences in the data produced by the sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in the literature included herein by reference (Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46).
[0031] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequence of individual nucleic acid molecules (e.g., Helicos BioSciences' HeliScope Gene Sequencing system and Pacific BioSciences' PacBio RS system). In other embodiments, sequencing methods such as high-density parallel short-read sequencing, which produces more bases of sequence per sequencing unit than other sequencing methods that produce fewer but longer reads (e.g., the Illumina Inc. Solexa sequencer located in San Diego, California), determine the nucleotide sequences of cloned proxies for individual nucleic acid molecules (e.g., the Illumina Inc. Solexa sequencer located in San Diego, California; Life Sciences (located in Branford, Connecticut) and Ion Torrent). Other methods or machines for next-generation sequencing are provided by, but are not limited to, 454 Life Sciences (located in Branford, Connecticut), Applied Biosystems (located in Foster City, California; SOLiD sequencers), Helicos Biosciences Corporation (located in Cambridge, Massachusetts), and emulsion and microfluidic sequencing techniques such as nanoinfusions (e.g., GnuBio infusions).
[0032] Platforms for next-generation sequencing include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX system, Illumina / Solexa's Genome Analyzer (GA), Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene Sequencing system, and Pacific Biosciences' PacBio RS system.
[0033] In the present invention, the sorting step in step (b) above is not limited to the above, but may be performed using the BWA algorithm and the Hg19 sequence.
[0034] In the present invention, the BWA algorithm may include, but is not limited to, BWA-ALN, BWA-SW, or Bowtie2.
[0035] In the present invention, the length of the sequence information (reads) in step (b) is 5 to 5000 bp, and the number of sequence information entries used may be 5,000 to 5 million, but is not limited to these values.
[0036] In the present invention, the filtering in step (c) above can be performed without limitation as long as it is a method that can distinguish between single nucleotide mutations that originated in healthy individuals and single nucleotide mutations that originated specifically in cancer. Preferably, the extracted single nucleotide mutations have a read depth of 3 or more in the mutation region and an average sequencing quality of 30 or more, but are not limited to these.
[0037] In the present invention, the mutation region refers to the precise location where a single nucleotide mutation occurs, and the read depth of the mutation region being 3 or more means that the number of reads aligned at that location is 3 or more.
[0038] In the present invention, the filtering in step (c) above may be further characterized by removing artifacts and germline mutations that occurred during the sequence analysis process, and the above process is i) A mutation detected in only one of the read pairs; ii) Mutations in which two or more types are detected at one location; iii) Mutations in which normal bases are not detected at each position; and iv) Mutations detected from healthy individual databases; This method can be characterized by removing one or more mutations selected from the group consisting of the following, but is not limited to these.
[0039] In the present invention, the healthy person database can be any database containing nucleotide sequence mutation information of healthy persons without limitation. Preferably, it may be a database containing cfDNA WGS data of healthy persons, WGS data of tissue samples, etc. More preferably, it may be a publicly available database such as dbSNP, 1000 Genome, Hapmap, ExAC, Gnomad, etc., but is not limited to these.
[0040] In the present invention, the interval in step (d) can be set arbitrarily to any interval in which the distribution of single nucleotide mutations can be calculated, preferably 100kb to 10Mb, more preferably 500kb to 5Mb, and most preferably 1Mb, but is not limited to these.
[0041] In the present invention, the step of calculating the regional mutation density (RMD) of the extracted single nucleotide variants in step (d) above can be characterized by being carried out by a method including the following steps: (di) A step of calculating the number of single nucleotide variants extracted for each interval, excluding intervals in which no mutations are detected above the reference value for the entire sample; and (d-ii) The next step is to normalize the calculated number by dividing it by the total number of variations in each interval.
[0042] In the present invention, the reference value can be used without limitation as long as it is a value that can significantly distinguish the extracted single base mutations, preferably 40-60%, more preferably 45-55%, and most preferably 50%, but is not limited to these values.
[0043] In the present invention, the interval excluding the interval in which no mutations are detected at or above the reference value of the entire sample means, when the reference value is 50%, the interval excluding the interval in which no single nucleotide mutations are present, extracted from 50% or more of the samples in the entire sample.
[0044] In the present invention, the interval can be characterized by being one or more selected from the intervals listed in Table 1.
[0045] In this invention, the regional mutation density (RMD) is used interchangeably with the background mutation rate, and refers to the mutation frequency calculated by dividing the entire genome into fixed intervals.
[0046] In this invention, the distribution of single-gene mutations by cancer type is a quantitative value indicating whether a region is prone to mutations or not in that particular cancer. Single-gene mutations in cancer are not uniformly distributed throughout the human genome. The amount of single-gene mutations accumulated differs across the entire genome, and the patterns of accumulation also differ significantly depending on the cancer type. Furthermore, epigenomic characteristics (Histone modification, replication time) are the main cause of the distribution of single-gene mutations by cancer type, and the distribution of single-gene mutations reflects the epigenomic characteristics of that particular cancer type.
[0047] The distribution of single-gene mutations differs across the entire genome and across cancer types, making it a useful indicator for cancer diagnosis and cancer type differentiation. The distribution of single-gene mutations can be used to determine whether a discovered mutation is located in a region with a high probability of occurrence in that particular cancer.
[0048] In the present invention, the step of calculating the frequency of each type of single nucleotide mutation (mutation signature) in step (e) above can be characterized by being carried out by a method including the following steps: (ei) The step of calculating the number of mutations for each type of mutation; and (1) A mutation in which cytosine (C) is replaced by thymine (T), adenine (A), or guanine (G); (2) Mutations in which thymine is replaced by cytosine, adenine, or guanine; (3) A mutation in (1) or (2) that includes an additional 5' direction base; (4) A mutation in (1) or (2) that includes an additional 3' direction base; and (5) Mutations further comprising one 5' directional base and one 3' directional base of a mutation in which adenine, guanine, cytosine, and thymine are substituted with different bases; (e-ii) The step of normalizing the calculated sum of mutation counts by dividing it by the grand total. In the present invention, the type of mutation can be characterized by being one or more selected from the mutations listed in Table 2.
[0049] In the present invention, the type of single nucleotide mutation (mutation signature) can be any mutation in which a normal base is mutated to another base and a functional abnormality of the gene occurs, and is preferably one or more selected from the group consisting of C->A, C->G, C->T, T->A, T->C, and T->G, but is not limited to these.
[0050] In this invention, C->A means confirming that the detected mutation is one in which the normal base C is mutated into the mutant base A, and C->G means confirming that the detected mutation is one in which the normal base C is mutated into the mutant base G, and the same applies to the others.
[0051] In the present invention, the reference value in step (f) above can be used without limitation as long as it is a value that can diagnose cancer, and is preferably 0.5, but is not limited thereto. If the reference value is 0.5, the present invention may be characterized in that cancer is determined to be present if the reference value is 0.5 or higher.
[0052] In this invention, the artificial intelligence model learns to produce an output result close to 1 when cancer is present, and an output result close to 0 when cancer is absent. Using 0.5 as a baseline, it is determined that cancer is present if the result is 0.5 or higher, and that cancer is absent if the result is 0.5 or lower, and performance is measured (accuracy of training, validation, and test evaluation).
[0053] It is obvious to any competent technician that the 0.5 threshold is a value that can be changed at any time. For example, to reduce false positives, a threshold higher than 0.5 can be set, making the criteria for determining the presence of cancer stricter. To reduce false negatives, a lower threshold can be measured, making the criteria for determining the presence of cancer slightly weaker.
[0054] In the present invention, (g) A step in which the distribution of single nucleotide variants and the frequency values of each type of single nucleotide variant in the sample identified as cancer are input into a second artificial intelligence model trained to distinguish between cancer types, and the output values are compared to predict the cancer type; It can be characterized by further including the following.
[0055] In the present invention, the comparison of the output values in step (g) above can be characterized by a method that includes a step of determining the cancer type showing the highest value among the output values as the cancer of the sample.
[0056] In the present invention, the artificial intelligence model can be used without limitation as long as it is a model capable of diagnosing cancer or identifying cancer types. Preferably, it may be an artificial neural network model, and more preferably, it may be selected from the group consisting of convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), and autoencoders. Most preferably, it may be a deep neural network, but is not limited to these.
[0057] In the present invention, when the artificial intelligence model is a DNN and learns binary classification, the loss function can be characterized by being a binary crossentropy, as shown in the following equation 1:
number
[0058] In the present invention, when the second artificial intelligence model is a DNN and learns multi-class classification, the loss function can be characterized by being the categorical cross-entropy shown in the following equation 2:
number
[0059] In the present invention, if the artificial intelligence model is a DNN, the learning process may be characterized by including the following steps: i) The stage of classifying the produced data into training, validation, and test data; In this process, the training data is used to train the DNN model, the validation data is used for hyperparameter tuning validation, and the performance evaluation data is used for performance evaluation after the optimal model has been produced. ii) The stage of constructing the optimal DNN model through hyperparameter tuning and the learning process; iii) The stage in which the performance of various models obtained through hyperparameter tuning is compared using validation data, and the model with the best performance on the validation data is determined to be the optimal model;
[0060] In the present invention, the hyperparameter tuning process is a process of optimizing the values of various parameters (number of layers, number of filters, etc.) that constitute the DNN model, and the hyperparameter tuning process is characterized by using hyperband optimization, Bayesian optimization, and grid search techniques.
[0061] In the present invention, the learning process is characterized by optimizing the intrinsic parameters (weights) of the DNN model using predetermined hyperparameters, and determining that the model is overfitting when the validation loss begins to increase compared to the learning loss, and interrupting model learning before that point.
[0062] In another aspect, the present invention is a decoding unit for extracting nucleic acids from a biological sample and decoding their sequence information; An alignment unit that sorts the decoded sequences into a standard chromosome sequence database; A mutation discovery unit that extracts single nucleotide variants from aligned sequences and filters them to extract cancer-specific single nucleotide variants; A single nucleotide mutation distribution calculation unit that divides a standard chromosome into fixed intervals and calculates the distribution of single nucleotide mutations (regional mutation density) extracted for each interval; A mutation frequency calculation unit that calculates the frequency of each type of single-nucleotide mutation (mutation signature) among the extracted mutations; A cancer diagnosis unit that inputs the calculated distribution values and mutation frequencies of single nucleotide variants into an artificial intelligence model trained to perform cancer diagnosis, and compares the output values with reference values to determine the presence or absence of cancer; and This invention relates to an artificial intelligence-based cancer diagnosis and cancer type prediction device that includes a cancer type prediction unit that inputs the distribution values and mutation frequencies of single nucleotide mutations in a sample diagnosed with cancer into a second artificial intelligence model trained to classify cancer types, and predicts the cancer type by comparing the output results.
[0063] In the present invention, the decoding unit may include a nucleic acid injection unit for injecting nucleic acids extracted from an independent device, and a sequence information analysis unit for analyzing the sequence information of the injected nucleic acids. Preferably, it may be an NGS analysis device, but is not limited thereto.
[0064] In the present invention, the decoding unit is characterized by receiving and decoding sequence information data generated by an independent device.
[0065] In another aspect, the present invention also includes a computer-readable recording medium and instructions configured to be executed by a processor that predicts cancer diagnosis and cancer type, (a) The step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) The step of extracting cancer-specific single nucleotide variants by searching for single nucleotide variants in the aligned sequence information (reads) and performing filtering; (d) A step of dividing the standard chromosome into certain intervals and calculating the distribution of single nucleotide mutations (regional mutation density) extracted for each interval; (e) A step of calculating the frequency of each type of single nucleotide variant (mutation signature) of the extracted mutations; (f) A step in which the distribution of single nucleotide mutations calculated in step (d) and the frequency values of each type of single nucleotide mutation calculated in step (e) are input into an artificial intelligence model trained to perform cancer diagnosis, and the output value is compared with a reference value to determine the presence or absence of cancer; and (g) A step in which the distribution of single nucleotide mutations and the frequency values of each type of single nucleotide mutation in the sample determined to be cancerous in step (f) above are input into a second artificial intelligence model trained to classify cancer types, and the output values are compared to predict the cancer type; The present invention relates to a computer-readable recording medium containing instructions configured to be executed by a processor that predicts the presence and type of cancer.
[0066] In other embodiments, the method according to the present invention can be implemented using a computer. In one embodiment, the computer includes one or more processors connected to a chipset. The chipset is also connected to memory, a storage device, a keyboard, a graphics adapter, a pointing device, and a network adapter, etc. In one embodiment, the performance of the chipset is enabled by a memory controller hub and an I / O controller hub. In other embodiments, the memory can be used by being directly connected to the processor instead of the chipset. The storage device is any device capable of storing data, including a hard drive, a CD-ROM (Compact Disk Read-Only Memory), a DVD, or other memory device. The memory is involved with the data and instructions used by the processor. The pointing device may be a mouse, a trackball, or other type of pointing device, and is used in combination with a keyboard to transfer input data to the computer system. The graphics adapter displays images and other information on a display. The network adapter is connected to the computer system via a short-range or long-range communication network. However, the computer used in this application is not limited to the configuration described above, and may lack some configurations, include additional configurations, and may be part of a Storage Area Network (SAN). The computer of this application may be configured to accommodate the execution of modules in a program for performing the method according to this application.
[0067] In this application, "module" can mean a functional and structural combination of hardware for implementing the technical concept of this application and software for driving said hardware. For example, the module can mean a logical unit of predetermined code and hardware resources for which said predetermined code is performed, and it will be obvious to those skilled in the art that it does not necessarily mean physically connected code or a single type of hardware.
[0068] Examples The present invention will be described in more detail below through examples. It will be obvious to those of ordinary skill in the art that these examples are solely for illustrative purposes and that the scope of the present invention should not be construed as being limited by these examples.
[0069] Example 1. DNA is extracted from blood and next-generation sequencing is performed. Blood samples (10 mL each) were collected from 471 healthy individuals, 151 ovarian cancer patients, and 131 liver cancer patients and stored in EDTA tubes. Within 2 hours of collection, the plasma portion was primarily centrifuged at 1200 g, 4°C, and 15 minutes. The purified plasma was then subjected to secondary centrifugation at 16000 g, 4°C, and 10 minutes, and the upper plasma layer was separated after removing the precipitate. Cell-free DNA was extracted from the separated plasma using the Tiangen micro DNA kit (Tiangen), and library preparation was performed using the MGIAsy cell-free DNA library prep set kit. Sequences were then performed using a DNBseq G400 equipped with (MGI) in 100-base paired-end mode. As a result, it was confirmed that approximately 170 million sequence readings were produced per sample.
[0070] Example 2. Extraction of single nucleotide variants, distribution of single nucleotide variants, and extraction of frequency features by type. 2-1. Filtering for cancer-specific mutation extraction The NGS data obtained in Example 1 was aligned to the reference chromosome (hg 19), and the resulting bam file was processed using the GATK pipeline. To obtain sample-specific mutation profiles, mutations were discovered using varscan (mutation caller).
[0071] The Varscan variant discovery criteria were applied very generously. Variant calling was performed with generous criteria, including the presence of at least one variant read, a total depth of 3 or more in the variant region, an average base quality of 30 or more, removal of the minimum variant allele frequency criterion, removal of the strand filter, and removal of the Varscan variant P value criterion (variant allele frequency represents the ratio of the number of reads in which a variant was discovered to the total number of reads in which the variant was discovered at the variant site).
[0072] After identifying all possible cancer-derived mutations using generous criteria, artifacts and germline mutations were removed using various standards. Four different methods were used to remove mutations in inaccurate locations.
[0073] Firstly, when sequencing was performed on locations where mutations were present in both the forward and reverse reads of a fragment, if a mutation was found in only one of the reads, it was removed. Secondly, if there were two or more mutations at a single location, it was removed. Thirdly, if the variant allele frequency was 1, it meant that the mutation was present in all DNA in the blood, so it was assumed that there was no probability that it was a tumor-derived mutation and it was removed.
[0074] Fourth, we removed mutations from various healthy individual mutation databases and blacklist regions. Blacklist regions are areas that are highly likely to be misaligned during alignment, such as repeats and centromeres. The blacklist regions used were those compiled in Haley M amemiya et al., Scientific report Vol. 9, no. 9354, 2019. In addition, to remove mutations that are highly likely to be healthy individual mutations, we used a public database of healthy individual mutations. The dbSNP (https: / / data.amerigeoss.org / ko_KR / dataset / dbsnp), 1000 Genome (https: / / www.internationalgenome.org / ), Hapmap (https: / / ftp.ncbi.nlm.nih.gov / hapmap / ), ExAC (https: / / gnomad.broadinstitute.org / downloads#exac-variants), and Gnomad (https: / / gnomad.broadinstitute.org / ) databases were used.
[0075] Furthermore, mutations in the cfDNA WGS database of 20,000 healthy individuals produced by Green Cross Corporation (South Korea) were filtered out because they were unlikely to be tumor-derived mutations. In addition, for the input values of the algorithm classifying cancer types, mutations found in the cell-free DNA WGS of 412 healthy individuals in Example 1 were also removed.
[0076] 2-2. Calculation of the distribution of single nucleotide variants The entire genome was segmented into 1Mb segments, and the regional mutation density (RMD) for each segment was calculated. The mutations extracted in Example 2-1 accounted for more than 50% of the total sample. Excluding segments where no mutations were present, a total of 2726 segments were used as input values for the algorithm. The number of mutations in each segment was calculated and divided by the total number of mutations across all 2726 segments for normalization. Finally, features for the distribution of 2726 single-gene mutations were generated, and the feature list is shown in Table 1 below. [Table 1] TIFF0007927096000006.tif255148TIFF0007927096000007.tif255148TIFF0007927096 000008.tif255148TIFF0007927096000009.tif255149TIFF0007927096000010.tif25514 8TIFF0007927096000011.tif255148TIFF0007927096000012.tif255148TIFF0007927096 000013.tif255148TIFF0007927096000014.tif255147TIFF0007927096000015.tif25514 9TIFF0007927096000016.tif255148TIFF0007927096000017.tif255148TIFF0007927096 000018.tif255148TIFF0007927096000019.tif255148TIFF0007927096000020.tif25514 9TIFF0007927096000021.tif255148TIFF0007927096000022.tif255148TIFF0007927096 000023.tif255148TIFF0007927096000024.tif255148TIFF0007927096000025.tif79151
[0077] 2-3. Calculation of the frequency of each type of single nucleotide variant The frequency of each type of single-gene mutation (mutation signature) was calculated across the entire genome. Four criteria were defined to classify the types of mutations.
[0078] Firstly, we categorized the types of mutations using the standard base and the type of altered base, defining a total of six basic mutation types (C>A, C>G, C>T, T>A, T>C, T>G). Secondly, we further considered one 5' base in the basic mutation types, defining 24 (4×6) mutation types. Thirdly, we further considered one 3' base in the basic mutation types, defining 24 (6×4) mutation types. Finally, we further considered one 5' base and one 3' base in the basic mutation types, determining 96 (4×6×4) mutation types, which are commonly used in mutation signature analysis.
[0079] The frequency of occurrence was calculated for each of the 150 types of mutations thus separated. Then, the total number of mutations for each of the four mutation classification methods was calculated and normalized by dividing it by the total sum of all mutations that occurred in all bases.
[0080] The defined types of mutations are as shown in Table 2 below. [Table 2] TIFF0007927096000027.tif20170
[0081] Ultimately, a total of 2,876 features were used as input values for the algorithm: 2,726 features representing the distribution of single-gene mutations and 150 features representing the types of single-gene mutations.
[0082] Example 3. DNN model construction and training process To develop algorithms for cancer diagnosis and cancer type classification from cfDNA, we used a total of 2876 features related to the distribution and types of single-gene mutations, which were previously obtained through analysis. Two artificial intelligence algorithms were developed.
[0083] First, a binary classification model was constructed to diagnose whether a person was healthy or had cancer. Second, a multiple classification model was constructed to classify different types of cancer. For algorithm training, the binary classification model used binary classification as the loss function, while the multiple classification model used category cross-entropy. A deep neural network artificial intelligence model was used for algorithm training.
[0084] The entire dataset was divided into training, validation, and performance evaluation datasets, and a model was trained using Bayesian optimization and hyperparameter tuning. The entire dataset was divided into five training, validation, and performance evaluation sets, and training was performed five times to create five algorithmic models. Then, predictions were made for each of the five performance evaluation datasets using these five algorithmic models, so that the entire dataset could be used as a performance evaluation dataset once each. The performance of the models was then evaluated using the prediction probability when the entire dataset was used as a performance evaluation dataset.
[0085] Example 4. Construction and performance verification of a deep learning model for cancer diagnosis and cancer type classification. To test the performance of the deep learning model constructed using the reads obtained in Example 1, we applied the methods of a conventional, known artificial intelligence model used for cancer diagnosis and cancer type discrimination (Cristiano, S. et al., Nature, Vol. 570(7761), pp. 385-389. 2019) to cfDNA and constructed a fragmentation pattern and copy number variation (CNV) based comparative cancer diagnosis and cancer type discrimination model based on the dataset from Example 1.
[0086] More specifically, the Fragment Pattern method involved GC correction of the entire genome, dividing it into 5Mb segments, and then normalizing the ratio of short fragments in each segment to the total number of fragments using Z-score normalization. Here, short fragments refer to fragments with lengths between 100bp and 150bp. The CNV method involved GC correction of the entire genome by dividing it into non-overlapping 50KB segments, calculating the depth for each segment, and then converting it to log2 values for use as input. XGBoost was used to train both the Fragment Pattern and CNV models.
[0087] To compare the performance of cancer diagnostic models, we examined the sensitivity at predictive probability thresholds when specificity was 95%, 98%, and 99%.
[0088] As a result, as shown in Figure 2, we confirmed that the performance of the cancer diagnostic model constructed in this invention is superior to that of conventional methods. Furthermore, as shown in Figure 3, we confirmed that the cancer diagnostic model constructed in this invention not only performs better in cancer diagnosis in all accuracy aspects, but also, as shown in Figure 3(B), while the performance of conventional methods is hindered in early cancer diagnosis (stage I), the cancer diagnostic model constructed in this invention shows superior performance even in early cancer diagnosis.
[0089] Furthermore, as a result of comparing the performance of cancer type discrimination models, as shown in Figure 4, it was confirmed that the cancer type discrimination model constructed in this invention is superior to conventional methods in cancer type discrimination performance at all stages.
[0090] Example 5. Confirmation of the effect of filtering conditions 5-1. Verification of the effectiveness of filtering criteria The inventors extracted all possible cancer-derived mutations using a generous criterion, and then removed artifacts and germline mutations using various criteria. The inventors compared the performance of three mutation extraction methods: a strict method, a less strict method, and a lenient method. The strict method extracted a mutation if a variant read was present in both the forward and reverse reads, while the less strict method extracted a mutation even if the variant read contained two or more mutations. The lenient method was the same as the method described in Example 2-1. After mutation extraction, the models were trained using the same filtering and learning processes, and then their performance was compared.
[0091] As a result, as shown in Figure 5, we confirmed that the best performance was achieved when all possible cancer-derived mutations were identified using generous criteria, and then filtered.
[0092] 5-2. Verification of the effectiveness of the filtering database In this invention, a method was used to filter mutations that appear in the cfDNA and tissues of healthy individuals, along with the discovery of generous mutations. By using mutations discovered in a large-scale healthy-person cfDNA / tissue WGS for mutation filtering, it was expected that artifacts and germ cell mutations that may occur in cfDNA could be effectively removed.
[0093] Since there was no large-scale database of healthy individual cfDNA mutations in public databases, we used the WGS of 20,000 healthy individuals produced by GreenCross.
[0094] As a result, as shown in Figure 6, we confirmed that using healthy individual cfDNA and tissue mutations improved performance compared to not using them. Therefore, the prediction model of the present invention utilized all public databases related to healthy individual cfDNA mutations and healthy individual tissue mutations.
[0095] Example 6. RMD distribution of cfDNA mutations in cancer-specific mutation regions When cfDNA mutations were discovered using the cfDNA mutation discovery method developed in this invention and RMD values were calculated, it was confirmed that the method accurately reflected the actual characteristics and distribution of the cancer in question.
[0096] In the PCAWG, a large-scale cohort of cancer genes, tumor mutations were discovered sample by sample using cancer tissue WGS for ovarian cancer and liver cancer. After calculating RMD values in 1 Mbp bin units, edgeR was used to search for regions with a high and low incidence of mutations specific to each cancer type. Then, it was confirmed whether regions with a high incidence of mutations specific to the cancer tissue also had high RMD values in cfDNA, and regions with low mutations specific to the cancer tissue also had low cfDNA RMD values.
[0097] As a result, as shown in Figure 7, it was confirmed that cfDNA reflected the characteristics of the RMD region of both ovarian and liver cancers in the same way as the tissue samples. The liver and ovaries on the X axis of the figure represent the actual cancers in the cfDNA samples. The region type refers to the region with a high / low mutation rate specific to that cancer, as defined using PCAWG data.
[0098] Although specific parts of the present invention have been described in detail above, it will be clear to those with ordinary skill in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the invention. Therefore, the substantial scope of the invention is defined by the appended claims and their equivalents.
Claims
1. (a) The step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) The step of extracting cancer-specific single nucleotide variants by searching for single nucleotide variants in the sorted sequence information (reads) and performing filtering; (d) A step of dividing the standard chromosome into certain intervals and calculating the distribution of single nucleotide mutations (regional mutation density, RMD) extracted for each interval; (e) A step of calculating the mutation signature frequency of each type of single nucleotide variant of the extracted mutation; and (f) The distribution of single nucleotide mutations calculated in step (d) and the frequency values of each type of single nucleotide mutation calculated in step (e) are input into an artificial intelligence model trained to perform cancer diagnosis, and the output probability values are compared with reference values; Includes, The step of performing the aforementioned filtering is, (ci) A stage in which single nucleotide variants are extracted in which the read depth of the mutant region containing the discovered single nucleotide variant is 3 or greater and the average sequencing quality is 30 or greater; and (c-ii) The step of removing artifacts and germline mutations that occurred during the sequence analysis process; Includes, The step of calculating the distribution of the extracted single nucleotide variants (regional mutation density) is as follows: (di) A step of calculating the number of single nucleotide variants extracted for each interval, excluding intervals in which no mutations are detected above the reference value for the entire sample; and (d-ii) The step of normalizing the calculated number by dividing it by the total number of variations in each interval; This is done in a way that includes, The step of calculating the frequency of each type of single nucleotide mutation (mutation signature) extracted above is: (ei) The step of calculating the number of mutations for each type of mutation below; and (1) Mutations in which cytosine (C) is replaced by thymine (T), adenine (A), or guanine (G); (2) Mutations in which thymine is replaced by cytosine, adenine, or guanine; (3) A mutation in (1) or (2) that includes an additional 5' base; (4) A mutation in (1) or (2) that includes an additional 3' direction base; and (5) Mutations further comprising one 5' directional base and one 3' directional base of a mutation in which adenine, guanine, cytosine, and thymine are substituted with different bases; (e-ii) The step of normalization by dividing the total number of calculated mutations by the total sum of all mutations that occurred in all bases; A computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized by being carried out by a method including the above.
2. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations according to claim 1, characterized in that step (a) is carried out by a method including the following steps: (ai) The stage of obtaining nucleic acids from biological samples; (a-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using the salting-out method, column chromatography method, or beads method to obtain purified nucleic acids; (a-iii) The step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method; (a-iv) The step of reacting the prepared library with a next-generation sequencer; and (av) The stage in which nucleic acid sequence information (reads) is obtained using a next-generation gene sequencing machine.
3. The step of removing the aforementioned artifact and germline mutation is: i) Mutations detected in only one of the read pairs; ii) Mutations in which two or more types are detected at one location; iii) Mutations in which normal bases are not detected at each position; and iv) Variations detected from healthy individual databases; A computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized by removing one or more mutations selected from the group consisting of the following.
4. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized in that the interval of step (d) is 100kb to 10Mb, as described in claim 1.
5. A computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized in that the reference value is 40 to 60%.
6. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized in that the interval is one or more selected from the intervals listed in Table 1 below, as described in claim 1. Table 1
7. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations according to claim 1, characterized in that the type of mutation is one or more selected from the mutations listed in Table 2 below. Table 2
8. A computer implementation method for providing information for cancer diagnosis and cancer type prediction using a single nucleotide mutation, characterized in that the reference value for the (f) stage is 0.5, and if it is 0.5 or higher, it is determined to be cancer.
9. (g) A step in which the distribution of single nucleotide mutations calculated in step (d) above and the frequency values of each type of single nucleotide mutation calculated in step (e) above are input into a second artificial intelligence model trained to distinguish between cancer types, and the output probability values are compared to predict the cancer type; A computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations according to claim 1, further comprising:
10. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, as described in claim 9, characterized in that the comparison of the output values in step (g) above is performed by a method that includes a step of determining the cancer type showing the highest value among the output values as the cancer of the sample.
11. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, as described in claim 1, characterized in that the artificial intelligence model in step (f) is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and an autoencoder.
12. A computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized in that, when the artificial intelligence model is a DNN and learns binary classification, the loss function is binary crossentropy as shown in the following formula 1: [Math 1] Here, in binary classification, N is the total number of samples. 【number】 is the probability value that the model predicted the i-th input value would be close to class 1, and y i This is the actual class of the i-th input value.
13. The computer implementation method for providing information for cancer diagnosis and cancer type prediction using single nucleotide mutations, characterized in that, when the second artificial intelligence model is a DNN and learns multi-class classification, the loss function is the categorical cross-entropy shown in the following formula 2: [Math 2] Here, in the category cross-entropy, N is the total number of samples, J is the total number of classes, and y j This value indicates the actual class of the sample; it is 1 if the actual class is j, and 0 if the actual class is not j. 【number】 This value represents the probability that the sample is predicted to be in class j, with a value closer to 1 indicating a higher probability that the sample belongs to that class.
14. A decoding unit that extracts nucleic acids from biological samples and decodes their sequence information; Alignment unit that sorts the decoded sequences into a standard chromosome sequence database; A mutation discovery unit that extracts single nucleotide variants from aligned sequences and filters them to extract cancer-specific single nucleotide variants; A single nucleotide mutation distribution calculation unit that divides a standard chromosome into fixed intervals and calculates the distribution of single nucleotide mutations (regional mutation density) extracted for each interval; A mutation frequency calculation unit that calculates the frequency of each type of single-nucleotide mutation (mutation signature) of the extracted mutations; A cancer diagnosis unit that inputs the calculated distribution values and mutation frequencies of single nucleotide variants into an artificial intelligence model trained to perform cancer diagnosis, and compares the output probability values with a reference value to determine the presence or absence of cancer; and A cancer type prediction unit inputs the distribution values and mutation frequencies of single nucleotide mutations in samples diagnosed as cancer into a second artificial intelligence model trained to classify cancer types, and predicts the cancer type by comparing the output probability values; An artificial intelligence-based cancer diagnosis and cancer type prediction device, including The step of performing filtering in the mutation discovery unit is as follows: (1-1) The stage of extracting single nucleotide mutations in which the read depth of the mutant region containing the discovered single nucleotide mutation is 3 or more and the average sequencing quality is 30 or more; and (1-2) The step of removing artifacts and germline mutations that occurred during the sequence analysis process; Includes, In the unit for calculating the distribution of single nucleotide mutations, the step of calculating the distribution of extracted single nucleotide mutations (regional mutation density) is as follows: (2-1) The step of calculating the number of single nucleotide variants extracted for each interval, excluding intervals in which no mutations are detected above the reference value for the entire sample; and (2-2) The next step is to normalize the calculated number by dividing it by the total number of mutations in each interval; This is done in a way that includes, The step in the mutation frequency calculation unit to calculate the frequency of each type of extracted single nucleotide mutation (mutation signature) is as follows: (3-1) The step of calculating the number of mutations for each type of mutation; and (1) Mutations in which cytosine (C) is replaced by thymine (T), adenine (A), or guanine (G); (2) Mutations in which thymine is replaced by cytosine, adenine, or guanine; (3) A mutation in (1) or (2) that includes an additional 5' base; (4) A mutation in (1) or (2) that includes an additional 3' direction base; and (5) Mutations further comprising one 5' directional base and one 3' directional base of a mutation in which adenine, guanine, cytosine, and thymine are substituted with different bases; (3-2) The step of normalization, in which the total number of calculated mutations is divided by the total sum of all mutations that occurred in all bases; The artificial intelligence-based cancer diagnosis and cancer type prediction device, characterized by being carried out by a method including the following.
15. A computer-readable recording medium, comprising instructions configured to be executed by a processor for cancer diagnosis and cancer type prediction, (a) The step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) The step of extracting cancer-specific single nucleotide variants by searching for single nucleotide variants in the sorted sequence information (reads) and performing filtering; (d) A step of dividing the standard chromosome into certain intervals and calculating the distribution of single nucleotide mutations (regional mutation density) extracted for each interval; (e) A step of calculating the mutation signature frequency of each type of single nucleotide variant of the extracted mutation; (f) A step in which the distribution of single nucleotide mutations calculated in step (d) and the frequency values of each type of single nucleotide mutation calculated in step (e) are input into an artificial intelligence model trained to perform cancer diagnosis, and the output probability values are compared with reference values to determine the presence or absence of cancer; and (g) A step in which the distribution of single nucleotide mutations and the frequency values of each type of single nucleotide mutation in the sample determined to be cancerous in step (f) above are input into a second artificial intelligence model trained to classify cancer types, and the output probability values are compared to predict the cancer type; A computer-readable recording medium comprising instructions configured to be executed by a processor that predicts the presence and type of cancer, The step of performing the aforementioned filtering is, (ci) A stage in which single nucleotide variants are extracted in which the read depth of the mutant region containing the discovered single nucleotide variant is 3 or greater and the average sequencing quality is 30 or greater; and (c-ii) The step of removing artifacts and germline mutations that occurred during the sequence analysis process; Includes, The step of calculating the distribution of the extracted single nucleotide variants (regional mutation density) is as follows: (di) A step of calculating the number of single nucleotide variants extracted for each interval, excluding intervals in which no mutations are detected above the reference value for the entire sample; and (d-ii) The step of normalizing the calculated number by dividing it by the total number of variations in each interval; This is done in a way that includes, The step of calculating the frequency of each type of single nucleotide mutation (mutation signature) extracted above is: (ei) The step of calculating the number of mutations for each type of mutation below; and (1) Mutations in which cytosine (C) is replaced by thymine (T), adenine (A), or guanine (G); (2) Mutations in which thymine is replaced by cytosine, adenine, or guanine; (3) A mutation in (1) or (2) that includes an additional 5' base; (4) A mutation in (1) or (2) that includes an additional 3' direction base; and (5) Mutations further comprising one 5' directional base and one 3' directional base of a mutation in which adenine, guanine, cytosine, and thymine are substituted with different bases; (e-ii) The step of normalization by dividing the total number of calculated mutations by the total sum of all mutations that occurred in all bases; The computer-readable recording medium, characterized by being operated in a manner that includes the above.
Citation Information
Patent Citations
Ultrasound-sensitive detection of circulating tumor DNA by genome-wide integration
JP2021519607A
Multi-Assay Prediction Model for Cancer Detection
US20190316209A1
Ultra-sensitive detection of circulating tumor DNA through genome-wide integration
US20210043275A1
Methods and systems for detecting microsatellite instability of a cancer in a liquid biopsy assay
US20210098078A1
Methods and Systems for Analyzing Nucleic Acid Molecules
US20210172022A1