Method for early diagnosis of cancer based on artificial intelligence using cell-free DNA distribution of tissue-specific regulatory region
The AI-based method addresses the limitations of existing cancer diagnosis techniques by imaging and analyzing tissue-specific regulatory regions in cfDNA, enabling accurate early cancer detection with high sensitivity and specificity.
Patent Information
- Application Number
- JP2025188274
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-28
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-25
AI Technical Summary
Existing methods for early cancer diagnosis using liquid biopsies, such as analyzing cfDNA mutations and transcription factor binding patterns, face challenges with low reliability and the need for large data volumes, making it difficult to accurately detect early-stage cancers.
An artificial intelligence-based method that images and analyzes the distribution of cell-free nucleic acids in tissue-specific regulatory regions, using an AI model to distinguish between normal and cancerous images by generating image data from selected nucleic acid fragments.
Achieves high sensitivity and accuracy in early cancer diagnosis, particularly for solid and blood cancers, by leveraging tissue-specific regulatory regions and AI models to differentiate between normal and cancerous samples.
Smart Images

Figure 2026032004000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an artificial intelligence-based method for early cancer diagnosis, and more particularly to an artificial intelligence-based method for early cancer diagnosis using a method of inputting information on cell-free DNA distribution of tissue-specific regulatory regions into an artificial intelligence model trained to perform early cancer diagnosis and analyzing the information. [Background technology]
[0002] Liquid biopsy techniques are being used to detect chromosomal abnormalities using cell-free DNA (cfDNA; cell-free DNA) present in plasma due to cell necrosis, apoptosis, and secretion. In particular, circulating cfDNA derived from tumor cells contains tumor-specific chromosomal abnormalities and mutations not present in normal cells. Its short half-life of approximately 2 hours allows it to accurately reflect the current state of the tumor. Furthermore, because it is noninvasive and can be collected repeatedly, cfDNA has emerged as a potential tumor-specific biomarker for various cancer-related fields, including cancer diagnosis, monitoring, and prognosis.
[0003] Taking advantage of the ability to diagnose cancer with a simple blood test, many researchers are working to use liquid biopsies for early diagnosis. Because cancer is a disease caused by the gradual accumulation of mutations in DNA, cancer-derived cfDNA contains mutations that differ from those in normal individuals. Using these characteristics, cancer can be diagnosed by detecting DNA containing mutations. However, among the 3 billion human genomes, only a very small number of mutations are commonly found in cancer cells. Furthermore, many people develop cancer without these mutations, making early cancer diagnosis using mutations difficult.
[0004] Recently, methods have been developed that obtain full-length genomic data from cfDNA, derive transcription start site profiles based on read depth, and then use SVM to learn whether each gene is expressed (Ulz, P., Thallinger, G., Auer, M. et al. Nat. Genet. Vol. 48, pp. 1273-1278, 2016). Other methods include analyzing transcription factor binding patterns based on cfDNA fragmentation patterns to diagnose cancer early or classify cancer types (Ulz, P. et al. Nat. Commun. Vol. 10, 4666, 2019). However, these methods have drawbacks, such as relatively low reliability or the need for large amounts of data.
[0005] Against this technical background, the present inventors have made intensive efforts to develop an artificial intelligence-based method for early cancer diagnosis. As a result, they have confirmed that early cancer diagnosis can be achieved with high sensitivity and accuracy when the distribution of cell-free nucleic acids in tissue-specific regulatory regions is imaged and input into an artificial intelligence model trained to perform early cancer diagnosis, thereby completing the present invention. Summary of the Invention
[0006] An object of the present invention is to provide a method for providing information for early cancer diagnosis based on artificial intelligence.
[0007] Another object of the present invention is to provide an information providing device for early cancer diagnosis based on artificial intelligence.
[0008] Another object of the present invention is to provide a computer readable medium containing instructions adapted to be executed by a processor that provides information for early cancer diagnosis in said method.
[0009] Another object of the present invention is to provide an artificial intelligence-based method for early cancer diagnosis.
[0010] Another object of the present invention is to provide an artificial intelligence-based early cancer diagnosis device.
[0011] Another object of the present invention is to provide a computer readable medium containing instructions adapted to be executed by a processor to perform early cancer diagnosis in the above-described manner.
[0012] To achieve the above object, the present invention provides a method for providing information for early cancer diagnosis based on artificial intelligence, including the steps of: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) selecting nucleic acid fragments of a regulatory region based on the aligned sequence information (reads); (d) generating the selected nucleic acid fragments as image data; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancer images, analyzing the data, and comparing it with a cut-off value to determine the presence or absence of cancer.
[0013] The present invention also provides an information providing device for early cancer diagnosis based on artificial intelligence, including a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment selection unit that selects nucleic acid fragments of regulatory regions from nucleic acid fragments based on the aligned sequences; a data generation unit that generates the selected nucleic acid fragments as image data; and an information providing unit that inputs the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzes the data, and provides information for early cancer diagnosis.
[0014] The present invention also provides a computer-readable storage medium including instructions configured to be executed by a processor that provides information for the early diagnosis of cancer, the instructions being configured to be executed by a processor, the computer providing information for the early diagnosis of cancer through the steps of: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) selecting nucleic acid fragments of a regulatory region based on the aligned sequence information (reads); (d) generating the selected nucleic acid fragments as image data; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancer images, analyzing the image data, and comparing it with a cut-off value to determine the presence or absence of cancer.
[0015] The present invention also provides a method for early cancer diagnosis, comprising the steps of: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) selecting nucleic acid fragments of regulatory regions based on the aligned sequence information (reads); (d) generating image data from the selected nucleic acid fragments; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancer images, analyzing the image data, and comparing it with a cut-off value to determine the presence of cancer if the cut-off value is exceeded.
[0016] The present invention also provides an AI-based early cancer diagnosis device, including a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment selection unit that selects nucleic acid fragments of regulatory regions from nucleic acid fragments based on the aligned sequences; a data generation unit that generates the selected nucleic acid fragments as image data; and a cancer diagnosis unit that inputs the generated image data into an AI model trained to distinguish between normal images and cancer images, and determines the presence of cancer if the image data exceeds a reference value.
[0017] The present invention also provides a computer-readable storage medium containing instructions configured to be executed by a processor for performing early cancer diagnosis, the instructions being configured to be executed by a processor for performing early cancer diagnosis, through the steps of: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) selecting nucleic acid fragments of a regulatory region based on the aligned sequence information (reads); (d) generating the selected nucleic acid fragments as image data; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancer images, analyzing the input data, and determining that cancer is present if the cut-off value is exceeded. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is an overall flowchart for carrying out the method of the present invention. [Figure 2] Schematic diagram and actual examples showing differences in nucleosome positioning in regulatory regions by tissue. [Figure 3] FIG. 1 is a schematic representation of regulator data for various tissues. [Figure 4] FIG. 1 is a schematic diagram showing a method for discovering tissue-specific regulatory factors. [Figure 5]FIG. 1 shows the principle of generating an image of the cfDNA distribution of regulatory regions obtained by one embodiment of the present invention to input into an artificial intelligence model. [Figure 6] FIG. 1 is a diagram illustrating an algorithm of an artificial intelligence model constructed according to one embodiment of the present invention. [Figure 7] 1 shows the results of confirming the performance of a liver cancer prediction model constructed according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs. Generally, the nomenclature used herein and the laboratory procedures described below are well known and commonly used in the art.
[0020] Terms such as "first," "second," "A," and "B" may be used to describe various components, but the components are not limited by these terms and are used solely to distinguish one component from another. For example, a first component can be designated a "second component," and similarly, a second component can be designated a "first component," without departing from the scope of the technology described below. The term "and / or" includes a combination of multiple related listed items or any of multiple related listed items.
[0021] In terms used in this specification, singular terms should be understood to include plural terms unless the context clearly dictates otherwise, and terms such as "comprise" should be understood to mean the presence of a stated feature, number, step, operation, component, part, or combination thereof, but not to exclude the possible presence or addition of one or more other features, numbers, step operations, components, parts, or combinations thereof.
[0022] Before proceeding to a detailed description of the drawings, it should be made clear that the division of components in this specification is merely a division according to the main function of each component. That is, two or more components described below may be combined into one component, or one component may be divided into two or more components according to more specific functions. Furthermore, each of the components described below may perform some or all of the functions performed by other components in addition to its own main function, and of course, some of the main functions performed by each component may be exclusively performed by other components.
[0023] Furthermore, in performing a method or method of operation, the steps making up the method may be performed in an order different from that stated, unless the context clearly dictates a particular order. That is, the steps may be performed in the same order as stated, substantially simultaneously, or in the reverse order.
[0024] In the present invention, we have attempted to confirm that early cancer diagnosis can be achieved with high sensitivity and accuracy by aligning sequence analysis data obtained from a sample to a reference genome, selecting nucleic acid fragments of regulatory regions from the aligned nucleic acid fragments, generating image data, and inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancer images.
[0025] That is, in one embodiment of the present invention, nucleic acids were extracted from liquid biopsies obtained from 187 normal individuals, 12 patients with early-stage liver cancer, and 150 patients with late-stage liver cancer. cfDNA sequencing was then performed to select nucleic acid fragments corresponding to liver-specific regulatory regions, which were then generated as image data. An artificial intelligence learning model for early diagnosis of liver cancer was constructed using the data from 187 normal individuals and 150 patients with late-stage liver cancer. The performance of the learning model was confirmed using data from 12 patients with early-stage liver cancer, and it was confirmed that the constructed learning model could distinguish with high accuracy between images of normal individuals, liver cancer patients, and patients with early-stage liver cancer (Figures 7 and 8).
[0026] The term "reads" as used herein refers to a single nucleic acid fragment whose sequence information has been analyzed using various methods known in the art. Therefore, the terms "sequence information" and "read" as used herein have the same meaning in that they are the result of obtaining sequence information through a sequencing process.
[0027] The term "regulatory region" as used herein refers to any location on a chromosome that can regulate gene expression, and refers to a region where RNA polymerase and regulatory proteins bind for RNA synthesis. Preferably, the region may include, but is not limited to, a promoter, an enhancer, a silencer, and an insulator.
[0028] In the present invention, the term "NFR (Nucleosome Free Region)" refers to the same region as the regulatory region, but specifically refers to a region within the regulatory region that is free of nucleosomes. For example, in the case of an enhancer region consisting of 1-147 bp of the first nucleosome, 148-346 bp of internucleosomal nucleic acid, 347-493 bp of the second nucleosome, 494-692 bp of internucleosomal nucleic acid, 693-839 bp of the third nucleosome, and 840-1039 bp of internucleosomal nucleic acid, if, upon transcription initiation, the second nucleosome is released and a transcription regulatory protein can bind, the NFR would be the 148-692 bp region.
[0029] Furthermore, while transcription proceeds in the manner described above in normal samples, NFRs may not be present in cancer samples, or nucleosomes in other regions may be separated, generating other NFRs, and NFRs that are not present in normal samples may be newly generated in cancer samples.
[0030] Furthermore, while transcription proceeds in the manner described above in blood cells, NFRs may not be present in other tissues (e.g., the liver), or nucleosomes in other regions may be separated to generate different NFRs, and NFRs that are not present in blood cells may be newly generated in cancer samples.
[0031] Therefore, from one aspect, the present invention provides: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A method for providing information for early cancer diagnosis based on artificial intelligence, including a step of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzing the image data, and comparing the image data with a cut-off value to determine whether or not cancer is present.
[0032] In the present invention, the cancer may be a solid cancer or a blood cancer, and is preferably selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphoid leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colorectal / rectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, liver cancer, thyroid cancer, gastric cancer, gallbladder cancer, biliary tract cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary site, kidney cancer, esophageal cancer, and mesothelioma, and more preferably, is liver cancer, but is not limited to these.
[0033] In the present invention, the step of obtaining sequence information in step (a) may be performed by a method including the following steps: (ai) obtaining nucleic acids from a biological sample; (a-ii) removing proteins, fats, and other residues from the collected nucleic acid using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid; (a-iii) preparing a single-end sequencing or paired-end sequencing library from the purified nucleic acids or the nucleic acids randomly fragmented by enzymatic cleavage, crushing, or hydroshear method; (a-iv) subjecting the prepared library to a next-generation sequencer; and (av) A step of obtaining nucleic acid sequence information (reads) using a next-generation sequencer.
[0034] In the present invention, the step of obtaining sequence information in step (a) may be characterized by obtaining the sequence information by full-length genome sequencing of the separated cell-free DNA at a depth of 1 million to 100 million reads.
[0035] In the present invention, the biological sample refers to any substance, biological fluid, tissue, or cell obtained from or derived from an individual, such as whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic juice, etc. The fluid may include, but is not limited to, blood, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, semen, hair, saliva, urine, buccal cells, placental cells, cerebrospinal fluid, and mixtures thereof.
[0036] The term "reference population" as used herein refers to a reference population that can be compared, such as a standard nucleotide sequence database, and refers to a group of people who are currently free of a specific disease or condition. In the present invention, the standard nucleotide sequence in the standard chromosome sequence database of the reference population may be a reference chromosome registered with a public health organization such as NCBI.
[0037] In the present invention, the nucleic acid in step (a) may be, but is not limited to, cell-free DNA, more preferably circulating tumor cell DNA.
[0038] In the present invention, the next-generation sequencer may be used in any sequencing method known in the art. Sequencing of nucleic acids separated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of individual nucleic acid molecules or one of their clonally extended proxies in a highly similar manner (e.g., 10 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species in a library can be estimated by measuring the relative occurrence of their cognate sequences in the data generated by the sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference.
[0039] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequence of each nucleic acid molecule (for example, Helicos BioSciences' HeliScope Gene Sequencing system and Pacific Biosciences' PacBio RS system).In another embodiment, sequencing, for example, massively parallel short-read sequencing (for example, Illumina Inc.'s Solexa sequencer, San Diego, California) method, which generates more bases of sequence per sequencing unit than other sequencing methods that generate fewer but longer reads, determines the nucleotide sequence of the clone-extended proxy for each nucleic acid molecule (for example, Illumina Inc.'s Solexa sequencer, San Diego, California; 454 Life Sciences (Branford, Connecticut) and Ion Torrent). Other methods or machines for next-generation sequencing include, but are not limited to, those provided by 454 Life Sciences (Branford, Connecticut), Applied Biosystems (Foster City, California; SOLiD sequencer), Helicos Biosciences Corporation (Cambridge, Massachusetts), and emulsion and microflow sequencing nano-infusion (e.g., GnuBio infusion).
[0040] Platforms for next-generation sequencing include, but are not limited to, the Roche / 454 Genome Sequencer (GS) FLX system, the Illumina / Solexa Genome Analyzer (GA), the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, the Polonator G.007 system, the Helicos BioSciences HeliScope Gene Sequencing system, and the Pacific Biosciences PacBio RS system.
[0041] In the present invention, the alignment step of step (b) may be performed using, but is not limited to, the BWA algorithm and the Hg19 sequence.
[0042] In the present invention, the BWA algorithm may include, but is not limited to, BWA-ALN, BWA-SW, or Bowtie2.
[0043] In the present invention, prior to performing step (c), the method may further include a step of selecting reads whose mapping quality score of the aligned nucleic acid fragments is equal to or greater than a reference value, and the reference value may be any value that can confirm the quality of the aligned nucleic acid fragments, and may be preferably 50 to 70 points, more preferably 60 points, but is not limited thereto.
[0044] In the present invention, the regulatory region in step (c) may be a tissue-specific regulatory region.
[0045] In the present invention, the tissue-specific regulatory region may be characterized in that the length and / or amount of cell-free DNA detected varies depending on the tissue.
[0046] In the present invention, the tissue-specific regulatory region may be such that the length and / or amount of cell-free DNA detected in a particular tissue, for example, the liver, differs from the length and / or amount of cell-free DNA detected in tissues such as blood, brain, stomach, and heart, or the length and / or amount of cell-free DNA detected in solid tissues (such as the brain, liver, stomach, lung, and heart) differs from that detected in blood tissues (such as blood cells and bone marrow).
[0047] In the present invention, the tissue-specific regulatory region may more specifically refer to a region in the regulatory region where no nucleosomes exist, i.e., a nucleosome-free region (NFR), but is not limited thereto.
[0048] In the present invention, the tissue-specific regulatory regions may be used in any number that can generate image data to be input into an artificial intelligence model, and preferably may be, but are not limited to, 10, 100, 1,000, 10,000, 20,000 to 50,000.
[0049] In the present invention, the image in step (d) may be any image that can be used for learning by the artificial intelligence model, and may preferably be a one-dimensional image in which the x-axis is composed of the number of reads for each sequence position of the selected nucleic acid fragment, but is not limited thereto.
[0050] In the present invention, the image in step (d) is an arrangement of values obtained by stacking cfDNA reads for each base pair, and may show a structure such as [0.91, 0.93, ~~, 0.73, 0.86]. If a total of 2000 bp is used, ±1000 bp from the position selected as the tissue-specific regulatory region, the number in [ ] will be 2000.
[0051] In the present invention, the artificial intelligence model in step (e) may be any learning model that can learn to distinguish between normal images and cancer images, and may be preferably an artificial neural network, and more preferably a convolutional neural network (CNN) or a recurrent neural network (RNN), but is not limited to these.
[0052] In the present invention, the reference value in step (e) may be any value that allows early diagnosis of cancer, and may preferably be 0.5, but is not limited thereto. If the reference value is 0.5, it may be characterized in that a value of 0.5 or more is determined to be cancer.
[0053] In the present invention, the artificial intelligence model was trained so that the output result was close to 1 if there was cancer and close to 0 if there was no cancer, and performance was measured (Training, validation, test accuracy) by determining that there was cancer if the result was 0.5 or higher, and that there was no cancer if the result was 0.5 or lower, based on a standard of 0.5.
[0054] It is obvious to any skilled technician that the 0.5 reference value can be changed at any time. For example, if you want to reduce false positives, you can set a reference value higher than 0.5 and set a stricter standard for determining whether cancer is present. If you want to reduce false negatives, you can measure a lower reference value and set a slightly weaker standard for determining whether cancer is present.
[0055] In the present invention, when the artificial intelligence model is a CNN, the loss function may be expressed by the following Equation 1:
number
[0056] In the present invention, when the artificial intelligence model is a CNN, the learning may be characterized by including the following steps: i) Classifying the produced image data into training, validation, and test data; In this case, the training data is used to train the CNN model, the validation data is used to verify hyper-parameter tuning, and the test data is used to evaluate the performance of the optimal model after it is produced. ii) constructing an optimal CNN model through hyper-parameter tuning and learning process; iii) comparing the performance of multiple models obtained through hyper-parameter tuning using validation data, and determining the model with the best validation data performance as the optimal model;
[0057] In the present invention, the hyper-parameter tuning process is a process of optimizing values of a plurality of parameters (such as the number of convolution layers, the number of dense layers, and the number of convolution filters) that constitute a CNN model, and may be characterized by using hyperband optimization, Bayesian optimization, and grid search methods as the hyper-parameter tuning process.
[0058] In the present invention, the learning process may be characterized by optimizing internal parameters (weights) of the CNN model using predetermined hyper-parameters, and determining that the model is overfitting when the validation loss begins to increase relative to the training loss, and interrupting model learning before that occurs.
[0059] In the present invention, the result value of the analysis of the image data input to the artificial intelligence model in step (e) may be any specific score or real number, and is preferably characterized as being a real number, but is not limited thereto.
[0060] In the present invention, a real value means a value expressed as a probability value by adjusting the output of an artificial intelligence model to a scale of 0 to 1 using a sigmoid function or a softmax function in the last layer of the artificial intelligence model.
[0061] From another viewpoint, the present invention provides: A decoding unit that extracts nucleic acids from biological samples and decodes the sequence information; an alignment section that aligns the decoded sequences to a standard chromosome sequence database; a nucleic acid fragment selection unit for selecting nucleic acid fragments of regulatory regions from the nucleic acid fragments based on the aligned sequences; a data generating unit that generates image data of the selected nucleic acid fragments; and an information providing unit that provides information for early cancer diagnosis by inputting the generated image data into an artificial intelligence model that has been trained to distinguish between normal images and cancer images; This invention relates to an information providing device for early cancer diagnosis based on artificial intelligence.
[0062] In the present invention, the decoding unit may include a nucleic acid injection unit that injects nucleic acid extracted from an independent device; and a sequence information analysis unit that analyzes the sequence information of the injected nucleic acid, and may preferably be, but is not limited to, an NGS analysis device.
[0063] In the present invention, the decoding unit may receive and decode sequence information data generated by an independent device.
[0064] From another viewpoint, the present invention provides: 1. A computer-readable storage medium comprising instructions configured to be executed by a processor to provide information for early cancer diagnosis; (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A computer-readable storage medium including instructions configured to be executed by a processor to provide information for early cancer diagnosis by inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancerous images, analyzing the image data, and comparing the image data with a cut-off value to determine whether or not cancer is present.
[0065] In another aspect, the method of the present invention can be implemented using a computer. In one embodiment, the computer includes one or more processors coupled to a chipset. The chipset is also coupled to a memory, a storage device, a keyboard, a graphics adapter, a pointing device, and a network adapter. In one embodiment, the chipset's capabilities are enabled by a memory controller hub and an I / O controller hub. In another embodiment, the memory may be directly coupled to the processor instead of the chipset. The storage device may be any device capable of retaining data, including a hard drive, a compact disk read-only memory (CD-ROM), a DVD, or other memory device. The memory contains data and instructions used by the processor. The pointing device may be a mouse, trackball, or other type of pointing device, and is used in combination with a keyboard to send input data to the computer system. The graphics adapter displays images and other information on a display. The network adapter is coupled to the computer system via a short- or long-distance communication network. The computer used in the present application is not, however, limited to the configuration described above, and may include some or all of the components, or may be part of a storage area network (SAN), and the computer of the present application may be configured to execute modules in a program for performing the method of the present application.
[0066] The term "module" as used herein may refer to a functional and structural combination of hardware for executing the technical idea of the present application and software for driving the hardware. For example, the term "module" may refer to a logical unit of predetermined code and hardware resources for executing the predetermined code, and it is obvious to those skilled in the art that the term "module" does not necessarily refer to physically connected code or a single type of hardware.
[0067] In another aspect, the present invention provides a method for producing a method for producing a nucleic acid molecule comprising the steps of: (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A method for early cancer diagnosis, comprising the steps of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzing the image data, comparing the image data with a cut-off value, and determining that cancer is present if the cut-off value is exceeded.
[0068] From another aspect, the present invention relates to a method for treating a cancer patient, comprising: (a) inputting image data of a nucleic acid fragment of a regulatory region into an artificial intelligence model for analysis using the above method; (b) determining that the patient has cancer if the value output by the artificial intelligence model exceeds a reference value; and (c) treating the patient determined to have cancer.
[0069] In the present invention, the cancer therapeutic agent may be used without limitation as long as it is a method capable of treating cancer or minimal residual cancer. Preferably, the cancer therapeutic agent may be performed by one or more methods selected from the group consisting of surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, and combinations thereof. More preferably, the cancer therapeutic agent may be administered to treat the cancer. Most preferably, the cancer therapeutic agent may be administered to treat the cancer by one or more anticancer agents selected from the group consisting of chemical anticancer agents, targeted anticancer agents, and immunological anticancer agents, but is not limited to these.
[0070] From another perspective, the present invention relates to an AI-based early cancer diagnosis device, including a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment selection unit that selects nucleic acid fragments of regulatory regions from nucleic acid fragments based on the aligned sequences; a data generation unit that generates the selected nucleic acid fragments as image data; and a cancer diagnosis unit that inputs the generated image data into an AI model trained to distinguish between normal images and cancer images, and determines that there is cancer if the image data exceeds a reference value.
[0071] In another aspect, the present invention relates to a computer-readable storage medium including instructions configured to be executed by a processor for performing early cancer diagnosis, the instructions being configured to be executed by a processor for performing early cancer diagnosis, the computer-readable storage medium including instructions configured to be executed by a processor for performing early cancer diagnosis, the instructions including the steps of: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) selecting nucleic acid fragments of a regulatory region based on the aligned sequence information (reads); (d) generating the selected nucleic acid fragments as image data; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzing the input data, and determining that cancer is present if a cut-off value is exceeded. [Example]
[0072] The present invention will be described in more detail below with reference to examples. It will be obvious to those skilled in the art that these examples are intended solely to illustrate the present invention and should not be construed as limiting the scope of the present invention.
[0073] Example 1. Securing adjustment area Regulatory regions can be identified using next-generation sequencing (NGS) techniques such as ATAC-seq, DNase-seq, and FAIRE-seq. The present inventors used TCGA data, which produced regulatory region data for over 400 patients with 23 types of cancer, and regulatory region profiling data for 16 blood cells (DOI: 10.1126 / science.aav1898, DOI: https: / / doi.org / 10.1038 / ng.3646).
[0074] The regulatory data was used to search for nucleosome-free regions (NFRs) using the HMMRATAC tool, and to search for regulatory regions using the MACS2 tool. HMMRATAC was used to find NFRs in the genome using the default options, and MACS2 was used to find regulatory regions using the options "--shift -75 --extsize 150 --nomodel --nolambda --call-summits -q 0.05 -B -SPMR".
[0075] First, using HMMRATAC as a blood-related cell type, we found 39,604 NFRs in B cells, 40,795 in CD4 T cells, 44,687 in CD8 T cells, 36,342 in monocytes, and 42,458 in NK cells. Of these, CD8 T cells, which had the highest number of NFRs, were used as a representative blood cell type, and we used Bedtools' intersectBed to calculate whether the NFR regions found in CD8 T cells overlapped with the regulatory regions of 17 liver cancer patients. Since we used the basic option of intersectBed, overlap was calculated if more than half of the two regions overlapped.
[0076] Similarly, using the ATAC-seq data of 17 liver cancer patients, we searched for NFR regions using the HMMRATAC program, identifying a minimum of 13,712 to 62,344 NFR regions per sample. Of these, the sample with the highest number of NFRs (62,344) was used as the liver cancer representative. It was then overlaid with a total of five blood cell types and calculated using the same method. When the NFRs of CD8 T cells, a representative blood cell type, were overlapped with the liver cancer regulatory region, regions with no overlap were designated as blood-specific NFRs, and NFRs that completely overlapped with the liver cancer regulatory region were designated as blood-common NFRs. Conversely, when the NFRs of liver cancer control samples were overlapped with the regulatory regions of blood cells, regions with no overlap were designated as liver cancer-specific NFRs, and regions that completely overlapped with the blood-specific NFRs were designated as liver cancer-common NFRs.
[0077] Using this method, we selected 8,806 blood cell-specific NFRs, 17,508 blood-generated NFRs, 24,642 liver cancer-specific NFRs, and 19,134 liver cancer-generated NFRs, and constructed a deep learning image based on the distribution of cfDNA reads accumulated in these regions (Figure 4).
[0078] Example 2. Building an artificial intelligence model A distribution of cfDNA at regulatory region positions was generated as shown in Figure 5 to be used as input for deep learning models.
[0079] In other words, NGS provides information about millions of cfDNA fragments floating in the blood, and to create an image using each cfDNA fragment located in a regulatory region, we stacked information about the cfDNA fragment's location within the genome on the x-axis to create a 1D image (Figure 5).
[0080] Tissue-specific regulatory regions were created as input images for deep learning. Because this model is designed to distinguish between normal individuals and liver cancer, two input images were created: one for blood cell-specific regulatory regions and one for liver cancer-specific regulatory regions, and the two were combined to create the final input image.
[0081] The x-axis represents the cfDNA position, which is based on the NFRs called by HMMRATAC, ±1000 bp, for a total of 2000 bp. In other words, the stacked values of the cfDNA reads at each bp constitute a 1D image.
[0082] Therefore, the final input image consists of 2000 (x-axis, cfDNA position) x 4 (blood cell-specific, common regulatory region, liver cancer-specific, common regulatory region).
[0083] The convolutional neural network (CNN) model is a model that exhibits good performance in image classification because it captures local features well through the kernel. Therefore, we created a model that generates the cfDNA distribution as image data, learns the pattern using the CNN model, and determines whether the image is cancerous or normal based on this learned pattern.
[0084] Experimental example 1. Construction of an early diagnosis model for liver cancer To confirm whether this model can be used to diagnose liver cancer, blood samples were collected from 187 healthy individuals and 64 liver cancer patients and stored in stretch tubes. After centrifugation, plasma was separated from the upper portion of the blood. cfDNA was extracted using the Tiangen kit and sequenced using the MGI DNB-seq system.
[0085] A total of 251 people, including both terminal liver cancer and healthy people, were used for model learning, with 150 people used for training and 49 people used for validation, and performance was confirmed in 52 people through testing.
[0086] Deep learning works better the more data there is to learn from, so to increase the number of samples that can be learned, we performed down-sampling for each sample and randomly selected 1.7 x 10^7 reads 10 times to increase the number of samples and perform learning.
[0087] [Table 1]
[0088] Experimental Example 2: Confirmation of the performance of the liver cancer early diagnosis model Using 2020 training sets, 670 validation sets, and 680 test sets, we tuned various hyperparameters using Hyperband, and ultimately confirmed high performance with an AUC of 0.98 for training, 0.94 for validation, and 0.86 for test (Figure 7).
[0089] Furthermore, when randomly selected regions were used instead of the selected tissue-specific NFRs, the AUC was confirmed to be 0.83 in training, 0.79 in validation, and 0.70 in test. This confirmed that the selected tissue-specific NFRs are important in distinguishing between normal individuals and liver cancer patients, and that liver cancer patients can be accurately identified through the selected regions.
[0090] Although certain parts of the present invention have been described in detail above, it is clear to those skilled in the art that these specific descriptions are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the true scope of the present invention is to be defined by the appended claims and their equivalents. [Industrial Applicability]
[0091] The method for early cancer diagnosis according to the present invention diagnoses cancer early based on artificial intelligence using cell-free nucleic acid distribution of tissue-specific regulatory regions obtained by next generation sequencing (NGS). Since the method has high accuracy and sensitivity and is commercially applicable, the method is useful for early cancer diagnosis.
Claims
1. (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A method for providing information for early cancer diagnosis based on artificial intelligence, comprising the steps of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzing the image data, and comparing the image data with a cut-off value to determine whether or not cancer is present.
2. (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A method for early cancer diagnosis, comprising the steps of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzing the image data, comparing the image data with a cut-off value, and determining that cancer is present if the cut-off value is exceeded.
3. The method according to claim 1 or 2, wherein the step of acquiring sequence information in step (a) is carried out by a method comprising the steps of: (ai) obtaining nucleic acids from the biological sample; (a-ii) removing proteins, fats, and other residues from the collected nucleic acid by salting-out, column chromatography, or beads to obtain purified nucleic acid; (a-iii) preparing a single-end sequencing or paired-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, crushing, or hydroshear method; (a-iv) subjecting the prepared library to a next-generation sequencer; and (av) A step of obtaining nucleic acid sequence information (reads) using a next-generation sequencer.
4. The method according to claim 1 or 2, characterized in that the nucleic acid in step (a) is cell-free DNA.
5. The method of claim 1 or 2, further comprising a step of selecting reads whose mapping quality score of aligned nucleic acid fragments is equal to or greater than a reference value before performing step (c).
6. The method according to claim 5, wherein the reference value is 50 to 70 points.
7. The method according to claim 1 or 2, wherein the regulatory region in step (c) is a tissue-specific regulatory region.
8. The method of claim 7, wherein the tissue-specific regulatory region differs in length and / or amount of cell-free DNA detected in different tissues.
9. The method according to claim 1 or 2, wherein the image in step (d) is a one-dimensional image whose x-axis is composed of the number of reads for each alignment position of the selected nucleic acid fragments.
10. 3. The method according to claim 1 or 2, wherein the artificial intelligence model in step (e) is an artificial neural network.
11. 11. The method according to claim 10, wherein the artificial neural network is a convolutional neural network (CNN) or a recurrent neural network (RNN).
12. a decoding unit that extracts nucleic acids from the biological sample and decodes the sequence information; an alignment section that aligns the decoded sequences to a standard chromosome sequence database; a nucleic acid fragment selection unit for selecting nucleic acid fragments of regulatory regions from the nucleic acid fragments based on the aligned sequences; A data generating unit that generates image data of the selected nucleic acid fragments; and an information providing unit that inputs the generated image data into an artificial intelligence model that has been trained to distinguish between normal images and cancer images, analyzes the data, and provides information for early cancer diagnosis; An information providing device for early cancer diagnosis based on artificial intelligence, including:
13. 1. A computer-readable storage medium comprising instructions configured to be executed by a processor to provide information for early cancer diagnosis; (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A computer-readable storage medium containing instructions configured to be executed by a processor to provide information for early cancer diagnosis by inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancerous images, analyzing the image data, and comparing the image data with a cut-off value to determine whether or not cancer is present.
14. a decoding unit that extracts nucleic acids from the biological sample and decodes the sequence information; an alignment section that aligns the decoded sequences to a standard chromosome sequence database; a nucleic acid fragment selection unit for selecting nucleic acid fragments of regulatory regions from the nucleic acid fragments based on the aligned sequences; A data generating unit that generates image data of the selected nucleic acid fragments; and a cancer diagnosis unit that inputs the generated image data into an artificial intelligence model that has been trained to distinguish between normal images and cancer images, and determines that cancer is present if the image data exceeds a reference value; An artificial intelligence-based early cancer diagnosis device.
15. 1. A computer-readable storage medium comprising instructions configured to be executed by a processor for performing early cancer diagnosis; (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence reads to a reference genome database; (c) selecting nucleic acid fragments of the regulatory region based on the aligned sequence reads; (d) generating image data of the selected nucleic acid fragments; and (e) A computer-readable storage medium containing instructions configured to be executed by a processor to perform early cancer diagnosis by inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images, analyzing the image data, and determining that cancer is present if a cut-off value is exceeded.