An AI-based early cancer diagnosis method using cell-free DNA distribution of tissue-specific regulatory regions.

An AI-based method analyzing cell-free nucleic acids in tissue-specific regulatory regions addresses the limitations of existing cfDNA methods by providing accurate early cancer diagnosis through high sensitivity and specificity.

JP7830520B2Active Publication Date: 2026-03-16GREEN CROSS GENOME CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2026-03-16

AI Technical Summary

Technical Problem

Existing methods for early cancer diagnosis using cell-free DNA (cfDNA) in liquid biopsies suffer from low reliability and require large amounts of data, making it challenging to accurately detect cancer mutations.

Method used

An AI-based method that analyzes the distribution of cell-free nucleic acids in tissue-specific regulatory regions by extracting and imaging nucleic acid fragments from a biological sample, aligning them with a reference genome, and inputting the data into an AI model trained to distinguish between normal and cancerous images.

Benefits of technology

The method achieves high sensitivity and accuracy in early cancer diagnosis, enabling precise differentiation between normal and cancerous samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007830520000003
    Figure 0007830520000003
  • Figure 0007830520000004
    Figure 0007830520000004
  • Figure 0007830520000005
    Figure 0007830520000005
Patent Text Reader

Abstract

The present invention relates to an artificial intelligence-based method for early cancer diagnosis, more particularly, to an artificial intelligence-based method for early cancer diagnosis using a method of inputting information on cell-free DNA distribution of tissue-specific regulatory regions into an artificial intelligence model trained to diagnose cancer early, and analyzing the information. The method for early cancer diagnosis according to the present invention diagnoses cancer early based on artificial intelligence using cell-free nucleic acid distribution of tissue-specific regulatory regions obtained by next generation sequencing (NGS), and has high accuracy and sensitivity and is commercially applicable, so that the method of the present invention is useful for early cancer diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for early cancer diagnosis based on artificial intelligence, and more specifically, to a method for early cancer diagnosis based on artificial intelligence using a method of inputting and analyzing information on the cell-free DNA distribution in tissue-specific regulatory regions into an artificial intelligence model trained to diagnose early cancer.

Background Art

[0002] Research is underway to detect chromosomal abnormalities using cell-free DNA (cfDNA; cell-free DNA) present in plasma due to cell necrosis, apoptosis, and secretion by using liquid biopsy technology. In particular, blood cell-free DNA derived from tumor cells contains tumor-specific chromosomal abnormalities and mutations that do not appear in normal cells, and has the advantage of a short half-life of about 2 hours and reflecting the current state of the tumor. In addition, because it is non-invasive and can be repeatedly collected, blood cell-free DNA has attracted attention as a tumor-specific biomarker in various fields related to cancer, such as cancer diagnosis, monitoring, and prognosis observation.

[0003] Many researchers are working hard to use liquid biopsy for early diagnosis by taking advantage of the fact that cancer can be diagnosed by just a blood test. Since cancer is a disease that occurs as mutations gradually accumulate in DNA, cfDNA derived from cancer has mutations different from those of normal people, and if DNA containing mutations is discovered using such characteristics, it can be diagnosed as cancer. However, among the 3 billion human genomes, the mutations commonly found in cancer cells by humans are very few, and furthermore, many people with cancer do not have such mutations, so early cancer diagnosis using mutations has not yet shown good performance.

[0004] Recently, methods have been developed to obtain full-length genome data of cfDNA, derive a transcription start site profile based on read depth, and learn the expression status of each gene using SVM (Ulz, P., Thallinger, G., Auer, M. et al. Nat. Genet. Vol. 48, pp. 1273-1278, 2016), and to analyze transcription factor binding patterns based on cfDNA fragmentation patterns to diagnose cancer early or classify cancer types (Ulz, P. et al., Nat. Commun. Vol. 10, 4666, 2019). However, these methods have the drawbacks of relatively low reliability or the need for large amounts of data.

[0005] Against this technological backdrop, the inventors diligently worked to develop an artificial intelligence-based method for early cancer diagnosis. As a result, they confirmed that by imaging the distribution of cell-free nucleic acids in tissue-specific regulatory regions and inputting this data into an artificial intelligence model trained to diagnose cancer early, cancer can be diagnosed early with high sensitivity and accuracy, thus completing the present invention. [Overview of the project]

[0006] The objective of this invention is to provide an artificial intelligence-based method for providing information for early cancer diagnosis.

[0007] Another object of the present invention is to provide an artificial intelligence-based information provision device for early cancer diagnosis.

[0008] Another object of the present invention is to provide a computer-readable medium that includes instructions configured to be executed by a processor that provides information for early cancer diagnosis in the manner described above.

[0009] Another object of the present invention is to provide an artificial intelligence-based method for early cancer diagnosis.

[0010] Another object of the present invention is to provide an artificial intelligence-based cancer early diagnosis device.

[0011] Another object of the present invention is to provide a computer-readable medium containing instructions configured to be executed by a processor that performs early cancer diagnosis in the manner described above.

[0012] To achieve the above objective, the present invention provides an artificial intelligence-based method for providing information for early cancer diagnosis, comprising the steps of: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) in a reference genome database; (c) selecting nucleic acid fragments of regulatory regions based on the aligned sequence information (reads); (d) generating image data from the selected nucleic acid fragments; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis and comparing it with a cut-off value to determine the presence or absence of cancer.

[0013] The present invention also provides an artificial intelligence-based information provision device for early cancer diagnosis, which includes: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence in a standard chromosome sequence database; a nucleic acid fragment selection unit that selects nucleic acid fragments of regulatory regions from nucleic acid fragments based on the aligned sequence; a data generation unit that generates the selected nucleic acid fragments as image data; and an information provision unit that inputs the generated image data into an artificial intelligence model trained to distinguish between normal images and cancer images for analysis and provides information for early cancer diagnosis.

[0014] The present invention also provides a computer-readable storage medium comprising instructions configured to be executed by a processor that provides information for early cancer diagnosis, the instructions comprising: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) in a reference genome database; (c) sorting nucleic acid fragments of regulatory regions based on the aligned sequence information (reads); (d) generating image data from the sorted nucleic acid fragments; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis and comparing it to a cut-off value to determine the presence or absence of cancer.

[0015] The present invention also provides a method for early cancer diagnosis, comprising the steps of (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) in a reference genome database; (c) selecting nucleic acid fragments of regulatory regions based on the aligned sequence information (reads); (d) generating image data from the selected nucleic acid fragments; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis, and comparing it to a cut-off value. If the cut-off value is exceeded, the method determines that cancer is present.

[0016] The present invention also provides an artificial intelligence-based early cancer diagnostic device, which includes: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence in a standard chromosome sequence database; a nucleic acid fragment selection unit that selects nucleic acid fragments of regulatory regions from nucleic acid fragments based on the aligned sequence; a data generation unit that generates the selected nucleic acid fragments as image data; and a cancer diagnosis unit that inputs the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images and determines that cancer is present if the values ​​exceed a certain threshold.

[0017] The present invention also provides a computer-readable storage medium comprising instructions configured to be executed by a processor for early cancer diagnosis, the instructions comprising: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) in a reference genome database; (c) selecting nucleic acid fragments of regulatory regions based on the aligned sequence information (reads); (d) generating image data from the selected nucleic acid fragments; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis, and determining the presence of cancer if the value exceeds a cut-off value. [Brief explanation of the drawing]

[0018] [Figure 1] This is an overall flowchart for carrying out the method of the present invention. [Figure 2] This diagram shows the differences in nucleosome location within regulatory regions across different tissues, along with actual examples. [Figure 3] This is a schematic diagram of regulatory factor data for various tissues. [Figure 4] This is a schematic diagram illustrating a method for identifying tissue-specific regulatory factors. [Figure 5]This figure shows the principle for generating an image of the cfDNA distribution of the regulatory region obtained by one embodiment of the present invention for input into an artificial intelligence model. [Figure 6] This figure shows the algorithm of an artificial intelligence model constructed according to one embodiment of the present invention. [Figure 7] This is the result of verifying the performance of a liver cancer prediction model constructed according to one embodiment of the present invention. [Modes for carrying out the invention]

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by skilled experts in the art to which this invention pertains. In general, the nomenclature used herein and the experimental methods described below are well known and commonly used in the art.

[0020] Terms such as First, Second, A, B, etc., may be used to describe various components, but the components are not limited by such terms and are used solely for the purpose of distinguishing one component from another. For example, within the scope of the rights of the technology described below, the First component may be named the Second component, and similarly, the Second component may be named the First component. The terms and / or include combinations of multiple related descriptions or any of multiple related descriptions.

[0021] In the terminology used herein, singular expressions should be understood to include plural expressions unless the context clearly indicates otherwise, and terms such as “includes” should be understood to mean that the described features, quantities, steps, actions, components, parts, or combinations thereof exist, and not to exclude the possibility of the existence or addition of one or more other features, quantities, steps, actions, components, parts, or combinations thereof.

[0022] Prior to providing a detailed description of the drawings, it is to be made clear that the classification of components in this specification is merely based on the main functions that each component undertakes. That is, two or more components described below may be combined into one component, or one component may be divided into two or more components according to more refined functions. And each of the components described below may further perform part or all of the functions of other components in addition to its own main function. Of course, it is also possible that part of the main function that each component undertakes is solely performed by other components.

[0023] In addition, when implementing a method or an operation method, each process constituting the method may be performed in an order different from the specified order, unless a specific order is clearly described in the context. That is, each process may be performed in the same order as the specified order, may be performed substantially simultaneously, or may be performed in the reverse order.

[0024] In the present invention, after aligning the sequence analysis data obtained from a sample with a reference genome, a nucleic acid fragment of a regulatory region is selected from the aligned nucleic acid fragments to generate image data, and when the generated image data is input into an artificial intelligence model trained to distinguish normal images from cancer images, it was attempted to confirm that cancer can be diagnosed at an early stage with high sensitivity and accuracy.

[0025] That is, in one embodiment of the present invention, nucleic acids were extracted from liquid biopsies obtained from 187 normal people, 12 early-stage liver cancer patients, and 150 late-stage liver cancer patients. After obtaining cfDNA sequencing, nucleic acid fragments corresponding to liver-specific regulatory regions were selected, and these were generated as image data. An artificial intelligence learning model for early diagnosis of liver cancer was constructed using the data of 187 normal people and 150 late-stage liver cancer patients, and the performance of the learning model was confirmed using the data of 12 early-stage liver cancer patients. As a result, it was confirmed that the constructed learning model with high accuracy can distinguish images of normal people, liver cancer patients, and early-stage liver cancer patients (FIGS. 7 and 8).

[0026] In this invention, the term "reads" refers to a single nucleic acid fragment whose sequence information has been analyzed using various methods known in the art. Therefore, the terms "sequence information" and "reads" in this specification have the same meaning in that they are the results obtained through the sequencing process.

[0027] In this invention, the term "regulatory region" refers to all locations on a chromosome where gene expression can be regulated, and specifically to regions to which RNA polymerase and electronically regulatory proteins bind for RNA synthesis. Preferably, it may include, but is not limited to, promoters, enhancers, silencers, and insulators.

[0028] In this invention, the term "NFR (Nucleosome Free Region)" refers to the same region as the regulatory region, but specifically refers to a region within the regulatory region where nucleosomes are absent. For example, if there is an enhancer region consisting of the first nucleosome (1-147 bp), internucleonucleotides (148-346 bp), the second nucleosome (347-493 bp), internucleonucleotides (494-692 bp), the third nucleosome (693-839 bp), and internucleonucleotides (840-1039 bp), and when transcription begins, the second nucleosome detaches and a transcription regulatory protein binds to it, then the NFR will be the region from 148 to 692 bp.

[0029] Furthermore, while transcription proceeds in normal samples using the method described above, cancer samples may lack NFRs, or nucleosomes in other regions may detach, resulting in the generation of other NFRs. This means that NFRs that are not present in normal samples may be newly generated in cancer samples.

[0030] Furthermore, while transcription proceeds in blood cells in the manner described above, NFRs may not exist in other tissues (e.g., the liver), or nucleosomes in other regions may detach, generating different NFRs. Additionally, NFRs that do not exist in blood cells may be newly generated in cancer samples.

[0031] Therefore, from one perspective, the present invention is, (a) A step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) A step of selecting nucleic acid fragments of the regulatory region based on the aligned sequence information (reads); (d) A step of generating selected nucleic acid fragments as image data; and (e) A method for providing information for early cancer diagnosis on an artificial intelligence basis, which includes the step of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancerous images for analysis and comparing it with a cut-off value to determine the presence or absence of cancer.

[0032] In the present invention, the cancer may be a solid tumor or a hematological cancer, and may be preferably selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphoblastic leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, liver cancer, thyroid cancer, gastric cancer, gallbladder cancer, biliary tract cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary site, kidney cancer, esophageal cancer, and mesothelioma, and more preferably liver cancer, but is not limited to these.

[0033] In the present invention, the step of obtaining sequence information in step (a) above may be characterized by being carried out in a manner that includes the following steps: (ai) Steps to obtain nucleic acids from biological samples; (a-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using the salting-out method, column chromatography method, or beads method to obtain purified nucleic acids; (a-iii) A step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, disruption, or hydroshear method; (a-iv) The step of reacting the prepared library with a next-generation sequencer; and (av) Steps to obtain nucleic acid sequence information (reads) using a next-generation sequencer.

[0034] In the present invention, the step of obtaining sequence information in step (a) above may be characterized by obtaining the isolated cell-free DNA by full-length genome sequencing at a depth of 1 million to 100 million reads.

[0035] In the present invention, the term "biological sample" means any substance, biological fluid, tissue, or cell obtained from or derived from an individual, such as whole blood, leukocytes, peripheral blood mononuclear cells, leukocyte buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, and pancreatic fluid. This may include, but is not limited to, fluids such as lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extract, semen, hair, saliva, urine, oral cells, placental cells, cerebrospinal fluid, and mixtures thereof.

[0036] In this invention, the term "reference population" refers to a comparable reference population, such as a standard nucleotide sequence database, that currently does not have a specific disease or condition. In this invention, the standard nucleotide sequences in the standard chromosome sequence database of the reference population may be reference chromosomes registered with public health organizations such as the NCBI.

[0037] In the present invention, the nucleic acid in step (a) may be cell-free DNA, and more preferably circulating tumor DNA, but is not limited to these.

[0038] In the present invention, the next-generation sequencer may be used with any sequencing method known in the art. Sequencing of nucleic acids isolated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the single nucleotide sequence of individual nucleic acid molecules or, in a very similar manner, a cloned proxie for individual nucleic acid molecules (e.g., 10⁵ or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species in a library can be estimated by measuring the relative occurrence of their congeneral sequences in the data produced by the sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in the literature included herein by reference (Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46).

[0039] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequences of individual nucleic acid molecules (e.g., HeliScope Gene Sequencing system from Helicos BioSciences and PacBio RS system from Pacific Biosciences). In another embodiment, sequencing, for example, large-scale parallel short-read sequencing (e.g., Solexa sequencer from Illumina Inc. in San Diego, California), which produces more bases of sequence per sequencing unit than other sequencing methods that produce fewer but longer reads, determines the nucleotide sequences of cloned proxies for individual nucleic acid molecules (e.g., Solexa sequencer from Illumina Inc. in San Diego, California; 454 Life Sciences (located in Branford, Connecticut) and Ion Torrent). Other methods or machines for next-generation sequencing are provided by, but are not limited to, 454 Life Sciences (located in Branford, Connecticut), Applied Biosystems (located in Foster City, California; SOLiD sequencers), Helicos Biosciences Corporation (located in Cambridge, Massachusetts), and emulsion and microflow sequencing methods such as nanoinfusions (e.g., GnuBio infusions).

[0040] Platforms for next-generation sequencing include, but are not limited to, the Roche / 454 Genome Sequencer (GS) FLX system, the Illumina / Solexa Genome Analyzer (GA), the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, the Polonator G.007 system, the Helicos BioSciences HeliScope Gene Sequencing system, and the Pacific Biosciences PacBio RS system.

[0041] In the present invention, the sorting step of step (b) above is not limited to the above, but may be performed using the BWA algorithm and the Hg19 array.

[0042] In the present invention, the BWA algorithm may include, but is not limited to, BWA-ALN, BWA-SW, or Bowtie2.

[0043] The present invention may further include a step of selecting reads whose mapping quality score of the aligned nucleic acid fragments is equal to or greater than a reference value, prior to performing step (c) above, and the reference value may be any value that can confirm the quality of the aligned nucleic acid fragments, preferably 50 to 70 points, more preferably 60 points, but is not limited thereto.

[0044] In the present invention, the regulatory region in step (c) may be characterized by being a tissue-specific regulatory region.

[0045] In the present invention, the tissue-specific regulatory region may be characterized by having different lengths and / or amounts of cell-free DNA detected depending on the tissue.

[0046] In the present invention, the tissue-specific regulatory region may be different in length and / or amount of cell-free DNA detected in specific tissues, such as the liver, or different in length and / or amount of cell-free DNA detected in blood, brain, stomach, and heart, or different in length and / or amount of cell-free DNA detected in solid tissues (such as the brain, liver, stomach, lungs, and heart) and blood tissues (such as blood cells and bone marrow).

[0047] In the present invention, the tissue-specific regulatory region may more specifically mean a region within the regulatory region where nucleosomes are absent, i.e., an NFR (Nucleosome Free Region), but is not limited thereto.

[0048] In the present invention, the number of tissue-specific regulatory regions may be unlimited as long as it is sufficient to generate image data to be input to the artificial intelligence model. Preferably, there may be 10, 100, 1,000, 10,000, or 20,000 to 50,000 tissue-specific regulatory regions, but the number is not limited to these.

[0049] In the present invention, the image in step (d) may be any image that can be used for training by the artificial intelligence model, and may preferably be a one-dimensional image in which the x-axis is composed of the number of reads for each sequence position of the selected nucleic acid fragments, but is not limited thereto.

[0050] In the present invention, the image in step (d) above is a sequence of values ​​obtained by accumulating cfDNA reads for each base pair, and may show a structure of the form, for example, [0.91, 0.93, ~~, 0.73, 0.86], and if ±1000 bp centered on the position selected as a tissue-specific regulatory region, totaling 2000 bp, then the number in [ ] will be 2000.

[0051] In the present invention, the artificial intelligence model in step (e) above may be any learning model that learns to distinguish between normal images and cancerous images, preferably an artificial neural network, and more preferably a convolutional neural network (CNN) or a recurrent neural network (RNN), but is not limited to these.

[0052] In the present invention, the reference value in step (e) above may be any value that allows for the early diagnosis of cancer, and may preferably be 0.5, but is not limited thereto. If the reference value is 0.5, the invention may be characterized by determining that cancer is present if the value is 0.5 or higher.

[0053] In this invention, the artificial intelligence model learns to produce an output result close to 1 if cancer is present, and an output result close to 0 if cancer is not present. Based on a baseline of 0.5, it is determined that cancer is present if the output is 0.5 or higher, and that cancer is not present if the output is 0.5 or lower, and performance is measured accordingly (Training, validation, test accuracy).

[0054] It is obvious to any engineer that the 0.5 threshold can be changed at any time. For example, to reduce false positives, a higher threshold can be set than 0.5, making the criteria for determining the presence of cancer stricter. To reduce false negatives, a lower threshold can be measured, making the criteria for determining the presence of cancer slightly weaker.

[0055] In the present invention, if the artificial intelligence model is a CNN, the loss function may be characterized by being expressed by the following formula 1.

number

[0056] In the present invention, if the artificial intelligence model is a CNN, the learning process may be characterized by including the following steps: i) A step to classify the produced image data into training, validation, and test data; In this process, the training data is used to train the CNN model, the validation data is used for hyper-parameter tuning verification, and the test data is used to evaluate the performance after the optimal model has been produced. ii) Steps to build an optimal CNN model through hyperparameter tuning and the learning process; iii) A step in which the performance of multiple models obtained through hyper-parameter tuning is compared using validation data, and the model with the best validation data performance is determined to be the optimal model;

[0057] In the present invention, the hyper-parameter tuning process is a process of optimizing the values ​​of multiple parameters (such as the number of convolutional layers, dense layers, and convolutional filters) that constitute the CNN model, and the hyper-parameter tuning process may be characterized by using hyperband optimization, Bayesian optimization, and grid search methods.

[0058] In the present invention, the learning process may be characterized by optimizing the internal parameters (weights) of the CNN model using predetermined hyper-parameters, and determining that the model has overfitted when the validation loss begins to increase relative to the training loss, and interrupting model learning before that point.

[0059] In the present invention, the result value analyzed by the artificial intelligence model from the input image data in step (e) above may be any specific score or real number without limitation, and may preferably be characterized by being a real number, but is not limited thereto.

[0060] In this invention, real values ​​refer to values ​​expressed as probability values ​​obtained by adjusting the output of the artificial intelligence model to a scale of 0 to 1 using a sigmoid function or softmax function in the last layer of the artificial intelligence model.

[0061] From another perspective, this invention A decoding unit that extracts nucleic acids from biological samples and decodes their sequence information; An alignment unit that aligns the decoded sequences into a standard chromosome sequence database; A nucleic acid fragment sorting unit that sorts regulatory region nucleic acid fragments from aligned sequence-based nucleic acid fragments; A data generation unit that generates selected nucleic acid fragments as image data; and An information provision unit that inputs the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images, thereby providing information for the early diagnosis of cancer; This relates to an artificial intelligence-based information provision device for early cancer diagnosis.

[0062] In the present invention, the decoding unit may include a nucleic acid injection unit for injecting nucleic acids extracted from an independent device, and a sequence information analysis unit for analyzing the sequence information of the injected nucleic acids, and may preferably be an NGS analyzer, but is not limited thereto.

[0063] In the present invention, the decoding unit may be characterized by receiving and decoding sequence information data generated by an independent device.

[0064] From another perspective, this invention A computer-readable storage medium comprising instructions configured to be executed by a processor that provides information for the early diagnosis of cancer, (a) A step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) A step of selecting nucleic acid fragments of the regulatory region based on the aligned sequence information (reads); (d) A step of generating selected nucleic acid fragments as image data; and (e) relating to a computer-readable storage medium that includes instructions configured to be executed by a processor that provides information for early cancer diagnosis by inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis and comparing it to a cut-off value to determine the presence or absence of cancer.

[0065] In other embodiments, the methods according to the present invention can be carried out using a computer. In one embodiment, the computer includes one or more processors coupled to a chipset. The chipset is also coupled to memory, a storage device, a keyboard, a graphics adapter, a pointing device, and a network adapter, etc. In one embodiment, the performance of the chipset is enabled by a memory controller hub and an I / O controller hub. In other embodiments, the memory may be used by being directly coupled to the processor instead of the chipset. The storage device is any device capable of holding data, including a hard drive, a CD-ROM (Compact Disk Read-Only Memory), a DVD, or other memory device. The memory is involved with the data and instructions used by the processor. The pointing device may be a mouse, a trackball, or other type of pointing device, and is used in combination with a keyboard to transmit input data to the computer system. The graphics adapter displays images and other information on a display. The network adapter is connected to the computer system by a short-range or long-range communication network. The computer used in this application is not limited to the configuration described above, and may be missing some of the configurations or may include additional configurations, and may be part of a Storage Area Network (SAN), and the computer of this application may be configured to be suitable for executing modules of a program for performing the method according to this application.

[0066] In this application, "module" may mean a functional and structural combination of hardware for implementing the technical concept of this application and software for driving said hardware. For example, the module may mean a logical unit of a predetermined code and a hardware resource for which said code is performed, and it will be obvious to those skilled in the art that it does not necessarily mean physically connected code or a single type of hardware.

[0067] From another perspective, the present invention includes the step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) A step of selecting nucleic acid fragments of the regulatory region based on the aligned sequence information (reads); (d) A step of generating selected nucleic acid fragments as image data; and (e) The present invention relates to a method for early cancer diagnosis, which includes the step of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancerous images for analysis, and comparing it to a cut-off value, and determining if cancer is present if the cut-off value is exceeded.

[0068] In other words, the present invention relates to a method for treating a cancer patient, comprising the steps of (a) inputting image data of a nucleic acid fragment of a regulatory region into an artificial intelligence model for analysis using the method described above; (b) determining that cancer is present if the value output by the artificial intelligence model exceeds a reference value; and (c) treating the patient who has been determined to have cancer.

[0069] In the present invention, the cancer treatment agent may be used without limitation as long as it is a method that can treat cancer or minimal residual cancer, preferably characterized by one or more methods selected from the group consisting of surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adaptive T cell therapy, targeted therapy, and combinations thereof, more preferably characterized by treatment by administering the cancer treatment agent, most preferably characterized by treatment by administering one or more anticancer agents selected from the group consisting of chemoanticancer agents, targeted anticancer agents, and immunoanticancer agents, but is not limited thereto.

[0070] From another perspective, the present invention relates to an artificial intelligence-based early cancer diagnostic device, which includes: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence in a standard chromosome sequence database; a nucleic acid fragment selection unit that selects nucleic acid fragments of regulatory regions from nucleic acid fragments based on the aligned sequence; a data generation unit that generates the selected nucleic acid fragments as image data; and a cancer diagnosis unit that inputs the generated image data into an artificial intelligence model trained to distinguish between normal images and cancerous images, and determines that cancer is present if the values ​​exceed a reference value.

[0071] In other words, the present invention relates to a computer-readable storage medium comprising instructions configured to be executed by a processor for early cancer diagnosis, the instructions comprising: (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) in a reference genome database; (c) selecting nucleic acid fragments of regulatory regions based on the aligned sequence information (reads); (d) generating image data from the selected nucleic acid fragments; and (e) inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis, and determining the presence of cancer if the value exceeds a cut-off value. [Examples]

[0072] The present invention will be described in more detail below with reference to examples. These examples are solely for illustrative purposes and it will be obvious to those with ordinary skill in the art that the scope of the present invention is not to be limited by these examples.

[0073] Example 1. Securing the adjustment area The regulatory regions can be identified using next-generation sequencing (NGS) technologies such as ATAC-seq, DNase-seq, and FAIRE-seq. However, the inventors used TCGA data, which produced regulatory region data from over 400 patients with 23 types of cancer, and data profiling regulatory regions for 16 blood cell types (DOI: 10.1126 / science.aav1898, DOI:https: / / doi.org / 10.1038 / ng.3646).

[0074] Using the regulatory factor data, we used the HMMRATAC tool to search for Nucleosome Free Regions (NFRs) and the MACS2 tool to search for regulatory regions. HMMRATAC found NFRs on the genome using the default option, and MACS2 found regulatory regions using the "--shift -75 --extsize 150 --nomodel --nolambda --call-summits -q 0.05 -B -SPMR" option.

[0075] First, using HMMRATAC as a blood-related cell type, we found 39,604 NFRs in B cells, 40,795 in CD4 T cells, 44,687 in CD8 T cells, 36,342 in Monocytes, and 42,458 in NK cells. Of these, CD8 T cells, which had the highest number of NFRs, were used as the representative of the blood cell type, and we used Bedtools' intersectBed to calculate whether the NFR regions found in CD8 T cells overlapped with the regulatory regions of 17 liver cancer patients. At this time, we proceeded with the basic options of intersectBed, so if more than half of the two regions overlapped, it was calculated as overlap.

[0076] Similarly, using ATAC-seq data from a total of 17 liver cancer patients, the HMMRATAC program was used to search for NFR regions, securing a minimum of 13,712 to 62,344 NFR regions per sample. Of these, the sample with the most NFRs (62,344) was used as the liver cancer representative, and its values ​​were calculated by overlaying it with a total of five blood cell types using the method described above. When the NFRs of the blood representative cell type, CD8 T cells, overlapped with the liver cancer regulatory regions, regions that did not overlap were named blood-specific NFRs, and NFRs that completely overlapped with the liver cancer regulatory regions were named blood-common NFRs. Conversely, when the NFRs of the liver cancer sample overlapped with the regulatory regions of blood cells, regions that did not overlap were named liver cancer-specific NFRs, and regions that completely overlapped with the blood-specific NFRs were named liver cancer-common NFRs.

[0077] Using this method, 8,806 blood cell-specific NFRs, 17,508 blood-common NFRs, 24,642 liver cancer-specific NFRs, and 19,134 liver cancer-common NFRs were selected, and deep learning images were constructed from the distribution of cfDNA reads accumulated in these regions (Figure 4).

[0078] Example 2. Construction of an artificial intelligence model To use the distribution of cfDNA at the regulatory region as input for a deep learning model, a cfDNA distribution was created as shown in Figure 5.

[0079] In other words, we were able to obtain information about millions of cfDNA fragments floating in the blood via NGS, and to create an image like the one shown, we stacked information about the genomic location of the cfDNA fragments on the x-axis to create a 1D image (Figure 5).

[0080] Tissue-specific regulatory regions were created as input images for deep learning. Since this model distinguishes between normal individuals and liver cancer patients, two input images were created: one for blood cell-specific regulatory regions and another for liver cancer-specific regulatory regions. These two images were then combined to create the final input image.

[0081] The cfDNA positions corresponding to the x-axis were primarily based on NFRs called by HMMRATAC, with a range of ±1000 bp, totaling 2000 bp. In other words, the 1D image was constructed by accumulating the values ​​of cfDNA reads for each bp.

[0082] Therefore, the final input image consists of 2000 (x-axis, cfDNA location) x 4 (blood cell specific, common regulatory region, liver cancer specific, common regulatory region).

[0083] Convolutional neural network (CNN) models are good at image classification because they capture local features well through the kernel. Therefore, we generated the aforementioned cfDNA distribution as image data, trained a CNN model to learn the pattern, and created a model that determines whether the image is cancerous or normal based on this learned pattern.

[0084] Experimental Example 1: Construction of an Early Diagnosis Model for Liver Cancer To determine if this model could be used for liver cancer diagnosis, blood samples were collected from 187 healthy individuals and 64 liver cancer patients and stored in stretch tubes. After centrifugation, plasma was separated from the top of the blood, and cfDNA was extracted using the Tiangen Kit and sequenced using MGI DNB-seq.

[0085] The model was trained using a total of 251 people, including those with terminal liver cancer and healthy individuals. 150 people were used for training, 49 for validation, and the performance was confirmed with 52 people in the test.

[0086] Deep learning performs better with a larger amount of training data. To increase the number of samples available for training, we performed down-sampling on each sample, randomly selecting 1.7 × 10^7 reads 10 times to increase the sample size for training.

[0087] [Table 1]

[0088] Experimental Example 2. Verification of the performance of the early diagnosis model for liver cancer. Using 2020 training sets, 670 validation sets, and 680 test sets, various hyperparameters were tuned using hyperband, and high performance was confirmed, with an AUC of 0.98 in training, 0.94 in validation, and 0.86 in test (Figure 7).

[0089] Furthermore, when using randomly selected regions instead of the selected tissue-specific NFRs, we confirmed that the AUC was 0.83 in training, 0.79 in validation, and 0.70 in test. This confirmed that the selected tissue-specific NFRs are important for distinguishing between normal individuals and liver cancer patients, and that liver cancer patients can be accurately selected through the selected regions.

[0090] Although specific parts of the present invention have been described in detail above, it is clear to those with ordinary skill in the art that these specific descriptions are merely preferred embodiments and do not limit the scope of the invention. Therefore, the substantial scope of the invention is defined by the appended claims and their equivalents. [Industrial applicability]

[0091] The cancer early diagnosis method according to the present invention uses cell-free nucleic acid distribution of tissue-specific regulatory regions obtained by next-generation sequencing (NGS) to diagnose cancer early on an artificial intelligence basis. Because it has high accuracy and sensitivity and high commercial applicability, the method of the present invention is useful for early cancer diagnosis.

Claims

1. (a) A step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) A step of selecting nucleic acid fragments of tissue-specific nucleosome-free regions (NFRs) based on the aligned sequence information (reads); (d) A step of generating a one-dimensional image, (di) A step of superimposing selected nucleic acid fragments located within ±1000 bp from the center of the tissue-specific nucleosome-free region; (d-ii) A step of calculating the number of reads for each bp; and (d-iii) Step to generate a 2000bp one-dimensional image. The steps including the above; and (e) A step of inputting the generated image data into an artificial intelligence model trained to distinguish between normal images and cancerous images for analysis and comparing it with a cut-off value to determine the presence or absence of cancer. Includes, A computer-based method for providing information for early cancer diagnosis, characterized in that the amount of cell-free DNA detected differs depending on the tissue in the tissue-specific nucleosome-free region.

2. The method according to claim 1, characterized in that the step of obtaining the sequence information in step (a) above is performed in a manner that includes the following steps: (ai) Step of obtaining nucleic acids from a biological sample; (a-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using the salting-out method, column chromatography method, or beads method to obtain purified nucleic acids; (a-iii) A step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, disruption, or hydroshear method; (a-iv) The step of reacting the prepared library with a next-generation sequencer; and (av) Steps to obtain nucleic acid sequence information (reads) using a next-generation sequencer.

3. The method according to claim 1, characterized in that the nucleic acid in step (a) is cell-free DNA.

4. The method according to claim 1, further comprising the step of selecting reads whose mapping quality score of aligned nucleic acid fragments is equal to or greater than a reference value before performing step (c) above.

5. The method according to claim 4, characterized in that the aforementioned reference value is 50 to 70 points.

6. The method according to claim 1, characterized in that the artificial intelligence model in step (e) is an artificial neural network.

7. The method according to claim 6, characterized in that the artificial neural network is a convolutional neural network (CNN) or a recurrent neural network (RNN).

8. A decoding unit that extracts nucleic acids from biological samples and decodes their sequence information; Alignment unit that sorts the decoded sequences into a standard chromosome sequence database; A nucleic acid fragment sorting unit that sorts nucleic acid fragments from tissue-specific nucleosome-free regions from aligned sequence-based nucleic acid fragments; A data generation unit that generates a one-dimensional image, wherein the generation is (i) A step of superimposing selected nucleic acid fragments located within ±1000 bp from the center of the tissue-specific nucleosome-free region; (ii) A step of calculating the number of reads for each bp; and (iii) Step of generating a one-dimensional image of size 2000 bp The data generation unit including; and An information provision unit that inputs the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis, and provides information for early cancer diagnosis; Includes, The aforementioned tissue-specific nucleosome-free region is characterized by different amounts of cell-free DNA detected depending on the tissue, and is an artificial intelligence-based information provision device for early cancer diagnosis.

9. A computer-readable storage medium comprising instructions configured to be executed by a processor that provides information for the early diagnosis of cancer, (a) A step of extracting nucleic acids from a biological sample and obtaining sequence information; (b) The step of aligning the acquired sequence information (reads) with the reference genome database; (c) A step of selecting nucleic acid fragments of tissue-specific nucleosome-free regions based on the aligned sequence information (reads); (d) A step of generating a one-dimensional image, (di) A step of superimposing selected nucleic acid fragments located within ±1000 bp from the center of the tissue-specific nucleosome-free region; (d-ii) A step of calculating the number of reads for each bp; and (d-iii) Step to generate a 2000bp one-dimensional image. The steps including the above; and (e) Instructions configured to be executed by a processor that provides information for early cancer diagnosis through the step of inputting the generated image data into an artificial intelligence model trained to distinguish between normal and cancerous images for analysis and comparing it to a cut-off value to determine the presence or absence of cancer, The aforementioned tissue-specific nucleosome-free region is characterized by a computer-readable storage medium in which the amount of cell-free DNA detected differs depending on the tissue.

Citation Information

Patent Citations

  • Deep Learning-Based Variant Classifier

    JP2020525893A

  • Transcription factor profiling

    WO2020076772A1