Method for diagnosing cancer and predicting cancer type using frequency and size of terminal sequence motifs of cell-free nucleic acid fragments

By extracting and analyzing nucleic acid sequence information from biological samples, derive the frequency and size of terminal sequence templates, and using artificial intelligence models for analysis, the problem of insufficient invasiveness and accuracy of existing cancer diagnosis methods is solved, and high sensitivity and specific diagnosis and type prediction are achieved.

JP7674525B2Active Publication Date: 2025-05-09GREEN CROSS GENOME CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023573426
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-28
Filing Date
2022-05-30
Publication Date
2025-05-09
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

Existing cancer diagnosis methods are highly invasive, sensitive and specific, making it difficult to detect and accurately predict cancer types early.

Method used

By extracting nucleic acids from biological samples, obtaining sequence information, normalizing and aligning them, the frequency of the terminal sequence template and the size of nucleic acid fragments are derived, vectorized data are generated, and analysed using a learning artificial intelligence model to achieve cancer diagnosis and type prediction.

Benefits of technology

High sensitivity and high specificity of cancer diagnosis and type prediction are achieved, avoiding the invasiveness and insufficient accuracy of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674525000006
    Figure 0007674525000006
  • Figure 0007674525000007
    Figure 0007674525000007
  • Figure 0007674525000008
    Figure 0007674525000008
Patent Text Reader

Abstract

The present invention relates to a method for diagnosing cancer and predicting cancer type using the frequency and size of terminal sequence motifs of cell-free nucleic acid fragments, more specifically, a method for diagnosing cancer and predicting cancer type using a method of extracting nucleic acid from a biological sample, obtaining sequence information, deriving the frequency of terminal sequence motifs of nucleic acid fragments and the size of nucleic acid fragments based on aligned reads, generating the vectorized data, inputting the vectorized data into a trained artificial intelligence model, and analyzing the calculated values. The method for diagnosing cancer and predicting cancer type using the frequency and size of terminal sequence motifs of cell-free nucleic acid fragments according to the present invention is useful because it shows high sensitivity and accuracy even when read coverage is low, since it generates vectorized data and analyzes it using an AI algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for diagnosing cancer and predicting cancer type using the frequency and size of terminal sequence motifs of cell-free nucleic acid fragments, and more specifically, to a method for diagnosing cancer and predicting cancer type using a method in which nucleic acid is extracted from a biological sample, sequence information is obtained, and the frequency of terminal sequence motifs of the nucleic acid fragments and the size of the nucleic acid fragments are derived based on aligned reads, which are then generated as vectorized data, which are then input into a trained artificial intelligence model, and the calculated values ​​are analyzed. [Background technology]

[0002] Clinical cancer diagnosis is usually confirmed by tissue biopsy after medical history, physical examination and clinical evaluation. Clinical cancer diagnosis is possible only when the number of cancer cells is more than 1 billion and the diameter of the cancer is more than 1 cm. In this case, the cancer cells already have the ability to metastasize, and at least half of them have already metastasized. In addition, tissue biopsy is invasive, causing considerable discomfort to the patient, and there are many cases where tissue biopsy is not possible when treating cancer patients. In addition, tumor markers are used in cancer screening to monitor substances produced directly or indirectly by cancer, but their accuracy is limited because more than half of the results of tumor marker screening are normal even when cancer is present, and frequently positive even when cancer is not present.

[0003] Due to the demand for a relatively simple, non-invasive, highly sensitive and specific cancer diagnostic method that can overcome the problems of the conventional cancer diagnostic methods, liquid biopsy, which utilizes the patient's body fluids, has been widely used recently for cancer diagnosis and follow-up testing. Liquid biopsy is a non-invasive method and is a diagnostic technology that is attracting attention as an alternative to conventional invasive diagnostic and testing methods.

[0004] Recently, methods have been developed for diagnosing cancer and differentiating cancer types using cell-free DNA obtained by liquid biopsy (US 10975431, Zhou, Xionghui et al., bioRxiv, 2020.07.16.201350), and in particular, methods are known for analyzing motif frequency information of the terminal sequences of cell-free nucleic acids for use in cancer diagnosis, prenatal diagnosis, or organ transplant monitoring (WO 2020-125709, Peiyong Jiang et al., cancer discovery, Vol. 10, 2020, pp. 664-673).

[0005] Meanwhile, an artificial neural network is a computational model realized by software or hardware that uses a large number of artificial neurons connected by connection lines to mimic the computational capabilities of biological systems. An artificial neural network uses artificial neurons that simplify the functions of biological neurons. They are then interconnected through connection lines with connection strengths to perform human cognition and learning processes. The connection strength is a specific value that a connection line has, and is also called a connection weight value. Learning in an artificial neural network can be divided into supervised learning and unsupervised learning. Supervised learning refers to a method in which input data and corresponding output data are input together into a neural network, and the connection strength of the connection lines is updated so that output data corresponding to the input data is output. Representative learning algorithms include the Delta Rule and back propagation learning. Unsupervised learning refers to a method in which an artificial neural network learns its own connection strengths using only input data without a target value. Unsupervised learning is a method in which the connection weight value is updated according to the correlation between input patterns.

[0006] Much of the data applied in machine learning becomes complex, and as the dimensions increase, the curse of dimensionality problem occurs. In other words, this means that the more infinite the dimensions of the required data, the more the distance between any two points diverges infinitely, and the amount of data, i.e., density, becomes somewhat lower in high-dimensional space, making it impossible to properly reflect the features of the data (Richard Bellman, Dynamic Programming, 2003, chapter 1). Recently, the development of deep neural networks (deep learning), which have a structure with a hidden layer between the input layer and the output layer, has been reported to have significantly improved the performance of classifiers for high-dimensional data such as images, videos, and signal data by processing linear combinations of variable values ​​transmitted from the input layer with nonlinear functions (Hinton, Geoffrey, et al., IEEE Signal Processing Magazine Vol. 29.6, pp. 82-97, 2012).

[0007] There are various patents (KR 10-2018-0124550, KR 10-2019-7038076, KR 10-2019-0003676, KR 10-2019-0001741) that use such artificial neural networks in the bio field, but there is a lack of research on methods to predict cancer types through artificial neural network analysis based on sequence analysis information of cell-free DNA (cfDNA) in the blood.

[0008] Therefore, the present inventors have made extensive efforts to solve the above problems and develop an AI-based method for cancer diagnosis and cancer type prediction with high sensitivity and accuracy. As a result, they have found that when vectorized data is generated based on the terminal sequence motifs of cell-free nucleic acid fragments and the length information of the nucleic acid fragments, and the vectorized data is analyzed using a trained AI model, cancer can be diagnosed and the type of cancer can be predicted with high sensitivity and accuracy, thereby completing the present invention. Summary of the Invention [Problem to be solved by the invention]

[0009] An object of the present invention is to provide a method for diagnosing cancer and predicting cancer type using the frequency and size of terminal sequence motifs of cell-free nucleic acid fragments. Another object of the present invention is to provide a cancer diagnosis and cancer type prediction device using the frequency and size of terminal sequence motifs of cell-released nucleic acid fragments.

[0010] Another object of the present invention is to provide a computer readable storage medium comprising instructions adapted to be executed by a processor for the method of cancer diagnosis and cancer type prediction. [Means for solving the problem]

[0011] In order to achieve the above object, the present invention provides a method for providing information for cancer diagnosis and cancer type prediction, comprising: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) using the aligned sequence information (reads) to derive the frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments; (d) generating vectorized data using the derived frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the output result value, and comparing it with a cut-off value to determine the presence or absence of cancer; and (f) predicting the type of cancer through the comparison of the output result values.

[0012] The present invention also provides a method for diagnosing and predicting cancer type, comprising: (a) extracting nucleic acid from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) deriving the frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments using the aligned sequence information (reads); (d) generating vectorized data using the derived frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the output result value, and comparing it with a cut-off value to determine the presence or absence of cancer; and (f) predicting the type of cancer through the comparison of the output result values.

[0013] The present invention also provides a cancer diagnosis and cancer type prediction device, including a decoding unit that extracts nucleic acid from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment analysis unit that derives the frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments based on the aligned sequences; a data generation unit that generates vectorized data using the derived frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments; a cancer diagnosis unit that inputs the generated vectorized data into a trained artificial intelligence model for analysis and compares it with a reference value to determine the presence or absence of cancer; and a cancer type prediction unit that analyzes the output result value and predicts the type of cancer.

[0014] The present invention also provides a computer-readable storage medium, comprising instructions configured to be executed by a processor for cancer diagnosis and cancer type prediction, the computer-readable storage medium comprising instructions configured to be executed by a processor for predicting the presence or absence of cancer and the type of cancer through the steps of: (a) extracting nucleic acid from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) using the aligned sequence information (reads) to derive the frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments; (d) generating vectorized data using the derived frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the output result value, and comparing it with a cut-off value to determine the presence or absence of cancer; and (f) predicting the type of cancer through the comparison of the output result values. [Brief description of the drawings]

[0015] [Figure 1] 1 is an overall flow chart for carrying out the method for diagnosing cancer and predicting cancer type using the frequency and size of terminal sequence motifs of cell-free nucleic acid fragments according to the present invention. [Diagram 2] 1 shows an example of a process for selecting motifs whose expression frequencies differ between healthy subjects and cancer patients or between different types of cancer in one embodiment of the present invention. [Diagram 3] 1 is a graph showing the size distribution of nucleic acid fragments selected in one embodiment of the present invention. [Figure 4] The left panel shows an example of a FEMS table produced in accordance with one embodiment of the present invention using a single nucleic acid fragment, and the right panel shows an example using the entire nucleic acid fragment. [Diagram 5] The left panel shows an example of a FEMS table created by further performing Edge summary in one embodiment of the present invention, and the right panel shows the visualization result. [Figure 6] 1 is a visualization example of a FEMS table created based on data of a healthy person, a liver cancer patient, and an esophageal cancer patient used in an embodiment of the present invention. [Figure 7] (A) shows the results of checking the performance of the CNN model constructed in one embodiment of the present invention using Accuracy and micro AUC, and (B) shows the confusion matrix. [Figure 8] The results are obtained by checking the degree to which the probability values ​​of healthy individuals, liver cancer patients, and esophageal cancer patients predicted by the CNN model constructed in one embodiment of the present invention match with the actual patients through the distribution of DPI values ​​output by the CNN model. [Figure 9] FIG. 1 is a schematic diagram showing the configuration of a CNN model constructed in an embodiment of the present invention. MODE FOR CARRYING OUT THEINVENTION

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs. Generally, the nomenclature used herein and the laboratory procedures described below are those well known and commonly used in the art.

[0017] Terms such as first, second, A, B, etc. may be used to describe various elements, but the elements are not limited by the terms and are used only for the purpose of distinguishing one element from another element. For example, the first element may be named the second element, and similarly, the second element may be named the first element, without departing from the scope of the technology described below. The term "and / or" includes any combination of multiple related items or multiple related items.

[0018] In the terms used in this specification, the singular term "a" or "an" should be understood to include the plural term unless the context clearly dictates otherwise, and terms such as "comprise" or "comprise" should be understood to mean the presence of a stated feature, number, step, operation, component, or combination thereof, but not to exclude the possible presence or addition of one or more other features, numbers, steps, operations, components, components, or combinations thereof.

[0019] Before describing the drawings in detail, it should be made clear that the components in this specification are merely classified according to the main function of each component. That is, two or more components described below may be combined into one component, or one component may be divided into two or more components according to more specific functions. Furthermore, each of the components described below may perform some or all of the functions of the other components in addition to its own main function, and some of the main functions of each component may be exclusively performed by the other components.

[0020] Additionally, in performing a method or method of operation, the steps of the method may be performed in an order different from that stated, unless the context clearly dictates a particular order, i.e., the steps may be performed in the same order as stated, substantially simultaneously, or in the reverse order.

[0021] In the present invention, it was confirmed that cancer diagnosis and cancer type prediction can be performed with high sensitivity and accuracy by aligning sequence analysis data obtained from a sample to a reference genome, deriving the frequency of terminal sequence motifs of nucleic acid fragments and the size of nucleic acid fragments based on the aligned sequence information, generating vectorized data using the derived frequency of terminal sequence motifs of nucleic acid fragments and the size of nucleic acid fragments, and then calculating and analyzing the DPI value using a trained artificial intelligence model.

[0022] That is, in one embodiment of the present invention, DNA extracted from blood is sequenced and aligned to a reference chromosome, and then the frequency of the terminal sequence motif of the nucleic acid fragment and the size of the nucleic acid fragment are derived using this to generate vectorized data with the frequency of the terminal sequence motif of the nucleic acid fragment on the X-axis and the size of the nucleic acid fragment on the Y-axis. This is then trained into a deep learning model to calculate a DPI value, which is compared with a reference value to perform cancer diagnosis, and a method has been developed in which the cancer type with the highest DPI value among the DPI values ​​calculated for each cancer type is determined to be the cancer type of the sample (Figure 1).

[0023] Therefore, in one aspect, the present invention provides a method for producing a method for manufacturing a semiconductor device comprising the steps of: (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive frequencies of terminal sequence motifs of nucleic acid fragments and sizes of nucleic acid fragments; (d) generating vectorized data using the frequencies of the terminal sequence motifs of the derived nucleic acid fragments and the sizes of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the data, and comparing the output result value with a cut-off value to determine the presence or absence of cancer; and (f) A method for providing information for cancer diagnosis and cancer type prediction, comprising the step of predicting a cancer type through comparison of the output result values.

[0024] In the present invention, the nucleic acid fragment may be any fragment of nucleic acid extracted from a biological sample, and may be, but is not limited to, a fragment of cell-free nucleic acid or intracellular nucleic acid.

[0025] In the present invention, the nucleic acid fragment may be obtained by any method known to those skilled in the art, and preferably, but is not limited to, direct sequence analysis, sequence analysis through next generation sequencing analysis, sequence analysis through non-specific whole genome amplification, or probe-based sequence analysis.

[0026] In the present invention, the cancer may be a solid cancer or a blood cancer, and may be preferably selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphoid leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colon / rectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, liver cancer, thyroid cancer, gastric cancer, gallbladder cancer, biliary tract cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary site, kidney cancer, esophageal cancer, and mesothelioma, and more preferably may be, but is not limited to, liver cancer or esophageal cancer.

[0027] In the present invention, The step (a) (ai) obtaining nucleic acid from the biological sample; (a-ii) removing proteins, fats, and other residues from the collected nucleic acid by a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid; (a-iii) preparing a single-end sequencing or pair-end sequencing library from the purified nucleic acids or the nucleic acids randomly fragmented by enzymatic cleavage, crushing, or hydroshear method; (a-iv) subjecting the produced library to a next-generation sequencer; and (av) obtaining sequence information (reads) of the nucleic acid using a next-generation sequencer; The method can be characterized in that it includes:

[0028] In the present invention, the step of acquiring sequence information in step (a) may be characterized in that the separated cell-free DNA is acquired by full-length genome sequencing at a depth of 1 million to 100 million reads.

[0029] In the present invention, the biological sample means any substance, biological fluid, tissue or cell obtained from or derived from an individual, and includes, for example, whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic juice, and the like. The fluid may include, but is not limited to, tissue, fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, semen, hair, saliva, urine, buccal cells, placental cells, cerebrospinal fluid, and mixtures thereof.

[0030] In the present invention, the next-generation sequencer may be used in any sequencing method known in the art. Sequencing of the nucleic acids isolated by the selection method is usually performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of an individual nucleic acid molecule or one of the clonally extended proxies for an individual nucleic acid molecule in a very similar manner (e.g., 10 5(More than one molecule are sequenced simultaneously). In one embodiment, the relative abundance of a nucleic acid species in a library can be estimated by measuring the relative occurrence of its cognate sequence in the data generated by the sequencing experiment. Next generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference.

[0031] In one embodiment, next generation sequencing is performed to determine the nucleotide sequence of individual nucleic acid molecules (eg, the HeliScope Gene Sequencing system from Helicos BioSciences and the PacBio RS system from Pacific Biosciences). In other embodiments, sequencing, e.g., massively parallel short-read sequencing (e.g., Solexa sequencer, Illumina Inc., San Diego, Calif.), which generates more bases of sequence per sequencing unit than other sequencing methods that generate fewer but longer reads, determines the nucleotide sequence of clonally extended proxies for individual nucleic acid molecules (e.g., Solexa sequencer, Illumina Inc., San Diego, Calif.; 454 Life Sciences (Branford, Conn.) and Ion Torrent. Other methods or machines for next-generation sequencing include, but are not limited to, those provided by 454 Life Sciences (Branford, Conn.), Applied Biosystems (Foster City, Calif.; SOLiD sequencer), Helicos Biosciences Corporation (Cambridge, Massachusetts), and emulsion and microflow sequencing methods nanoinfusion (e.g., GnuBio infusion).

[0032] Platforms for next generation sequencing include, but are not limited to, the Genome Sequencer (GS) FLX system from Roche / 454, the Genome Analyzer (GA) from Illumina / Solexa, the Support Oligonucleotide Ligation Detection (SOLiD) system from Life / APG, the G.007 system from Polonator, the HeliScope Gene Sequencing system from Helicos BioSciences, and the PacBio RS system from Pacific Biosciences. NGS techniques may include, for example, at least one of the steps of template preparation, sequencing and imaging, and data analysis.

[0033] Template preparation. The method of preparing template may include steps such as randomly breaking nucleic acid (e.g., genomic DNA or cDNA) into small sizes and making sequencing template (e.g., fragment template or mate pair template). Spatially separated template may be attached or fixed to a solid surface or support, which allows large-scale sequencing reaction to be performed simultaneously. The types of template that can be used for NGS reaction include, for example, clone amplified template derived from a single DNA molecule and single DNA molecule template. Methods for producing clonally amplified templates include, for example, emulsion PCR (emPCR) and solid-phase amplification.

[0034] EmPCR may be used to produce templates for NGS. Typically, a library of nucleic acid fragments is created and adapters containing universal priming sites are ligated to the ends of the fragments. The fragments are then denatured into single strands and captured by beads. Each bead captures a single nucleic acid molecule. After amplification and enrichment of the emPCR beads, a quantity of templates may be attached, immobilized in a polyacrylamide gel on a standard microscope slide (e.g., Polonator), chemically crosslinked to an amino-coated glass surface (e.g., Life / APG; Polonator), or deposited on individual PicoTiterPlate (PTP) wells (e.g., Roche / 454), at which point the NGS reaction may be performed.

[0035] Solid-phase amplification may also be used to generate templates for NGS. Typically, forward and reverse primers are covalently attached to a solid support. The surface density of the amplified fragments is defined as the ratio of primers to templates on the support. Solid-phase amplification can generate millions of spatially separated template clusters (e.g., Illumina / Solexa). The ends of the template clusters can be hybridized to universal primers for the NGS reaction.

[0036] Other methods for the production of clonally amplified templates include, for example, Multiple Displacement Amplification (MDA) (Lasken RS Curr Opin Microbiol:510-6). MDA is a non-PCR-based DNA amplification method. The reaction involves annealing random hexamer primers to the template and synthesizing DNA with a high-fidelity enzyme, usually Φ29, at a constant temperature. MDA can produce products of large size with a lower error frequency.

[0037] Template amplification methods such as PCR can couple NGS platforms to targets or enrich specific regions of the genome (e.g., exons). Representative template enrichment methods include, for example, microdroplet PCR (Tewhey R. et al., Nature Biotech. 2009, 27:1025-1031), custom-designed oligonucleotide microarrays (e.g., Roche / NimbleGen oligonucleotide microarrays) and solution-based hybridization methods (e.g., molecular inversion probes (MIPs)) (Porreca GJ et al., Nature Methods, 2007, 4:931-936; Krishnakumar S. et al. USA, 2008, 105:9296-9310; Turner EH et al., Nature Methods, 2009, 6:315-316) and biotinylated RNA capture sequences (Gnirke A. et al.:182-9).

[0038] Unimolecular templates are another type of template that can be used for NGS reactions. Spatially separated unimolecular templates may be immobilized on a solid support by a variety of methods. In one approach, individual primer molecules are covalently attached to a solid support. Adapters are added to the template, and the template is then hybridized to the immobilized primer. In another approach, the unimolecular template is covalently attached to a solid support by priming and extending the single-stranded unimolecular template from an immobilized primer. A universal primer is then hybridized to the template. In another approach, a single polymerase molecule is attached to the solid support to which the primed template is attached.

[0039] Sequencing and Imaging. Exemplary sequencing and imaging methods for NGS include, but are not limited to, cyclic reversible termination (CRT), sequencing by ligation (SBL), single molecule addition (pyrosequencing), and real-time sequencing.

[0040] CRT uses reversible terminators in a cyclic method that involves minimal nucleotide inclusion, fluorescent imaging, and cleavage steps. Typically, DNA polymerase incorporates a single fluorescently modified nucleotide into the primer that is complementary to the template base's complementary nucleotide. DNA synthesis is terminated after the addition of the single nucleotide, and the unincorporated nucleotide is washed away. Imaging is performed to determine the identity of the incorporated labeled nucleotide. A cleavage step then removes the terminator / inhibitor and fluorescent dye. Representative NGS platforms that use the CRT method include, but are not limited to, the Illumina / Solexa Genome Analyzer (GA), which uses a clonally amplified template method coupled with a four-color CRT method detected by total internal reflection fluorescence (TIRF), and the Helicos BioSciences / HeliScope, which uses a single-molecule template method coupled with a one-color CRT method detected by TIRF. SBL uses DNA ligase and either single- or double-encoded probes for sequencing.

[0041] Typically, a fluorescently labeled probe hybridizes to a complementary sequence adjacent to the primed template. DNA ligase is used to ligate the dye-labeled probe to the primer. After washing away the unligated probe, fluorescent imaging is performed to determine the identity of the ligated probe. The fluorescent dye may be removed using a cleavable probe that regenerates the 5'-PO4 group for the subsequent ligation cycle. Alternatively, new primers may be hybridized to the template after removing the old primer. Exemplary SBL platforms include, but are not limited to, Life / APG / SOLiD (supported oligonucleotide ligation detection), which uses a two-base coded probe.

[0042] Pyrosequencing methods are based on detecting the activity of DNA polymerase with different chemiluminescent enzymes. Typically, this method sequences a single strand of DNA by synthesizing the complementary strand along one base pair at a time and detecting the base actually added at each step. The template DNA is fixed, and solutions of A, C, G, and T nucleotides are added sequentially and removed from the reaction. Light is only produced when a nucleotide solution replenishes an unpaired base in the template. The sequence of the solution that produces a chemiluminescent signal will determine the sequence of the template. Representative pyrosequencing platforms include, but are not limited to, Roche / 454, which uses DNA templates produced by emPCR with 1-2 million beads deposited in PTP wells.

[0043] Real-time sequencing involves imaging the sequential inclusion of dye-labeled nucleotides during DNA synthesis. Representative real-time sequencing platforms include, but are not limited to, the Pacific Biosciences platform, which uses DNA polymerase molecules attached to the surface of individual zero-mode waveguide (ZMW) detectors to obtain sequence information as phosphate-linked nucleotides are included in the growing primer strand; the Life / VisiGen platform, which uses genetically engineered DNA polymerases with attached fluorescent dyes to generate enhanced signals after nucleotide inclusion by fluorescence resonance energy transfer (FRET); and the LI-COR Biosciences platform, which uses dye-quencher nucleotides in the sequencing reaction.

[0044] Other sequencing methods of NGS include, but are not limited to, nanopore sequencing, sequencing by hybridization, nanotransistor array-based sequencing, polony sequencing, scanning tunneling microscopy (STM)-based sequencing, and nanowire molecular sensor-based sequencing.

[0045] Nanopore sequencing involves electrophoresis of nucleic acid molecules in solution through nanoscale pores that provide a highly confined space that allows single nucleic acid polymers to be analyzed. Representative methods of nanopore sequencing are described, for example, in [Branton D. et al., Nat Biotechnol. 2008; 26(10): 1146-53].

[0046] Sequencing by hybridization is a non-enzymatic method using DNA microarrays. Usually, a single pool of DNA is fluorescently labeled and hybridized to an array containing a known sequence. Hybridization signals can identify the DNA sequence from a given drop on the array. The binding of a single strand of DNA to its complementary strand in a DNA duplex is sensitive to single base mismatches if the hybridization region is short or if a embodied mismatch detection protein is present. Representative methods of sequencing by hybridization are described in the literature, for example (Hanna GJ et al. J. Clin.Microbiol. 2000; 38(7):2715-21; and Edwards JR et al.2005; 573(1-2):3-12).

[0047] Polony sequencing is based on sequencing by polony amplification and multiple single base extension (FISSEQ). Polony amplification is a method for amplifying DNA in situ on a polyacrylamide film. A representative polony sequencing method is described, for example, in U.S. Patent Application Publication No. 2007 / 0087362.

[0048] Nanotransistor array-based devices such as Carbon NanoTube Field Effect Transistors (CNTFETs) can also be used for NGS. For example, a DNA molecule is stretched and driven across a nanotube by microfabricated electrodes. The DNA molecule is sequentially brought into contact with the surface of the carbon nanotube, and charge transfer between the DNA molecule and the nanotube results in differences in current flow from each base. The DNA is sequenced by recording these differences. Exemplary nanotransistor array-based sequencing methods are described, for example, in U.S. Patent Publication No. 2006 / 0246497.

[0049] Scanning electron tunneling microscopes (STMs) can also be used for NGS. STMs use a piezoelectrically controlled probe that raster scans the sample to form an image of the sample surface. STMs may be used, for example, to image the physical properties of single DNA molecules, resulting in consistent electron tunneling imaging and spectroscopy, by integrating a scanning electron tunneling microscope with an actuator-driven flexible gap. Exemplary sequencing methods using STMs are described, for example, in U.S. Patent Application Publication No. 2007 / 0194225.

[0050] Molecular analysis devices consisting of nanowire molecular sensors may also be used in NGS. Such devices can detect the interaction of nitrogenous substances disposed on the nanowires with nucleic acid molecules, such as DNA. Molecular guides are disposed to guide molecules in the vicinity of the molecular sensor to allow for interaction and subsequent detection. Exemplary sequencing methods using nanowire molecular sensors are described, for example, in U.S. Patent Application Publication No. 2006 / 0275779.

[0051] For NGS, double-end sequencing methods may be used.Double-end sequencing uses blocked and unblocked primers to sequence both the sense and antisense strands of DNA.Usually, these methods include: annealing an unblocked primer to a first strand of nucleic acid; annealing a second blocked primer to a second strand of nucleic acid; extending the nucleic acid along the first strand with a polymerase; terminating the first sequencing primer; deblocking the second primer; and extending the nucleic acid along the second strand.Representative double-stranded sequencing methods are described, for example, in U.S. Patent No. 7,244,567.

[0052] After the NGS reads are generated, they are aligned or de novo assembled against a known reference sequence. For example, genetic variations such as single nucleotide polymorphisms and structural variations in a sample (e.g., a tumor sample) can be identified by aligning the NGS reads against a reference sequence (e.g., a wild-type sequence). NGS sequence alignment methods are described, for example, in the literature (Trapnell C. and Salzberg SL Nature Biotech., 2009, 27:455-457).

[0053] Examples of de novo assembly are described, for example, in the literature (Warren R. et al., Bioinformatics, 2007, 23:500-501; Butler J. et al., Genome Res., 2008, 18:810-820; and Zerbino DR and Birney E., Genome Res., 2008, 18:821-829).

[0054] Sequence alignment or assembly may be performed using read data from one or more NGS platforms, for example, mixing Roche / 454 and Illumina / Solexa read data. In the present invention, the alignment step may be performed using, but is not limited to, the BWA algorithm and hg19 sequence.

[0055] In the present invention, the sequence alignment in step (b) includes a computational method or approach used as a computer algorithm to identify the lead sequence in a genome (e.g., a short lead sequence from next generation sequencing) from which most may be derived by evaluating the similarity between the lead sequence and a reference sequence. Various algorithms can be applied to the sequence alignment problem. Some algorithms are relatively slow but allow relatively high specificity. These include, for example, algorithms based on dynamic programming. Dynamic programming is a method of solving complex problems by dividing them into simpler steps. Other approaches are relatively efficient but generally not exhaustive. These include, for example, heuristic algorithms and probabilistic methods designed for large-scale database searching.

[0056] Typically, there may be two steps in the alignment process: candidate inspection and sequence alignment. Candidate inspection reduces the search space for sequence alignment from the entire genome to a shorter enumeration of possible alignment positions. As the term suggests, sequence alignment involves aligning sequences with sequences provided in the candidate inspection step. This may be done using global alignment (e.g., Needleman-Wunsch alignment) or local alignment (e.g., Smith-Waterman alignment).

[0057] Most attribute alignment algorithms can be characterized as one of three types based on indexing schemes: algorithms based on hash tables (e.g. BLAST, ELAND, SOAP), suffix trees (e.g. Bowtie, BWA) and merge sorts (e.g. Slider). Short read sequences are usually used for alignment. Examples of sequence alignment algorithms / programs for short read sequences include, but are not limited to, BFAST (Homer N. et al., PLoS One.2009; 4(11): e7767), BLASTN (from blast.ncbi.nlm.nih.gov on the World Wide Web), BLAT (Kent WJ Genome Res.2002; 12(4): 656-64), Bowtie (Langmead B. et al., Genome Biol. 2009; 10(3): R25), BWA (Li H. and Durbin R. Bioinformatics, 2009, 25: 1754-60), BWA-SW (Li H. and Durbin R. Bioinformatics, 2010; 26(5): 589-95), CloudBurst (Schatz MC Bioinformatics.2009;25(11):1363-9), Corona Lite (Applied Biosystems, Carlsbad, California, USA), CASHX (Fahlgren N. et al., RNA, 2009; 15, 992-1002), CUDA-EC (Shi H. et al, J Comput Biol. 2010;17(4):603-15), ELAND (from bioit.dbi.udel.edu / howto / eland on the World Wide Web), GNUMAP (Clement NL et al.2010;26(1):38-45), GMAP (Wu TD and Watanabe CK Bioinformatics.2005;21(9):1859-75), GSNAP(Wu TD and Nacu S., Bioinformatics.2010;26(7):873-81)Geneious Assembler(Copyright © Biomatters Ltd.) LAST, MAQ(Li H. et al., Genome Res.2008;18(11):1851-8), Mega-BLAST(World Wide Web of ncbi.nlm.nih.gov / blast / megablast.shtml), MOM(Eaves HL and Gao Y). Bioinformatics.2009;25(7):969-70)、MOSAIC(Bioinformatics.bc.edu / marthlab / Mosaic)、Novoalign(Novocraft.com / main / index.phpより), PALMapper(World Wide Web of fml.tuebingen.mpg.de / raetsch / suppl / palmapper), PASS(Campagna D. et al, Bioinformatics.2009;25(7):967-8), PatMaN(Prufer K. et al.2008; 24(13):1530-1), PerM(Chen Y. et al., Bioinformatics, 2009, 25(19):2514-2521), ProbeMatch(Kim YJ et al.2009;25(11):1424-5), QPalma(de Bona F. et al., Bioinformatics, 2008, 2008). 24(16):i174); 2008;24:2395-2396.) Shrec(Salmela L., Bioinformatics.2010;26(10):1284-90), SHRiMP(Rumble SM et al., PLoS Comput. Biol., 2009, 5(5):e1000386), SLIDER (Malhis N. et al., Bioinformatics, 2009, 25 (1):6-13), SLIM Search (Muller T. et al., Bioinformatics. 2001;17 Suppl 1:S182-9), SOAP (Li R. et al., Bioinformatics.2008;24(5):713-4), SOAP2 (Li R. et al., Bioinformatics.2009;25(15):1966-7), SOCS (Ondov BD et al., Bioinformatics, 2008; 24(23):2776-7), SSAHA (Ning Z. et al. al.2001;11(10):1725-9), SSAHA2 (Ning Z. et al.et al.2001;11(10):1725-9), Stampy (Lunter G. and Goodson M. Genome Res. 2010, epub ahead of print), Taipan (on the World Wide Web at taipan.sourceforge.net), UGENE (on the World Wide Web at ugene.unipro.ru), XpressAlign (on the World Wide Web at bcgsc.ca / platform / bioinfo / software / XpressAlign), and ZOOM (Bioinformatics Solutions Inc., Waterloo, Ontario, Canada).

[0058] A sequence alignment algorithm may be selected based on multiple factors, such as, for example, the sequencing method, the length of the read, the number of reads, the available computational data, and the sensitivity / scoring requirements. Different sequence alignment algorithms can achieve different speed levels, alignment sensitivity, and alignment specificity. Alignment specificity refers to the percentage of target sequence residues that are aligned as found in the normal submission that are correctly aligned compared to the predicted alignment. Alignment sensitivity also refers to the percentage of target sequence residues that are aligned as found in the normal predicted alignment that are correctly aligned in the submission.

[0059] Alignment algorithms such as ELAND and SOAP may be used to align short reads (e.g., from Illumina / Solexa sequencers) against a reference genome when speed is the primary factor. Alignment algorithms such as BLAST and Mega-BLAST may be used to search for similarity using short reads (e.g., from RocheFLX) when specificity is the most important factor, even though these methods are relatively slow. Alignment algorithms such as MAQ and Novoalign take into account quality scores and may be used on single-end or paired-end data when accuracy is important (e.g., in fast and large-scale SNP searches). Alignment algorithms such as Bowtie and BWA require a relatively small memory footprint because they use block sorting (Burrows-Wheeler Transform: BWT). Alignment algorithms such as BFAST, Perm, SHRiMP, SOCS, and ZOOM may be used with ABI's SOLiD platform to map color space reads. In some applications, the results of two or more alignment algorithms may be combined.

[0060] In the present invention, the length of the sequence information (reads) in step (b) is 5 to 5,000 bp, and the number of pieces of sequence information used may be, but is not limited to, 5,000 to 5,000,000.

[0061] In the present invention, the terminal sequence motif of the nucleic acid fragment in step (c) can be characterized as being a pattern of 2 to 30 base sequences at both ends of the nucleic acid fragment. That is, if there is a nucleic acid fragment sequenced by paired-end sequencing as follows: Forward strand: 5'-TACAGACTTTGGAAT-3' (SEQ ID NO: 1) Reverse strand: 3`-ATGACTGAAACCTTA-5` (SEQ ID NO: 2) TACA read in sequence from the 5' end of the forward strand and ATTC read in sequence from the 5' end of the reverse strand are the terminal sequence motif values ​​of this nucleic acid fragment.

[0062] In the present invention, the frequency of the terminal sequence motif of the nucleic acid fragment in the step (c) may be characterized as the number of each motif detected from the entire nucleic acid fragment.

[0063] In other words, when analyzing the terminal sequence motif of a nucleic acid fragment based on the four bases at both ends (4-mer motif), there are four possible combinations of bases, A, T, G, and C, at the 1st, 2nd, 3rd, and 4th positions, respectively, so a total of 256 combinations of motif values ​​(4*4*4*4) can be analyzed.

[0064] The number of times each motif is observed among all the nucleic acid fragments generated by sequencing is the motif frequency, and the relative frequency of each motif is calculated by dividing this value by the total number of nucleic acid fragments generated.

[0065] [Table 1]

[0066] As shown in Table 1, the total number of nucleic acid fragments is 126,430,124, and the number of nucleic acid fragments analyzed for AAAA as the nucleic acid fragment terminal sequence motif is 125,071. Therefore, the frequency of the AAAA nucleic acid fragment terminal sequence motif is 125,071, and the relative frequency of the nucleic acid fragment terminal sequence motif calculated by dividing this by the total number of nucleic acid fragments is 0.00099. In the present invention, the size of the nucleic acid fragment in step (c) can be characterized as the number of bases from the 5' end to the 3' end of the nucleic acid fragment. For example, the size of the nucleic acid fragment analyzed in SEQ ID NO: 1 and SEQ ID NO: 2 is 15.

[0067] In the present invention, the size of the nucleic acid fragment may be 1 to 10,000, preferably 10 to 1,000, more preferably 50 to 500, and most preferably 90 to 250, but is not limited thereto.

[0068] In the present invention, the vectorized data in the step (d) can be characterized in that the type of terminal sequence motif of the nucleic acid fragment is represented on the X-axis, and the size of the nucleic acid fragment is represented on the Y-axis. In other words, assuming there is one nucleic acid fragment like this: Forward strand: 5'-TACAGACTAGT … TTGGAAT-3' (SEQ ID NO: 3) Reverse strand: 3'-ATGACTGATCA … AACCTTA-5' (SEQ ID NO: 4) Fragment Size:176

[0069] This nucleic acid fragment may be represented as a two-dimensional vector as shown in the left panel of Figure 4, and when this process is extended and accumulated to the entire nucleic acid fragment, it will generate a two-dimensional vector as shown in the right panel of Figure 4.

[0070] In the present invention, the vectorized data may further include a sum of frequencies of terminal motifs of nucleic acid fragments and a sum of frequencies of nucleic acid fragments by size.

[0071] That is, in order to add frequency information for each fragment end motif unrelated to the fragment size, a column sum value is added four times to the bottom of the two-dimensional vector of FIG. 4, and in order to add fragment size information unrelated to the fragment end motif, a row sum value is added four times to the rightmost part of the two-dimensional vector of FIG. 4 by additionally executing Edge Summary, thereby generating a two-dimensional vector as shown in the left panel of FIG. 5.

[0072] In the present invention, the two-dimensional vector is defined as the Fragment End Motif frequency and Size (FEMS) table. The FEMS table can be visualized as shown in the right panel of FIG. 5 and FIG. 6.

[0073] In the present invention, the vectorized data is preferably, but not limited to, an image. An image is basically composed of pixels, and when an image composed of pixels is vectorized, it may be represented as a one-dimensional 2D vector (black and white), a three-dimensional 2D vector (color (RGB)), or a four-dimensional 2D vector (color (CMYK)) depending on the type of image.

[0074] The vectorized data of the present invention is not limited to images, and may be, for example, an n-dimensional 2D vector (multi-dimensional vector) obtained by stacking multiple n black-and-white images, and used as input data for an artificial intelligence model.

[0075] In the present invention, the method may further comprise a step of classifying nucleic acid fragments that satisfy a mapping quality score of the aligned nucleic acid fragments before performing the (c) step.

[0076] In the present invention, the mapping quality score may vary depending on the desired criteria, but may be preferably 15 to 70 points, more preferably 50 to 70 points, and most preferably 60 points.

[0077] In the present invention, the artificial intelligence model in step (e) may be any model that can learn to distinguish images according to cancer type, and is preferably characterized as being a deep learning model.

[0078] In the present invention, the artificial intelligence model may be an artificial neural network algorithm capable of analyzing vectorized data based on an artificial neural network without limitation, and may be preferably selected from the group consisting of a Convolutional Neural Network (CNN), a Deep Neural Network (DNN), and a Recurrent Neural Network (RNN), but is not limited thereto.

[0079] In the present invention, the recurrent neural network may be selected from the group consisting of a long-short term memory (LSTM) neural network, a gated recurrent unit (GRU) neural network, a vanilla recurrent neural network, and an attentive recurrent neural network.

[0080] In the present invention, when the artificial intelligence model is a CNN, the loss function for binary classification may be characterized by being expressed by the following Equation 1, and the loss function for multi-class classification may be characterized by being expressed by the following Equation 2.

[0081]

number

[0082]

number

[0083] In the present invention, when the artificial intelligence model is a CNN, the learning can be characterized by including the following steps: i) classifying the produced vector data into training, validation, and test data; In this case, the training data is used to train the CNN model, the validation data is used to verify hyper-parameter tuning, and the test data is used to evaluate the performance of the optimal model after it is produced. ii) constructing an optimal CNN model through hyperparameter tuning and learning process; iii) comparing the performances of the multiple models obtained through hyperparameter tuning using validation data, and determining the model with the best performance on the validation data as the optimal model; In the present invention, the hyperparameter tuning process is a process of optimizing values ​​of a plurality of parameters (such as the number of convolutional layers, the number of dense layers, the number of convolutional filters, etc.) constituting a CNN model, and the hyperparameter tuning process may be characterized by using a Bayesian optimization and a grid search method.

[0084] In the present invention, the learning process can be characterized in that the internal parameters (weights) of the CNN model are optimized using the determined hyperparameters, and when the validation loss starts to increase relative to the training loss, it is determined that the model is overfitting, and model learning is interrupted before that happens.

[0085] In the present invention, the result value analyzed from the vectorized data input to the artificial intelligence model in step e) may be any specific score or real number without limitation, and is preferably, but is not limited to, a DPI (Deep Probability Index) value.

[0086] In the present invention, the deep probability index refers to a value expressed as a probability value obtained by adjusting the output of the artificial intelligence model on a scale of 0 to 1 using a sigmoid function in the case of binary classification, or a softmax function in the case of multi-class classification, in the last layer of the artificial intelligence model.

[0087] In the case of binary classification, a sigmoid function is used to train the model so that in the case of cancer, the DPI value is 1. For example, when a breast cancer sample and a normal sample are input, the model trains the model so that the DPI value of the breast cancer sample approaches 1.

[0088] In the case of multi-class classification, a softmax function is used to extract the DPI values ​​for the number of classes. The sum of the DPI values ​​for the number of classes is 1, and the system is trained so that the DPI value for the actual corresponding cancer type is 1. For example, if there are three classes: breast cancer, liver cancer, and normal, and a breast cancer sample is entered, the system will train to bring the breast cancer class closer to 1.

[0089] In the present invention, the output result value of the step (e) may be derived for each cancer type.

[0090] In the present invention, the artificial intelligence model learns so that the output result is close to 1 if there is cancer and close to 0 if there is no cancer, and measures performance based on this (Training, validation, test accuracy) by using 0.5 as the standard and judging that there is cancer if the standard is 0.5 or above and judging that there is no cancer if the standard is 0.5 or below.

[0091] It is obvious to ordinary engineers that the standard value of 0.5 can be changed at any time. For example, if you want to reduce false positives, you can set a standard value higher than 0.5 and set a stricter standard for determining whether cancer is present. If you want to reduce false negatives, you can measure a lower standard and set a slightly weaker standard for determining whether cancer is present.

[0092] Most preferably, a trained artificial intelligence model can be used to apply unseen data (data that knows the answer and has not been trained on) to ascertain the probability of DPI values ​​and determine a baseline value.

[0093] In the present invention, the step of predicting the type of cancer by comparing the output result values ​​in step (f) can be characterized as being carried out by a method including a step of determining that the cancer type showing the highest output result value is cancer of the sample.

[0094] From another aspect, the present invention provides a decoding unit for extracting nucleic acid from a biological sample and decoding sequence information; an alignment section for aligning the decoded sequence to a standard chromosome sequence database; and a nucleic acid fragment analysis unit for deriving the frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments based on the aligned sequences; a data generating unit for generating vectorized data using the derived frequencies of terminal sequence motifs of nucleic acid fragments and sizes of the nucleic acid fragments; A cancer diagnosis unit that inputs the generated vectorized data into a trained artificial intelligence model for analysis and compares it with a reference value to determine the presence or absence of cancer; and The present invention relates to a cancer diagnosis and cancer type prediction device including a cancer type prediction unit that analyzes the output result values ​​and predicts the cancer type.

[0095] In the present invention, the decoding unit may include a nucleic acid injection unit that injects nucleic acid extracted from an independent device; and a sequence information analysis unit that analyzes the sequence information of the injected nucleic acid, and may be preferably, but is not limited to, an NGS analysis device.

[0096] In the present invention, the decoding unit may receive and decode sequence information data generated by an independent device.

[0097] In another aspect, the present invention is a computer readable storage medium including instructions configured to be executed by a processor for cancer diagnosis and predicting cancer type, the instructions comprising: (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive frequencies of terminal sequence motifs of nucleic acid fragments and sizes of nucleic acid fragments; (d) generating vectorized data using the frequencies of the terminal sequence motifs of the derived nucleic acid fragments and the sizes of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the data, and comparing the output result value with a cut-off value to determine the presence or absence of cancer; and (f) a computer readable storage medium comprising instructions configured to be executed by a processor to predict the presence or absence and type of cancer through the step of predicting the type of cancer through a comparison of the output result values.

[0098] In another aspect, the method according to the present invention can be implemented using a computer. In one embodiment, the computer includes one or more processors coupled to a chipset. Also coupled to the chipset are memory, storage, a keyboard, a graphics adapter, a pointing device, and a network adapter. In one embodiment, the capabilities of the chipset are enabled by a memory controller hub and an I / O controller hub. In another embodiment, the memory can be directly coupled to the processor instead of the chipset. The storage is any device capable of maintaining data, including a hard drive, a compact disk read-only memory (CD-ROM), a DVD, or other memory device. The memory is responsible for the data and instructions used by the processor. The pointing device may be a mouse, a track ball, or other type of pointing device, and is used in combination with a keyboard to send input data to the computer system. The graphics adapter shows images and other information on a display. The network adapter is coupled to the computer system by a local or long-distance communication network. The computer used in the present application is not, however, limited to the above configurations, and may include none or additional configurations, and may be part of a storage area network (SAN), and the computer of the present application may be configured to execute program modules for performing the method of the present application.

[0099] In the present application, a module may refer to a functional and structural combination of hardware for carrying out the technical idea of ​​the present application and software for driving the hardware. For example, the module may refer to a logical unit of a certain code and a hardware resource for executing the certain code, and it is obvious to those skilled in the art that the module does not necessarily refer to physically connected code or to one type of hardware.

[0100] From another viewpoint, the present invention provides: (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive frequencies of terminal sequence motifs of nucleic acid fragments and sizes of nucleic acid fragments; (d) generating vectorized data using the frequencies of the terminal sequence motifs of the derived nucleic acid fragments and the sizes of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the data, and comparing the output value with a cut-off value to determine the presence or absence of cancer; and (f) A method for diagnosing and predicting a type of cancer, comprising predicting a type of cancer through comparison of the output result values. EXAMPLES

[0101] The present invention will be described in more detail below with reference to examples. It will be obvious to those skilled in the art that these examples are intended solely to illustrate the present invention and are not intended to limit the scope of the present invention.

[0102] Example 1. Extract DNA from blood and perform next-generation sequencing analysis. 10mL of blood was collected from 349 healthy subjects, 51 liver cancer patients, and 108 esophageal cancer patients, and stored in EDTA tubes. Within 2 hours of collection, the plasma portion was centrifuged at 1200g, 4℃, and 15 minutes, and the plasma supernatant was separated by centrifuging the plasma at 16000g, 4℃, and 10 minutes to remove precipitates. Cell-free DNA was extracted from the separated plasma using the Tiangenmicro DNA kit (Tiangen), and library preparation was performed using the MGI Easy cell-free DNA library prep set kit, followed by sequencing using the DNBseq G400 equipment (MGI) in 100 base paired end mode. As a result, it was confirmed that approximately 170 million reads were generated per sample.

[0103] Example 2. Selection of terminal motifs and size of nucleic acid fragments 2-1. Selection of terminal motifs of nucleic acid fragments The terminal motifs of nucleic acid fragments are set to four bases (A, T, G, C), and among the 256 motifs (4*4*4*4), there are motifs whose relative frequency does not differ between the Normal, HCC, and EC groups. If such motifs with no difference are included in the FEMS table, no meaningful information for classification is provided, and they become noise that only increases the amount of calculations required for the model. Therefore, to exclude such meaningless motifs, we selected only specific motifs whose relative frequency differed significantly between the three groups.

[0104] In addition, to prevent the problem of overfitting the model during the size and motif selection process, only the training set is used in the size and motif selection process. That is, using the NGS data generated in Example 1, the terminal motifs of the nucleic acid fragments were set to four bases (A, T, G, C), and some motifs that showed a statistically significant level (Kruskal-Wallis Test, FDR-adjusted p<0.05) of relative frequency difference between healthy subjects (Normal), liver cancer (HCC), and esophageal cancer (EC) patient groups were selected from a total of 256 motifs (4*4*4*4) (Figure 2).

[0105] In addition, to prevent overfitting, motifs selected in the above process were further selected that had an average frequency in the healthy subject group higher than the random baseline (1 / 256, 0.004). As a result, a total of 84 motifs were selected, and the detailed motif information is as follows: CTGG, ACTT, CCTA, TGGA, TGGG, CAGG, TATA, CCTT, CAGC, TAGA, AGAA, AGAG, CATA, CAGT, CAGA, ACCT, CTGT, ACAT, GCTT, GCTA, TCAG, CTTA, GGCC, ATTT, CCCA, TATC, CCTG, TCTA, GCCT, ACTG, TGAG, GGTA, CATT, TATT, CCAT, CCTC, CCAA, CTTT, TAAG, GCTG, CCCT, TGAA, ACCA, GTTT, TGTA, CTCA, GCCA, TATG, GCAT, AAAG, AAAA, GGCT, TGAC, AGCA, TCTT, CTGA, CATC, ACAA, GACA, AACA, CCCC, CACT, GGAG, GGCA, TCAA, CAAG, TAAA, AAAT, TGCC, GGTT, GGGA, CCAC, TGTG, CATG, TGCA, GAAT, TGTC, TGCT, CAAT, GGAA, AGTG, TACT, CACA, TCCC

[0106] 2-2. Size selection of nucleic acid fragments In the case of size selection of nucleic acid fragments, most of the nucleic acid fragments that have undergone quality confirmation have sizes in the range of 90 to 250, as shown in Figure 3. Therefore, if an FEMS table is generated including areas outside this size range, most of the areas will be filled with zero values, and only meaningless noise will increase, so the above sizes were selected.

[0107] Example 3. Creation of a Fragment End Motif frequency and Size (FEMS) table A two-dimensional vector was generated by placing the motif type on the X-axis and the fragment size on the Y-axis so that the fragment end motif frequency value and size information of the nucleic acid fragments selected in Example 2 could be simultaneously represented. More specifically, as shown in the left panel of Figure 4, the type and size of the nucleic acid motif at both ends of one nucleic acid fragment were expressed as a frequency number, and this was expanded and accumulated to the entire nucleic acid fragment to generate a two-dimensional vector as shown in Figure 4.

[0108] Additionally, to add frequency information for each Fragment End Motif that is unrelated to Fragment Size, a column sum value was added four times to the bottom of the two-dimensional vector, and an Edge Summary step was performed to add row sum values ​​four times to the rightmost part of the two-dimensional vector to add Fragment Size information that is unrelated to Fragment End Motif, ultimately generating a two-dimensional vector as shown in Figure 5. This two-dimensional vector was defined as the Fragment End Motif frequency and Size (FEMS) table, and an example of its visualization is shown in Figure 5.

[0109] Example 3. CNN model construction and learning process Using the FEMS table two-dimensional vector as input, a CNN artificial intelligence model was trained to distinguish between healthy individuals, liver cancer patients, and esophageal cancer patients. All samples were divided into Training, Validation, and Test datasets, and the Training dataset was used for model training, the Validation dataset for hyperparameter tuning, and the Test dataset for evaluating the performance of the final model. The number of samples in each set is as follows.

[0110] [Table 2]

[0111] The basic configuration of the CNN model is shown in Figure 9. The activation function used was ReLU (Rectified Linear Unit), one convolution layer was used, and five 10*10 patches were used. The pooling method used was max, and 2x2 patches were used. One fully connected layer was used, and it contained 512 hidden nodes. Finally, the final DPI value was calculated using the softmax function value.

[0112] The hyperparameter tuning process optimizes the values ​​of multiple parameters that make up the CNN model (number of convolutional layers, number of dense layers, number of convolutional filters, etc.). Bayesian optimization and grid search methods were used for the hyperparameter tuning process, and when the validation loss began to increase relative to the training loss, it was determined that the model was overfitting, and model training was interrupted.

[0113] The performance of multiple models obtained through hyperparameter tuning was compared using the validation dataset, and the model with the best performance on the validation dataset was determined to be the optimal model, and final performance evaluation was performed on the test dataset.

[0114] When the FEMS table 2D vector of any sample is input into the model created through the above process, the probability that the sample is a healthy person, a liver cancer patient, or an esophageal cancer patient is calculated through the softmax function, which is the last layer of the CNN model, and this probability value is defined as the Deep Probability Index (DPI).

[0115] A given sample is judged to belong to the group with the highest of the three DPI values. For example, if the DPI values ​​calculated for a given sample for a healthy person, a liver cancer patient, and an esophageal cancer patient are 0.6, 0.3, and 0.1, respectively, the sample is judged to belong to a healthy person.

[0116] Example 4. Performance verification of the constructed deep learning model 4-1.Performance check The performance of the DPI value output by the deep learning model constructed in Example 3 was tested. All samples were divided into Train, Validation, and Test groups, and a model was constructed using the Train samples, and then the performance of the model constructed using the Train samples was confirmed using samples in the Validation and Test groups.

[0117] [Table 3]

[0118] As a result, as shown in Table 3 and Figure 7, the accuracy was confirmed to be 91.3%, 92.7%, and 89.5% for the Train, Valid, and Test groups, respectively, and the micro AUC values, which are the results of the multi-class ROC analysis, were confirmed to be 0.991, 0.990, and 0.955 for the Train, Valid, and Test groups, respectively. Figure 7 (A) shows the performance of the CNN model confirmed by accuracy and microAUC for the Train, Validation, and Test groups, and Figure 7 (B) shows the performance of the CNN model confirmed by confusion matrix for the Train, Validation, and Test groups.

[0119] 4-2. Check the DPI distribution We confirmed how well the DPI values, which are output values ​​of the deep learning model constructed in Example 3, matched those of actual patients. The X-axis in Fig. 8 indicates the group (True label) information of the actual samples, and the Y-axis indicates, from the left, the DPI values ​​of healthy subjects (Normal), liver cancer patients (HCC), and esophageal cancer patients (EC) calculated by the CNN model.

[0120] As a result, as shown in Figure 8, it was confirmed that the DPI distribution was distributed in all of the Train, Validation, and Test datasets, with healthy subject samples having the highest probability of being healthy subjects, liver cancer patient samples having the highest probability of being liver cancer patients, and esophageal cancer patient samples having the highest probability of being esophageal cancer patients.

[0121] Although the specific parts of the present invention have been described in detail above, it is clear to those skilled in the art that these specific techniques are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents. [Industrial Applicability]

[0122] The method for cancer diagnosis and cancer type prediction using the frequency and size of terminal sequence motifs of cell-free nucleic acid fragments according to the present invention generates vectorized data and analyzes it using an AI algorithm, and therefore is useful because it shows high sensitivity and accuracy even when read coverage is low.

Claims

1. (a) extracting cell-free nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive frequencies of terminal sequence motifs of nucleic acid fragments and sizes of nucleic acid fragments; (d) generating vectorized data using the frequencies of the terminal sequence motifs of the derived nucleic acid fragments and the sizes of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the data, and comparing the output result value with a cut-off value to determine the presence or absence of cancer; and (f) A method for providing information for cancer diagnosis and cancer type prediction, comprising a step of predicting a cancer type through comparison of the output result values.

2. The method according to claim 1, characterized in that step (a) is carried out in a manner comprising the steps of: (ai) obtaining nucleic acid from blood, semen, vaginal cells, hair, saliva, urine, buccal cells, amniotic fluid containing placental or fetal cells, tissue cells or mixtures thereof; (a-ii) removing proteins, fats, and other residues from the collected nucleic acid by a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid; (a-iii) preparing a single-end sequencing or pair-end sequencing library from the purified nucleic acid or the nucleic acid randomly fragmented by enzymatic cleavage, crushing, or hydroshear method; (a-iv) subjecting the prepared library to a next-generation sequencer; and (av) A step of obtaining sequence information (reads) of the nucleic acid using a next-generation sequencer.

3. The method according to claim 1, wherein the terminal sequence motif in step (c) is a pattern of 2 to 30 base sequences at both ends of the nucleic acid fragment.

4. The method according to claim 1, wherein the frequency of the terminal sequence motif in step (c) is the number of each motif detected from the entire nucleic acid fragment.

5. The method according to claim 1, wherein the size of the nucleic acid fragment in step (c) is the number of bases from the 5' end to the 3' end of the nucleic acid fragment.

6. The method according to claim 1, wherein the vectorized data in step (d) has the type of terminal sequence motif of the nucleic acid fragment on the X-axis and the size of the nucleic acid fragment on the Y-axis.

7. The method of claim 6, wherein the vectorized data further includes a sum of the frequencies of the nucleic acid fragments by terminal motif and a sum of the frequencies of the nucleic acid fragments by size.

8. The method of claim 1 , wherein the artificial intelligence model in step (e) is trained to distinguish between vectorized data of healthy individuals and vectorized data of patients with cancer.

9. 9. The method of claim 8, wherein the artificial intelligence model is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN) and a recurrent neural network (RNN).

10. 2. The method according to claim 1, wherein the result value output by the artificial intelligence model in step (e) after analyzing the input vectorized data is a DPI (Deep Probability Index) value.

11. The method according to claim 1, wherein the reference value in step (e) is 0.5, and when the reference value is 0.5 or more, the cancer is judged to be present.

12. The method according to claim 1, characterized in that the step of predicting the type of cancer by comparing the output result values ​​in step (f) is performed by a method including a step of determining that the cancer type showing the highest output result value is cancer of the sample.

13. a decoding unit for extracting cell-free nucleic acids from the biological sample and decoding sequence information; an alignment section which aligns the decoded sequence to a standard chromosomal sequence database; a nucleic acid fragment analysis unit for deriving the frequency of terminal sequence motifs of nucleic acid fragments and the size of the nucleic acid fragments based on the aligned sequences; a data generating unit for generating vectorized data using the derived frequencies of terminal sequence motifs of nucleic acid fragments and sizes of the nucleic acid fragments; A cancer diagnosis unit that inputs the generated vectorized data into a trained artificial intelligence model for analysis and compares the data with a reference value to determine the presence or absence of cancer; and The cancer diagnosis and cancer type prediction device includes a cancer type prediction unit that analyzes the output result value and predicts the cancer type.

14. 1. A computer-readable storage medium comprising instructions configured to be executed by a processor for cancer diagnosis and predicting cancer type, (a) extracting cell-free nucleic acids from a biological sample and obtaining sequence information; (b) aligning the sequence reads to a reference genome database; (c) using the aligned sequence reads to derive frequencies of terminal sequence motifs of nucleic acid fragments and sizes of nucleic acid fragments; (d) generating vectorized data using the frequencies of the terminal sequence motifs of the derived nucleic acid fragments and the sizes of the nucleic acid fragments; (e) inputting the generated vectorized data into a trained artificial intelligence model, analyzing the data, and comparing the output result value with a cut-off value to determine the presence or absence of cancer; and (f) a computer-readable storage medium comprising instructions configured to be executed by a processor to predict the presence or absence and type of cancer through a step of predicting the type of cancer through a comparison of the output result values.

Citation Information

Patent Citations

  • Method for detecting cancer using fragment end sequence frequency and size by position of cell-free nucleic acid

    JP2024544749A

  • Cell-free DNA damage analysis and its clinical applications

    WO2020020174A1

  • Cell-free DNA end characteristics

    WO2020125709A1