Method for diagnosing and predicting cancer type based on artificial intelligence based on artificial intelligence
An AI-based cancer diagnosis method using nucleic acid sequence analysis and vectorized data processing enhances cancer detection and classification accuracy, addressing the invasiveness and accuracy issues of existing techniques.
Patent Information
- Application Number
- JP2023532698
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-27
- Filing Date
- 2021-11-15
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2041-11-15
AI Technical Summary
Current cancer diagnosis methods, including tissue biopsies and liquid biopsies, are invasive, have low accuracy, and struggle to detect early-stage cancers or differentiate between cancer types accurately.
An artificial intelligence-based method that extracts nucleic acids from a biological sample, aligns sequence information with a reference genome, generates vectorized data, and uses a learned artificial intelligence model to predict cancer types with high sensitivity and accuracy.
The method enables non-invasive, high-sensitivity, and high-accuracy cancer diagnosis and type prediction, overcoming limitations of conventional methods by leveraging deep learning models to analyze nucleic acid fragment patterns.
Smart Images

Figure 0007710042000012 
Figure 0007710042000013 
Figure 0007710042000014
Abstract
Description
Technical Field
[0001] The present invention relates to an artificial intelligence-based cancer diagnosis and cancer type prediction method. More specifically, after extracting nucleic acids from a biological sample, obtaining sequence information, generating vectorized data based on the aligned reads, and then analyzing the values calculated by inputting the data into a learned artificial intelligence model, the present invention relates to an artificial intelligence-based cancer diagnosis and cancer type prediction method using such a method.
Background Art
[0002] In clinical cancer diagnosis, it is usually confirmed by performing a tissue biopsy after a medical history investigation, physical examination, and clinical evaluation. Cancer diagnosis by clinical experiments is possible only when the number of cancer cells is 1 billion or more and the diameter of the cancer is 1 cm or more. In this case, the cancer cells already have the ability to metastasize, and at least half of them are already in a metastasized state. In addition, tissue biopsy is invasive, causes considerable discomfort to the patient, and there is often a problem that tissue biopsy cannot be performed when treating cancer patients. In addition, in cancer screening, tumor markers for monitoring substances produced directly or indirectly from cancer are used. However, even when cancer is present, more than half of the results of tumor marker screening are normal, and even when there is no cancer, they are frequently positive, so there are limitations in its accuracy.
[0003] Due to the need for a cancer diagnosis method that is relatively simple, non-invasive, and has high sensitivity and specificity to complement the problems of such conventional cancer diagnosis methods, recently, liquid biopsy that utilizes a patient's body fluid for cancer diagnosis and follow-up examinations has been widely used. Liquid biopsy is a non-invasive method and is a diagnostic technique attracting attention as an alternative to conventional invasive diagnosis and examination methods. However, there are still no large-scale research results confirming the effectiveness of liquid biopsy in cancer diagnosis methods, and there are no research results on diagnosing ambiguous cancers or differentiating ambiguous cancer types through liquid biopsy.
[0004] To mitigate the health impact of cancer, a significant amount of research effort has been dedicated to cancer diagnosis and treatment technologies. Among these, SMCT (Somatic mutation Based Cancer Typing) is one of the most important research themes. SMCT enables the determination of cancer types and subtypes based on somatic gene mutations in patients, thereby allowing for the formulation of treatment plans. In recent years, with the decrease in DNA sequencing costs, DNA sequencing data has increased rapidly, greatly promoting the development of SMCT. Different from conventional cancer typing methods based on the morphological appearance of tumors or gene expression levels (i.e., mRNA profiles or protein profiles), SMCT can distinguish tumors with similar histopathological appearances and is advantageous for providing accurate cancer subtype classification results that better reflect the cancer microenvironment (Sun, Y. et al. Sci Rep Vol. 9, 17256, 2019).
[0005] In recent years, not only SMCT, but also methods using the three-dimensional structure of chromosomes or copy number abnormalities have been reported in cancer subtype prediction (Yuan et al. BMC Genomics, Vol. 19(Suppl 6), pp. 565, 2018, 10-2019-0036494).
[0006] On the other hand, as a method for solving the problem of classifying frequently encountered input patterns in the field of engineering into specific groups, research is actively underway to apply the efficient pattern recognition methods possessed by humans to actual computers.
[0007] Among various computer application studies, there is research on artificial neural networks that engineer a model of the human brain cell structure where efficient pattern recognition occurs. To solve the problem of classifying input patterns into specific groups, artificial neural networks use an algorithm that mimics the learning ability of humans. Through this algorithm, the artificial neural network can generate a mapping between the input pattern and the output pattern, which is expressed as the artificial neural network having a learning ability. In addition, the artificial neural network has a generalization ability to generate relatively correct outputs for input patterns that were not used in learning, based on the learned results. Due to these two typical performances of learning and generalization, artificial neural networks are applied to problems that are difficult to solve with conventional sequential programming methods. Artificial neural networks have a wide range of uses and are actively applied in fields such as pattern classification problems, continuous mapping, non-linear system identification, non-linear control, and robot control.
[0008] An artificial neural network refers to a computational model implemented in software or hardware that mimics the computational ability of biological systems using a large number of artificial neurons connected by connection lines. In an artificial neural network, artificial neurons that simplify the functions of biological neurons are used. Then, they are interconnected via connection lines with connection strengths, and are capable of performing human cognitive functions and learning processes. The connection strength is a specific value that a connection line has and is also called the connection weight. The learning of an artificial neural network can be divided into supervised learning and unsupervised learning. Supervised learning is a method in which input data and the corresponding output data are input into the neural network together, and the connection strength of the connection lines is updated so that the output data corresponding to the input data is output. Typical learning algorithms include the Delta Rule and Back propagation Learning. Unsupervised learning is a method in which the artificial neural network itself learns the connection strength using only the input data without a target value. Unsupervised learning is a method of updating the connection weights based on the correlation between input patterns.
[0009] Much of the data applied in machine learning becomes complex, and as the dimensions increase, a problem known as the curse of dimensionality occurs. This is because the greater the dimensionality of the required data, the more the distance between any two points diverges to infinity, the amount of data present, that is, the density becomes somewhat lower in a high-dimensional space, and the characteristics (features) of the data cannot be appropriately reflected (Richard Bellman, Dynamic Programming, 2003, chapter 1). The recent development of deep neural networks (deep learning) has a structure with hidden layers between the input layer and the output layer, and it has been reported that while processing the linear combination of variable values transmitted from the input layer with a non-linear function, it has significantly improved the performance of classifiers in high-dimensional data such as images, videos, and signal data (Hinton, Geoffrey, et al., IEEE Signal Processing Magazine Vol. 29.6, pp. 82-97, 2012).
[0010] Although there are various patents (KR 10-2017-0185041, KR 10-2017-0144237, KR 10-2018-124550) that utilize such artificial neural networks in the bio field, there is currently a lack of research on methods for predicting cancer types through artificial neural network analysis based on sequence analysis information of cell-free DNA (cfDNA) in the blood.
[0011] Therefore, the inventors made earnest efforts to solve the above problems and develop an artificial intelligence-based diagnosis and cancer type prediction method with high sensitivity and accuracy. As a result, when generating vectorized data based on reads aligned to chromosomal regions and analyzing this with a learned artificial intelligence model, it was confirmed that cancer diagnosis and cancer types can be predicted with high sensitivity and accuracy, and the present invention was completed.
Summary of the Invention
[0012] An object of the present invention is to provide an artificial intelligence-based cancer diagnosis and cancer type prediction method.
[0013] Another object of the present invention is to provide an artificial intelligence-based cancer diagnosis and cancer type prediction apparatus.
[0014] Another object of the present invention is to provide a computer-readable storage medium including commands configured to be executed by a processor that diagnoses cancer and predicts cancer types by the above method.
[0015] To achieve the above object, the present invention provides a method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction, including: (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) generating vectorized data using nucleic acid fragments based on the aligned sequence information (reads); (d) inputting the generated vectorized data into a learned artificial intelligence model, comparing the output result value obtained by analysis with a cut-off value, and determining the presence or absence of cancer; and (e) predicting the cancer type through the comparison of the output result values.
[0016] The present invention also provides an artificial intelligence-based cancer diagnosis and cancer type prediction method including the steps of: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) generating vectorized data using nucleic acid fragments based on the aligned sequence information (reads); (d) inputting the generated vectorized data into a learned artificial intelligence model, analyzing the output result value, comparing it with a cut-off value, and determining the presence or absence of cancer; and (e) predicting the cancer type through the comparison of the output result value.
[0017] The present invention also provides an artificial intelligence-based cancer diagnosis and cancer type prediction apparatus including: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequence with a reference genome database; a data generation unit that generates vectorized data using nucleic acid fragments based on the aligned sequence; a cancer diagnosis unit that inputs the generated vectorized data into a learned artificial intelligence model, analyzes it, compares it with a reference value, and determines the presence or absence of cancer; and a cancer type prediction unit that analyzes the output result value and predicts the cancer type.
[0018] The present invention also provides a computer-readable storage medium including commands configured to be executed by a processor for diagnosing cancer and predicting cancer types, the commands including: (a) extracting nucleic acids from a biological sample to obtain sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) generating vectorized data using nucleic acid fragments based on the aligned sequence information (reads); (d) inputting the generated vectorized data into a learned artificial intelligence model for analysis, comparing with a reference value to determine the presence or absence of cancer; and (e) analyzing the output result value to predict cancer types.
Brief Description of the Drawings
[0019] FIG. 1 is an overall flowchart for determining artificial intelligence-based chromosomal abnormalities of the present invention.
[0020] FIG. 2 is an example of GCplot, an image obtained by vectorizing NGS data.
[0021] FIG. 3 is a schematic diagram showing the configuration of a CNN model constructed in an embodiment of the present invention.
[0022] (A) of FIG. 4 shows the result of confirming the accuracy of cancer presence / absence determination for a deep learning model trained with the generated GC plot image data according to an embodiment of the present invention, and (B) shows the result of showing the probability distribution for each dataset.
[0023] (A) of FIG. 5 shows the result of confirming the accuracy of cancer type prediction for a deep learning model trained with the generated GC plot image data according to an embodiment of the present invention, and (B) shows the result of showing the probability distribution for each dataset.
Modes for Carrying Out the Invention
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In general, the nomenclature used herein and the experimental methods described below are well known and commonly used in the art.
[0025] In the present invention, after aligning the sequence analysis data obtained from a sample to a reference gene, vectorized data is generated based on the aligned nucleic acid fragments, and when calculating and analyzing the DPI value with a learned artificial intelligence model, it was attempted to confirm that cancer diagnosis and cancer types can be predicted with high sensitivity and accuracy.
[0026] That is, in one embodiment of the present invention, after sequencing the DNA extracted from blood and aligning it to a reference chromosome, the distance or amount between nucleic acid fragments is calculated for each fixed chromosomal region, while each genetic region is on the X-axis and the distance or amount between nucleic acid fragments is on the Y-axis. Vectorized data is generated, and this is learned by a deep learning model to calculate the DPI value, which is compared with a reference value to perform cancer diagnosis. Among the DPI values calculated for each cancer type, a method was developed to determine the cancer type with the highest DPI value as the cancer type of the sample (Figure 1).
[0027] Therefore, from one aspect, the present invention (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) generating vectorized data using nucleic acid fragments (fragments) based on the aligned sequence information (reads); (d) inputting the generated vectorized data into a learned artificial intelligence model, analyzing the output result value, comparing it with a reference value (cut-off value), and determining the presence or absence of cancer, and (e) predicting a cancer type through comparison of the output result values relates to a method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction including the same.
[0028] In the present invention, the nucleic acid fragment can be used without limitation as long as it is a fragment of nucleic acid extracted from a biological sample, preferably including, but not limited to, a fragment of cell-free nucleic acid or intracellular nucleic acid.
[0029] In the present invention, the nucleic acid fragment is characterized in that it is obtained by directly performing sequence analysis, performing sequence analysis through next-generation sequencing analysis, or performing sequence analysis through non-specific whole genome amplification.
[0030] In the present invention, when using next-generation sequence analysis, the nucleic acid fragment means a read.
[0031] In the present invention, the cancer includes solid cancer or blood cancer, preferably selected from the group consisting of non-Hodgkin lymphoma, Hodgkin lymphoma, acute myeloid leukemia, acute lymphoid leukemia, multiple myeloma, head and neck cancer, lung cancer, glioblastoma, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, melanoma, prostate cancer, thyroid cancer, gastric cancer, gallbladder cancer, biliary tract cancer, bladder cancer, small intestine cancer, cervical cancer, cancer of unknown primary site, kidney cancer, and mesothelioma, but not limited thereto.
[0032] In the present invention, the step (a) is (a-i) The step of obtaining nucleic acids from amniotic fluid, tissue cells, and mixtures thereof, including blood, semen, vaginal cells, hair, saliva, urine, oral cells, placental cells, or fetal cells; (a-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using the salting-out method, column chromatography method, or beads method to obtain purified nucleic acids; (a-iii) The step of creating a single-end sequencing or pair-end sequencing library for the purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, disruption, or the hydroshear method; (a-iv) The step of reacting the produced library with a next-generation sequencer, and (a-v) The step of obtaining nucleic acid sequence information (reads) from the next-generation sequencer, characterized by including the above steps.
[0033] In the present invention, the next-generation sequencer can be used by any sequencing method known in the art. Sequencing of nucleic acids isolated by a selected method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of one proxy cloned relative to an individual nucleic acid molecule or in a very similar way to an individual nucleic acid molecule (e.g., 105 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species within a library can be estimated by measuring the relative number of occurrences of its cognate sequences in the data generated by the sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in the literature incorporated herein by reference (Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46).
[0034] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequence of individual nucleic acid molecules (e.g., the HeliScope Gene Sequencing system of Helicos BioSciences and the PacBio RS system of Pacific Biosciences). In other embodiments, a massively parallel short-read sequencing method that generates more bases of sequence per sequencing unit than other sequencing methods, such as other sequencing methods that generate fewer but longer reads, for example, the Solexa sequencer of Illumina Inc. in San Diego, California, determines the nucleotide sequence of a cloned proxy for an individual nucleic acid molecule (e.g., the Solexa sequencer of Illumina Inc. in San Diego, California, 454 Life Sciences (Branford, Connecticut), and Ion Torrent). Other methods or machines for next-generation sequencing include, but are not limited to, 454 Life Sciences (Branford, Connecticut), Applied Biosystems (Foster City, California, SOLiD sequencer), Helicos Biosciences Corporation (Cambridge, Massachusetts), and emulsion and microfluidic sequencing technologies nanopore (e.g., GnuBio nanopore).
[0035] Platforms for next-generation sequencing include, but are not limited to, the Roche / 454 Genome Sequencer (GS) FLX system, the Illumina / Solexa Genome Analyzer (GA), the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, the Polonator G.007 system, the Helicos BioSciences HeliScope Gene Sequencing system, and the Pacific Biosciences PacBio RS system.
[0036] NGS technology includes, for example, one or more steps of template preparation, sequencing, imaging, and data analysis.
[0037] Methods for preparing templates include steps such as randomly fragmenting nucleic acids (e.g., genomic DNA or cDNA) into small sizes and creating sequencing templates (e.g., single-fragment templates or mate-pair templates). Spatially separated templates are attached or immobilized to a solid surface or support to enable a large number of sequencing reactions to be performed simultaneously. Types of templates that can be used for NGS reactions include, for example, templates in which clones derived from a single DNA molecule are amplified and single DNA molecule templates.
[0038] Methods for preparing templates in which clones are amplified include, for example, emulsion PCR (emulsion PCR: emPCR) and solid-phase amplification.
[0039] EmPCR can be used to produce templates for NGS. Typically, a library of nucleic acid fragments is created, and adapters containing universal priming sites are ligated to the ends of the fragments. The fragments are then denatured to single strands and captured by beads. Each bead captures a single nucleic acid molecule. After amplification and enrichment of the emPCR beads, a large amount of template may be attached, fixed to a polyacrylamide gel on a standard microscope slide (e.g., Polonator), chemically cross-linked to an amino-coated glass surface (e.g., Life / APG; Polonator), or deposited onto individual picotiter plate (PicoTiterPlate: PTP) wells (e.g., Roche / 454)), at which point the NGS reaction can be performed.
[0040] Solid-phase amplification can also be used to create templates for NGS. Typically, the forward primer and the reverse primer covalently bind to a solid support. The surface density of the amplified fragments is defined as the ratio of primers to template on the support. Solid-phase amplification can generate millions of spatially separated template clusters (e.g., Illumina / Solexa). The ends of the template clusters can hybridize to universal primers for the NGS reaction.
[0041] Other methods for the production of cloned amplified templates include, for example, Multiple Displacement Amplification (MDA) (Lasken R. S. Curr Opin Microbiol:510-6). MDA is a non-PCR-based DNA amplification technique. The reaction involves annealing random hexamer primers to the template and synthesizing DNA by a high-fidelity enzyme, typically Φ29, at a constant temperature. MDA can produce products of huge size with a lower error frequency.
[0042] Template amplification methods such as PCR can bind the NGS platform to the target or enrich specific regions of the genome (e.g., exons). Representative template enrichment methods include, for example, microdroplet PCR techniques (Tewhey R. et al, Nature Biotech. 2009, 27:1025-1031), custom-designed oligonucleotide microarrays (e.g., Roche / NimbleGen oligonucleotide microarrays), and solution-based hybridization methods (e.g., molecular inversion probe (MIP) (Porreca G. J. et al, Nature Methods, 2007, 4:931-936; Krishnakumar S. et al. USA, 2008, 105:9296-9310; Turner E. H. et al., Nature Methods, 2009, 6:315-316) and biotinylated RNA capture sequences (Gnirke A. et al., Nat. Biotechnol. 2009; 27(2): 182-9).
[0043] Single-molecule templates are another type of template that can be used for NGS reactions. Spatially separated single-molecule templates can be immobilized on a solid support by various methods. In one approach, individual primer molecules covalently bind to the solid support. Adapters are added to the template, and the template is then hybridized to the immobilized primer. In another approach, single-molecule templates covalently bind to the solid support by priming and extending a single-stranded single-molecule template from an immobilized primer. Then, a universal primer is hybridized to the template. In another approach, a single polymerase molecule attaches to the solid support to which the primed template is bound.
[0044] Representative sequencing and imaging methods for NGS include, but are not limited to, cyclic reversible termination (CRT), sequencing by ligation (SBL), single molecule addition (pyrosequencing), and real-time sequencing.
[0045] CRT uses a cyclic method with reversible terminators that contain nucleotides and minimize fluorescence imaging and cleavage steps. Typically, DNA polymerase incorporates a single fluorescently modified nucleotide complementary to the template base into the primer. DNA synthesis terminates after the addition of a single nucleotide, and unincorporated nucleotides are washed away. Imaging is performed to determine the identity of the incorporated labeled nucleotide. Subsequently, in the cleavage step, the terminator / inhibitor and the fluorescent dye are removed. Representative NGS platforms that use the CRT method include, but are not limited to, the Illumina / Solexa Genome Analyzer (GA), which uses a clonally amplified template method combined with a 4-color CRT method detected by total internal reflection fluorescence (TIRF), and Helicos BioSciences / HeliScope, which uses a single molecule template method combined with a 1-color CRT method detected by TIRF. SBL uses either DNA ligase and a one-base-encoded probe or a two-base-encoded probe for sequencing.
[0046] Typically, a fluorescently labeled probe hybridizes to a complementary sequence adjacent to the primed template. DNA ligase is used to ligate a probe labeled with a dye to the primer. After unligated probes are washed away, fluorescence imaging is performed to determine the identity of the ligated probes. The fluorescent dye can be removed using a cleavable probe that regenerates the 5'-PO4 group for the next ligation cycle. Alternatively, after removing the old primer, a new primer may be hybridized to the template. Representative SBL platforms include, but are not limited to, Life / APG / SOLiD (Supported Oligonucleotide Ligation Detection), which uses two-base-encoded probes.
[0047] Pyrosequencing is based on the step of detecting the activity of DNA polymerase with different chemiluminescent enzymes. Typically, this method sequences a single strand of DNA by synthesizing the complementary strand along one base pair at a time and detecting the base that is actually added at each step. The template DNA is immobilized, and solutions of A, C, G, and T nucleotides are added sequentially and removed from the reaction. Light is generated only when the nucleotide solution fills in a base that is not paired to the template. The sequence of the solution that generates the chemiluminescent signal will determine the sequence of the template. Representative pyrosequencing platforms include, but are not limited to, Roche / 454, which uses DNA templates produced by emPCR with one to two million beads deposited in PTP wells.
[0048] Real-time sequencing involves imaging the continuous incorporation of dye-labeled nucleotides during DNA synthesis. Representative real-time sequencing platforms include, but are not limited to, individual zero-mode waveguides (ZMWs) for obtaining sequence information when phosphate-linked nucleotides are included in the growing primer strand.
[0049] The Pacific Biosciences platform that uses DNA polymerase molecules attached to the surface of a detector, the Life / VisiGen platform that uses recombinant DNA polymerase along with a fluorescent dye attached to create an enhanced signal after incorporating nucleotides by fluorescence resonance energy transfer (FRET), and the LI-COR Biosciences platform that uses dye-labeled nucleotides in a sequencing reaction are included.
[0050] Other sequencing methods for NGS include, but are not limited to, nanopore sequencing, sequencing by hybridization, nanoparticle transistor array-based sequencing, polony sequencing, scanning tunneling microscopy (STM)-based sequencing, and nanowire molecular sensor-based sequencing.
[0051] Nanopore sequencing involves the electrophoresis of nucleic acid molecules in solution through a nanoscale pore that provides a highly confined space where single nucleic acid polymers can be analyzed. Representative methods of nanopore sequencing are described, for example, in the literature [Branton D. et al., Nat Biotechnol. 2008; 26(10): 1146-53].
[0052] Sequencing by hybridization is a non-enzymatic method that uses DNA microarrays. Typically, a single pool of DNA is fluorescently labeled and hybridized to an array containing known sequences. The hybridization signal from a given spot on the array can identify the DNA sequence. In a DNA double strand, the binding of one strand of DNA to its complementary strand is sensitive to single-base mismatches if the hybrid region is short or if there are embodied mismatch-detecting proteins. Representative methods of sequencing by hybridization are described, for example, in the literature (Hanna G.J. et al. J. Clin. Microbiol. 2000; 38(7):2715-21; and Edwards J.R. et al. 2005; 573(1-2):3-12).
[0053] Polony sequencing is based on sequencing by polony amplification and multiplex single base extension (FISSEQ). Polony amplification is a method of in situ amplifying DNA on a polyacrylamide film. Representative polony sequencing methods are described, for example, in U.S. Patent Application Publication No. 2007 / 0087362.
[0054] Nanotransistor array-based devices such as carbon nanotube field effect transistors (CNTFET) can also be used in NGS. For example, DNA molecules are stretched and driven across the nanotubes by microfabricated electrodes. The DNA molecules are sequentially contacted with the surface of the carbon nanotubes, and differences in the flow of current from each base occur due to charge transfer between the DNA molecules and the nanotubes. The DNA is sequenced by recording these differences. Representative nanotransistor array-based sequencing methods are described, for example, in U.S. Patent Publication No. 2006 / 0246497.
[0055] A scanning tunneling microscope (STM) can also be used in NGS. The STM forms an image of the surface using a piezo-electrically controlled probe that raster scans the sample. The STM is used to image the physical properties of single DNA molecules, for example, by integrating an actuator-driven flexible gap and a scanning tunneling microscope, and by creating consistent electron tunneling imaging and spectroscopy. Representative sequencing methods using STM are described, for example, in U.S. Patent Application Publication No. 2007 / 0194225.
[0056] A molecular analysis device consisting of a nanowire molecular sensor can also be used in NGS. Such devices can detect the interaction between a nanowire such as DNA and a nitrogenous substance placed on a nucleic acid molecule. A molecular guide is placed to guide molecules near the molecular sensor to allow interaction and subsequent detection. Representative sequencing methods using nanowire molecular sensors are described, for example, in U.S. Patent Application Publication No. 2006 / 0275779.
[0057] A double-ended sequencing method can be used for NGS. Double-ended sequencing uses blocking and non-blocking primers to sequence both the sense and antisense strands of DNA. Typically, these methods include the steps of annealing an unblocking primer to the first strand of nucleic acid, annealing a second blocking primer to the second strand of nucleic acid, extending the nucleic acid along the first strand with polymerase, terminating the first sequencing primer, deblocking the second primer, and extending the nucleic acid along the second strand. Representative double-stranded sequencing methods are described, for example, in U.S. Patent No. 7,244,567.
[0058] In the data analysis stage, after the NGS reads are created, they are aligned or de novo assembled against a known reference sequence.
[0059] For example, confirmation of genetic variations such as single nucleotide polymorphisms and structural variants in a sample (e.g., a tumor sample) may be performed by aligning NGS reads against a reference sequence (e.g., a wild-type sequence). A sequence alignment method for NGS is described, for example, in the literature (Trapnell C. and Salzberg S.L. Nature Biotech., 2009, 27:455-457).
[0060] Examples of de novo assemblies are described, for example, in the literature (Warren R. et al., Bioinformatics, 2007, 23:500-501; Butler J. et al., Genome Res., 2008, 18:810-820, and Zerbino D.R. and Birney E., Genome Res., 2008, 18:821-829).
[0061] Sequence alignment or assembly may be performed using read data from one or more NGS platforms, for example, by mixing Roche / 454 and Illumina / Solexa read data. In the present invention, the alignment step may be performed using, but not limited to, the BWA algorithm and the hg19 sequence.
[0062] In the present invention, the sequence alignment in the step (b) is, as a computer algorithm, a computer-based method or approach that may be derived mainly by evaluating the similarity between a read sequence (e.g., from next-generation sequencing, e.g., a short read sequence) in a genome and a reference sequence, and is used for identity when there is a possibility that it may be derived from evaluating the similarity between the read sequence and the reference sequence. Various algorithms can be applied to the sequence alignment problem. Some algorithms are relatively slow but allow relatively high specificity. These include, for example, dynamic programming-based algorithms. Dynamic programming is a method of dividing a complex problem into simpler steps for solution. Other approaches are relatively efficient but generally not thorough. These include, for example, heuristic algorithms and probabilistic methods designed for large-scale database searches.
[0063] Typically, the alignment process has two stages: candidate inspection and sequence alignment. Candidate inspection reduces the search space for sequence alignment from the entire genome for a shorter enumeration of possible alignment positions. As the term suggests, sequence alignment includes the stage of aligning sequences with the sequences provided in the candidate inspection stage. This can be done using global alignment (e.g., Needleman-Wunsch alignment) or local alignment (e.g., Smith-Waterman alignment).
[0064] Most attribute alignment algorithms are characterized by one of three types based on the indexing method. Algorithms based on hash tables (e.g., BLAST, ELAND, SOAP), suffix trees (e.g., Bowtie, BWA), and merge sort (e.g., Slider). Short read arrays are typically used for alignment. Examples of sequence alignment algorithms / programs for short read arrays include, but are not limited to, BFAST (Homer N. et al., PLoS One. 2009; 4(11): e7767), BLASTN (from blast.ncbi.nlm.nih.gov on the World Wide Web), BLAT (Kent W.J. Genome Res. 2002;12(4):656-64), Bowtie (Langmead B. et al., Genome Biol. 2009;10(3):R25), BWA (Li H. and Durbin R. Bioinformatics, 2009, 25:1754-60), BWA-SW (Li H. and Durbin R. Bioinformatics, 2010;26(5):589-95), CloudBurst (Schatz M.C. Bioinformatics. 2009;25(11):1363-9), Corona Lite (Applied Biosystems, Carlsbad, California, USA), CASHX (Fahlgren N. et al., RNA, 2009; 15, 992-1002), CUDA-EC (Shi H. et al, J Comput Biol. 2010;17(4):603-15), ELAND (from bioit.dbi.udel.edu / howto / eland on the World Wide Web), GNUMAP (Clement N.L. et al. 2010;26(1):38-45), GMAP (Wu T.D. and Watanabe C.K. Bioinformatics. 2005;21(9):1859-75), GSNAP (Wu T.D. and Nacu S., Bioinformatics.2010;26(7):873-81), Geneious Assembler (Biomatters Ltd. in Auckland, New Zealand), LAST, MAQ (Li H. et al., Genome Res. 2008;18(11):1851-8), Mega-BLAST (from ncbi.nlm.nih.gov / blast / megablast.shtml on the World Wide Web), MOM (Eaves H.L. and Gao Y. Bioinformatics. 2009;25(7):969-70), MOSAIK (from bioinformatics.bc.edu / marthlab / Mosaik on the World Wide Web), Novoalign (from novocraft.com / main / index.php on the World Wide Web), PALMapper (from fml.tuebingen.mpg.de / raetsch / suppl / palmapper on the World Wide Web), PASS (Campagna D. et al, Bioinformatics. 2009;25(7):967-8), PatMaN (Prufer K. et al. 2008; 24(13):1530-1), PerM (Chen Y. et al., Bioinformatics, 2009, 25(19):2514-2521), ProbeMatch (Kim Y.J. et al. 2009;25(11):1424-5), QPalma (de Bona F. et al., Bioinformatics, 2008, 24(16):i174), RazerS (Weese D. et al., Genome Research, 2009, 19:1646-1654), RMAP (Smith A.D. et al., Bioinformatics. 2009;25(21):2841-2), SeqMap (Jiang H. et al. Bioinformatics. 2008;24:2395-2396.), Shrec (Salmela L., Bioinformatics. 2010;26(10):1284-90), SHRiMP (Rumble S.M. et al., PLoS Comput. Biol., 2009, 5(5):e1000386), SLIDER (Malhis N. et al., Bioinformatics, 2009, 25(1):6-13), SLIM Search (Muller T. et al., Bioinformatics. 2001;17 Suppl 1:S182-9), SOAP (Li R. et al., Bioinformatics.2008;24(5):713-4), SOAP2 (Li R. et al., Bioinformatics.2009;25(15):1966-7), SOCS (Ondov B.D. et al, Bioinformatics, 2008; 24(23):2776-7), SSAHA (Ning Z.et al.2001;11(10):1725-9), SSAHA2 (Ning Z. et al.et al.2001;11(10):1725-9), Stampy (Lunter G. and Goodson M. Genome Res. 2010, epub ahead of print), Taipan (from taipan.sourceforge.net on the World Wide Web), UGENE (from ugene.unipro.ru on the World Wide Web), XpressAlign (from bcgsc.ca / platform / bioinfo / software / XpressAlign on the World Wide Web), and ZOOM (Bioinformatics Solutions Inc. in Waterloo, Ontario, Canada).
[0065] The alignment algorithm is selected based on multiple factors including, for example, sequence technology, read length, number of reads, available computing resources, and sensitivity / scoring requirements. Different alignment algorithms achieve different speed levels, alignment sensitivities, and alignment specificities. Alignment specificity refers to the percentage of target sequence residues that are correctly aligned as seen in a typical submission compared to the predicted alignment. Also, alignment sensitivity refers to the percentage of target sequence residues that are correctly aligned in a typical submission as seen in the predicted alignment.
[0066] Alignment algorithms such as ELAND and SOAP can be used for the purpose of aligning short reads (e.g., Illumina / Solexa sequencers) to a reference genome when speed is the first factor to be considered. Alignment algorithms such as BLAST and Mega-BLAST can be used for the purpose of examining similarity using short reads (e.g., Roche's FLX, etc.) when specificity is the most important factor, although these methods are relatively slow. Alignment algorithms such as MAQ and Novoalign can be used for single-end or paired-end data when accuracy is important because quality scores are considered (e.g., in high-throughput SNP searches). Alignment algorithms such as Bowtie or BWA require a relatively small memory footprint because they use the Burrows-Wheeler Transform (BWT). Alignment algorithms such as BFAST, Perm, SHRiMP, SOCS, ZOOM, etc. may be used with ABI's SOLiD platform for mapping color space reads. In some applications, the results from two or more alignment algorithms can be combined.
[0067] In the present invention, the length of the sequence information (reads) in the step (b) is 5 to 5,000 bp, and the number of sequence information used is 5,000 to 5 million, but is not limited thereto.
[0068] In the present invention, the vectorized data in the step (c) can be used without limitation as long as it is vectorized data using aligned read-based nucleic acid fragments, but is preferably a Grand Canyon plot (GC plot), but is not limited thereto.
[0069] In the present invention, the vectorized data is preferably characterized by being imaged, although not limited thereto. An image basically consists of pixels, but when an image composed of pixels is vectorized, it can be expressed as a one-dimensional 2D vector (black and white), a three-dimensional 2D vector (color (RGB)), or a four-dimensional 2D vector (color (CMYK)) depending on the type of the image.
[0070] In the present invention, the vectorized data is not limited to an image. For example, multiple black and white images of n pieces can be stacked and used as input data for an artificial intelligence model using an n-dimensional 2D vector (Multi-dimensional Vector).
[0071] In the present invention, the GC plot is a plot in which a specific section (a constant bin or bins of different sizes) is placed on the X-axis and a numerical value that can be expressed by nucleic acid fragments such as the distance or number between nucleic acid fragments is generated on the Y-axis. In the present invention, the bin is 1 kb to 10 Mbp, but is not limited thereto.
[0072] In the present invention, it further includes a step of separately classifying nucleic acid fragments that satisfy the mapping quality score of the aligned nucleic acid fragments aligned before the execution of the step (c).
[0073] In the present invention, the mapping quality score varies depending on a desired standard, but is preferably 15 to 70 points, more preferably 50 to 70 points, and most preferably 60 points.
[0074] In the present invention, the GC plot in the step (c) is generated as vectorized data by calculating the number of nucleic acid fragments per chromosomal interval or the distance between nucleic acid fragments with respect to the chromosomal interval distribution of the aligned nucleic acid fragments.
[0075] In the present invention, the method for vectorizing the calculated value of the number of nucleic acid fragments or the distance between nucleic acid fragments can be used without limitation as long as the technique for vectorizing the calculated value is a known technique.
[0076] In the present invention, calculating the chromosomal interval distribution of the aligned sequence information by the number of nucleic acid fragments is characterized by performing the following steps. i) A step of dividing the chromosome into fixed intervals (bins); ii) A step of determining the number of nucleic acid fragments arranged in each interval; iii) A step of dividing the number of nucleic acid fragments determined in each interval by the total number of nucleic acid fragments of the sample and normalizing; and iv) A step of generating a GC plot with the order of each interval as the value on the X-axis and the normalized value calculated in the step iii) as the value on the Y-axis.
[0077] In the present invention, calculating the chromosomal interval distribution of the aligned sequence information by the distance between nucleic acid fragments is characterized by performing the following steps. i) A step of dividing the chromosome into fixed intervals (bins); ii) A step of calculating the distance (Fragments Distance, FD) between nucleic acid fragments aligned in each interval; iii) A step of determining a representative value (RepFD) of the distance for each interval based on the distance value calculated for each interval; iv) A step of dividing the representative value calculated in the step iii) by the representative value of the total nucleic acid fragment distance and normalizing; and v) A step of generating a GC plot with the order of each interval as the value on the X-axis and the normalized value calculated in the step iv) as the value on the Y-axis.
[0078] In the present invention, the fixed interval (bin) is characterized by being 1 Kb to 3 Gb, but is not limited thereto.
[0079] In the present invention, the step of grouping nucleic acid fragments can be additionally used. In this case, the grouping criterion can be based on the adapter sequences of the aligned nucleic acid fragments. The nucleic acid fragments aligned in the forward direction and those aligned in the reverse direction can be separately classified, and the distance between nucleic acid fragments can be calculated for the selected sequence information.
[0080] In the present invention, the FD value is defined as the distance between the i-th nucleic acid fragment and the reference value of any one or more nucleic acid fragments selected from the i + 1-th to n-th nucleic acid fragments for the obtained n nucleic acid fragments.
[0081] In the present invention, for the obtained n nucleic acid fragments, the FD value is calculated as the distance between the first nucleic acid fragment and the reference value of any one or more nucleic acid fragments selected from the group consisting of the second to n-th nucleic acid fragments, and one or more values selected from the group consisting of their sum, difference, product, average, logarithm of the product, logarithm of the sum, median, quantile, minimum value, maximum value, variance, standard deviation, median absolute deviation, and coefficient of variation and / or one or more reciprocal values of these, and the calculation result including the weighted value and statistical values not limited thereto can be used as the FD value, but is not limited thereto.
[0082] In the present invention, the description “one or more values and / or one or more reciprocal values of these” is interpreted to mean that one or a combination of two or more of the previously described numerical values can be used.
[0083] In the present invention, the “reference value of the nucleic acid fragment” is characterized by being a value obtained by adding or subtracting an arbitrary value from the median of the nucleic acid fragment.
[0084] The FD value can be defined as follows for the obtained n nucleic acid fragments. FD = Dist(Ri~Rj) (1 < i < j < n), Here, the Dist function calculates one or more values selected from the group consisting of the sum, difference, product, average, logarithm of the product, logarithm of the sum, median, quantile, minimum value, maximum value, variance, standard deviation, mean absolute deviation, and coefficient of variation of the sequence position values of all nucleic acid fragments contained between two selected nucleic acid fragments Ri and Rj, and / or one or more reciprocal values of these, the calculation result including the weighted value, and statistical values not limited thereto.
[0085] That is, in the present invention, the FD value (Fragment Distance Value) means the distance between aligned nucleic acid fragments. Here, the number of cases of selecting nucleic acid fragments for distance calculation can be defined as follows. When there are a total of N nucleic acid fragments, The number of combinations of distances between 10,170 nucleic acid fragments in JPEG0007710042000001.jpg is possible. That is, when i is 1, i + 1 is 2, and the distance from any one or more nucleic acid fragments selected from the 2nd to the nth nucleic acid fragments can be defined.
[0086] In the present invention, the FD value is characterized by calculating the distance between a specific position inside the i-th nucleic acid fragment and a specific position inside any one or more nucleic acid fragments from i + 1 to n.
[0087] For example, if the length of a certain nucleic acid fragment is 50 bp and it is aligned at position 4,183 on chromosome 1, the genetic position values that can be used for distance calculation of this nucleic acid fragment are 4,183 to 4,232 on chromosome 1.
[0088] When the adjacent 50-bp-long nucleic acid fragment to the said nucleic acid fragment is aligned at the 4,232nd position on chromosome 1, the genetic position values that can be used for distance calculation of this nucleic acid fragment are 4,232 to 4,281 on chromosome 1, and the FD value between the two nucleic acid fragments is from 1 to 99.
[0089] When another adjacent nucleic acid fragment with a length of 50 bp is aligned at the 4123rd position of chromosome 1, the genetic position values that can be used for calculating the distance of this nucleic acid fragment are 4,123 to 4,172 on chromosome 1, the FD value between the two nucleic acid fragments is 61 to 159, and the FD value with the first exemplary nucleic acid fragment is 12 to 110. The calculation results including one or more values selected from the group consisting of the sum, difference, product, average, logarithm of the product, logarithm of the sum, median, quantile, minimum value, maximum value, variance, standard deviation, mean absolute deviation, and coefficient of variation of any value in either of the two FD value ranges and / or one or more reciprocal values of these, and weighted values, and statistical values not limited thereto can be used as the FD value. Preferably, it is the reciprocal value of one value in the two FD value ranges, but is not limited thereto.
[0090] Preferably, in the present invention, the FD value is characterized in that it is a value obtained by adding or subtracting an arbitrary value from the median of the nucleic acid fragment.
[0091] In the present invention, the median of FD means the value located at the most central position when the calculated FD values are arranged in ascending order of magnitude. For example, when there are three values such as 1, 2, and 100, 2 is at the most central position, so 2 becomes the median. If there are an even number of FD values, the median is determined by the average of the two central values. For example, when there are FD values of 1, 10, 90, and 200, the median is 50, which is the average of 10 and 90.
[0092] In the present invention, the arbitrary value can be used without limitation as long as it can indicate the position of the nucleic acid fragment. Preferably, it is 0 to 5 kbp or 0 to 300% of the nucleic acid fragment length, 0 to 3 kbp or 0 to 200% of the nucleic acid fragment length, 0 to 1 kbp or 0 to 100% of the nucleic acid fragment length, and more preferably 0 to 500 bp or 0 to 50% of the nucleic acid fragment length, but is not limited thereto.
[0093] In the present invention, in the case of paired-end sequencing, the FD value is derived based on the position values of the forward and reverse sequence information (reads).
[0094] For example, in the case of a paired-end read pair with a length of 50 bp, if the forward read is aligned at the 4183rd position on chromosome 1 and the reverse read is aligned at the 4349th position, the two ends of this nucleic acid fragment are 4183 and 4349, and the reference value that can be used for the nucleic acid fragment distance is 4183 - 4349. At this time, for another paired-end read pair adjacent to this nucleic acid fragment, if the forward read is aligned at the 4349th position on chromosome 1 and the reverse read is aligned at the 4515th position, the position value of this nucleic acid fragment is 4349 - 4515. The distance between these two nucleic acid fragments is 0 - 333, and most preferably, it is 166, which is the distance of the median value of each nucleic acid fragment.
[0095] In the present invention, when obtaining sequence information by the paired-end sequence, in the case of a nucleic acid fragment where the number of alignment points of the sequence information (reads) is less than the reference value, it further includes the step of excluding it in the calculation process.
[0096] In the present invention, in the case of single-end sequencing, the FD value is derived based on one of the position values of the forward or reverse sequence information (read).
[0097] In the present invention, in the case of the single-end sequencing, when deriving a position value based on the array information aligned in the forward direction, an arbitrary value is added, and when deriving a position value based on the array information aligned in the reverse direction, an arbitrary value is subtracted. The arbitrary value can be used without limitation as long as the FD value clearly indicates the position of the nucleic acid fragment. Preferably, it is 0 to 5 kbp or 0 to 300% of the nucleic acid fragment length, 0 to 3 kbp or 0 to 200% of the nucleic acid fragment length, 0 to 1 kbp or 0 to 100% of the nucleic acid fragment length, and more preferably 0 to 500 bp or 0 to 50% of the nucleic acid fragment length, but is not limited thereto.
[0098] In the present invention, the nucleic acid to be analyzed is sequenced and expressed in units of reads. These reads can be divided into single end sequencing reads (SE) and paired end sequencing reads (PE) according to the sequencing method. The SE mode read means that either the 5' or 3' of the nucleic acid molecule is sequenced in a random direction for a certain length, and the PE mode read means that both the 5' and 3' are sequenced for a certain length. Due to such differences, it is a well-known fact to ordinary technicians that when sequencing in the SE mode, one read is generated from one nucleic acid fragment, and in the PE mode, two reads are generated in pairs from one nucleic acid fragment.
[0099] The most ideal method for calculating the exact distance between nucleic acid fragments is to sequence the nucleic acid molecule from start to end, align its reads, and utilize the median (center) of the aligned values. However, technically, the above method is currently restricted in terms of the limitations of sequencing technology and cost. Therefore, sequencing will be performed using methods such as SE or PE. In the case of the PE method, since the start and end positions of the nucleic acid molecule can be known, the exact position (median) of the nucleic acid fragment can be grasped through combinations of these values. In the case of the SE method, however, since only information about one end of the nucleic acid fragment can be utilized, there are limitations in calculating the exact position (median).
[0100] Also, when calculating the distance of a nucleic acid molecule using the end information of all reads sequenced (aligned) in both the forward and reverse directions, inaccurate values may be calculated due to the element of the sequencing direction.
[0101] Therefore, due to technical reasons of the sequencing method, the 5` end of the forward read will have a position value smaller than the central position of the nucleic acid molecule, and the 3` end of the reverse read will have a larger value. Utilizing such characteristics, in the case of the forward read, by adding an arbitrary value (Extended bp), and for the reverse read, subtracting it, a value close to the central position of the nucleic acid molecule can be estimated.
[0102] That is, the arbitrary value (Extended bp) may vary depending on the sample used. In the case of cell-free nucleic acids, since the average length of the nucleic acid is said to be about 166 bp, it can be set to about 80 bp. If the experiment is performed through a fragmentation device (for example, sonication), about half of the target length set in the fragmentation process can be set as the extended bp.
[0103] In the present invention, the representative value (RepFD) is one or more values selected from the group consisting of the sum, difference, product, average, median, quantile, minimum value, maximum value, variance, standard deviation, mean absolute deviation, and coefficient of variation of FD values and / or one or more reciprocals thereof, and is preferably the median, average value of FD values, or their reciprocals, but is not limited thereto.
[0104] In the present invention, the vectorized data is characterized in that a plurality of chromosome-specific plots are included in one image.
[0105] In the present invention, the artificial intelligence model in the step (d) can be used without limitation as long as it is a model that can be trained to distinguish images by cancer type, and is preferably a deep learning model.
[0106] In the present invention, the artificial intelligence model can be used without limitation as long as it is an artificial neural network algorithm that can analyze vectorized data based on an artificial neural network. Preferably, it is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and an autoencoder, but is not limited thereto.
[0107] In the present invention, the recurrent neural network is characterized by being selected from the group consisting of a long short-term memory (LSTM) neural network, a gated recurrent unit (GRU) neural network, a vanilla recurrent neural network, and an attentive recurrent neural network.
[0108] In the present invention, when the artificial intelligence model is a CNN, the loss function for performing binary classification is characterized by being represented by the following formula 1, and the loss function for performing multi-class classification is characterized by being represented by the following formula 2. Formula 1: Binary classification
Number
Number
[0109] In the present invention, the binary classification means that the artificial intelligence model is trained to determine the presence or absence of cancer, and the multi-class classification means that the artificial intelligence model is trained to determine two or more types of cancer.
[0110] In the present invention, when the artificial intelligence model is a CNN, the learning is characterized by including the following steps. i) Classifying the produced GC plot into training, validation, and test data. At this time, the training data is used when learning the CNN model, the validation data is used for hyper-parameter tuning verification, and the test data is used for performance evaluation after the optimal model is produced. ii) Constructing an optimal CNN model through hyper-parameter tuning and the learning process. (iii) A step of comparing the performances of a plurality of models obtained through hyperparameter tuning by using validation data, and determining the model with the best performance on the validation data as the optimal model.
[0111] In the present invention, the hyperparameter tuning process is a process of optimizing the values of a plurality of parameters (the number of convolution layers, the number of dense layers, the number of convolution filters, etc.) constituting the CNN model. As the hyperparameter tuning process, Bayesian optimization and grid search techniques are used.
[0112] In the present invention, the learning process optimizes the internal parameters (weights) of the CNN model by using the defined hyperparameters. When the validation loss starts to increase with respect to the training loss, it is determined that the model is overfitting, and the model learning is interrupted before that.
[0113] In the present invention, the result value analyzed from the vectorized data input to the artificial intelligence model in the step d) can be used without limitation if it is a specific score or a real number, preferably a DPI (Deep Probability Index) value, but is not limited thereto.
[0114] In the present invention, the Deep probability Index means a value obtained by adjusting the output of the artificial intelligence on a 0-1 scale using the sigmoid function for binary classification and the softmax function for multi-class classification for the last layer of the artificial intelligence model and expressing it as a probability value.
[0115] In the case of binary classification, the sigmoid function is used for training such that the DPI value becomes 1 in the case of cancer. For example, when breast cancer samples and normal samples are input, training is performed so that the DPI value of the breast cancer samples approaches 1.
[0116] In the case of multi-class classification, the softmax function is used to extract DPI values for the number of classes. The sum of the DPI values for the number of classes is 1, and training is performed such that the DPI value of the actual corresponding cancer type becomes 1. For example, if there are three classes of breast cancer, liver cancer, and normal, and a breast cancer sample is input, training is performed so that the breast cancer class approaches 1.
[0117] In the present invention, it is characterized in that the output result value in the step (d) is derived for each cancer type.
[0118] In the present invention, when the artificial intelligence model is being trained, if there is cancer, the output result is trained to be close to 1, and if there is no cancer, the output result is trained to be close to 0. Based on 0.5, if it is 0.5 or more, it is determined that there is cancer, and if it is less than 0.5, it is determined that there is no cancer, and performance measurement is performed (Training, validation, test accuracy). Here, it is obvious to ordinary technicians that the reference value of 0.5 can be changed at any time. For example, if one wants to reduce false positives, a reference value higher than 0.5 can be set, and the criterion for determining the presence of cancer can be made stricter. If one wants to reduce false negatives, the reference value can be measured lower, and the criterion for determining the presence of cancer can be made looser.
[0119] Most preferably, the learned artificial intelligence model can be used to apply unseen data (data for which the answer is known but has not been used for training) to confirm the probability of the DPI value and determine the reference value.
[0120] In the present invention, the step of predicting the cancer type by comparing the output result values in the step (e) is performed by a method including a step of determining, as the cancer of the sample, the cancer type showing the highest value among the output result values.
[0121] From another aspect, the present invention relates to an artificial intelligence-based cancer diagnosis and cancer type prediction apparatus including a decoding unit that extracts nucleic acid from a biological sample and decodes sequence information, an alignment unit that aligns the decoded sequence with a standard chromosomal sequence database, a data generation unit that generates vectorized data using the aligned sequence-based nucleic acid fragments, and a cancer diagnosis unit that inputs the generated vectorized data into a learned artificial intelligence model for analysis, compares it with a reference value, and determines the presence or absence of cancer, and a cancer type prediction unit that analyzes the output result value to predict the cancer type.
[0122] In the present invention, the decoding unit includes a nucleic acid injection unit that injects nucleic acid extracted from an independent device, and a sequence information analysis unit that analyzes the sequence information of the injected nucleic acid, and is preferably an NGS analyzer, but is not limited thereto.
[0123] In the present invention, the decoding unit is characterized by receiving and decoding sequence information data generated by an independent device.
[0124] In the present invention, the vectorized data of the data generation unit is a Grand Canyon plot (GC plot).
[0125] The GC plot in the present invention is a plot in which a specific section (constant bin or bins with different sizes) is placed on the X-axis, and a numerical value that can be expressed by nucleic acid fragments such as the distance or number between nucleic acid fragments is generated on the Y-axis. In the present invention, the bin is 1 kb to 10 Mbp, but is not limited thereto.
[0126] In the present invention, the data generation unit further includes a nucleic acid fragment classification unit that separately classifies nucleic acid fragments that satisfy the mapping quality score of the aligned nucleic acid fragments before generating the vectorized data.
[0127] In the present invention, the mapping quality score varies depending on a desired standard, but is preferably 15 to 70 points, more preferably 50 to 70 points, and most preferably 60 points.
[0128] In the present invention, the GC plot of the data generation unit is generated with vectorized data by calculating the number of nucleic acid fragments per interval or the distance between nucleic acid fragments with respect to the chromosomal interval distribution of the aligned nucleic acid fragments.
[0129] In the method of vectorizing the number of nucleic acid fragments or the calculated value of the distance between nucleic acid fragments in the present invention, vectorizing the calculated value can be used without limitation as long as it is a known technique.
[0130] In the present invention, calculating the chromosomal interval distribution of the aligned sequence information by the number of nucleic acid fragments is characterized by including the following steps. i) The step of dividing the chromosome into fixed intervals (bins), ii) The step of determining the number of nucleic acid fragments arranged in each interval, iii) The step of dividing the number of nucleic acid fragments determined in each interval by the total number of nucleic acid fragments in the sample and normalizing, and iv) The step of generating a GC plot with the order of each interval as the value on the X-axis and the normalized value calculated in the step iii) as the value on the Y-axis.
[0131] In the present invention, calculating the chromosomal interval distribution of the aligned sequence information by the distance between nucleic acid fragments is characterized by including the following steps. i) The step of dividing the chromosome into fixed intervals (bins), ii) calculating the distance between nucleic acid fragments (Fragments Distance, FD) arranged in each interval; iii) determining a representative value (RepFD) of the distance for each interval based on the distance values calculated for each interval; iv) normalizing by dividing the representative value calculated in step iii) by the representative value of the total nucleic acid fragment distance, and v) generating a GC plot with the order of each interval as the value on the X-axis and the normalized value calculated in step iv) as the value on the Y-axis.
[0132] In the present invention, the fixed interval (bin) is characterized by being 1 Kb to 3 Gb, but is not limited thereto.
[0133] In the present invention, the step of grouping nucleic acid fragments can be additionally used. In this case, the grouping criterion can be based on the adapter sequences of the aligned nucleic acid fragments. The distance between nucleic acid fragments can be calculated for the sequence information separately selected and sorted for the nucleic acid fragments aligned in the forward direction and the nucleic acid fragments aligned in the reverse direction.
[0134] In the present invention, the FD value is defined as the distance between the i-th nucleic acid fragment and the reference value of any one or more nucleic acid fragments selected from the i+1-th to n-th nucleic acid fragments for the obtained n nucleic acid fragments.
[0135] In the present invention, for the obtained n nucleic acid fragments, the FD value is calculated as the distance between the first nucleic acid fragment and the reference value of any one or more nucleic acid fragments selected from the group consisting of the second to n-th nucleic acid fragments, and one or more values selected from the group consisting of their sum, difference, product, average, logarithm of the product, logarithm of the sum, median, quantile, minimum value, maximum value, variance, standard deviation, median absolute deviation, and coefficient of variation and / or one or more reciprocal values thereof, and calculation results including weights and statistical values not limited thereto can be used as the FD value, but are not limited thereto.
[0136] In the present invention, the description "one or more values and / or one or more reciprocal values thereof" is interpreted to mean that one or two or more of the previously described numerical values can be used in combination.
[0137] In the present invention, the "reference value of the nucleic acid fragment" is characterized in that it is a value obtained by adding or subtracting an arbitrary value from the median value of the nucleic acid fragment.
[0138] In the present invention, the artificial intelligence model of the cancer diagnosis unit can be used without limitation as long as it is a model that can be trained to distinguish images by cancer type, and is preferably a deep learning model.
[0139] In the present invention, the artificial intelligence model can be used without limitation as long as it is an artificial neural network algorithm that can analyze data vectorized based on an artificial neural network. Preferably, it is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and an autoencoder, but is not limited thereto.
[0140] In the present invention, the recurrent neural network is characterized in that it is selected from the group consisting of an LSTM (Long-short term memory) neural network, a GRU (Gated Recurrent Unit) neural network, a vanilla recurrent neural network, and an attentive recurrent neural network.
[0141] In the present invention, when the artificial intelligence model is a CNN, the loss function for binary classification is characterized by being represented by the following formula 1, and the loss function for multi-class classification is characterized by being represented by the following formula 2. Formula 1: Binary classification
Number
Number
[0142] In the present invention, the binary classification means that the artificial intelligence model is trained to determine the presence or absence of cancer, and the multi-class classification means that the artificial intelligence model is trained to determine two or more types of cancer.
[0143] In the present invention, when the artificial intelligence model is a CNN, the learning is characterized by including the following steps. i) Classifying the produced GC plot into training, validation, and test data. At this time, the training data is used when learning the CNN model, the validation data is used for hyper-parameter tuning verification, and the test data is used for performance evaluation after producing the optimal model. ii) Constructing an optimal CNN model through hyper-parameter tuning and the learning process. iii) Comparing the performance of multiple models obtained through hyper-parameter tuning using the validation data, and determining the model with the best performance on the validation data as the optimal model.
[0144] In the present invention, the Hyper-parameter tuning process is a process of optimizing the values of a plurality of parameters (the number of convolution layers, the number of dense layers, the number of convolution filters, etc.) that make up the CNN model. As the Hyper-parameter tuning process, it is characterized by using Bayesian optimization and grid search techniques.
[0145] In the present invention, the learning process optimizes the internal parameters (weights) of the CNN model using the defined hyper-parameters. When the validation loss starts to increase with respect to the training loss, it is determined that the model is overfitting, and the model learning is interrupted before that.
[0146] In the present invention, in the cancer diagnosis unit, the result value analyzed from the vectorized data input to the artificial intelligence model can be used without limitation as long as it is a specific score or real number, preferably the DPI (Deep Probability Index) value, but is not limited thereto.
[0147] In the present invention, the Deep probability Index means a value obtained by adjusting the output of the artificial intelligence on a 0-1 scale using the sigmoid function in the case of binary classification and the softmax function in the case of multi-class classification in the last layer of the artificial intelligence model and expressing it as a probability value.
[0148] In the case of binary classification, the sigmoid function is used for learning so that the DPI value becomes 1 in the case of cancer. For example, when breast cancer samples and normal samples are input, learning is performed so that the DPI value of the breast cancer samples approaches 1.
[0149] In the case of multi-class classification, the softmax function is used to extract DPI values for the number of classes. The sum of the DPI values for the number of classes becomes 1, and learning is performed so that the DPI value of the actually corresponding cancer type becomes 1. For example, if there are three classes: breast cancer, liver cancer, and normal, when a breast cancer sample is input, learning is performed to make the breast cancer class approach 1.
[0150] In the present invention, it is characterized in that the output result value of the cancer diagnosis unit is derived for each cancer type.
[0151] In the present invention, when the artificial intelligence model is being trained, if there is cancer, it is trained so that the output result is close to 1, and if there is no cancer, it is trained so that the output result is close to 0. Based on 0.5, if it is 0.5 or more, it is determined that there is cancer, and if it is less than 0.5, it is determined that there is no cancer, and performance measurement is performed (Training, validation, test accuracy). Here, it is obvious to ordinary technicians that the reference value of 0.5 can be changed at any time. For example, if you want to reduce False positive, you can set a reference value higher than 0.5 and take a strict criterion for determining that there is cancer. If you want to reduce False Negative, you can measure the reference value lower and take a lenient criterion for determining that there is cancer.
[0152] Most preferably, the learned artificial intelligence model can be used to apply unseen data (data for which the answers are known but not used in training) to confirm the probability of DPI values and determine the reference value.
[0153] In the present invention, the cancer type prediction unit predicts the cancer type through comparison of output result values, and is characterized in that it is performed by a method including a step of determining the cancer type showing the highest value among the output result values as the cancer of the sample.
[0154] From another perspective, the present invention relates to a computer-readable storage medium including commands configured to be executed by a processor that diagnoses cancer and predicts cancer types, (a) extracting nucleic acids from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) generating vectorized data using nucleic acid fragments based on the aligned sequence information (reads); (d) inputting the generated vectorized data into a learned artificial intelligence model for analysis, comparing with a reference value to determine the presence or absence of cancer; and (e) analyzing the output result value to predict cancer types, and relates to a computer-readable storage medium including commands configured to be executed by a processor that predicts the presence or absence of cancer and cancer types.
[0155] In the present invention, the step (a) is characterized by obtaining pre-generated sequence information, and at this time, the pre-generated sequence information is generated by extracting nucleic acids from a biological sample using an NGS device or the like.
[0156] In another aspect, the method according to the present invention can be implemented using a computer. In one embodiment, the computer includes one or more processors connected to a chipset. Also, connected to the chipset are a memory, a storage device, a keyboard, a Graphics Adapter, a Pointing Device, and a Network Adapter, etc. In one embodiment, the performance of the chipset is realized by a Memory Controller Hub and an I / O Controller Hub. In other embodiments, the memory can be directly connected to the processor instead of the chipset for use. The storage device is any device capable of holding data, including a hard drive, a CD-ROM (Compact Disk Read-Only Memory), a DVD, or other memory devices. The memory is involved in the data and commands used by the processor. The Pointing Device is a mouse, a Track Ball, or other types of pointing devices, and is used in combination with the keyboard to send input data to the computer system. The Graphics Adapter displays images and other information on the display. The Network Adapter connects the computer system to a short-distance or long-distance communication network. However, the computer used in the present invention is not limited to the above configuration, may have some components missing or include additional components, and may also be part of a Storage Area Network (SAN). The computer of the present invention can be configured to be compatible with the execution of the modules in the program for implementing the method according to the present invention.
[0157] In the present invention, a module means a functional and structural combination of hardware for implementing the technical idea of the present invention and software for driving the hardware. For example, the module means a logical unit of a predetermined code and hardware resources for executing the predetermined code, and it is obvious to those skilled in the technical field of the present invention that it does not necessarily mean physically connected code or a certain type of hardware.
[0158] The method according to the present invention can be implemented by hardware, firmware, software, or a combination thereof. When implemented by software, the storage medium includes any medium that can be stored or transmitted in a form readable by a device such as a computer. For example, computer-readable media include ROM (Read Only Memory), RAM (Random Access Memory), magnetic disk storage media, optical storage media, flash memory devices, and other electrical, optical, or acoustic signal transmission media.
[0159] From this perspective, the present invention also relates to a computer-readable medium including an execution module for causing a processor to execute operations including the steps according to the present invention described above.
Examples
[0160] Hereinafter, the present invention will be described in more detail with reference to examples. It will be obvious to those skilled in the art that these examples are for illustrative purposes of the present invention and should not be construed as limiting the scope of the present invention by these examples.
[0161] Example 1. Extract DNA from blood and perform next-generation sequencing analysis Blood samples of 184 normal individuals and 580 cancer patients were collected at 10 mL each and stored in EDTA tubes. Within 2 hours after collection, only the plasma was centrifuged once at 1200 g, 4 °C for 15 minutes, and the plasma obtained from the first centrifugation was centrifuged a second time at 16000 g, 4 °C for 10 minutes to separate the supernatant plasma layer excluding the precipitate. Cell-free DNA was extracted from the separated plasma using the Tiangen micro DNA kit (Tiangen), and after performing the library preparation process using the MGIEasy cell-free DNA library prep set kit, sequencing was performed on the DNBseq G400 equipment (MGI) in 100 base Paired end mode. As a result, it was confirmed that approximately 170 million reads were generated per sample.
[0162] Example 2. Generate a GC plot based on nucleic acid fragment distance Using the NGS data generated in Example 1 above, GCplot was generated (vectorized). The hg19 reference chromosome was divided based on a bin size of 100 k bases, and the generated NGS reads were assigned to each bin. Then, the reciprocal value of the median of the FD (Fragment Distance) values was calculated for each bin, and an image was generated with the X-axis representing the position of each bin and the Y-axis representing the reciprocal value of the median of the FD values calculated earlier (Figure 2).
[0163] Example 3. Construction and learning process of the CNN model The basic structure of the CNN model is as shown in Figure 3. The ReLU (Rectified Linear Unit) activation function was used, and each convolution layer used 20 10*10 patches. The max pooling method was utilized with 2x2 patches. Five fully connected layers were used, and each layer contained 175 hidden nodes. Finally, the sigmoid function value was used to calculate the final DPI value. The hyperparameter values used in the CNN model were obtained by the Bayesian Optimization method, and the model structure may change depending on the data used and the optimization of the model.
[0164] Example 4. Construction and performance verification of a cancer diagnosis deep learning model using a GC plot based on nucleic acid fragment distance The performance of the DPI value output by the deep learning model constructed using the GC plot based on the distance between nucleic acid fragments was tested using the leads obtained in Example 1. All samples were divided into Train, Validation, and Test groups and processed. After constructing the model using the Train samples, the performance of the model created using the Train samples was confirmed using the samples in the Validation and Test groups.
[0165]
Table 1
[0166]
Table 2
[0167] As a result, as shown in Table 2 and Figure 4, the Accuracy was confirmed to be 100%, 88.7%, and 90% in the Train, Valid, and Test groups respectively, and the AUC value, which is the result of the ROC analysis, was 1.00, 0.95, and 0.938 in the Train, Valid, and Test groups respectively.
[0168] In (A) of FIG. 4, among the methods for measuring accuracy, in the analysis using the ROC (Receiver Operating Characteristic) curve, it is interpreted that the higher the AUC (Area Under the Curve) value, which is the area under the curve, the higher the accuracy. The AUC value has a value between 0 and 1. When randomly predicting the label value (baseline), the expected AUC value is 0.5, and when predicting completely accurately, the expected AUC value is 1.
[0169] In (B) of FIG. 4, the probability value (DPI value) of chromosomal aneuploidy calculated by the artificial intelligence model of the present invention is shown by a boxplot for the normal sample and the cancer patient sample group, and the red line indicates 0.5 which is the DPI cutoff.
[0170] Example 5. Construction and performance verification of a cancer diagnosis deep learning model using a GC plot based on the number of nucleic acid fragments Using the leads obtained in Example 1, the performance of the DPI value output by the deep learning model was tested using a GC plot based on the number of nucleic acid fragments. All samples were divided into Train, Validation, and Test groups and proceeded. After constructing a model using the Train samples, the performance of the model created using the Train samples was confirmed using the samples of the Validation group and the Test group.
[0171] [Table 3]
[0172] [Table 4]
[0173] As a result, as described in Table 4 and Figure 5, Accuracy was confirmed to be 100%, 91%, and 86.8% in the Train, Valid, and Test groups, respectively, and the AUC value, which is the result of the ROC analysis, was confirmed to be 1.00, 0.968, and 0.936 in the Train, Valid, and Test groups, respectively.
[0174] (A) in Figure 5 is an analysis using the ROC (Receiver Operating Characteristic) curve among the methods for measuring accuracy, and it is interpreted that the higher the AUC (Area Under the Curve) value, which is the area under the curve, the higher the accuracy. The AUC value has a value between 0 and 1. When randomly predicting the label value (baseline), the expected AUC value is 0.5, and when predicting completely accurately, the expected AUC value is 1.
[0175] (B) in Figure 5 shows, in a boxplot, the probability values (DPI values) of chromosomal aneuploidy calculated by the artificial intelligence model of the present invention for normal sample and cancer patient sample groups, and the red line indicates 0.5, which is the DPI cutoff.
[0176] As described above, specific parts of the content of the present invention have been described in detail. However, it will be apparent to those skilled in the art that these specific techniques are merely preferred embodiments and do not limit the scope of the present invention. Therefore, it can be said that the substantial scope of the present invention is defined by the appended claims and their equivalents.
Industrial Applicability
[0177] The artificial intelligence-based cancer diagnosis and cancer type prediction method according to the present invention is useful because, compared to using the method of determining the chromosomal amount based on the conventional read count and using each value related to the read as a standardized value, it generates vectorized data and analyzes it using an AI algorithm, so it can exhibit the same effect even with low read coverage.
Claims
1. (a) A step of extracting nucleic acid from a biological sample and obtaining sequence information; (b) A step of aligning the obtained sequence information (reads) with a reference genome database; (c) A step of generating vectorized data using nucleic acid fragments (fragments) based on the aligned sequence information (reads); (d) A step of comparing the output result value analyzed by inputting the generated vectorized data into a learned artificial intelligence model with a cut-off value to determine the presence or absence of cancer, and (e) A method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction, including a step of predicting a cancer type through comparison of the output result value, wherein the vectorized data in step (c) is a Grand Canyon plot (GC plot), the GC plot is generated by calculating the distance between nucleic acid fragments for the distribution of aligned nucleic acid fragments by chromosomal interval and vectorizing the data, and the step of predicting a cancer type through comparison of the output result value in step (e) is performed by a method including a step of determining the cancer type showing the highest value among the output result values as the cancer of the sample. Method.
2. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to Claim 1, wherein the step (a) is performed by a method including the following steps: (a-i) A step of obtaining nucleic acid from blood, semen, vaginal cells, hair, saliva, urine, oral cells, amniotic fluid containing placental cells or fetal cells, tissue cells, and mixtures thereof; (a-ii) A step of removing proteins, fats, and other residues from the collected nucleic acid using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid. (a-iii) For the purified nucleic acid or the nucleic acid randomly fragmented by enzymatic cleavage, disruption, or the hydroshear method, the step of creating a single-end sequencing or pair-end sequencing library; (a-iv) The step of reacting the produced library with a next-generation sequencer, and (a-v) The step of obtaining sequence information (reads) of the nucleic acid with the next-generation sequencer.
3. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 1, wherein calculating the distribution by chromosomal interval based on the distance between nucleic acid fragments is performed including the following steps: i) The step of dividing a chromosome into certain intervals (bins); ii) The step of calculating the distance between nucleic acid fragments arranged in each interval; iii) The step of determining a representative value of the distance (RepFD) for each interval based on the distance values calculated for each interval; iv) The step of normalizing by dividing the representative value calculated in step (iii) by the representative value of the distance between all nucleic acid fragments; and v) The step of generating a GC plot with the order of each interval as the value on the X-axis and the normalized value calculated in step (iv) as the value on the Y-axis.
4. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 3, wherein the representative value is one or more selected from the group consisting of the sum, difference, product, average, median, quantile, minimum value, maximum value, variance, standard deviation, median absolute deviation, coefficient of variation, the reciprocal values thereof, and combinations thereof.
5. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 1, wherein the artificial intelligence model in step (d) is trained to be able to distinguish between vectorized data with normal chromosomal states and vectorized data with chromosomal abnormalities.
6. The artificial intelligence model is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and an autoencoder, and the method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 5 is characterized in that.
7. When the artificial intelligence model is a CNN and learns binary classification, the loss function is represented by Equation 1 below. When the artificial intelligence model is a CNN and learns multi-class classification, the loss function is represented by Equation 2 below. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 6 is characterized in that. Equation 1: 【Number 1】 Equation 2: 【Number 2】
8. The result value output by analyzing the vectorized data input to the artificial intelligence model in step (d) is a DPI (Deep Probability Index) value. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 1 is characterized in that.
9. The reference value in step (d) is 0.
5. When it is 0.5 or more, it is determined that the subject has cancer. The method for providing information for artificial intelligence-based cancer diagnosis and cancer type prediction according to claim 1 is characterized in that.
10. A decoding unit that extracts nucleic acids from a biological sample and decodes sequence information, An alignment unit that aligns the decoded sequences with a standard chromosomal sequence database, and A data generation unit that generates vectorized data using the aligned sequence-based nucleic acid fragments, A cancer diagnosis unit that inputs the generated vectorized data into a learned artificial intelligence model for analysis, compares it with a reference value, and determines the presence or absence of cancer, and A cancer type prediction unit that analyzes the output result value to predict the cancer type, The vectorized data generated by the data generation unit is a Grand Canyon plot (GC plot), The GC plot is generated from vectorized data obtained by calculating the distance between nucleic acid fragments with respect to the distribution of aligned nucleic acid fragments by chromosomal region, and the cancer type prediction unit predicts the cancer type through comparison of the output result values, and determines the cancer type indicated by the highest value among the output result values as the cancer of the sample. An artificial intelligence-based cancer diagnosis and cancer type prediction device.
11. A computer-readable storage medium including commands configured to be executed by a processor that diagnoses cancer and predicts the cancer type, (a) extracting nucleic acid from a biological sample and obtaining sequence information; (b) aligning the obtained sequence information (reads) with a reference genome database; (c) generating vectorized data using nucleic acid fragments based on the aligned sequence information (reads); (d) inputting the generated vectorized data into a learned artificial intelligence model for analysis, comparing with a reference value, and determining the presence or absence of cancer; and (e) analyzing the output result value to predict the cancer type including commands configured to be executed by a processor that predicts the presence or absence of cancer and the cancer type through the above steps, the vectorized data in step (c) is a Grand Canyon plot (GC plot), the GC plot is generated from vectorized data obtained by calculating the distance between nucleic acid fragments with respect to the distribution of aligned nucleic acid fragments by chromosomal region, and the step of predicting the cancer type through comparison of the output result values in step (e) is performed by a method including the step of determining the cancer type indicated by the highest value among the output result values as the cancer of the sample. A computer-readable storage medium.
Citation Information
Patent Citations
Prediction methods for gastrointestinal and pancreatic neuroendocrine neoplasms (GEP-NENs)
JP2014512172A
Artificial intelligence-based chromosomal abnormality detection method
JP2023504139A
Model-based featurization and classification
WO2020232109A1