Method for improving cancer patient detection accuracy by correcting experimental batch effects in cell-free whole genome sequencing fragmentomic data

By deriving terminal sequence motif frequencies and sizes from nucleic acid fragments and correcting batch effects, the method enhances cancer diagnosis and type prediction using artificial intelligence, addressing the limitations of existing techniques.

WO2025174151A1PCT designated stage Publication Date: 2025-08-21GREEN CROSS GENOME CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/002274
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-02-17
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing cancer diagnostic methods, particularly those using cell-free nucleic acid sequencing, suffer from low accuracy and are invasive, while conventional tumor markers and tissue biopsies have limited sensitivity and specificity, and there is a lack of effective methods for predicting cancer types using artificial neural networks based on cell-free DNA analysis.

Method used

A method involving the extraction of nucleic acids from a biological sample, deriving terminal sequence motif frequencies and sizes, correcting for batch effects using a global reference value, and inputting the data into a trained artificial intelligence model for cancer diagnosis and type prediction.

Benefits of technology

The method achieves high sensitivity and accuracy in cancer diagnosis and type prediction by generating vectorized data from nucleic acid fragment information, effectively overcoming batch effects and improving diagnostic performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025002274_21082025_PF_FP_ABST
    Figure KR2025002274_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for cancer diagnosis and cancer type prediction using the frequency and size of cell-free nucleic acid fragment end sequence motifs, and more specifically, to a method for cancer diagnosis and cancer type prediction by analyzing a value obtained by extracting nucleic acids from a biological sample, deriving the frequencies of end sequence motifs and the fragment sizes of the nucleic acids on the basis of reads aligned by acquiring sequence information, then creating same as vectorized data and post-processing same, and then inputting the post-processed vectorized data to a trained artificial intelligence model. The method for cancer diagnosis and cancer type prediction using the frequency and size of end sequence motifs of cell-free nucleic acid fragments according to the present invention creates vectorized data and analyzes same using an AI algorithm, and thus exhibits high sensitivity and accuracy even with low read coverage, and is useful.
Need to check novelty before this filing date? Find Prior Art

Description

A method for improving the detection accuracy of cancer patients by correcting for experimental batch effects in whole-genome sequencing data of cell-free nucleic acid fragments.

[0001] The present invention relates to a method for improving the detection accuracy of cancer patients by correcting for experimental batch effects in cell-free nucleic acid fragment whole-genome sequencing data, and more specifically, to a method for diagnosing cancer and predicting cancer types using a method of extracting nucleic acids from a biological sample, obtaining sequence information, deriving the terminal sequence motif frequency and the size of the nucleic acid fragment based on aligned reads, generating the vectorized data, correcting for batch effects, post-processing the data, and then inputting the calculated values ​​into a trained artificial intelligence model to analyze the data.

[0002]

[0003] In clinical practice, a cancer diagnosis is typically confirmed by performing a tissue biopsy following a medical history, physical examination, and clinical evaluation. A clinical diagnosis of cancer is only possible if the number of cancer cells exceeds 1 billion and the tumor diameter exceeds 1 cm. In this case, the cancer cells already have the ability to metastasize, and at least half of them have already metastasized. Furthermore, tissue biopsies are invasive, causing significant discomfort to the patient, and often become unavailable during cancer treatment. Furthermore, tumor markers are used in cancer screening to monitor substances produced directly or indirectly by cancer. However, their accuracy is limited, as more than half of tumor marker screening results are normal even in the presence of cancer, and positive results are frequently found even in the absence of cancer.

[0004] To address the challenges of conventional cancer diagnostic methods, the demand for a relatively simple, noninvasive, and highly sensitive and specific cancer diagnostic method has led to the recent proliferation of liquid biopsy, a diagnostic technique utilizing a patient's bodily fluids for cancer diagnosis and follow-up testing. Liquid biopsy, a noninvasive method, is gaining attention as an alternative to existing invasive diagnostic and testing methods.

[0005] Recently, a method for diagnosing cancer and differentiating cancer types using cell-free DNA obtained from liquid biopsy has been developed (US 10975431), and in particular, a method for analyzing motif frequency information of cell-free nucleic acid terminal sequences and using it for cancer diagnosis, prenatal diagnosis, or organ transplant monitoring is known (WO 2020-125709, Peiyong Jiang et al., cancer discovery, Vol. 10, 2020, pp. 664-673).

[0006] In addition, a method for diagnosing cancer using the ends of cell-free nucleic acids has been known (US 2020-0199656 A1), but it has the disadvantage of low accuracy.

[0007] Meanwhile, an artificial neural network (ANN) is a computational model implemented in software or hardware that mimics the computational capabilities of biological systems using a large number of artificial neurons connected by connections. An ANN uses artificial neurons that simplify the function of biological neurons. These neurons are interconnected through connections with strengths, mimicking human cognitive functions and learning processes. A strength is a specific value associated with a connection, also known as a connection weight. ANN learning can be divided into supervised and unsupervised learning. Supervised learning involves input data and corresponding output data, and updating the connection strengths of the connections so that the output data corresponds to the input data. Representative learning algorithms include the delta rule and backpropagation learning. Unsupervised learning involves allowing an ANN to learn connection strengths on its own using only input data, without a target value. Unsupervised learning updates connection weights based on correlations between input patterns.

[0008] In machine learning, as the data becomes more complex and dimensionless, the curse of dimensionality arises. This means that as the required data dimensionality becomes infinite, the distance between any two points diverges infinitely, and the amount of data, or density, becomes somewhat lower in high-dimensional space, failing to properly reflect the data's features (Richard Bellman, Dynamic Programming, 2003, chapter 1). The recent development of deep neural networks (deep learning) has been reported to significantly improve the performance of classifiers in high-dimensional data such as images, videos, and signal data by processing linear combinations of variable values ​​transmitted from the input layer as nonlinear functions, with a structure that includes a hidden layer between the input and output layers (Hinton, Geoffrey, et al., IEEE Signal Processing Magazine Vol. 29.6, pp. 82-97, 2012).

[0009] There are various patents (KR KR 10-2018-124550, KR 10-2019-7038076, KR 10-2019-0003676, KR 10-2019-0001741) that utilize these artificial neural networks in the bio field, but there is a lack of research on methods to predict cancer types through artificial neural network analysis based on sequence analysis information of cell-free DNA (cfDNA) in blood.

[0010] Accordingly, the inventors of the present invention have made great efforts to solve the above problems and develop an artificial intelligence-based cancer diagnosis and cancer type prediction method with high sensitivity and accuracy. As a result, they have confirmed that cancer diagnosis and cancer type can be predicted with high sensitivity and accuracy when vectorized data is generated based on the terminal sequence motif of cell-free nucleic acid fragments and the length information of the nucleic acid fragments, the batch effect is corrected, and then the data is post-processed and analyzed with a learned artificial intelligence model, thereby completing the present invention.

[0011]

[0012] Summary of the invention

[0013] The purpose of the present invention is to provide a method for diagnosing cancer and predicting cancer types using the frequency and size of cell-free nucleic acid fragment terminal sequence motifs.

[0014] Another object of the present invention is to provide a device for diagnosing cancer and predicting cancer types using the frequency and size of cell-free nucleic acid fragment terminal sequence motifs.

[0015] Another object of the present invention is to provide a computer-readable storage medium comprising instructions configured to be executed by a processor for diagnosing cancer and predicting cancer types by the above method.

[0016] In order to achieve the above object, the present invention provides a method for providing information for cancer diagnosis and cancer type prediction, including the steps of: (a) obtaining sequence information of nucleic acids extracted from a biological sample; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) deriving terminal sequence motif frequencies and sizes of nucleic acid fragments using the aligned sequence information (reads); (d) generating vectorized data using the derived terminal sequence motif frequencies and sizes of nucleic acid fragments; (e) correcting batch effects of the vectorized data using a global reference value of a negative control (NC); (f) post-processing the corrected data; (g) inputting the post-processed data into a trained artificial intelligence model and comparing the analyzed output result values ​​with a cut-off value to determine the presence or absence of cancer; and (h) predicting the cancer type through the comparison of the output result values.

[0017] The present invention also provides a method for diagnosing and predicting cancer, comprising the steps of: (a) obtaining sequence information of nucleic acids extracted from a biological sample; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) deriving terminal sequence motif frequencies and sizes of nucleic acid fragments using the aligned sequence information (reads); (d) generating vectorized data using the derived terminal sequence motif frequencies and sizes of nucleic acid fragments; (e) correcting batch effects of the vectorized data using a global reference value of a negative control (NC); (f) post-processing the corrected data; (g) inputting the post-processed data into a trained artificial intelligence model and comparing the analyzed output results with a cut-off value to determine the presence or absence of cancer; and (h) predicting the type of cancer through the comparison of the output results.

[0018] The present invention also provides a cancer diagnosis and cancer prediction device, including: a decoding unit that decodes sequence information of nucleic acids extracted from a biological sample; an alignment unit that aligns the decoded sequence to a standard chromosome sequence database; a nucleic acid fragment analysis unit that derives terminal sequence motif frequencies and sizes of nucleic acid fragments based on the aligned sequences; a data generation unit that generates vectorized data using the derived terminal sequence motif frequencies and sizes of nucleic acid fragments, and then performs post-processing by correcting the batch effect of the vectorized data using reference values ​​of a negative control (NC); a cancer diagnosis unit that inputs the generated post-processed vectorized data into a learned artificial intelligence model for analysis and compares it with a reference value to determine the presence or absence of cancer; and a cancer prediction unit that analyzes the output result value to predict the cancer type.

[0019] The present invention also provides a computer-readable storage medium, comprising instructions configured to be executed by a processor for diagnosing cancer and predicting a cancer type, comprising the steps of: (a) obtaining sequence information of nucleic acids extracted from a biological sample; (b) aligning the obtained sequence information (reads) to a reference genome database; (c) deriving terminal sequence motif frequencies and sizes of nucleic acid fragments using the aligned sequence information (reads); (d) generating vectorized data using the derived terminal sequence motif frequencies and sizes of nucleic acid fragments; (e) correcting a batch effect of the vectorized data using a reference value of a negative control (NC); (f) post-processing the corrected data; (g) inputting the post-processed data into a learned artificial intelligence model and comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; And (h) a computer-readable storage medium for cancer diagnosis and cancer type prediction is provided, including instructions configured to be executed by a processor for cancer diagnosis and cancer type prediction through a step of predicting cancer type by comparing the output result values.

[0020]

[0021] Figure 1 is an overall flowchart for performing a method for improving the detection accuracy of cancer patients by correcting experimental batch effects in cell-free nucleic acid fragment whole-genome sequencing data of the present invention.

[0022] The left panel of Fig. 2 is an example of a FEMS table produced in one embodiment of the present invention written as a single nucleic acid fragment, and the right panel is an example of a FEMS table produced as a whole nucleic acid fragment.

[0023] FIG. 3 is a diagram illustrating a batch effect correction step of the present invention, wherein (A) is an example showing that the measured values ​​for each feature are different for each batch, (B) shows the configuration of learning samples included in a batch performed according to the method of the present invention, (C) shows a method for calculating a global reference value for each feature according to the method of the present invention, (D) shows a method for calculating a correction coefficient to correct the batch effect according to the method of the present invention, and (E) shows the result value of (A) in which the batch effect is corrected according to the method of the present invention.

[0024] FIG. 4 is a graph showing the frequency distribution by feature before (upper panel) and after (lower panel) correction for the placement effect according to one embodiment of the present invention.

[0025] FIG. 5 is a drawing explaining the difference in frequency values ​​by zone of a FEMS table manufactured in one embodiment of the present invention.

[0026] Figure 6 is a schematic diagram showing the manufacturing process of a FEMS_BC_Z table manufactured in one embodiment of the present invention.

[0027] Figure 7 illustrates a one-dimensional conversion process of the FEMS table and FEMS_BC_Z table used in one embodiment of the present invention.

[0028] FIG. 8 shows the results of verifying the performance of an MLP model (Before_BC, left graph) using a one-dimensional conversion of a FEMS table constructed in one embodiment of the present invention and an MLP model (After_BC, right graph) using a one-dimensional conversion of a FEMS_BC_Z table using a discovery data set used for learning.

[0029] Figure 9 shows the results of verifying the performance of an MLP model (Before_BC, left graph) that uses a one-dimensional conversion of the FEMS table constructed in one embodiment of the present invention and an MLP model (After_BC, right graph) that uses a one-dimensional conversion of the FEMS_BC_Z table using an external data set.

[0030] FIG. 10 shows the results of verifying the performance of an MLP model (Before_BC, upper panel) that uses a FEMS table constructed in one embodiment of the present invention after converting it into one dimension, and an MLP model (After_BC, lower panel) that uses a FEMS_BC_Z table after converting it into one dimension.

[0031] FIG. 11 shows the results of confirming the distribution of output values ​​(DPI) of an MLP model (Before_BC, upper panel) that uses a FEMS table constructed in one embodiment of the present invention after converting it into one dimension and an MLP model (After_BC, lower panel) that uses a FEMS_BC_Z table after converting it into one dimension.

[0032]

[0033] Detailed description of the invention and preferred embodiments

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the nomenclature used herein and the experimental methods described below are well known and commonly used in the art.

[0035] Terms such as first, second, A, and B may be used to describe various components, but these components are not limited by these terms and are used solely to distinguish one component from another. For example, without departing from the scope of the technology described below, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0036] In the terms used herein, the singular expression should be understood to include the plural expression unless the context clearly dictates otherwise, and the term "comprises" and the like should be understood to mean that a stated feature, number, step, operation, component, part, or combination thereof is present, but does not exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0037] Before going into a detailed description of the drawings, it should be made clear that the division of components in this specification is merely a division based on the main function of each component. In other words, two or more components described below may be combined into a single component, or a single component may be further subdivided into two or more components with more detailed functions. In addition to its own main function, each component described below may additionally perform some or all of the functions of other components, and of course, some of the main functions of each component may be exclusively performed by other components.

[0038] Additionally, in performing a method or method of operation, each process constituting the method may occur in a different order than the stated order, unless the context clearly indicates a specific order. That is, each process may occur in the same order as the stated order, may be performed substantially simultaneously, or may be performed in the opposite order.

[0039]

[0040] In the present invention, when sequence analysis data obtained from a sample is aligned to a reference genome, and then the terminal sequence motif frequency and the size of the nucleic acid fragment are derived based on the aligned sequence information, vectorized data is generated using the terminal sequence motif frequency and the size of the nucleic acid fragment derived above, and then the batch effect is corrected, post-processed, and then the DPI value is calculated and analyzed using the learned artificial intelligence model, it was confirmed that cancer diagnosis and cancer type can be predicted with high sensitivity and accuracy.

[0041] That is, in one embodiment of the present invention, after sequencing DNA extracted from blood, it is aligned to a reference chromosome, and then the terminal sequence motif frequency of the nucleic acid fragment and the size of the nucleic acid fragment are derived using this, and vectorized data with the terminal sequence motif frequency of the nucleic acid fragment as the X-axis and the size of the nucleic acid fragment as the Y-axis are generated to correct the batch effect, and then post-process it, and then input it into an artificial intelligence model trained to perform cancer diagnosis and cancer type classification to output a DPI value, and then perform cancer diagnosis by comparing this with a reference value, and then a method was developed to determine the cancer type showing the highest DPI value among the DPI values ​​output for each cancer type as the cancer type of the sample (Fig. 1).

[0042] Therefore, the present invention, from a consistent perspective,

[0043] (a) A step of obtaining sequence information of nucleic acid extracted from a biological sample;

[0044] (b) a step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database);

[0045] (c) a step of deriving the terminal sequence motif frequency and the size of the nucleic acid fragments using the aligned sequence information (reads);

[0046] (d) a step of generating vectorized data using the terminal sequence motif frequency of the nucleic acid fragment derived above and the size of the nucleic acid fragment;

[0047] (e) a step of correcting the batch effect of the vectorized data using the global reference value of the negative control (NC);

[0048] (f) a step of post-processing the above corrected data;

[0049] (g) a step of inputting the above post-processed data into a learned artificial intelligence model, comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; and

[0050] (h) It relates to a method for providing information for cancer diagnosis and cancer type prediction, including a step of predicting cancer type by comparing the above output result values.

[0051] The present invention also provides:

[0052] (a) A step of obtaining sequence information of nucleic acid extracted from a biological sample;

[0053] (b) a step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database);

[0054] (c) a step of deriving the terminal sequence motif frequency and the size of the nucleic acid fragments using the aligned sequence information (reads);

[0055] (d) a step of generating vectorized data using the terminal sequence motif frequency of the nucleic acid fragment derived above and the size of the nucleic acid fragment;

[0056] (e) a step of correcting the batch effect of the vectorized data using the global reference value of the negative control (NC);

[0057] (f) a step of post-processing the above corrected data;

[0058] (g) a step of inputting the above post-processed data into a learned artificial intelligence model, comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; and

[0059] (h) It relates to a method for diagnosing cancer and predicting cancer types, including a step of predicting cancer types by comparing the above output result values.

[0060]

[0061] In the present invention, the nucleic acid fragment may be used without limitation as long as it is a fragment of nucleic acid extracted from a biological sample, and may preferably be a fragment of cell-free nucleic acid or intracellular nucleic acid, but is not limited thereto.

[0062] In the present invention, the nucleic acid fragment can be obtained by any method known to a person skilled in the art, preferably, by direct sequencing, by next-generation sequencing, by non-specific whole genome amplification, or by probe-based sequencing, but is not limited thereto.

[0063]

[0064] In the present invention, the cancer may be a solid cancer or a blood cancer, and preferably ovarian cancer, soft tissue sarcoma, peripheral T-cell cancer, colon cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T-cell lymphoma, non-Hodgkin's lymphoma, urinary tract cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lung cancer, Hodgkin's lymphoma, renal cell cancer, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancer, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, blood cancer, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, Hemangiosarcoma, osteosarcoma, refractory cervical cancer, cholangiocarcinoma, gastroesophageal adenocarcinoma, rhabdomyosarcoma, carcinoma, non-muscle-invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open-angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumor, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancy, endometrial cancer, mucinous / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tendon sheath giant cell tumor, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal cancer, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, stomach cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B-cell lymphoma, relapsed / refractory It may be selected from the group consisting of lymphoma, metastatic colorectal cancer, advanced malignant tumor and acute lymphoblastic leukemia, and more preferably, lung cancer, but is not limited thereto.

[0065]

[0066] In the present invention,

[0067] Step (a) above

[0068] (ai) a step of obtaining nucleic acid from a biological sample;

[0069] (a-ii) a step of removing proteins, fats, and other residues from the collected nucleic acid using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid;

[0070] (a-iii) a step of producing a single-end sequencing or pair-end sequencing library for purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method;

[0071] (a-iv) a step of reacting the produced library with a next-generation sequencer; and

[0072] (av) A step of obtaining sequence information (reads) of nucleic acids from a next-generation genetic sequencer;

[0073] It may be characterized by including .

[0074]

[0075] In the present invention, the step of obtaining sequence information in step (a) may be characterized by obtaining the separated cell-free DNA through whole-genome sequencing at a depth of 1 million to 100 million reads.

[0076] In the present invention, the biological sample means any material, biological fluid, tissue or cell obtained from or derived from an individual, for example, whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, blood (including plasma and serum), sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, pelvic fluids, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple It may include, but is not limited to, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, semen, hair, saliva, urine, buccal cells, placental cells, cerebrospinal fluid, and mixtures thereof.

[0077]

[0078] In the present invention, the next-generation sequencer can be used with any sequencing method known in the art. Sequencing of nucleic acids isolated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of an individual nucleic acid molecule or a clonally expanded proxy for an individual nucleic acid molecule in a highly similar manner (e.g., 105 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of a nucleic acid species in a library can be estimated by counting the relative occurrence of its cognate sequence in data generated by a sequencing experiment. Next-generation sequencing methods are well known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference.

[0079] Platforms for next-generation sequencing include, but are not limited to, the Roche / 454 Genome Sequencer (GS) FLX system, the Illumina / Solexa Genome Analyzer (GA), the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, the Polonator G. 007 system, the Helicos BioSciences' HeliScope Gene Sequencing system, and the Pacific Biosciences PacBio RS system.

[0080]

[0081] In the present invention, the length of the sequence information (reads) of step (b) is 5 to 5,000 bp, and the number of sequence information used may be 5,000 to 5,000,000, but is not limited thereto.

[0082]

[0083] In the present invention, the nucleic acid fragment terminal sequence motif of step (c) may be characterized as a pattern of 2 to 30 base sequences at both ends of the nucleic acid fragment.

[0084] That is, when there is a nucleic acid fragment sequenced by paired-end sequencing as follows,

[0085] Forward strand: 5`-TACAGACTTTGGAAT-3` (SEQ ID NO: 1)

[0086] Reverse strand: 3`-ATGACTGAAACCTTA-5` (SEQ ID NO: 2)

[0087] TACA, read sequentially from the 5' end of the forward strand, and ATTC, read sequentially from the 5' end of the reverse strand, become the terminal sequence motif values ​​of this nucleic acid fragment.

[0088]

[0089] In the present invention, the frequency of the nucleic acid fragment terminal sequence motif in step (c) may be characterized as being the number of each motif detected in the entire nucleic acid fragment.

[0090] That is, when analyzing the nucleic acid fragment terminal sequence motif based on the four bases at both ends (4-mer motif), since there are four types of base combinations of A, T, G, and C possible at the 1st, 2nd, 3rd, and 4th positions, a total of 256 (4*4*4*4) combinations of motif values ​​are subject to analysis.

[0091] The motif frequency is the number of times each motif is observed in the entire nucleic acid fragments produced by sequencing, and the relative frequency of each motif is calculated by dividing this value by the total number of nucleic acid fragments produced.

[0092]

[0093] As described in Table 1 above, the total number of nucleic acid fragments is 126,430,124, and the number of nucleic acid fragments in which AAAA is analyzed as a nucleic acid fragment terminal sequence motif is 125,071, so the frequency of the AAAA nucleic acid fragment terminal sequence motif is 125,071, and the relative frequency of the nucleic acid fragment terminal sequence motif calculated by dividing this by the total number of nucleic acid fragments is 0.00099.

[0094]

[0095] In the present invention, the size of the nucleic acid fragment in step (c) may be characterized by being the number of bases from the 5' end to the 3' end of the nucleic acid fragment.

[0096] For example, the size of the nucleic acid fragments analyzed by the above sequence numbers 1 and 2 is 15.

[0097] In the present invention, the size of the nucleic acid fragment may be 1 to 10,000, preferably 10 to 1,000, more preferably 50 to 500, and most preferably 100 to 250, but is not limited thereto.

[0098]

[0099] In the present invention, the vectorized data of step (d) may be characterized in that the type of nucleic acid fragment terminal sequence motif is on the X-axis and the size of the nucleic acid fragment is on the Y-axis.

[0100] That is, assuming there is one nucleic acid fragment as follows,

[0101] Forward strand: 5`-TACAGACTAGT … TTGGAAT-3` (SEQ ID NO: 3)

[0102] Reverse strand: 3`-ATGACTGATCA … AACCTTA-5` (SEQ ID NO: 4)

[0103] Fragment Size: 176

[0104] This nucleic acid fragment can be expressed as a two-dimensional vector as in the left panel of Fig. 2, and when this process is extended to the entire nucleic acid fragment and accumulated, a two-dimensional vector as in the right panel of Fig. 2 is generated.

[0105]

[0106] In the present invention, the batch effect refers to the effect that data mixed with noise is produced depending on the slight environmental difference of each experimental batch, even if multiple experiments are performed repeatedly using the same protocol with the same blood extracted from the same sample, and as a result, as described in (A) of FIG. 3, for one feature constituting the FEMS table, the feature value of the cancer sample (Cancer, C) is higher than that of the negative control (NC) and normal (N) within the batch, but noise occurs in which the normal value of batch 2 is distributed higher than the cancer sample value of batch 1, and if this is not corrected, it becomes vulnerable to overfitting during the artificial intelligence model learning process, and a significant performance degradation phenomenon occurs during external validation, which must be overcome.

[0107]

[0108] To this end, in the present invention, for each batch, a negative control group was included, and NGS was performed by configuring samples to include NC for each batch as described in (B) of Fig. 3, and the median value of each NC feature obtained through this was used as a global reference value to correct the batch effect.

[0109] That is, in (B) of Fig. 3, 15 NC data are produced from 3 batches of 5 NC samples each, so the median value is calculated for each feature in the 15 NC data and used as a global reference value. For example, a feature with a length of 110 bp and an end of AAAA is 110_AAAA, and the median value of the frequency values ​​of the feature in the 15 NC samples becomes the global reference value of the feature, and the median value of the frequency values ​​of the 150_TTTT feature in the 15 NC samples becomes the global reference value of the feature, and a global reference value is calculated for all features as described in (C) of Fig. 3.

[0110] Thereafter, in order to correct the batch effect, as described in (D) of Fig. 3, the median value of the frequency value of each feature of NC existing in each batch is compared with the global reference value to calculate the correction coefficient using the following equation 1.

[0111] Equation 1: Normalization Factor feature = (Global Reference feature / Batch NCs Median feature )

[0112] For example, when the global reference value of the 110_AAAA feature is G, and the feature values ​​of the 5 NCs produced in Batch 1 are B1.1, B.12, …, B1.5, NF = G / Median(B1.1, B1.2, …, B1.5).

[0113] Using the calculated NF value, the correction value is obtained by multiplying the feature frequency values ​​of the remaining samples by NF.

[0114] For example, when the value of the Normal 1 sample of the 110_AAAA feature in Batch 1 is Raw1 and the Normalization Factor value is NF, the Bias Corrected value is calculated as Raw1 * NF.

[0115] After going through the above process, as described in (E) of Fig. 3, the values ​​for each batch become closer to the global reference value, and the batch effect is corrected.

[0116]

[0117] In the present invention, the step (e) may be characterized in that it is performed by a method including the following steps:

[0118] (ei) a step of calculating a global reference value for the entire negative control group;

[0119] (e-ii) a step of calculating a normalization factor (NF) by comparing the median value of the negative control group within the batch with the global reference value; and

[0120] (e-iii) A step of obtaining a correction value by multiplying the vectorized data by a correction coefficient.

[0121]

[0122] In the present invention, the negative control group can be used without limitation as long as it is a sample determined not to be cancerous, and is preferably characterized as being a normal sample.

[0123] In the present invention, the negative control group may be characterized by being used together in the step of obtaining sequence information in step (a) to perform analysis in the same batch as the biological sample. That is, in the present invention, when obtaining read data using the nucleic acid sample of the biological sample to be analyzed by NGS, the nucleic acid sample of the negative control group may be introduced together in the same batch to obtain read data together.

[0124]

[0125] In the present invention, the global reference value may be characterized as being the median of vectorized data of the entire negative control group.

[0126] In the present invention, the step (e-ii) is characterized in that the method is calculated using the following formula:

[0127] Equation 1: Normalization Factor feature = (Global Reference feature / Batch NCs Median feature )

[0128] Here, the feature can be characterized by the type of each nucleic acid fragment terminal sequence motif and the frequency value according to the size of the nucleic acid fragment.

[0129]

[0130] In the present invention, the step (f) may be characterized in that it is performed by a method including the following steps:

[0131] (fi) a step of calculating the average and standard deviation of the frequency values ​​of each nucleic acid fragment terminal sequence motif type and nucleic acid fragment size in a normal group;

[0132] (f-ii) a step of performing Z-standardization by subtracting the average of the frequency values ​​of the types of terminal sequence motifs of each nucleic acid fragment and the size of each nucleic acid fragment in the normal group from the frequency values ​​of the types of terminal sequence motifs of each nucleic acid fragment and the size of each nucleic acid fragment in the sample, and then dividing the average by the standard deviation of the frequency values ​​of the types of terminal sequence motifs of each nucleic acid fragment and the size of each nucleic acid fragment to derive a Z-standardization value; and

[0133] (f-iii) If the Z standardization value derived from (f-ii) above exceeds the reference range, a step of correcting it to the reference value.

[0134] In the present invention, the reference range may be -5 to 5, and the reference value may be -5 or 5.

[0135]

[0136] That is, the existing FEMS_BC table is characterized by a post-processing task to standardize the large difference in the distribution of values ​​calculated by area.

[0137] For example, the above post-processing operation can be performed through the following steps:

[0138] i) A step of selecting 99 healthy people included in the training data as a Z standardization reference group (Z Reference set);

[0139] ii) Step of calculating the mean and standard deviation of the values ​​observed at each position in the FEMS_BC table in the selected Z standardization reference group: For example, in the FEMS table of 99 Z standardization reference groups, the mean and standard deviation of the values ​​at position (a) having a nucleic acid fragment size of 180 and an AAAA motif are calculated and defined as Mean_180_AAAA and SD_180_AAAA, respectively.

[0140] iii) A step of performing Z-standardization using the average and standard deviation values ​​at each position in the FEMS table calculated in the above ii) process: Specifically, when the frequency value observed at a position having an AAAA motif and a nucleic acid fragment size of 180 is Value_180_AAAA, Z-standardization is performed using the formula Z_180_AAAA = (Value_180_AAAA - Mean_180_AAAA) / SD_180_AAAA.

[0141] iv) To exclude the influence of the Z standardization value being calculated outside the general range (-5 to 5) due to the standard deviation value being too small, a step of limiting the minimum and maximum ranges of the Z standardization value by setting the values ​​for Z < -5 to -5 and the values ​​for Z > 5 to 5.

[0142]

[0143] The FEMS_BC_Z table generated through the above steps is visualized as shown in Fig. 6.

[0144]

[0145] In the present invention, the vectorized data is not limited thereto, but may preferably be characterized as a 2D table or a 1D vector flattened therefrom.

[0146] The method of unrolling a 2D table into a 1D vector can be used without limitation by any method known to a person skilled in the art, and preferably, the method can be characterized by listing the frequency values ​​for each feature of the FEMS_BC_Z table of the present invention on the x-axis.

[0147] For example, the FEMS_BC_Z table of the present invention records the fragment end motif type on the x-axis, the fragment size on the Y-axis, and the frequency value within the table. When this is expanded in one dimension, as described in Fig. 8, the frequency value of a nucleic acid fragment having an AAAA terminal motif of 110 bp in size becomes the first value, the frequency value of a nucleic acid fragment having an AAAC terminal motif of 110 bp in size becomes the second value, and by repeating this, the frequency value of a nucleic acid fragment having a TTTT terminal motif of 230 bp in size becomes the 30,976th value.

[0148]

[0149] In the present invention, it may be characterized by further including a step of separately classifying nucleic acid fragments that satisfy the alignment match score (mapping quality score) of the aligned nucleic acid fragments prior to performing the step (c).

[0150] In the present invention, the alignment consistency score (mapping quality score) may vary depending on a desired criterion, but is preferably 15-70 points, more preferably 50-70 points, and most preferably 60 points.

[0151]

[0152] In the present invention, the artificial intelligence model of step (g) can be used without limitation as long as it is a model that can learn to distinguish images by cancer type, and is preferably characterized as being a deep learning model.

[0153]

[0154] In the present invention, the artificial intelligence model may be used without limitation as long as it is an artificial neural network algorithm that can analyze vectorized data based on an artificial neural network, but is preferably selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a multi-layer perceptron (MLP), and a recurrent neural network (RNN), but is not limited thereto.

[0155]

[0156] In the present invention, the recurrent neural network may be characterized by being selected from the group consisting of a long-short term memory (LSTM) neural network, a gated recurrent unit (GRU) neural network, a vanilla recurrent neural network, and an attentive recurrent neural network.

[0157]

[0158] In the present invention, when the artificial intelligence model is MLP, the loss function may be characterized by being expressed by the following equation 2 or equation 3.

[0159]

[0160]

[0161] In the present invention, when the artificial intelligence model is MLP, learning may be performed by including the following steps:

[0162] i) A step of classifying the produced vector data into training, validation, and test data;

[0163] At this time, the training data is used when learning the MLP model, the validation data is used for verifying hyper-parameter tuning, and the test data is used for performance evaluation after producing the optimal model.

[0164] ii) A step to build an optimal MLP model through hyper-parameter tuning and learning process;

[0165] iii) A step of comparing the performance of several models obtained through hyper-parameter tuning using validation data and determining the model with the best validation data performance as the optimal model;

[0166]

[0167] In the present invention, the hyper-parameter tuning process is a process of optimizing the values ​​of several parameters (number of dense layers, number of hidden nodes, etc.) that constitute the MLP model, and the hyper-parameter tuning process may be characterized by using Bayesian optimization and grid search techniques.

[0168] In the present invention, the learning process may be characterized in that the internal parameters (weights) of the MLP model are optimized using the determined hyper-parameters, and when the validation loss begins to increase compared to the training loss, the model is judged to be overfitted, and model learning is stopped before that.

[0169]

[0170] In the present invention, the result value analyzed from the vectorized data input by the artificial intelligence model in the step (g) can be used without limitation as long as it is a specific score or a real number, and is preferably characterized as a DPI (Deep Probability Index) value, but is not limited thereto.

[0171]

[0172] In the present invention, the Deep probability Index refers to a value expressed as a probability value by adjusting the output of the artificial intelligence to a scale of 0 to 1 using a sigmoid function in the case of binary classification or a softmax function in the case of multi-class classification in the last layer of the artificial intelligence model.

[0173] In the case of binary classification, the sigmoid function is used to train so that the DPI value becomes 1 if it is cancer. For example, if a breast cancer sample and a normal sample are input, the DPI value of the breast cancer sample is trained so that it is close to 1.

[0174] In the case of multi-class classification, the softmax function is used to extract as many DPI values ​​as the number of classes. The sum of the DPI values ​​for the number of classes is 1, and training is performed so that the DPI value of the actual corresponding cancer type is 1. For example, if there are three classes (breast cancer, liver cancer, and normal), and a breast cancer sample is input, the breast cancer class will be trained to be close to 1.

[0175] In the present invention, the output result value of step (g) may be characterized by determining the presence or absence of cancer.

[0176]

[0177] In the present invention, when the artificial intelligence model is trained, if there is cancer, the output result is trained to be closer to 1, and if there is no cancer, the output result is trained to be closer to 0, and if it is 0.5 or higher, it is determined that there is cancer, and if it is 0.5 or lower, it is determined that there is no cancer, and performance measurement is performed (Training, validation, test accuracy).

[0178] Here, it should be obvious to those skilled in the art that the cutoff value of 0.5 is a value that can change at any time. For example, to reduce false positives, a cutoff value higher than 0.5 can be set to strictly determine the presence of cancer. On the other hand, to reduce false negatives, a cutoff value can be set lower, making the cutoff value slightly more lenient.

[0179] Most preferably, the learned artificial intelligence model can be used to apply unseen data (data with known answers that were not trained) to determine the probability of the DPI value and set a reference value.

[0180]

[0181] In the present invention, the step of predicting the cancer type by comparing the output result values ​​of the step (h) may be characterized by being performed by a method including a step of determining the cancer type showing the highest value among the output result values ​​as the cancer of the sample.

[0182]

[0183] In another aspect, the present invention comprises a decoding unit for decoding sequence information of nucleic acid extracted from a biological sample;

[0184] An alignment unit that aligns the decoded sequence to a standard chromosome sequence database; and

[0185] A nucleic acid fragment analysis unit that derives the terminal sequence motif frequency and size of nucleic acid fragments based on aligned sequences;

[0186] A data generation unit that generates vectorized data using the terminal sequence motif frequency of the derived nucleic acid fragment and the size of the nucleic acid fragment, and then corrects the batch effect of the vectorized data using the reference value of the negative control (NC), and then performs post-processing;

[0187] A cancer diagnosis unit that inputs the generated post-processed vectorized data into a learned artificial intelligence model for analysis and compares it with a reference value to determine the presence or absence of cancer; and

[0188] The present invention relates to a cancer diagnosis and cancer type prediction device including a cancer type prediction unit that predicts cancer type by analyzing the output result values.

[0189] In the present invention, the decoding unit may include a nucleic acid injection unit that injects nucleic acid extracted from an independent device; and a sequence information analysis unit that analyzes sequence information of the injected nucleic acid, and may preferably be an NGS analysis device, but is not limited thereto.

[0190] In the present invention, the decoding unit may be characterized by receiving and decoding sequence information data generated in an independent device.

[0191]

[0192] In another aspect, the present invention comprises a computer-readable storage medium comprising instructions configured to be executed by a processor for diagnosing cancer and predicting cancer types,

[0193] (a) A step of obtaining sequence information of nucleic acid extracted from a biological sample;

[0194] (b) a step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database);

[0195] (c) a step of deriving the terminal sequence motif frequency and the size of the nucleic acid fragments using the aligned sequence information (reads);

[0196] (d) a step of generating vectorized data using the terminal sequence motif frequency of the nucleic acid fragment derived above and the size of the nucleic acid fragment;

[0197] (e) a step of correcting the batch effect of the vectorized data using the reference value of the negative control (NC);

[0198] (f) a step of post-processing the above corrected data;

[0199] (g) a step of inputting the above post-processed data into a learned artificial intelligence model, comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; and

[0200] (h) A computer-readable storage medium for cancer diagnosis and cancer type prediction, comprising instructions configured to be executed by a processor for cancer diagnosis and cancer type prediction, through a step of predicting cancer type by comparing the above output result values.

[0201]

[0202] In another aspect, the method according to the present invention can be implemented using a computer. In one embodiment, the computer includes one or more processors connected to a chipset. The chipset also includes memory, storage, a keyboard, a graphics adapter, a pointing device, and a network adapter. In one embodiment, the performance of the chipset is enabled by a memory controller hub and an I / O controller hub. In another embodiment, the memory may be directly connected to the processor instead of the chipset. The storage device is any device capable of retaining data, including a hard drive, a compact disk read-only memory (CD-ROM), a DVD, or other memory device. The memory stores data and instructions used by the processor. The pointing device may be a mouse, trackball, or other type of pointing device, and is used in combination with a keyboard to transmit input data to the computer system. The graphics adapter displays images and other information on a display. The network adapter connects the computer system to a local or wide area network. However, the computer used in the present invention is not limited to the above configuration, and may not have some configurations or may include additional configurations, and may also be part of a storage area network (SAN), and the computer in the present invention may be configured to be suitable for executing a module in a program for performing a method according to the present invention.

[0203]

[0204] The term "module" herein may refer to a functional and structural combination of hardware for implementing the technical concepts of the present invention and software for operating the hardware. For example, the module may refer to a logical unit of a given code and hardware resources for executing the given code. However, it is apparent to those skilled in the art that the module does not necessarily refer to physically connected code or a single type of hardware.

[0205]

[0206] Example

[0207] Hereinafter, the present invention will be described in more detail through examples. These examples are intended solely to illustrate the present invention, and it will be apparent to those skilled in the art that the scope of the present invention is not limited by these examples.

[0208]

[0209] Example 1. Extracting DNA from blood and performing next-generation base sequence analysis.

[0210] Blood samples were collected from 3,695 healthy individuals and 329 lung cancer patients (10 mL each) and stored in EDTA tubes. Within 2 hours of collection, the plasma portion was centrifuged for the first time at 1200 g, 4 °C, and 15 minutes. The first-centrifuged plasma was then centrifuged a second time at 16000 g, 4 °C, and 10 minutes to separate the plasma supernatant, excluding the precipitate. Cell-free DNA was extracted from the separated plasma using the Chemagic ccfNA 2K Kit (Chemagen), and library preparation was performed using the MGIEasy cell-free DNA library prep set kit. The DNA sequences were sequenced using a DNBseq G400 instrument (MGI) in 100-base paired end mode. As a result, approximately 170 million reads were produced per sample.

[0211] The generated data set is shown in Table 2 below.

[0212]

[0213]

[0214] Example 2. Selection of nucleic acid fragment size

[0215] In the case of nucleic acid fragment size selection, since most of the nucleic acid fragments that have completed quality verification have a size in the range of 110 to 230, when creating a FEMS table including an area outside this size range, most of the area is filled with 0 values, and only meaningless noise increases, so the above size was selected.

[0216]

[0217] Example 3. Creation of Fragment End Motif Frequency and Size (FEMS) table, bias correction, and FEMS_BC_Z table creation

[0218] 3-1 Creating a FEMS table

[0219] In order to simultaneously express the Fragment End Motif frequency value and Size information of the selected nucleic acid fragments in Example 2, a two-dimensional vector was created by arranging the motif type on the X-axis and the Fragment Size on the Y-axis. More specifically, as described in the left panel of Fig. 2, for one nucleic acid fragment, the nucleic acid motif type and size at both ends were expressed as frequencies, and this was expanded and accumulated to the entire nucleic acid fragment, thereby creating a two-dimensional vector as described in Fig. 2.

[0220]

[0221] 3-2 Bias correction

[0222] In the discovery data set described in Table 2, data from five cancer patients, normal subjects, and negative controls (NC) were produced together for each experimental batch.

[0223] That is, each time NGS was performed for each batch as in (b) of Fig. 3, it was performed together with 5 negative control samples, generating a total of 465 NC data.

[0224] The structure of the entire data is as shown in Table 3 below.

[0225]

[0226]

[0227]

[0228]

[0229]

[0230] The values ​​constituting the above FEMS table are the frequency values ​​of nucleic acid fragments having a specific size and terminal motif, and each nucleic acid fragment having a specific size and terminal motif is a featrue. A featrue means, for example, a nucleic acid fragment having a size of 110 bp and an AAAA terminal, and the frequency values ​​of this nucleic acid fragment are the values ​​constituting the FEMS table.

[0231] First, to calculate the global reference value, the median of the feature frequency values ​​of the NC data of the XX people was calculated. As a result, the global reference value was confirmed to be 0.127.

[0232] Afterwards, to correct for the experimental batch effect, the feature-specific Normalization Factor (NF) was calculated using Equation 1.

[0233] Equation 1: Normalization Factor feature = (Global Reference feature / Batch NCs Median feature )

[0234] For example, in the case of the 110_AAAA feature, the global reference value is 0.127, and the feature values ​​of the NCs produced in batch 1 were 0.12, 0.127, 0.128, 0.132, and 0.119, respectively, so the median value of these was 0.127, and the NF of the 110_AAAA feature was confirmed to be 1.000.

[0235] For Batch 2, the feature values ​​of NC were 0.237, 0.236, 0.231, 0.233, and 0.239, respectively, so the median value of these was 0.236, and the NF of the 110_AAAA feature was confirmed to be 0.538.

[0236] For each batch, NF was calculated for all features ((230-110+1)*256=30,976).

[0237] Finally, the calculated frequency values ​​for normal and cancer samples for each batch were multiplied by the above NF to obtain the corrected frequency values.

[0238] For example, in the case of batch 1, since the 110_AAAA feature value of normal 1 was 0.126, the corrected 0.126 was obtained by multiplying by the NF value of 1.000, and since the 110_AAAA feature value of lung cancer sample 1 was 0.135, the corrected 0.135 was obtained by multiplying by the NF value of 1.000.

[0239] For Batch 2, since the 110_AAAA feature value of normal person 1 was 0.236, 0.127 was obtained by multiplying it by the NF value of 0.538, and for lung cancer sample 1, since the 110_AAAA feature value was 0.244, 0.131 was obtained by multiplying it by the NF value of 0.538.

[0240] The above process was repeated for all features to generate the FEMS_BC table. Furthermore, after correcting for the placement effect as described in Figure 4, it was confirmed that the distribution of each feature value became consistent.

[0241]

[0242] 3-3 Create FEMS_BC_Z table

[0243] The frequency values ​​of Example 3-2 have the characteristic that the distribution difference between the values ​​calculated in the relatively high frequency areas (A, B) and the low frequency area (C) is large, as described in Fig. 5. For example, a difference of 100 units is observed in area A, a difference of 10,000 units is observed in area B, whereas a difference of only 1 unit is rarely observed in area C. If this FEMS_BC table is used as is, a problem occurs in that it becomes difficult for the CNN-based AI algorithm to learn parameters (weights). Therefore, an additional post-processing task was performed to create a FEMS_BC_Z table so that all areas within the FEMS_BC table have values ​​in a similar range.

[0244] Specifically, 1,499 healthy individuals included in the discovery data in Table 2 were selected as the Z Reference set, and then the mean and standard deviation of the values ​​observed at each location in the FEMS table were calculated in the selected Z Reference set.

[0245] That is, the mean and standard deviation of the values ​​at position (a) with a nucleic acid fragment size of 180 and an AAAA motif in the FEMS table of the Z-standardized reference group of 1,499 people were calculated and defined as Mean_180_AAAA and SD_180_AAAA, respectively.

[0246] Z-standardization was performed using the mean and standard deviation values ​​at each position in the FEMS table calculated in the above process. Specifically, when the frequency value observed at a position with a nucleic acid fragment size of 180 and an AAAA motif is called Value_180_AAAA, Z-standardization was performed using the formula Z_180_AAAA = (Value_180_AAAA - Mean_180_AAAA) / SD_180_AAAA (Fig. 7).

[0247] To exclude the influence of Z-standardized values ​​that are calculated outside the normal range (-5 to 5) due to their standard deviation values ​​being too small, the minimum and maximum ranges of Z-standardized values ​​were limited to -5 for values ​​with Z < -5 and 5 for values ​​with Z > 5.

[0248] Through the above process, a 2D vector that replaces the values ​​of all locations in the existing FEMS_BC table with Z-standardized values ​​was defined as the FEMS_BC_Z table.

[0249]

[0250] Example 4. MLP model construction and training process

[0251] 4-1 One-dimensional transformation of the FEMS_BC_Z table

[0252] The two-dimensional FEMS table and FEMS_BC_Z table were flattened into one dimension and used as input data by arranging the corrected values ​​corresponding to each feature on the X-axis as described in Fig. 7.

[0253] That is, for each sample, a one-dimensional vector listing the frequency values ​​of terminal motifs for each size of 30,976 specific nucleic acid fragments was used as input data.

[0254]

[0255] 4-2 Building and Training an MLP Model

[0256] An MLP artificial intelligence model was trained to distinguish between healthy people and lung cancer patients using the one-dimensional vectors of the FEMS_Z table and FEMS_BC_Z table as input.

[0257] The dataset in Table 2 was used, and the discovery dataset was divided into training, validation, and test datasets, respectively. The training dataset was used for model learning, the validation dataset was used for hyper-parameter tuning, and the test dataset was used for final model performance evaluation. The external validation data was used for additional performance evaluation on new data independent of model learning.

[0258]

[0259] The activation function used was sigmoid, there was one fully connected layer, and 511 hidden nodes were included. Finally, the final DPI value was calculated using the sigmoid function value.

[0260] The hyper-parameter tuning process is a process of optimizing the values ​​of various parameters (such as the number of dense layers and hidden nodes) that make up the MLP model. Bayesian optimization and grid search techniques were used in the hyper-parameter tuning process. When the validation loss began to increase compared to the training loss, the model was judged to be overfitting and model learning was stopped.

[0261] The performance of several models obtained through hyper-parameter tuning was compared using the validation data set, and the model with the best performance on the validation data set was judged to be the optimal model, and the final performance evaluation was performed using the test data set.

[0262] When a one-dimensional vector of the FEMS_BC_Z table of any sample is input into the model created through the above process, the probability of the sample being healthy and the probability of the sample being a lung cancer patient are calculated through the sigmoid function, which is the last layer of the MLP model, and this probability value is defined as the Deep Probability Index (DPI).

[0263]

[0264] Example 5. Performance verification of a deep learning model built using the FEMS_BC_Z table.

[0265] 5-1 Performance Check

[0266] The performance of the DPI values ​​output from the FEMS_Z table-based model and the FEMS_BC_Z table-based model constructed in Example 4 was tested.

[0267] Using the 85% percentile value of the predicted probability from the Normal value of the Discovery Dataset as a threshold, the model predicted cancer if the calculated DPI value was higher than this threshold. In other words, the model was designed to diagnose cancer patients at a level where 85% of all Normal samples could be predicted as Normal.

[0268] As a result, when checking the performance in the Discovery Dataset as described in Fig. 8, it was confirmed that there was no significant difference before (FEMS_Z table model) / after (FEMS_BC_Z table model) applying Bias Correction.

[0269] However, when measuring performance on an external validation dataset independent of the discovery dataset, as described in Fig. 9, the model showed significantly higher performance after applying bias correction.

[0270] When Bayes Correction is not applied, we confirmed that the performance degradation is severe on the external validation dataset, which is new data that has not been shown to the model during training. In particular, in the case of Specificity, it was confirmed that it was 76.1%, which is a significant drop from the target value of 85%. In other words, we confirmed that if BC is not applied, the model is vulnerable to false positive predictions on new validation data.

[0271] However, after applying Bias Correction, it was confirmed that the performance degradation was not significant even in the external validation dataset, as described in the lower panel of Fig. 10.

[0272]

[0273] 5-2. Checking DPI distribution

[0274] We checked how closely the DPI value, which is the output value of the deep learning model constructed in Example 5-1, matches the actual patient.

[0275] As a result, as described in Fig. 12, in the case of the FEMS_Z table model before applying Bias Correction, the DPI value was calculated higher in the External Validation data set than in the Discovery data set in normal samples, and was calculated the same in cancer samples, confirming that it was vulnerable to False Positive in new data. On the other hand, in the case of the FEMS_BC_Z table model after applying Bias Correction, the values ​​in the Discovery data set and the External Validation data set in normal samples were the same, confirming that Specificity reproduction in new data was excellent, and in the case of cancer samples, the values ​​in the Discovery data set were lower than or equal to the External Validation data set, confirming that Sensitivity reproduction in new data was excellent.

[0276]

[0277] While specific aspects of the present invention have been described in detail above, it will be apparent to those skilled in the art that these specific descriptions merely represent preferred embodiments and are not intended to limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.

[0278]

[0279] The method of improving the detection accuracy of cancer patients by correcting the experimental batch effect in cell-free nucleic acid fragment whole-genome sequencing data according to the present invention is useful because it drastically reduces the false positive effect that may occur due to noise signals appearing in each batch, corrects regional deviations in vectorized data, generates vectorized data that well reflects the characteristics of the sample itself, and analyzes it using an AI algorithm, thereby showing high sensitivity and accuracy compared to existing nucleic acid fragment-based methods.

[0280]

[0281] Electronic file attached.

Claims

1. (a) A step of obtaining sequence information of nucleic acid extracted from a biological sample; (b) a step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) a step of deriving the terminal sequence motif frequency and the size of the nucleic acid fragments using the aligned sequence information (reads); (d) a step of generating vectorized data using the terminal sequence motif frequency of the nucleic acid fragment derived above and the size of the nucleic acid fragment; (e) a step of correcting the batch effect of the vectorized data using the global reference value of the negative control (NC); (f) a step of post-processing the above corrected data; (g) a step of inputting the above post-processed data into a learned artificial intelligence model, comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; and (h) A method for providing information for cancer diagnosis and cancer type prediction, including a step of predicting cancer type by comparing the above output result values. 2.(a) A step of obtaining sequence information of nucleic acid extracted from a biological sample; (b) a step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) a step of deriving the terminal sequence motif frequency and the size of the nucleic acid fragments using the aligned sequence information (reads); (d) a step of generating vectorized data using the terminal sequence motif frequency of the nucleic acid fragment derived above and the size of the nucleic acid fragment; (e) a step of correcting the batch effect of the vectorized data using the global reference value of the negative control (NC); (f) a step of post-processing the above corrected data; (g) a step of inputting the above post-processed data into a learned artificial intelligence model, comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; and (h) A method for diagnosing and predicting cancer types, including a step of predicting cancer types by comparing the output result values.

3. A method according to claim 1 or 2, characterized in that step (a) is performed by a method including the following steps: (ai) a step of obtaining nucleic acid from a biological sample; (a-ii) a step of removing proteins, fats, and other residues from the collected nucleic acid using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid; (a-iii) a step of producing a single-end sequencing or pair-end sequencing library for purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method; (a-iv) a step of reacting the produced library with a next-generation sequencer; and (av) A step for obtaining sequence information (reads) of nucleic acids in a next-generation genetic sequencer.

4. A method according to claim 1 or 2, wherein the terminal sequence motif of step (c) is a pattern of 2 to 30 base sequences at both ends of the nucleic acid fragment.

5. A method according to claim 1 or 2, wherein the terminal sequence motif frequency of step (c) is the number of each motif detected in the entire nucleic acid fragment.

6. A method according to claim 1 or 2, characterized in that the size of the nucleic acid fragment in step (c) is the number of bases from the 5' end to the 3' end of the nucleic acid fragment.

7. A method according to claim 1 or 2, wherein the vectorized data of step (d) is characterized in that the type of nucleic acid fragment terminal sequence motif is represented on the X-axis and the size of the nucleic acid fragment is represented on the Y-axis.

8. A method according to claim 1 or 2, characterized in that step (e) is performed by a method including the following steps: (ei) a step of calculating a global reference value for the entire negative control group; (e-ii) a step of calculating a normalization factor (NF) by comparing the median value of the negative control group within the batch with the global reference value; and (e-iii) a step of obtaining a correction value by multiplying the vectorized data by a correction coefficient; 9. A method according to claim 8, characterized in that the global reference value is the median of vectorized data of the entire negative control group.

10. In the 8th paragraph, the method characterized in that the step (e-ii) is calculated using the following formula: 수식 1: Normalization Factor feature = (Global Reference feature / Batch NCs Median feature ) Here, the feature is characterized by the type of sequence motif at the end of each nucleic acid fragment and the frequency value according to the size of the nucleic acid fragment.

11. A method according to claim 1 or 2, characterized in that step (f) is performed by a method including the following steps: (fi) a step of calculating the average and standard deviation of the frequency values of each nucleic acid fragment terminal sequence motif type and nucleic acid fragment size in a normal group; (f-ii) a step of performing Z-standardization by subtracting the average of the frequency values of the types of terminal sequence motifs of each nucleic acid fragment and the size of each nucleic acid fragment in the normal group from the frequency values of the types of terminal sequence motifs of each nucleic acid fragment and the size of each nucleic acid fragment in the sample, and then dividing the average by the standard deviation of the frequency values of the types of terminal sequence motifs of each nucleic acid fragment and the size of each nucleic acid fragment to derive a Z-standardization value; and (f-iii) If the Z standardization value derived from (f-ii) above exceeds the reference range, a step of correcting it to the reference value.

12. In the 11th paragraph, the reference range is -5 to 5, and the reference value is -5 or 5, characterized in that the method 13. A method according to claim 1 or 2, characterized in that the artificial intelligence model of step (f) learns to distinguish between healthy vectorized data and cancer vectorized data.

14. A method according to claim 13, characterized in that the artificial intelligence model is selected from the group consisting of a convolutional neural network (CNN), a deep neural network (DNN), a multi-layer perceptron (MLP), and a recurrent neural network (RNN).

15. In the 14th paragraph, when the artificial intelligence model is MLP, the method is characterized in that the loss function is expressed by the following equation 2 or equation 3:

16. A method according to claim 1 or 2, characterized in that the result value output by analyzing the input vectorized data by the artificial intelligence model of step (g) is a DPI (Deep Probability Index) value.

17. A method characterized in that, in the first or second paragraph, the reference value of step (g) is 0.5, and if it is 0.5 or higher, it is determined to be cancer.

18. In paragraph 1 or 2, A method characterized in that the step of predicting the cancer type by comparing the output result values of the above step (h) is performed by a method including a step of determining the cancer type showing the highest value among the output result values as the cancer of the sample.

19. A decoding unit that decodes the sequence information of nucleic acids extracted from biological samples; An alignment unit that aligns the decoded sequence to a standard chromosome sequence database; A nucleic acid fragment analysis unit that derives the terminal sequence motif frequency and size of nucleic acid fragments based on aligned sequences; A data generation unit that generates vectorized data using the terminal sequence motif frequency of the derived nucleic acid fragment and the size of the nucleic acid fragment, and then corrects the batch effect of the vectorized data using the reference value of the negative control (NC), and then performs post-processing; A cancer diagnosis unit that inputs the generated post-processed vectorized data into a learned artificial intelligence model for analysis and compares it with a reference value to determine the presence or absence of cancer; and A cancer diagnosis and cancer type prediction device including a cancer type prediction unit that predicts cancer type by analyzing the output result values.

20. A computer-readable storage medium comprising instructions configured to be executed by a processor for diagnosing cancer and predicting cancer types, (a) A step of obtaining sequence information of nucleic acid extracted from a biological sample; (b) a step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) a step of deriving the terminal sequence motif frequency and the size of the nucleic acid fragments using the aligned sequence information (reads); (d) a step of generating vectorized data using the terminal sequence motif frequency of the nucleic acid fragment derived above and the size of the nucleic acid fragment; (e) a step of correcting the batch effect of the vectorized data using the reference value of the negative control (NC); (f) a step of post-processing the above corrected data; (g) a step of inputting the above post-processed data into a learned artificial intelligence model, comparing the analyzed output result value with a cut-off value to determine the presence or absence of cancer; and (h) A computer-readable storage medium for cancer diagnosis and cancer type prediction, comprising instructions configured to be executed by a processor for cancer diagnosis and cancer type prediction, through a step of predicting cancer type by comparing the above output result values.

Citation Information

Patent Citations

  • Colorectal cancer risk assessment model and system

    CN113160895A

  • Single-cell RNA-seq data clustering method based on deep noise reduction auto-encoder

    CN113889192A

  • Pallet tow device

    KR102703206B1

  • Systems and methods for automating RNA expression calls in a cancer prediction pipeline

    US20210272649A1

  • KR20220160806A