Method for distinguishing somatic mutations and artificial mutations using deep learning in DNA sequencing data generated from formalin-fixed paraffin-embedded samples and apparatus using the same
A deep learning model distinguishes somatic and artificial mutations in FFPE samples, addressing the challenge of artificial mutations introduced by formalin fixation, thereby improving cancer treatment predictions and enabling personalized medicine.
Patent Information
- Application Number
- JP2024570446
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-21
- Filing Date
- 2023-08-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-08-29
AI Technical Summary
FFPE samples, commonly used in cancer genome analysis, introduce artificial mutations due to formalin fixation, complicating the distinction between somatic and artificial mutations, leading to inaccuracies in Tumor Mutation Burden measurement and neoantigen discovery, which hampers advanced cancer treatments like immunotherapy and vaccine development.
A deep learning neural network model is employed to analyze DNA sequencing data from FFPE samples, utilizing cosine similarity and other metrics to distinguish between somatic and artificial mutations, thereby improving prediction accuracy.
The model effectively removes artificial mutations while preserving somatic mutations, enhancing the accuracy of cancer treatment predictions and enabling personalized medicine by accurately identifying somatic mutations in FFPE samples.
Smart Images

Figure 2025520108000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for distinguishing somatic mutations and artificial mutations using deep learning with DNA sequencing data generated from formalin-fixed paraffin-embedded samples.
Background Art
[0002] Factors affecting the onset of cancer include germline mutations including genetic predisposition and somatic mutations that occur during somatic cell division limited to specific organs or tissues.
[0003] Somatic mutations include various forms of mutations such as single nucleotide variations, structural variations, and aneuploidy. Specific cancer types are predicted by somatic single nucleotide mutations that occur when exposed to chemical drugs, ultraviolet rays, smoking, etc. and the repair mechanism fails. Large-scale structural variations affect the function of normal genes in cancer due to chromosome deletion, insertion, inversion, tandem duplication, translocation, complex rearrangement, etc.
[0004] Although major genes are involved in each cancer, understanding the mutations of these genes reveals the stage of cancer progression. Moreover, even cancers occurring in the same organ have different combinations of gene mutations (heterogeneity), so different genetic differences can be applied differently to the prognosis, treatment, recurrence, etc. of cancer.
[0005] Since the first analysis of the human genome in 2003, the development and utilization of genome analysis technologies have been diversely expanded. With the development of analysis technologies that can identify various gene mutations in a single test by using Next Generation Sequencing (NGS), cancer genome diagnosis has begun to play an important role in effective cancer treatment. Along with the success of developing targeted anticancer drugs against mutant proteins generated by somatic mutations in major genes such as EGFR and PIK3CA, the clinical demand for cancer genome diagnosis is increasing.
[0006] Ideally, for genome analysis, nucleic acids should be extracted immediately within 10 minutes after the blood supply is interrupted during surgery or frozen with liquid nitrogen, and then [Chemical Formula] stored at extremely low temperatures below and nucleic acids should be extracted when needed to proceed with DNA sequencing. However, it is impossible to apply the above method to the surgical tissues of all patients who are not research subjects in the hospital, and frozen tissues are not the optimal samples for pathological tissue examination. Therefore, all body-derived tissues are microtomed after being fixed with formalin and embedded in paraffin (formalin fixed paraffin embedded; FFPE) to perform pathological tissue examination. FFPE blocks are stored at room temperature for at least 10 years or permanently. The decision on the actual need for cancer genome analysis is made after the results of pathological tissue examination using FFPE tissues are obtained. Due to the size of cancer tissues that can be secured by surgery or biopsy, there is often no surplus tissue that can be cryopreserved after securing the necessary tissue for diagnostic FFPE blocks. Also, many hospitals do not have the infrastructure for cryopreservation. Therefore, most clinical genome analyses come to be carried out using FFPE samples.
[0007] FFPE samples have the advantage of being convenient for storage and optimal for pathological tissue examination, but they have weaknesses that are disadvantageous for genetic analysis. That is, chemical reactions occurring during the formalin fixation process cause artificial mutations (artifacts) unrelated to cancer, and when performing DNA sequencing using NGS techniques, more than 10 times as many mutations are found compared to the original mutations. When preparing a panel targeting only the hot spot driver mutations of some important cancer genes for genetic analysis, the presence or absence of driver mutations can be diagnosed very accurately. However, during exome or genome-level analysis, it is difficult to distinguish between actual mutations and artificial mutations. The presence of artificial mutations leads to inaccuracies in Tumor Mutation Burden measurement and neoantigen discovery, becoming a major obstacle to realizing advanced medicine such as immunotherapy agent response prediction and neoantigen vaccine development. In the meantime, methods for removing artificial mutations from FFPE NGS DNA sequencing data have been developed, but their clinical applicability is fragile because almost all actual mutations are removed at the same time as artificial mutations, and the need for a technology that removes almost all artificial mutations while preserving actual mutations has emerged.
[0008] The background art of the invention has been created to facilitate the understanding of the present invention. It should not be understood that the matters described in the background art of the invention exist as prior art.
Summary of the Invention
Problems to be Solved by the Invention
[0009] The inventors of the present invention recognized that the vast majority of general clinical samples are FFPE tissues, and in the case of old FFPE tissues or samples in which mutations occurred during the fixation process, NGS analysis is difficult due to the problems described above.
[0010] That is, when the inventors of the present invention discriminate somatic mutations using existing FFPE samples, artificial mutations appear due to various factors, resulting in low accuracy. After recognizing the need for a method that can accurately predict only somatic mutations, they noted that when using a deep learning neural network model as in the present invention, the accuracy can be improved.
[0011] In particular, the inventors of the present invention analyzed paired NGS data generated from frozen tissue and FFPE tissue derived from the same tissue, defined mutations found only in FFPE as artificial mutations, and recognized that somatic mutations can be predicted with high accuracy compared to artificial mutations by performing deep learning on learning data including cosine similarity.
[0012] As a result, the inventors of the present invention were able to confirm that the learned neural network model of the present invention can distinguish between artificial mutations and somatic mutations with high accuracy. As a result, the inventors of the present invention have developed an information - providing system for distinguishing artificial mutations and somatic mutations generated in FFPE based on the deep - learning neural network model of the present invention.
[0013] Therefore, the inventors of the present invention recognized that the introduction of the information - providing system for distinguishing artificial mutations and somatic mutations generated in FFPE based on the deep - learning neural network model of the present invention can complement the limitations of conventional diagnostic and information - providing systems, and thereby enable the presentation of an efficient treatment method.
[0014] Therefore, the problem to be solved by the present invention is to provide an information - providing method for predicting somatic mutations through artificial mutation removal and preservation of somatic mutation samples based on the deep learning of the present invention, and a device using the same.
[0015] The problems of the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the following description.
Means for Solving the Problems
[0016] In order to solve the problems as described above, there is provided a method for providing information for predicting somatic mutations based on deep learning according to an embodiment of the present invention. This information providing method is implemented by a processor, and includes steps of receiving data on mutations obtained from medical data and / or biological samples separated from an individual, calculating data necessary for discriminating artificial mutations or somatic mutations based on the received data, and using a neural network model configured to predict artificial mutations or somatic mutations with the data necessary for somatic mutation discrimination as an input, and predicting artificial mutations or somatic mutations for a sample obtained from an individual with the calculated data as an input.
[0017] At this time, the data obtained from medical data and / or biological samples separated from an individual are about mutations, specifically, may be data about somatic mutations. The medical data may be NGS (Next Generation Sequencing) data, specifically, may be WGS (Whole genome sequence), WES (whole exome sequencing), Targeted sequencing files, etc. that have been subjected to base sequence analysis for the entire genome. The biological sample separated from an individual may mean data about a sample in which a mutation has been observed, specifically, a sample in which a somatic mutation has been observed, but is not limited thereto.
[0018] According to the features of the present invention, the step of calculating data necessary for artificial mutation or somatic mutation discrimination based on the received data may further include the step of vectorizing the received data and the step of calculating a similarity based on the vectorized data. At this time, in the step of calculating a similarity based on the vectorized data, the similarity may be, but is not limited to, cosine similarity.
[0019] Further, using a neural network model configured to predict artificial mutation or somatic mutation with data necessary for somatic mutation discrimination as an input, predicting artificial mutation or somatic mutation for a sample obtained from an individual with the calculated data as an input. In this step, the neural network model may be, but is not limited to, a Deep Neural Network (DNN).
[0020] According to the features of the present invention, in the step of calculating data necessary for mutation or artificial mutation discrimination based on the received data, the calculated data includes: the number of reads in which the reference array is observed / the sequence depth at the corresponding position, the number of reads in which the alternate array is observed / the sequence depth at the corresponding position, the ratio of the difference between the number of sequences in which read 1 is mapped to the forward strand and the number of sequences in which read 2 is mapped to the reverse strand among the reads in which the reference array is observed to the reads in which the reference array is observed, the ratio of the difference between the number of sequences in which read 1 is mapped to the forward strand and the number of sequences in which read 2 is mapped to the reverse strand among the reads in which the alternate array is observed to the reads in which the alternate array is observed, the value obtained by dividing the median length of the insertions in the reads in which the alternate array is observed by the median length of the insertions in the reads in which the reference array is observed, the length of the sequences double-sequenced by read 1 and read 2 in the reference fragment, the number of times the sequences are double-sequenced by read 1 and read 2 in the reference fragment, the length of the sequences double-sequenced by read 1 and read 2 in the alternate fragment, the number of times the sequences are double-sequenced by read 1 and read 2 in the alternate fragment, the reference sequence (REF_3_BASES) including one base before and after centered on the mutation position, the alternate sequence (ALT_3_BASES) including one base before and after centered on the mutation position, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference array is observed to the reads in which the reference array is observed, and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the alternate array is observed to the reads in which the alternate array is observed, and can include any one or more selected from the group consisting of (in the reads in which the reference array is observed, read 1 is forwardThe cosine similarity between vectors defined by (the number of sequences mapped to the strand, the number of sequences mapped to the reverse strand by read 2), or between the vector defined by (the number of sequences mapped to the forward strand in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and the vector defined by (the number of sequences mapped to the forward strand in the reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand) can be further included, but is not limited thereto.
[0021] In one embodiment of the present invention, the cosine similarity can be expressed by the following formula.
[0022] [Number]
[0023] According to the features of the present invention, using a neural network model configured to predict artificial mutations or somatic mutations with the data required for somatic mutation discrimination as input, in the step of predicting artificial mutations or somatic mutations for a sample obtained from an individual with the calculated data as input, the sample obtained from the individual may be an FFPE sample, and the individual may be an individual with cancer. At this time, the cancer may be an individual in whom any one of liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma has developed, but is not limited thereto.
[0024] In order to solve the problems as described above, an information providing system including a device for providing information for somatic mutation prediction, a mutation prediction device, and a medical data providing server according to an embodiment of the present invention is provided.
[0025] Specific matters of other embodiments are included in the detailed description and the drawings. [Advantages of the Invention]
[0026] The present invention can provide information on somatic mutation prediction so that personalized treatment (Tailored personal medicine) can be achieved by providing information on prediction results for artificial mutations or somatic mutations using a deep - learned neural network model.
[0027] In particular, the present invention can calculate data related to somatic mutations and more accurately predict artificial mutations or somatic mutations by determining prediction results for artificial mutations or somatic mutations based on the data.
[0028] That is, the present invention can more easily set the treatment direction for individuals and contribute to the selection of personalized treatment methods by utilizing information on the prediction of artificial mutations or somatic mutations based on a deep - learned neural network model.
[0029] The present invention applies a deep - learned neural network model learned with learning data including cosine similarity for the prediction of artificial mutations or somatic mutations, so as to predict artificial mutations and somatic mutations with high accuracy, thus complementing the limitations when using conventional FFPE samples.
[0030] The effects according to the present invention are not limited to the contents exemplified above, and more various effects are included in this specification.
Brief Description of the Drawings
[0031]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5a
Figure 5b
Figure 5c
Figure 5d
Figure 6a
Figure 6b
Figure 6c
Figure 6d
Figure 6e
Figure 7a
Figure 7b
Figure 7c
Figure 7d
Figure 7e
Mode for Carrying Out the Invention
[0032] The advantages and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, but is embodied in a variety of different forms. Merely, these embodiments are provided so that the disclosure of the present invention is complete, and to fully inform those with ordinary knowledge in the technical field to which the present invention pertains of the scope of the invention. The present invention is only defined by the scope of the claims. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0033] In this document, expressions such as "have", "be able to have", "include", or "be able to include" refer to the presence of the corresponding features (e.g., components such as numerical values, functions, operations, or parts), and do not exclude the presence of further features.
[0034] In this document, expressions such as "A or B", "at least one of A or / and B", or "one or more of A or / and B" can all include all possible combinations of the items listed together. For example, "A or B", "at least one of A and B", or "at least one of A or B" can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0035] In this document, expressions such as "first", "second", "the first", or "the second" can modify various components regardless of order and / or importance, and are only used to distinguish one component from another without limiting the corresponding component. For example, the first user device and the second user device can indicate different user devices regardless of order or importance. For example, in order not to deviate from the scope of rights described in this document, the first component can be named the second component, and similarly, the second component can also be named the first component.
[0036] When a certain component (e.g., the first component) is referred to as being "(functionally or communicatively) coupled with / to" or "connected to" another component (e.g., the second component), it should be understood that the certain component can be directly coupled to the other component or can be coupled through another component (e.g., the third component). In contrast, when a certain component (e.g., the first component) is referred to as being "directly coupled" or "directly connected" to another component (e.g., the second component), it can be understood that there is no other component (e.g., the third component) between the certain component and the other component.
[0037] As used herein, the phrase "configured to" may, depending on the context, be used interchangeably with, for example, "suitable for", "having the capacity to", "designed to", "adapted to", "made to", or "capable of". The term "configured to" does not necessarily mean only something that is "specifically designed to" in hardware. Instead, in some situations, the expression "a device configured to" may mean that the device is "capable of" together with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing the corresponding operations, or a general-purpose processor (e.g., a CPU, GPU, or application processor) capable of performing the corresponding operations by executing one or more software programs stored in a memory device.
[0038] The terms used herein are merely used to describe specific embodiments and may not be intended to limit the scope of other embodiments. Singular expressions may include plural expressions unless the context clearly dictates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by one of ordinary skill in the art to which the technology described in this document pertains. Terms defined in a general dictionary among the terms used in this document may be interpreted to have the same or similar meaning as their meaning in the context of the related art, and will not be interpreted in an ideal or overly formal sense unless clearly defined herein. In some cases, even terms defined in this document may not be interpreted to exclude embodiments of this document.
[0039] The features of each of the various embodiments of the present invention can be partially or wholly combined or combined with each other, and various technical linkages and drives are possible as can be fully understood by those skilled in the art. Each embodiment may be implemented independently of each other or may be implemented together in a related relationship.
[0040] For the sake of clarity of interpretation of this specification, the terms used in this specification are defined below. As used herein, the term "individual" can mean all subjects for which it is desired to predict the presence or absence of mutations, specifically somatic mutations. For example, an individual may be an individual in which a somatic mutation has occurred. At this time, the individuals disclosed in this specification may be all mammals except humans, but are not limited thereto.
[0041] As used herein, the term "biological sample" refers to all biological samples that are separated from an individual and can be stored in FFPE form, and may include, but are not limited to, urine, tissue, cell lysate, whole blood, plasma, serum, saliva, ocular fluid, cerebrospinal fluid, sweat, milk, ascites fluid, synovial fluid, peritoneal fluid, etc. Preferably, the biological sample may be a tissue sample for which it is desired to confirm the presence or absence of somatic mutations separated from an individual, but is not limited thereto. Specifically, the biological sample of the present invention may be an FFPE sample (Formalin Fixed Paraffin Embedded), and the FFPE sample can be stored at room temperature by fixing the tissue with formalin and solidifying it in paraffin to produce a paraffin block, which is a standard method used when collecting and storing a patient's cancer tissue. It can mean a sample that is thinly cut from the paraffin block several times and used for various histological examinations.
[0042] As used herein, the term "somatic variant" means that one of the characteristics of cancer identified in cells is accompanied by genetic abnormalities, so it is possible to identify somatic variants to determine the occurrence and type classification of cancer. Therefore, in the present invention, the medical data and / or data obtained from biological samples isolated from an individual may be data about somatic mutations. The medical data may be NGS (Next Generation Sequencing) data, but is not limited thereto, and may mean a biological sample isolated from an individual in which somatic mutations are observed, but is not limited thereto.
[0043] As used herein, the term "artifact" may mean an artificial defect or the like that appears due to denaturation caused by a physical or chemical action and has no relation to cancer, but is not limited thereto.
[0044] As used herein, the term "NGS (Next Generation Sequencing)" is next-generation base sequence analysis and is one of the high-speed analysis methods for the base sequences of genomes.
[0045] At this time, NGS can be applied to genomes and whole-genome analysis to achieve various purposes including clinical research. On the other hand, most NGS analyses can be performed for simultaneous analysis of a large number of target samples.
[0046] As used herein, the term "NGS data" may mean a file containing base sequence data analyzed for a target sample. At this time, the NGS data can include a WGS (whole genome sequencing) file, a WES (whole exome sequencing) file, an RNA sequencing file, and a targeted sequencing file that has been sequenced for a specific region with respect to the entire genetic material. Specifically, it can be a WES (whole exome sequencing) file, but is not limited thereto.
[0047] Hereinafter, with reference to FIGS. 1 and 2, an information providing system for artificial mutation or somatic mutation prediction and an artificial mutation or somatic mutation prediction apparatus based on the artificial mutation or somatic mutation prediction apparatus according to an embodiment of the present invention will be described.
[0048] FIG. 1 shows an information providing system for artificial mutation or somatic mutation prediction using the artificial mutation or somatic mutation prediction apparatus according to an embodiment of the present invention. FIG. 2 exemplarily shows the configuration of the artificial mutation or somatic mutation prediction apparatus according to an embodiment of the present invention.
[0049] First, referring to FIG. 1, the information providing system 1000 may be a system configured to provide information related to the prediction of artificial mutations or somatic mutations by applying a deep learning neural network model based on medical data and / or data obtained from biological samples. At this time, the information providing system 1000 may include a data providing server / device 100 that provides data obtained from medical data and / or biological samples, a prediction device that predicts artificial mutations or somatic mutations based on the received data, and an information providing device 300 that visually provides information on the prediction of artificial mutations or somatic mutations for samples obtained from an individual based on the received prediction data.
[0050] Here, the information-providing device 300 is an electronic device that provides a user interface for presenting information related to the prediction of artificial mutations or somatic mutations, and can include at least one of a smartphone, a tablet PC (Personal Computer), a notebook computer, and / or a PC, etc.
[0051] For the information-providing device 300 for the prediction of somatic mutations, it can receive information related to the prediction results of artificial mutations or somatic mutations from the somatic mutation prediction device 200, and display the received results through a display unit described later.
[0052] The somatic mutation prediction device 200 can include a general-purpose computer, a laptop, and / or a data server, etc., which perform various operations for determining information related to the prediction of artificial mutations or somatic mutations based on various data (especially medical data and / or data about somatic mutations) that can be obtained from the data-providing server / device 100. At this time, the data-providing server 100 can be, but is not limited to, a device for accessing a web server that provides a web page or a mobile web server that provides a mobile web site.
[0053] More specifically, the somatic mutation prediction device 200 can receive medical data and / or data about somatic mutations from the medical data-providing server 100, predict artificial mutations or somatic mutations based on the received data, and provide information related thereto. At this time, the somatic mutation prediction device 200 can apply a deep-learned neural network model to perform the prediction of artificial mutations or somatic mutations based on the medical data received from the data-providing server 100 and / or the data about somatic mutations obtained from a biological sample.
[0054] Also, the somatic mutation prediction device 200 can provide the artificial mutation or somatic mutation prediction results to the information-providing device 300 for the prediction of somatic mutations. In this way, the information provided by the somatic mutation prediction device 200 can be provided to a web page through a web browser provided in the information providing device 300 for somatic mutation prediction, or in the form of an application or a program. In various embodiments, such data can be provided in a form included in a platform in a client-server environment.
[0055] In the following, it will be described that the somatic mutation prediction device 200 according to an embodiment of the present invention receives data obtained from medical data and / or biological samples from the data providing server 100 and performs operations. However, the information providing device 300 itself for somatic prediction can also perform all operations.
[0056] Next, with reference to FIG. 2, the components of the somatic mutation prediction device 200 of the present invention will be specifically described. Referring to FIG. 2, the somatic mutation prediction device 200 can include a communication interface 210, a memory 220, an I / O interface 230, and a processor 240, and each component can communicate with each other through one or more communication buses or signal lines.
[0057] The communication interface 210 can be connected to the information providing device 300 for somatic mutation prediction and the medical data providing server 100 through a wired / wireless communication network to exchange data. For example, the communication interface 210 can receive data about somatic mutations and / or individual data from the medical data providing server 100, and can transmit information about the determined artificial mutations or somatic mutation prediction results to the information providing device 300 for somatic mutation prediction.
[0058] On the one hand, the communication interface 210 that enables such data transmission and reception includes a wired communication port 211 and a wireless circuit 212. Here, the wired communication port 211 can include one or more wired interfaces, such as Ethernet (registered trademark), Universal Serial Bus (USB), FireWire, etc. Also, the wireless circuit 212 can transmit and receive data with an external device through an RF signal or an optical signal. Additionally, the wireless communication can use at least one of a plurality of communication standards, protocols, and technologies, such as GSM (registered trademark), EDGE, CDMA, TDMA, Bluetooth (registered trademark), Wi-Fi (registered trademark), VoIP, Wi-MAX (registered trademark), or any other suitable communication protocol.
[0059] The memory 220 can store various data used in the somatic mutation prediction device 200. For example, the memory 220 can store data about somatic mutations or store a deep learning model configured to predict artificial mutations or somatic mutations based on this.
[0060] In various embodiments, the memory 220 can include a volatile or non-volatile recording medium capable of storing various data, instructions, and information. For example, the memory 220 can include at least one type of storage medium such as a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (such as an SD or XD memory, etc.), RAM, SRAM, ROM, EEPROM, PROM, network storage, cloud, blockchain database.
[0061] In various embodiments, the memory 220 can store at least one configuration of the operation system 221, the communication module 222, the user interface module 223, and one or more applications 224.
[0062] The operating system 221 (for example, built-in operating systems such as LINUX (registered trademark), UNIX (registered trademark), MAC OS, WINDOWS (registered trademark), VxWorks (registered trademark), etc.) can include various software components and drivers for controlling and managing general system operations (for example, memory management, storage device control, power management, etc.), and can assist in communication between various hardware, firmware, and software components.
[0063] The communication module 223 can assist in communicating with other devices through the communication interface 210. The communication module 220 can include various software components for processing data received by the wired communication port 211 or the wireless circuit 212 of the communication interface 210.
[0064] The user interface module 223 can receive user requests or inputs from a keyboard, touch screen, microphone, etc. through the I / O interface 230 and provide a user interface on the display.
[0065] The application 224 can include a program or module configured to be executed by one or more processors 230. Here, an application for providing information related to somatic mutation prediction can be implemented on a server farm.
[0066] The I / O interface 230 can connect at least one of the input / output devices (not shown) of the somatic mutation prediction device 200, such as a display, keyboard, touch screen, and microphone, to the user interface module 223. The I / O interface 230 can receive user inputs (for example, voice input, keyboard input, touch input, etc.) together with the user interface module 223 and process commands based on the received inputs.
[0067] The processor 240 is connected to the communication interface 210, the memory 220, and the I / O interface 230 to control the overall operation of the somatic mutation prediction device 200, and can execute various instructions for information provision through an application or program stored in the memory 220.
[0068] The processor 240 may correspond to a computing device such as a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), or an AP (Application Processor). Further, the processor 240 may be embodied in the form of an integrated chip (Integrated Chip (IC)) such as a SoC (System on Chip) in which various computing devices are integrated. Or, the processor 240 may include a module for calculating an artificial neural network model, such as an NPU (Neural Processing Unit).
[0069] In various embodiments, the processor 240 may be configured to apply a deep - learned neural network model to calculate data necessary for artificial mutation or somatic mutation discrimination based on medical data and / or data obtained from biological samples, and predict and provide artificial mutations or somatic mutations.
[0070] In one embodiment, the somatic mutation prediction device 200 receives medical data and / or data obtained from a biological sample isolated from an individual from the medical data providing device 100, and uses the received data to apply a deep - learned neural network model to provide a prediction result for somatic mutations using a model trained to predict the presence or absence of artificial mutations or somatic mutations. At this time, the medical data and / or data obtained from a biological sample isolated from an individual may be data about somatic mutations. The medical data may be NGS (Next Generation Sequencing) data, but is not limited thereto, and may mean data about a biological sample isolated from an individual in which somatic mutations have been observed, but is not limited thereto. Specifically, the NGS data may be at least one of a WGS file, a WES file, an RNA sequencing file, and a target sequencing file. Specifically, it may be a WES file, but is not limited thereto.
[0071] Hereinafter, with reference to FIGS. 3 and 4, a method for providing information for predicting artificial mutations or somatic mutations according to an embodiment of the present invention will be specifically described. FIG. 3 is a block diagram of the procedure of a method for providing information for predicting artificial mutations or somatic mutations according to an embodiment of the present invention.
[0072] FIG. 4 illustratively shows the procedure of a method for providing information for artificial mutations or somatic mutations according to various embodiments of the present invention. First, referring to FIG. 3, the procedure of a method for providing information for predicting artificial mutations or somatic mutations according to an embodiment of the present invention is as follows.
[0073] First, data about mutations that have occurred in medical data and / or a biological sample isolated from an individual is received (S310). Thereafter, data necessary for discriminating artificial mutations or somatic mutations is calculated based on the received data (S320). Using a neural network model configured to predict artificial mutations or somatic mutations with the data required for somatic mutation discrimination as input, artificial mutations or somatic mutations for a sample obtained from an individual are predicted using the calculated data as input (S330).
[0074] According to the features of the present invention, in the step (S310) of receiving data on mutations that occurred in medical data and / or biological samples isolated from an individual, a WES (whole exome sequencing) file, an RNA sequencing file, or a targeted sequencing file may be received. Specifically, it may be a WES file, but is not limited thereto.
[0075] At this time, the WES file, the RNA sequencing file, and the targeted sequencing file may have the format of a BAM file, but the format of the NGS data is not limited thereto.
[0076] Next, the step (S320) of calculating the data required for discriminating artificial mutations or somatic mutations based on the received data may further include the step of vectorizing the received data and the step of calculating a similarity based on the vectorized data.
[0077] At this time, referring to both Table 1 and FIG. 4, the calculated data in the step of calculating the data necessary for mutation or artificial mutation discrimination based on the received data includes: the number of reads (reads) in which the reference array is observed / the sequence depth at the corresponding position, the number of reads in which the alternate array is observed / the sequence depth at the corresponding position, the ratio of the difference between the number of arrays in which read 1 is mapped to the forward strand and the number of arrays in which read 2 is mapped to the reverse strand among the reads in which the reference array is observed to the reads in which the reference array is observed, the ratio of the difference between the number of arrays in which read 1 is mapped to the forward strand and the number of arrays in which read 2 is mapped to the reverse strand among the reads in which the alternate array is observed to the reads in which the alternate array is observed, the value obtained by dividing the median of the length inserted in the reads in which the alternate array is observed by the median of the length inserted in the reads in which the reference array is observed, the length of the sequences (double-sequenced bases) read twice by read 1 and read 2 in the reference fragment, the number of times the sequences are read twice by read 1 and read 2 in the reference fragment, the length of the sequences (double-sequenced bases) read twice by read 1 and read 2 in the alternate fragment, the number of times the sequences are read twice by read 1 and read 2 in the alternate fragment, the reference sequence (REF_3_BASES) including one base before and after centered on the mutation position, the alternate sequence (ALT_3_BASES) including one base before and after centered on the mutation position, the ratio of the difference between the number of arrays mapped to the forward strand and the number of arrays mapped to the reverse strand among the reads in which the reference array is observed to the reads in which the reference array is observed, and the ratio of the difference between the number of arrays mapped to the forward strand and the number of arrays mapped to the reverse strand among the reads in which the alternate array is observed to the reads in which the alternate array is observed, and can include any one or more selected from the group consisting of (among the reads in which the reference array is observed, read 1 is forwardThe cosine similarity between the vectors defined as (the number of sequences mapped to the strand, the number of sequences where read 2 is mapped to the reverse strand), or the cosine similarity between the vector defined as (the number of sequences mapped to the forward strand in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and the vector defined as (the number of sequences mapped to the forward strand in the reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand) can be further included, but is not limited thereto.
[0078]
Table 1-1
[0079]
Table 1-2
[0080] As used herein, the term "read" may mean DNA including an adapter, and the term "fragment" may mean DNA excluding an adapter.
[0081] As used herein, the term "read1" means the one read in the 5' to 3' direction when the Paired End DNA entering the NGS equipment is read from both ends, and the term "read2" means the one read in the 3' to 5' direction.
[0082] As used herein, the term "depth" may be used interchangeably with the term "read-depth" and means the thickness or depth of the reads.
[0083] As used herein, the term "fr_ref" means the number of sequences in which read 1 maps to the forward strand in the reads where the reference sequence is observed. As used herein, the term "rf_ref" means the number of sequences in which read 2 maps to the reverse strand.
[0084] As used herein, the term "fr_alt" means the number of sequences in which read 1 maps to the forward strand in the reads where the alternative sequence is observed. As used herein, the term "rf_alt" means the number of sequences in which read 2 maps to the reverse strand.
[0085] As used herein, the term "fs_ref" means the number of sequences in the forward strand in the reads where the reference sequence is observed. As used herein, the term "rs_ref" means the number of sequences that map to the reverse strand.
[0086] As used herein, the term "fs_alt" means the number of sequences in the forward strand in the reads where the alternative sequence is observed. As used herein, the term "rs_alt" means the number of sequences that map to the reverse strand.
[0087] At this time, the model into which data on mutations that have occurred in medical data and / or biological samples isolated from an individual is input can be based on various algorithms such as a CNN (Convolution Neural Network) model, DNN (Deep Neural Network), DCNN (Deep Convolution Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), SSD (Single Shot Detector), and SVM (Support Vector Machine). Specifically, it may be a CNN (Convolution Neural Network) model or a DNN (Deep Neural Network) model. More specifically, it may be a DNN (Deep Neural Network) model, but it is not limited thereto.
[0088] Thereafter, using a neural network model configured to predict artificial mutations or somatic mutations with the data required for somatic mutation discrimination as input, and using the calculated data as input, in step (S330) of predicting artificial mutations or somatic mutations for a sample obtained from an individual, a prediction result for the artificial mutations or somatic mutations for the sample obtained from the individual is determined and provided. For example, the model can be configured to output whether the result value is an actual artificial mutation, a false artificial mutation, an actual somatic mutation, or a false somatic mutation. For example, the corresponding output can be shown as a percentage, and by setting a threshold value, it can be determined whether the case with a percentage equal to or higher than the threshold value is an actual artificial mutation, a false artificial mutation, an actual somatic mutation, or a false somatic mutation.
[0089] At this time, the sample obtained from the individual may be an FFPE sample, and the individual may be an individual in whom any one of liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma has developed, but it is not limited thereto.
[0090] On the other hand, the information providing procedure for predicting artificial mutations or somatic mutations is not limited to those described above. According to the above information providing method for predicting artificial mutations or somatic mutations, the present invention provides an information providing system for predicting artificial mutations or somatic mutations based on an algorithm, so that the setting of the treatment direction for an individual can be more easily carried out and contribute to the selection of an individualized treatment method.
[0091] That is, the present invention can complement the limitations of conventional diagnostic and information providing systems by predicting artificial mutations or somatic mutations with high accuracy by introducing an information providing system for predicting artificial mutations or somatic mutations based on a neural network model, and the setting of the treatment direction for an individual can be more easily carried out.
Example
[0092] Example 1: Prediction rate of artificial mutations or somatic mutations in FFPE samples For the generation of a trainable dataset, cancer tissue FFPE (Formalin-fixed paraffin-embedded) samples were collected, and mutations that also appeared in FF (Frozen Fresh) from 112,976 mutation calls were predicted.
[0093] As a result, there were 2,113 mutations that appeared in both FF samples and FFPE samples, and 110,863 mutations that appeared only in FFPE samples. It was confirmed that only about 1.8% showed somatic mutations that were not artificial mutations.
[0094] Thereafter, after calculating features highly related to somatic prediction using the obtained samples and the deep learning neural network model, the performance for predicting somatic mutations was evaluated. Example 2: Performance evaluation of artificial mutation or somatic mutation prediction of a deep learning neural network model Hereinafter, with reference to FIG. 5, the performance evaluation results of the deep - learned neural network model according to an embodiment of the present invention for predicting somatic mutations will be described.
[0095] FIGS. 5a to 5d are diagrams showing the results of evaluating the performance of the deep - learned neural network model according to an embodiment of the present invention and a comparative example. First, for the evaluation of the deep - learned neural network model according to an embodiment of the present invention, models that utilize various somatic mutation - related data were constructed as Comparative Examples 1 to 3. More specifically, the FFpolish (pre - trained) model was used as Comparative Example 1, the FFpolish (re - trained) model was used as Comparative Example 2, and the SOBDetector model was used as Comparative Example 3 for a comparative experiment with the neural network model (DEEPOMICS(R) FFPE) of the present invention.
[0096] Precision, recall, F1 score, accuracy, and specificity were calculated as follows.
[0097]
Equation
[0098] As a result, as shown in FIG. 5a, it was confirmed that the neural network model (DEEPOMICS (registered trademark) FFPE) according to an embodiment of the present invention removes 99.6% of artificial mutations and shows a preservation ability of 77.8% for actual somatic mutations.
[0099] On the other hand, as shown in FIG. 5b, it was confirmed that Comparative Example 1 (FFpolish (pre - trained)) removes 92.6% of artificial mutations and preserves 69.5% of actual somatic mutation samples, and it was confirmed that the ability to remove artificial mutations and the ability to preserve actual somatic mutation samples are lower than those of the neural network model according to an embodiment of the present invention.
[0100] As shown in Fig. 5c, in the case of Comparative Example 2 (FFpolish (re-trained)), the artificial mutation removal rate was 99.6%, indicating a high removal ability. However, it was confirmed that the actual somatic mutation preservation rate was 39.8%. It was confirmed that in the case of the model of Comparative Example 2, when removing artificial mutations, samples with actual somatic mutations also were removed together, affecting the prediction results.
[0101] As shown in Fig. 5d, in the case of Comparative Example 3 (SOBDetector), the artificial mutation removal rate was 72.1%, and the actual somatic mutation preservation rate was 84.5%, confirming that the artificial mutation removal ability was significantly low.
[0102] When using the neural network model according to an embodiment of the present invention, the F1 score was also 0.821, which was confirmed to be superior to other comparative examples. With the above results, it was confirmed that while having excellent artificial mutation removal ability, samples with actual somatic mutations could be preserved without being removed.
[0103] Evaluation 2: Measurement of the accuracy of artificial mutation or somatic mutation prediction of algorithms by cancer type Hereinafter, with reference to Figs. 6a to e and Figs. 7a to e, the performance evaluation results of the somatic mutation prediction of the neural network model according to an embodiment of the present invention in various cancer types will be described.
[0104] Figs. 6a to e compare and show the artificial mutation or somatic mutation prediction accuracies for liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma using the neural network model according to an embodiment of the present invention.
[0105] Figs. 7a to e compare and show the artificial mutation or somatic mutation prediction accuracies for liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma using the FFPolish model of Comparative Example 1 among the comparative examples of Evaluation 1 as a comparative example.
[0106] As a result, as shown in FIGS. 6a to 6e, when using the neural network model according to an embodiment of the present invention, the rates are 89.9%, 88.2%, 99.8%, 96.5% and 89.7% for liver cancer, colorectal cancer, breast cancer, lung cancer and fibroma samples respectively, indicating excellent artificial mutation removal ability. The somatic mutation preservation also appears as 88.5%, 89.3%, 63.9%, 76.4% and 90.2% respectively. It was confirmed that it is excellent in artificial mutation removal ability and at the same time excellent in the ability to preserve actual somatic mutation samples.
[0107] In contrast, in the case of the FFpolish model of the comparative example, as shown in FIGS. 7a to 7e, the rates are 85.3% for liver cancer, 63.4% for colorectal cancer, 92.8% for breast cancer, 91.2% for lung cancer and 71.8% for fibroma. It can be confirmed that the artificial mutation removal ability is lower than that of the model of the present invention in all samples. The somatic mutation sample preservation appears as 69.1% for liver cancer, 78.9% for colorectal cancer, 60.6% for breast cancer, 74.3% for lung cancer and 70.7% for fibroma respectively. It was confirmed that the probability of removing actual somatic mutation samples together during artificial mutation removal is higher than that of the neural network model according to an embodiment of the present invention.
[0108] As described above, the embodiments of the present invention have been described in more detail with reference to the accompanying drawings. However, the present invention is not necessarily limited to such embodiments, and can be variously modified within the scope not departing from the technical idea of the present invention. Therefore, the embodiments disclosed in the present invention are not for limiting the technical idea of the present invention, but for explaining it, and the scope of the technical idea of the present invention is not limited by such embodiments. Therefore, it should be understood that the embodiments described above are exemplary in all respects and not restrictive. The protection scope of the present invention should be interpreted by the following claims, and all technical ideas within the equivalent scope should be interpreted as being included in the scope of rights of the present invention.
Claims
**Claim 1** A method for providing information for predicting artificial mutations or somatic mutations implemented by a processor, comprising: receiving data obtained from medical data or data obtained from a biological sample isolated from an individual; calculating data necessary for discriminating artificial mutations or somatic mutations based on the received data; and using a neural network model configured to predict artificial mutations or somatic mutations with data necessary for discriminating somatic mutations as input, and predicting artificial mutations or somatic mutations for a sample obtained from an individual using the calculated data as input. A method for providing information for predicting artificial mutations or somatic mutations. **Claim 2** The method for providing information for predicting artificial mutations or somatic mutations according to claim 1, wherein the medical data is NGS (Next Generation Sequencing) data for somatic mutations. **Claim 3** The method for providing information for predicting artificial mutations or somatic mutations according to claim 2, wherein the NGS data is any one of a WGS (whole genome sequencing) file, a WES (whole exome sequencing) file, an RNA sequencing file, and a targeted sequencing file obtained by performing base sequence analysis on a specific region, which are subjected to base sequence analysis on the entire genome. **Claim 4** In the step of calculating data necessary for discriminating mutations or artificial mutations based on the received data, the calculated data is The number of reads in which the reference array is observed / the sequence depth at the corresponding position, the number of reads in which the alternative array is observed / the sequence depth at the corresponding position, the ratio of the difference between the number of arrays in which read 1 is mapped to the forward strand and the number of arrays in which read 2 is mapped to the reverse strand among the reads in which the reference array is observed to the reads in which the reference array is observed, the ratio of the difference between the number of arrays in which read 1 is mapped to the forward strand and the number of arrays in which read 2 is mapped to the reverse strand among the reads in which the alternative array is observed to the reads in which the alternative array is observed, the value obtained by dividing the median length of the insertions in the reads in which the alternative array is observed by the median length of the insertions in the reads in which the reference array is observed, the length of the double-sequenced bases read twice by read 1 and read 2 in the reference fragment, the number of times the double-sequenced bases are read twice by read 1 and read 2 in the reference fragment, the length of the double-sequenced bases read twice by read 1 and read 2 in the alternative fragment, the number of times the double-sequenced bases are read twice by read 1 and read 2 in the alternative fragment, the reference sequence (REF_3_BASES) including one base before and after centered on the mutation position, the alternative sequence (ALT_3_BASES) including one base before and after centered on the mutation position, the ratio of the difference between the number of arrays mapped to the forward strand and the number of arrays mapped to the reverse strand among the reads in which the reference array is observed to the reads in which the reference array is observed, the ratio of the difference between the number of arrays mapped to the forward strand and the number of arrays mapped to the reverse strand among the reads in which the alternative array is observed to the reads in which the alternative array is observed, and includes any one or more selected from the group consisting of: A method for providing information for predicting an artificial mutation or a somatic mutation according to claim 1.
5. The step of calculating data necessary for discriminating a mutation or an artificial mutation based on the received data is The method for providing information for predicting artificial mutations or somatic mutations according to claim 1, which includes the step of vectorizing the received data.
6. The step of calculating data necessary for mutation or artificial mutation discrimination based on the received data The method for providing information for predicting artificial mutations or somatic mutations according to claim 5, which further includes the step of calculating a similarity based on the vectorized data.
7. In the step of calculating data necessary for mutation or artificial mutation discrimination based on the received data, the calculated data is the cosine similarity between vectors defined as (the number of sequences mapped to the forward strand in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand in read 2), or (the number of sequences mapped to the forward strand in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and the cosine similarity between vectors defined as (the number of sequences mapped to the forward strand in the reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand), and further includes this. The method for providing information for predicting artificial mutations or somatic mutations according to claim 4.
8. In the step of determining a prediction for the presence of artificial mutations or somatic mutations for a sample obtained from an individual based on the calculated data The method for providing information for predicting artificial mutations or somatic mutations according to claim 1, wherein the sample obtained from the individual is an FFPE sample.
9. In the step of determining a prediction for the presence of artificial mutations or somatic mutations for a sample obtained from an individual based on the calculated data The individual is characterized in that it is an individual with cancer. The method for providing information for predicting artificial mutations or somatic mutations according to claim 1.
10. The method for providing information for predicting artificial mutations or somatic mutations according to claim 9, wherein the cancer is any one of liver cancer, colon cancer, breast cancer, lung cancer, and fibroma.
11. A communication unit that receives data on somatic mutations obtained from medical data or biological samples from a server for providing data, and including a processor electrically connected to the communication unit, the processor calculates data necessary for artificial mutation or somatic mutation discrimination based on the received data, using a neural network model configured to predict artificial mutations or somatic mutations with the data necessary for somatic mutation discrimination as input, and configured to predict artificial mutations or somatic mutations for a sample obtained from an individual with the calculated data as input, A device for providing information on the prediction of artificial mutations or somatic mutations.
12. The medical data is of NGS (Next Generation Sequencing) data for somatic mutations, The device for providing information on the prediction of artificial mutations or somatic mutations according to claim 11.
13. The NGS data is any one of a WGS (whole genome sequencing) file, a WES (whole exome sequencing) file, an RNA sequencing (RNA sequencing) file, and a targeted sequencing (targeted sequencing) in which the nucleotide sequence has been analyzed for a specific region, which have been subjected to nucleotide sequence analysis for the entire genome, The device for providing information on the prediction of artificial mutations or somatic mutations according to claim 12.
14. Data necessary for mutation or artificial mutation discrimination based on the received data includes the number of reads in which the reference sequence is observed / the sequence depth at the corresponding position, the number of reads in which the alternative sequence is observed / the sequence depth at the corresponding position, the ratio of the difference between the number of sequences in which read 1 is mapped to the forward strand and the number of sequences in which read 2 is mapped to the reverse strand among the reads in which the reference sequence is observed to the reads in which the reference sequence is observed, the ratio of the difference between the number of sequences in which read 1 is mapped to the forward strand and the number of sequences in which read 2 is mapped to the reverse strand among the reads in which the alternative sequence is observed to the reads in which the alternative sequence is observed, the value obtained by dividing the median length of the insertions in the reads in which the alternative sequence is observed by the median length of the insertions in the reads in which the reference sequence is observed, the length of the sequences (double-sequenced bases) read twice by read 1 and read 2 in the reference fragment, the number of times the sequences are read twice by read 1 and read 2 in the reference fragment, the length of the sequences (double-sequenced bases) read twice by read 1 and read 2 in the alternative fragment, the number of times the sequences are read twice by read 1 and read 2 in the alternative fragment, the reference sequence (REF_3_BASES) including one base before and after centered on the mutation position, the alternative sequence (ALT_3_BASES) including one base before and after centered on the mutation position, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference sequence is observed to the reads in which the reference sequence is observed, and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the alternative sequence is observed to the reads in which the alternative sequence is observed, and includes any one or more selected from the group consisting of these. A device for providing information for predicting artificial mutations or somatic mutations according to claim 11.
15. The processor is further configured to vectorize the received data The device for providing information for predicting artificial mutations or somatic mutations according to claim 11
16. The processor The device for providing information for predicting artificial mutations or somatic mutations according to claim 15, further configured to calculate a similarity based on the vectorized data
17. The data necessary for mutation or artificial mutation discrimination based on the received data is the cosine similarity between vectors defined as (the number of sequences mapped to the forward strand by read 1 in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand by read 2), or the cosine similarity between a vector defined as (the number of sequences mapped to the forward strand in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and a vector defined as (the number of sequences mapped to the forward strand in the reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand), and further includes The device for providing information for predicting artificial mutations or somatic mutations according to claim 14
18. The device for providing information for predicting artificial mutations or somatic mutations according to claim 11, characterized in that the sample obtained from the individual is an FFPE sample
19. The individual is characterized in that it is an individual with cancer The device for providing information for predicting artificial mutations or somatic mutations according to claim 11
20. The device for providing information for predicting artificial mutations or somatic mutations according to claim 19, wherein the cancer is any one of liver cancer, colon cancer, breast cancer, lung cancer, and fibroma
Citation Information
Patent Citations
Machine learning system and method for somatic mutation discovery
US20190189242A1