A method for distinguishing between somatic and artificial mutations using deep learning with DNA sequencing data generated from formalin-fixed, paraffin-embedded samples, and an apparatus utilizing the same.
A deep learning-based neural network model accurately distinguishes somatic mutations from artificial mutations in FFPE samples, enhancing the precision of somatic mutation prediction and enabling personalized treatment strategies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2026-04-08
AI Technical Summary
FFPE samples introduce artificial mutations during formalin fixation, making it difficult to accurately distinguish somatic mutations using NGS techniques, leading to inaccuracies in tumor mutation burden measurement and hindering advanced medical treatments like immunotherapy response prediction and new antigen vaccine development.
A deep learning-based neural network model is employed to analyze paired NGS data from frozen and FFPE tissues, utilizing cosine similarity to differentiate between artificial and somatic mutations, thereby improving accuracy in somatic mutation prediction.
The model enables high-accuracy distinction between artificial and somatic mutations, facilitating personalized treatment methods by providing accurate somatic mutation information, thus complementing the limitations of conventional diagnostic systems.
Smart Images

Figure 0007842487000005 
Figure 0007842487000006 
Figure 0007842487000007
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for distinguishing somatic mutations and artificial mutations using deep learning with DNA sequencing data generated from formalin-fixed paraffin-embedded samples.
Background Art
[0002] Factors affecting the onset of cancer include germline mutations (including genetic tendencies) and somatic mutations that occur during somatic cell division limited to specific organs or tissues.
[0003] Somatic mutations include various forms of mutations such as single nucleotide variations, structural variations, and aneuploidy. Specific cancer types are predicted by somatic single nucleotide mutations that occur when exposed to chemical drugs, ultraviolet rays, smoking, etc. and the repair mechanism fails. Large-scale structural variations affect the function of normal genes in cancer due to chromosome deletions, insertions, inversions, tandem duplications, translocations, complex rearrangements, etc. [[ID=第十八]]
[0004] [[ID=第十九]] Although major genes are involved in each cancer tumor, understanding the mutations of these genes reveals the stage of cancer progression. Not only that, but even cancers that occur in the same organ have different combinations of gene mutations (heterogeneity), so different genetic differences can be applied differently to the prognosis, treatment, recurrence, etc. of cancer.
[0005] Since the first human genetic analysis in 2003, the development and application of genetic analysis technology have expanded in diverse ways. With the development of analytical techniques that can identify various gene mutations in a single test using next-generation sequencing (NGS), cancer genetic diagnosis has begun to play an important role in effective anti-cancer treatment. Coupled with the success of developing targeted anti-cancer drugs against mutant proteins generated by somatic mutations in major genes such as EGFR and PIK3CA, the clinical demand for cancer genetic diagnosis is increasing.
[0006] For genetic analysis, ideally, nucleic acids should be extracted from cancer tissue immediately within 10 minutes of the interruption of vascular supply during surgery, or frozen with liquid nitrogen. -70℃ The tissue must be stored at extremely low temperatures below 10°C, and nucleic acids must be extracted as needed for DNA sequencing. However, it is impossible for hospitals to apply this method to the surgical tissue of all patients who are not research subjects, and frozen tissue is not the optimal sample for histopathological examination. Therefore, all body-derived tissue is formalin-fixed and then embedded in paraffin (FFPE), followed by microsection for histopathological examination. FFPE blocks are stored at room temperature for a minimum of 10 years or permanently. The decision on the necessity of actual cancer genetic analysis is made after the results of histopathological examination using FFPE tissue are available. Due to the size of cancer tissue that can be obtained through surgery or biopsy, there is often no surplus tissue that can be frozen after the tissue needed for diagnostic FFPE blocks has been secured. Furthermore, many hospitals lack the infrastructure for cryopreservation, so most clinical genetic analysis proceeds using FFPE samples.
[0007] While FFPE samples have advantages such as ease of storage and suitability for histopathological examination, they have weaknesses that make them unsuitable for genetic analysis. Specifically, the chemical reactions that occur during formalin fixation induce artificial mutations (artifacts) unrelated to cancer, and when DNA sequencing using NGS techniques, more than 10 times the original mutations can be found. When genetic analysis is performed by creating a panel targeting only hot-spot driver mutations of some important oncogenes, the presence or absence of driver mutations can be diagnosed with great accuracy. However, when analyzing at the exome or genome level, it is difficult to distinguish between actual mutations and artificial mutations. The presence of artificial mutations leads to inaccuracies in tumor mutation burden measurement and new antigen discovery, posing a major obstacle to realizing advanced medical advancements such as immunotherapy response prediction and new antigen vaccine development. Meanwhile, methods for removing artificial mutations from FFPE NGS DNA sequencing data have been developed, but since most actual mutations are also removed at the same time, their clinical applicability is weak. This has highlighted the need for a technology that removes most artificial mutations while preserving actual mutations.
[0008] The technical background information provided for this invention is intended to facilitate understanding of the present invention. It should not be understood that the matters described in the technical background information constitute prior art. [Overview of the Initiative] [Problems that the invention aims to solve]
[0009] The inventors of this invention recognized that the vast majority of common clinical samples are FFPE tissue, and that NGS analysis is difficult for older FFPE tissue or samples that have undergone mutations during the fixation process due to the aforementioned problems.
[0010] In other words, the inventors of the present invention recognized that when identifying somatic mutations using existing FFPE samples, various factors can cause artificial mutations to appear, resulting in low accuracy. Therefore, they recognized the need for a method that can accurately predict only somatic mutations, and then focused on the fact that using a deep learning-based neural network model, as in the present invention, can improve accuracy.
[0011] In particular, the inventors of this invention analyzed paired NGS data generated from frozen tissue and FFPE tissue derived from the same tissue, defined mutations found only in FFPE as artificial mutations, and recognized that by performing deep learning with training data including cosine similarity, it is possible to predict somatic mutations relative to artificial mutations with high accuracy.
[0012] As a result, the inventors of the present invention were able to confirm that it is possible to distinguish between artificial mutations and somatic mutations with high accuracy by applying the learned neural network model of the present invention. As a result, the inventors of the present invention have developed an information provision system that distinguishes between artificial mutations and somatic mutations generated by FFPE based on the deep learning-based neural network model of the present invention.
[0013] Therefore, the inventors of the present invention recognized that the introduction of an information provision system that distinguishes between artificial mutations and somatic mutations generated by FFPE based on the deep learning neural network model of the present invention can complement the limitations of conventional diagnostic and information provision systems, and thereby enable the presentation of efficient treatment methods.
[0014] Therefore, the problem that the present invention aims to solve is to provide a method for providing information for somatic mutation prediction through artificial mutation removal and preservation of somatic mutation samples based on the deep learning of the present invention, and a device that utilizes the same.
[0015] The problems addressed by the present invention are not limited to those mentioned above, and other problems not mentioned can be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0016] To solve the aforementioned problems, an information provision method for somatic mutation prediction based on deep learning according to one embodiment of the present invention is provided. This information provision method is embodied by a processor and includes the steps of: receiving data on mutations obtained from medical data and / or biological samples isolated from an individual; calculating data necessary for distinguishing between artificial mutations and somatic mutations based on the received data; and predicting artificial mutations or somatic mutations in a sample obtained from an individual using the calculated data as input, with the data necessary for somatic mutation discrimination as input and a neural network model configured to predict artificial mutations or somatic mutations as input.
[0017] In this context, the medical data and / or data obtained from biological samples isolated from an individual are of mutation, specifically data on somatic mutations; the medical data are NGS (Next Generation Sequencing) data, specifically WGS (Whole Genome Sequence), WES (Whole Exome Sequence), Targeted Sequence files, etc., obtained by sequencing the entire gene; and the biological samples isolated from an individual may, but are not limited to, data on samples in which mutations were observed, specifically data on samples in which somatic mutations were observed.
[0018] According to the features of the present invention, the step of calculating data necessary for determining artificial mutations or somatic mutations based on the received data may further include the step of vectorizing the received data and the step of calculating similarity based on the vectorized data, in which case the similarity in the step of calculating similarity based on the vectorized data may be, but is not limited to, cosine similarity.
[0019] Furthermore, in the step of predicting artificial mutations or somatic mutations in a sample obtained from an individual using the calculated data as input, the neural network model may be, but is not limited to, a deep neural network (DNN).
[0020] According to the features of the present invention, in the step of calculating data necessary for mutation or artificial mutation discrimination based on the received data, the calculated data includes: the number of reads in which the reference sequence is observed / the sequence depth at the relevant position; the number of reads in which the alternative sequence is observed / the sequence depth at the relevant position; the ratio of the difference between the number of sequences in which read 1 is mapped to the forward strand and the number of sequences in which read 2 is mapped to the reverse strand in the reads in which the reference sequence is observed to the number of reads in which the reference sequence is observed; the ratio of the difference between the number of sequences in which read 1 is mapped to the forward strand and the number of sequences in which read 2 is mapped to the reverse strand in the reads in which the alternative sequence is observed to the number of reads in which the alternative sequence is observed; the median length of the inserted sequences in the reads in which the alternative sequences are observed divided by the median length of the inserted sequences in the reads in which the reference sequence is observed; and the sequence in the reference fragment that has been read twice by read 1 and read 2 (double-sequenced The sequence can include one or more of the following: the length of bases, the number of times the sequence was read twice by read 1 and read 2 in the reference intercept, the length of the double-sequenced bases (sequenced sequences) read twice by read 1 and read 2 in the alternate intercept, the reference sequence containing one base before and after the mutation site (REF_3_BASES), the alternate sequence containing one base before and after the mutation site (ALT_3_BASES), the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in the reads in which the reference sequence is observed, and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in the reads in which the alternate sequence is observed, and (in the reads in which the reference sequence is observed, read 1 is forwardThe cosine similarity between vectors defined by (the number of sequences mapped to the strand, the number of sequences where read 2 is mapped to the reverse strand), or the cosine similarity between a vector defined by (the number of sequences mapped to the forward strand in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and a vector defined by (the number of sequences mapped to the forward strand in the reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand) can be further included, but is not limited thereto.
[0021] In one embodiment of the present invention, the cosine similarity can be expressed by the following formula.
[0022] [Number]
[0023] According to the features of the present invention, using a neural network model configured to predict artificial mutations or somatic mutations with the data required for somatic mutation discrimination as input, in the step of predicting artificial mutations or somatic mutations for a sample obtained from an individual using the calculated data as input, the sample obtained from the individual may be an FFPE sample, and the individual may be an individual with cancer. At this time, the cancer may be an individual in which any one of liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma has developed, but is not limited thereto.
[0024] In order to solve the problems as described above, an information providing system including a device for providing information for somatic mutation prediction, a mutation prediction device, and a medical data providing server according to an embodiment of the present invention is provided.
[0025] Specific matters of other embodiments are included in the detailed description and drawings. [Effects of the Invention]
[0026] The present invention can provide information on somatic mutation prediction so that personalized treatment (Tailored personal medicine) can be achieved by providing information on prediction results for artificial mutations or somatic mutations using a deep - learned neural network model.
[0027] In particular, the present invention can more accurately predict artificial mutations or somatic mutations by calculating data related to somatic mutations and determining prediction results for artificial mutations or somatic mutations based on the data.
[0028] That is, the present invention can more easily set the treatment direction for an individual and contribute to the selection of a personalized treatment method by utilizing information on the prediction of artificial mutations or somatic mutations based on a deep - learned neural network model.
[0029] The present invention applies a deep - learned neural network model learned with learning data including cosine similarity for the prediction of artificial mutations or somatic mutations, so as to predict artificial mutations and somatic mutations with high accuracy, thus complementing the limitations when using conventional FFPE samples.
[0030] The effects according to the present invention are not limited to the contents exemplified above, and more various effects are included in this specification.
Brief Description of the Drawings
[0031] [Figure 1] It shows an information - providing system for the prediction of artificial mutations or somatic mutations using a prediction device for artificial mutations or somatic mutations according to an embodiment of the present invention. [Figure 2] It exemplarily shows the configuration of a device for predicting artificial mutations or somatic mutations according to an embodiment of the present invention. [Figure 3]This is a block diagram of the procedure for a method of providing information for predicting artificial mutations or somatic mutations according to one embodiment of the present invention. [Figure 4] This document illustrates the procedure for providing information on artificial mutations or somatic mutations according to various embodiments of the present invention. [Figure 5a] This figure shows the results of evaluating the predictive performance of the deep learning-based neural network model of the present invention. [Figure 5b] This figure shows the results of evaluating the prediction performance of Comparative Example 1 (FFpolish pre-trained). [Figure 5c] This figure shows the results of evaluating the prediction performance of Comparative Example 2 (FFpolish re-trained). [Figure 5d] This figure shows the results of evaluating the prediction performance of Comparative Example 3 (SOBDetector). [Figure 6a] This shows a comparison of the accuracy of predicting artificial or somatic mutations in liver cancer samples using the deep learning neural network model of the present invention. [Figure 6b] This shows a comparison of the accuracy of predicting artificial or somatic mutations in colorectal cancer samples using the deep learning neural network model of the present invention. [Figure 6c] This paper compares and shows the accuracy of predicting artificial or somatic mutations in breast cancer samples using the deep learning neural network model of the present invention. [Figure 6d] This shows a comparison of the accuracy of predicting artificial or somatic mutations in lung cancer samples using the deep learning neural network model of the present invention. [Figure 6e] This shows a comparison of the accuracy of predicting artificial or somatic mutations in fibroma samples using the deep learning neural network model of the present invention. [Figure 7a] This shows a comparison of the accuracy of predicting artificial or somatic mutations in liver cancer samples using a comparative example (FFpolish pre-trained). [Figure 7b]This shows a comparison of the accuracy of predicting artificial or somatic mutations in colorectal cancer samples using a comparative example (FFpolish pre-trained). [Figure 7c] This shows a comparison of the accuracy of predicting artificial or somatic mutations in breast cancer samples using a comparative example (FFpolish pre-trained). [Figure 7d] This shows a comparison of the accuracy of predicting artificial or somatic mutations in lung cancer samples using a comparative example (FFpolish pre-trained). [Figure 7e] This shows a comparison of the accuracy of predicting artificial or somatic mutations in fibroma samples using a comparative example (FFpolish pre-trained). [Modes for carrying out the invention]
[0032] The advantages and features of the present invention, and the methods for achieving them, will become clear when you refer to the embodiments described in detail later with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, but can be embodied in a variety of different forms, and these embodiments are provided merely to complete the disclosure of the present invention and to fully inform a person ordinary skill in the art to which the invention belongs of the scope of the invention, and the present invention is defined only by the scope of the claims. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0033] In this document, expressions such as "have," "may have," "include," or "may include" refer to the existence of the relevant feature (e.g., numerical values, functions, operations, or components such as parts), and do not exclude the existence of further features.
[0034] In this document, expressions such as "A or B," "A or / and B, or one or more of A or / and B" can include all possible combinations of the items listed together. For example, "A or B," "A and B, or A or B, or at least one of A or B" can refer to all cases where (1) at least one A is included, (2) at least one B is included, or (3) at least one A and at least one B are included.
[0035] The terms "first," "second," "first," or "second" used in this document can modify various components regardless of their order and / or importance, and are used only to distinguish one component from another, without limiting it to that component. For example, the first user device and the second user device may refer to different user devices, regardless of their order or importance. For example, the first component may be named the second component in order to stay within the scope of rights described in this document, and similarly, the second component may be renamed in place of the first component.
[0036] When it is mentioned that a component (e.g., component 1) is "operally or communicatively coupled with / to" or "connected to" another component (e.g., component 2), it should be understood that the component is directly coupled to the other component or can be coupled through another component (e.g., component 3). Conversely, when it is mentioned that a component (e.g., component 1) is "directly coupled" or "directly connected" to another component (e.g., component 2), it should be understood that there is no other component (e.g., component 3) between the component and the other component.
[0037] The expression "configured to" as used in this document may be replaced with other expressions depending on the context, such as "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean only something that is "specifically designed to" in terms of hardware. Instead, in some situations, the expression "a device configured to" may mean that the device is "capable" of doing something together with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU, GPU, or application processor) that can perform those operations by running one or more software programs stored in a memory device.
[0038] The terms used herein are used solely to describe specific embodiments and are not intended to limit the scope of other embodiments. Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as those generally understood by someone with ordinary skill in the art described herein. Terms used herein that are defined in general dictionaries may be interpreted as having the same or similar meaning as they do in the context of the relevant art, and not as ideal or overly formal unless explicitly defined herein. In some cases, terms defined herein may not be interpreted as excluding the embodiments described herein.
[0039] The features of each of the various embodiments of the present invention are partially or entirely combinable with one another, and are technically capable of a wide range of interdependence and operation, as can be fully understood by those skilled in the art. Each embodiment may be implemented independently of the others or together in relation to one another.
[0040] For clarity of interpretation of this specification, the following definitions of terms used herein are provided. As used herein, the term "individual" may mean all subjects for which we are trying to predict the presence or absence of mutation, specifically somatic mutation. For example, an individual may be an individual in which a somatic mutation has occurred. In this case, the individual disclosed herein may be, but is not limited to, all mammals other than humans.
[0041] As used herein, the term "biological sample" refers to all biological samples isolated from an individual and storable in FFPE form, including but not limited to urine, tissue, cell lysates, whole blood, plasma, serum, saliva, ocular fluid, cerebrospinal fluid, sweat, milk, ascites fluid, synovial fluid, and peritoneal fluid. Preferably, the biological sample is a tissue sample isolated from an individual to confirm the presence or absence of somatic mutations, but is not limited to this. Specifically, the biological sample of the present invention may be an FFPE sample (Formalin Fixed Paraffin Embedded), and an FFPE sample may mean a sample in which tissue is fixed with formalin and solidified in paraffin in the standard method used when collecting and storing cancerous tissue from a patient, and manufactured into a paraffin block that can be stored at room temperature, and the paraffin block can be cut into thin slices several times and used for various histological examinations.
[0042] As used herein, the term "somatic variant" refers to a genetic abnormality, which is one of the characteristics of cancer observed in cells. Therefore, it is possible to identify somatic variants and classify cancer types. Accordingly, in this invention, medical data and / or data obtained from biological samples isolated from individuals may be data on somatic mutations. Medical data may be NGS (Next Generation Sequencing) data, but is not limited to this. It may also mean biological samples isolated from individuals in which somatic variants were observed, but is not limited to this.
[0043] As used herein, the term "artifact" may mean, but is not limited to, an artificial defect or other condition that appears as a result of degeneration caused by physical or chemical reactions, and is unrelated to cancer.
[0044] As used herein, the term "NGS (Next Generation Sequencing)" refers to next-generation nucleotide sequence analysis, which is one of the rapid methods for analyzing the nucleotide sequence of a gene.
[0045] In this context, NGS can be applied to genetic and whole-body analysis to achieve a variety of objectives, including clinical research. Furthermore, most NGS analyses can be performed simultaneously on a large number of target samples.
[0046] As used herein, the term "NGS data" may mean a file containing nucleotide sequence data analyzed for the sample in question. In this context, NGS data may include WGS (whole genome sequencing) files, WES (whole exome sequencing) files, RNA sequencing files, and targeted sequencing files, which are obtained by sequencing the entire gene. Specifically, it is limited to WES (whole exome sequencing) files.
[0047] In the following, with reference to Figures 1 and 2, an information provision system for predicting artificial mutations or somatic mutations and an artificial mutation or somatic mutation prediction device based on one embodiment of the present invention will be described.
[0048] Figure 1 shows an information provision system for predicting artificial mutations or somatic mutations using an artificial mutation or somatic mutation prediction device according to one embodiment of the present invention. Figure 2 illustrates the configuration of an artificial mutation or somatic mutation prediction device according to one embodiment of the present invention.
[0049] First, referring to Figure 1, the information provision system 1000 may be a system configured to provide information related to the prediction of artificial or somatic mutations by applying a deep-learned neural network model based on data obtained from medical data and / or biological samples. In this case, the information provision system 1000 may consist of a data provision server / device 100 that provides data obtained from medical data and / or biological samples, a prediction device that predicts artificial or somatic mutations based on the received data, and an information provision device 300 that visually provides information on the prediction of artificial or somatic mutations for samples obtained from individuals based on the received prediction data.
[0050] Here, the information-providing device 300 is an electronic device that provides a user interface for displaying information related to the prediction of artificial or somatic mutations, and may include at least one of a smartphone, tablet PC (Personal Computer), laptop computer, and / or PC.
[0051] The information-providing device 300 for predicting somatic mutations receives information related to the prediction results of artificial or somatic mutations from the somatic mutation prediction device 200, and can display the received results through a display unit described later.
[0052] The somatic mutation prediction device 200 may include a general-purpose computer, laptop, and / or data server that perform various calculations to determine information related to the prediction of artificial or somatic mutations based on various data (particularly medical data and / or data on somatic mutations) obtainable from the data provision server / device 100. In this case, the data provision server 100 may be, but is not limited to, a device for accessing a web server that provides web pages or a mobile web server that provides mobile websites.
[0053] More specifically, the somatic mutation prediction device 200 can receive medical data and / or data on somatic mutations from the medical data provision server 100, predict artificial mutations or somatic mutations based on the received data, and provide related information. In this case, the somatic mutation prediction device 200 can perform the prediction of artificial mutations or somatic mutations based on medical data and / or data on somatic mutations obtained from biological samples received from the data provision server 100 by applying a deep-learned neural network model.
[0054] Furthermore, the somatic mutation prediction device 200 can provide artificial mutation or somatic mutation prediction results to the information provision device 300 for somatic mutation prediction. The information provided by the somatic mutation prediction device 200 can be provided to a web page via a web browser provided on the somatic mutation prediction information provision device 300, or it can be provided in the form of an application or program. In various embodiments, such data can be provided in a form included in a platform in a client-server environment.
[0055] In the following description, the somatic cell mutation prediction device 200 according to one embodiment of the present invention will be described as receiving data obtained from medical data and / or biological samples from a data provision server 100 and performing its operations. However, the information provision device 300 for somatic cell prediction itself can also perform all operations.
[0056] Next, with reference to Figure 2, the components of the somatic mutation prediction device 200 of the present invention will be described in detail. Referring to Figure 2, the somatic mutation prediction device 200 may include a communication interface 210, memory 220, I / O interface 230, and processor 240, and each component can communicate with one or more communication buses or signal lines.
[0057] The communication interface 210 can connect to the information provision device 300 for somatic mutation prediction and the medical data provision server 100 via a wired / wireless network to exchange data. For example, the communication interface 210 can receive data on somatic mutations and / or individual data from the medical data provision server 100 and transmit information about the determined artificial mutation or somatic mutation prediction results to the information provision device 300 for somatic mutation prediction.
[0058] On the other hand, the communication interface 210 that enables the transmission and reception of such data includes a wired communication port 211 and a wireless circuit 212, where the wired communication port 211 may include one or more wired interfaces, such as Ethernet®, USB, FireWire, etc. The wireless circuit 212 can transmit and receive data with an external device via RF signals or optical signals. In addition, the wireless communication may use at least one of several communication standards, protocols and technologies, such as GSM®, EDGE, CDMA, TDMA, Bluetooth®, Wi-Fi®, VoIP, Wi-MAX®, or any other suitable communication protocol.
[0059] The memory 220 can store a variety of data used by the somatic mutation prediction device 200. For example, the memory 220 can store data about somatic mutations, or it can store deep learning models configured to predict artificial mutations or somatic mutations based on this data.
[0060] In various embodiments, the memory 220 may include volatile or non-volatile recording media capable of storing various types of data, instructions, and information. For example, the memory 220 may include at least one type of storage medium from among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), RAM, SRAM, ROM, EEPROM, PROM, network storage, cloud, and blockchain database.
[0061] In various embodiments, the memory 220 can store at least one of the following configurations: the operational system 221, the communication module 222, the user interface module 223, and one or more applications 224.
[0062] The operational system 221 (for example, the built-in operational systems of LINUX®, UNIX®, MAC OS, WINDOWS®, VxWorks®, etc.) can include a variety of software components and drivers for controlling and managing general system operations (e.g., memory management, storage device control, power management, etc.) and can support communication between various hardware, firmware, and software components.
[0063] Communication module 222 It can support communication with other devices through the communication interface 210. (Communication module) 222 This can include a variety of software components for processing data received by the wired communication port 211 or wireless circuit 212 of the communication interface 210.
[0064] The user interface module 223 can receive user requests or inputs from a keyboard, touchscreen, microphone, etc., via the I / O interface 230, and provide a user interface on the display.
[0065] Application 224 is a processor of one or more 240 This may include a program or module configured to be executed by a server farm. Here, an application for providing information related to somatic mutation prediction may be implemented on a server farm.
[0066] The I / O interface 230 can connect at least one of the input / output devices (not shown) of the somatic mutation prediction device 200, such as a display, keyboard, touchscreen, and microphone, to the user interface module 223. Together with the user interface module 223, the I / O interface 230 can receive user input (e.g., voice input, keyboard input, touch input, etc.) and process commands based on the received input.
[0067] The processor 240 is connected to the communication interface 210, memory 220, and I / O interface 230 to control the overall operation of the somatic cell mutation prediction device 200, and can execute a variety of instructions for providing information through applications or programs stored in memory 220.
[0068] Processor 240 could be a computing device such as a CPU (Central Processing Unit), GPU (Graphic Processing Unit), or AP (Application Processor). Alternatively, processor 240 could be embodied in the form of an integrated chip (IC), such as a System on Chip (SoC), which integrates various computing devices. Or, processor 240 could include a module for computing artificial neural network models, such as a Neural Processing Unit (NPU).
[0069] In various embodiments, the processor 240 may be configured to apply a deep-learned neural network model to calculate the data necessary for determining artificial or somatic mutations based on data obtained from medical data and / or biological samples, and to predict and provide artificial or somatic mutations.
[0070] In one embodiment, the somatic mutation prediction device 200 receives medical data and / or data obtained from biological samples isolated from individuals from the medical data provision device 100, and can provide prediction results for somatic mutations using a model trained to predict the presence or absence of artificial mutations or somatic mutations by applying a deep-learned neural network model using the received data. In this case, the medical data and / or data obtained from biological samples isolated from individuals may be data on somatic mutations, and the medical data may be, but is not limited to, NGS (Next Generation Sequencing) data, and may mean, but is not limited to, data on biological samples isolated from individuals in which somatic mutations were observed. Specifically, the NGS data may be at least one of WGS files, WES files, RNA sequencing files, and targeted sequencing files, and specifically may be, but is not limited to, WES files.
[0071] In the following, with reference to Figures 3 and 4, a method for providing information for predicting artificial mutations or somatic mutations according to one embodiment of the present invention will be specifically described. Figure 3 is a block diagram of the procedure for a method of providing information for predicting artificial or somatic mutations according to one embodiment of the present invention.
[0072] Figure 4 illustrates the procedure for providing information on artificial or somatic mutations according to various embodiments of the present invention. First, referring to Figure 3, the procedure for providing information regarding the prediction of artificial or somatic mutations according to one embodiment of the present invention is as follows.
[0073] First, medical data and / or data on mutations that have occurred in biological samples isolated from the individual are received (S310). Subsequently, data necessary for distinguishing between artificial mutations and somatic mutations is calculated based on the received data (S320), Using a neural network model configured to predict artificial mutations or somatic mutations with data necessary for somatic mutation discrimination as input, artificial mutations or somatic mutations are predicted for a sample obtained from an individual using the calculated data as input (S330).
[0074] According to the features of the present invention, in the step (S310) in which medical data and / or data on mutations occurring in a biological sample isolated from an individual are received, a whole exome sequencing (WES) file, an RNA sequencing file, or a targeted sequencing file may be received, specifically a WES file, but not limited thereto.
[0075] In this case, the WES file, RNA sequencing file, and target sequencing file may have the format of a BAM file, but the format of the NGS data is not limited to this.
[0076] Next, the step (S320) of calculating data necessary to distinguish between artificial mutations and somatic mutations based on the received data may further include the step of vectorizing the received data and the step of calculating similarity based on the vectorized data.
[0077] Referring to both Table 1 and Figure 4, the data calculated in the step of calculating the data necessary for mutation or artificial mutation discrimination based on the received data is: the number of reads in which the reference sequence is observed / the sequence depth at the relevant position; the number of reads in which the alternative sequence is observed / the sequence depth at the relevant position; the ratio of the difference between the number of sequences where read 1 is mapped to the forward strand and the number of sequences where read 2 is mapped to the reverse strand in the reads in which the reference sequence is observed to the number of reads in which the reference sequence is observed; the ratio of the difference between the number of sequences where read 1 is mapped to the forward strand and the number of sequences where read 2 is mapped to the reverse strand in the reads in which the alternative sequence is observed to the number of reads in which the alternative sequence is observed; the median length of the inserted sequences in the reads in which the alternative sequences are observed divided by the median length of the inserted sequences in the reads in which the reference sequence is observed; and the sequence in the reference fragment that has been read twice by read 1 and read 2 (double-sequenced The sequence can include one or more of the following: the length of bases, the number of times the sequence was read twice by read 1 and read 2 in the reference intercept, the length of the double-sequenced bases (sequenced sequences) read twice by read 1 and read 2 in the alternate intercept, the reference sequence containing one base before and after the mutation site (REF_3_BASES), the alternate sequence containing one base before and after the mutation site (ALT_3_BASES), the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in the reads in which the reference sequence is observed, and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in the reads in which the alternate sequence is observed, and (in the reads in which the reference sequence is observed, read 1 is forwardThis may further include, but is not limited to, the cosine similarity between vectors defined as (number of sequences mapped to the strand, number of sequences to which read 2 is mapped to the reverse strand), or the cosine similarity between a vector defined as (number of sequences mapped to the forward strand in the read where the reference sequence is observed, number of sequences mapped to the reverse strand) and a vector defined as (number of sequences mapped to the forward strand in the read where the alternative sequence is observed, number of sequences mapped to the reverse strand).
[0078] [Table 1-1]
[0079] [Table 1-2]
[0080] As used herein, the term "read" may mean DNA including the adapter, and the term "fragment" may mean DNA excluding the adapter.
[0081] In this specification, the terms "read 1" refer to the reading in the 5' to 3' direction when Paired End DNA entering the NGS equipment is read from both ends, and "read 2" refer to the reading in the 3' to 5' direction.
[0082] As used herein, the term "depth" may be used interchangeably with the term "read-depth" and refers to the thickness or depth of a lead.
[0083] As used herein, the term "fr_ref" refers to the number of sequences in which read 1 is mapped to the forward strand in the read where the reference sequence is observed. As used herein, the term "rf_ref" refers to the number of sequences to which lead 2 is mapped to the reverse strand.
[0084] As used herein, the term "fr_alt" refers to the number of sequences in which read 1 is mapped to the forward strand in reads where an alternative sequence is observed. In this specification, the term "rf_alt" refers to the number of sequences to which lead 2 is mapped to the reverse strand.
[0085] As used herein, the term "fs_ref" refers to the number of sequences mapped to the forward strand during a read in which a reference sequence is observed. As used herein, the term "rs_ref" refers to the number of arrays mapped to the reverse strand.
[0086] As used herein, the term "fs_alt" refers to the number of sequences mapped to the forward strand in a read where an alternative sequence is observed. As used herein, the term "rs_alt" refers to the number of sequences mapped to the reverse strand.
[0087] In this case, the model into which data on mutations occurring in medical data and / or biological samples isolated from individuals is input can be based on a variety of algorithms, such as CNN (Convolutional Neural Network) models, DNN (Deep Neural Network), DCNN (Deep Convolutional Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), SSD (Single Shot Detector), and SVM (Support Vector Machine). Specifically, it may be a CNN (Convolutional Neural Network) model or a DNN (Deep Neural Network) model, and more specifically, it may be a DNN (Deep Neural Network) model, but is not limited to these.
[0088] Subsequently, using a neural network model configured to predict artificial or somatic mutations with the data necessary for somatic mutation discrimination as input, the calculated data is used as input to predict artificial or somatic mutations in a sample obtained from an individual (S330). In this step, the prediction result for artificial or somatic mutations in the sample obtained from the individual is determined and provided. For example, the model may be configured to output whether the result value is an actual artificial mutation, a false artificial mutation, an actual somatic mutation, or a false somatic mutation. For example, the output may be shown as a percentage, and a threshold can be set to determine whether a percentage above the threshold is an actual artificial mutation, a false artificial mutation, an actual somatic mutation, or a false somatic mutation.
[0089] In this case, the sample obtained from the individual may be an FFPE sample, and the individual may be, but is not limited to, an individual that has developed one of the following: liver cancer, colorectal cancer, breast cancer, lung cancer, or fibroma.
[0090] On the other hand, the procedures for providing information regarding the prediction of artificial or somatic mutations are not limited to those described above. In accordance with the above-described method for providing information on the prediction of artificial or somatic mutations, the present invention provides an information provision system for the prediction of artificial or somatic mutations based on an algorithm, thereby making it easier to set treatment directions for individuals and contributing to the selection of personalized treatment methods.
[0091] In other words, the present invention can complement the limitations of conventional diagnostic and information provision systems by introducing an information provision system for predicting artificial or somatic mutations based on a neural network model, thereby predicting them with high accuracy, and making it easier to set treatment directions for individuals. [Examples]
[0092] Example 1: Prediction rate of artificial or somatic mutations in FFPE samples To generate a trainable dataset, we collected cancer tissue FFPE (Formalin-fixed paraffin-embedded) samples and predicted mutations that appeared in FF (Frozen Fresh) tissue from 112,976 mutation calls.
[0093] As a result, 2,113 mutations appeared in both FF and FFPE samples, while 110,863 mutations appeared only in FFPE samples. This confirmed that only about 1.8% were somatic mutations, not artificial mutations.
[0094] Subsequently, using the previously acquired samples and the deep-learned neural network model, features highly relevant to somatic cell prediction were calculated, and then the performance for somatic cell mutation prediction was evaluated. Example 2: Performance evaluation of artificial or somatic mutation prediction in a deep learning-based neural network model. In the following section, with reference to Figure 5, the performance evaluation results for somatic mutation prediction of a deep-learned neural network model according to one embodiment of the present invention will be explained.
[0095] Figures 5a to 5d show the results of evaluating the performance of a deep-learned neural network model according to one embodiment of the present invention and a comparative example. First, to evaluate the deep-learned neural network model according to one embodiment of the present invention, models utilizing diverse somatic mutation-related data were constructed in Comparative Examples 1 to 3. More specifically, the FFpolish (pre-trained) model was used in Comparative Example 1, the FFpolish (re-trained) model in Comparative Example 2, and the SOBDetector model in Comparative Example 3, and these were compared with the neural network model of the present invention (DEEPOMICS(R) FFPE).
[0096] Precision, recall, F1 score, accuracy, and specificity were calculated as follows.
[0097]
number
[0098] As a result, as shown in Figure 5a, the neural network model according to one embodiment of the present invention (DEEPOMICS® FFPE) was confirmed to remove 99.6% of artificial mutations while simultaneously showing 77.8% conservation of actual somatic mutations.
[0099] In contrast, as shown in Figure 5b, Comparative Example 1 (FFpolish (pre-trained)) was found to have 92.6% of artificial mutations removed and 69.5% of actual somatic mutation samples preserved, confirming that its ability to remove artificial mutations and preserve actual somatic mutation samples was lower than that of the present invention compared to the neural network model according to one embodiment of the present invention.
[0100] As shown in Figure 5c, in the case of Comparative Example 2 (FFpolish (re-trained)), artificial mutation removal was 99.6%, demonstrating high removal efficiency, but actual somatic mutation preservation was observed in 39.8% of cases. This confirms that, in the case of the Comparative Example 2 model, when artificial mutations are removed, samples in which actual somatic mutations have occurred are also removed, affecting the prediction results.
[0101] As shown in Figure 5d, in the case of Comparative Example 3 (SOBDetector), the artificial mutation removal rate was 72.1%, while the actual somatic mutation preservation rate was 84.5%, confirming that the artificial mutation removal ability was significantly low.
[0102] When using the neural network model according to one embodiment of the present invention, the F1 score was also 0.821, confirming its superiority over other comparative examples. As described above, it was confirmed that while the artificial mutation removal ability is excellent, samples in which actual somatic mutations appear can be preserved without being removed.
[0103] Evaluation 2: Measurement of the accuracy of artificial or somatic mutation prediction using cancer type-specific algorithms. In the following, with reference to Figures 6a to e and 7a to e, the performance evaluation results for predicting somatic mutations in various cancer types of the neural network model according to one embodiment of the present invention will be explained.
[0104] Figures 6a to 6e show a comparison of the accuracy of predicting artificial or somatic mutations for liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma using a neural network model according to one embodiment of the present invention.
[0105] Figures 7a to 7e show a comparison of the accuracy of the FFPolish model of Comparative Example 1 from the comparative examples of Evaluation 1, comparing its accuracy in predicting artificial or somatic mutations for liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma.
[0106] As a result, as shown in Figures 6a to 6e, when using the neural network model according to one embodiment of the present invention, the success rates for liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma samples were 89.9%, 88.2%, 99.8%, 96.5%, and 89.7%, respectively, demonstrating excellent artificial mutation removal ability. Somatic mutation preservation also showed success rates of 88.5%, 89.3%, 63.9%, 76.4%, and 90.2%, respectively, confirming that the model is excellent not only in artificial mutation removal ability but also in preserving actual somatic mutation samples.
[0107] In contrast, in the case of the comparative FFpolish model, as shown in Figures 7a to e, the percentages were 85.3% for liver cancer, 63.4% for colorectal cancer, 92.8% for breast cancer, 91.2% for lung cancer, and 71.8% for fibroma. It was confirmed that the ability to remove artificial mutations was lower in all samples compared to the model of the present invention. The preservation of somatic cell mutation samples was 69.1% for liver cancer, 78.9% for colorectal cancer, 60.6% for breast cancer, 74.3% for lung cancer, and 70.7% for fibroma, respectively. It was confirmed that the probability of removing actual somatic cell mutation samples along with artificial mutations was higher in the neural network model according to one embodiment of the present invention.
[0108] Although embodiments of the present invention have been described in more detail above with reference to the attached drawings, the present invention is not necessarily limited to these embodiments and can be modified and implemented in various ways within the scope of the technical concept of the present invention. Accordingly, the embodiments disclosed herein are for illustrative purposes only, not to limit the technical concept of the present invention, and the scope of the technical concept of the present invention is not limited by such embodiments. Therefore, the embodiments described above should be understood as illustrative in all respects and not limiting. The scope of protection of the present invention should be interpreted by the following claims, and all technical concepts within an equivalent scope should be interpreted as being included in the scope of the rights of the present invention.
Claims
1. A method for providing information regarding the prediction of artificial or somatic mutations embodied by a processor, A step of receiving medical data or data obtained from biological samples isolated from an individual; A step of calculating data necessary for determining artificial mutations or somatic mutations based on the received data; and A step of predicting artificial or somatic mutations in a sample obtained from an individual, using a neural network model configured to predict artificial or somatic mutations, with the data necessary for somatic mutation discrimination as input. Includes, In the step of calculating data necessary for determining artificial mutations or somatic mutations based on the received data, the calculated data is: The number of reads in which the reference sequence is observed / the sequence depth at the relevant position; the number of reads in which the alternative sequence is observed / the sequence depth at the relevant position; the ratio of the difference between the number of sequences where read 1 maps to the forward strand and the number of sequences where read 2 maps to the reverse strand in reads in which the reference sequence is observed; the ratio of the number of sequences where read 1 maps to the forward strand and the number of sequences where read 2 maps to the reverse strand in reads in which the alternative sequence is observed. The ratio of the difference in the number of sequences mapped to the strand to the reads in which the alternate sequence is observed, the median length of the inserted sequences in the reads in which the alternate sequence is observed divided by the median length of the inserted sequences in the reads in which the reference sequence is observed, the length of the double-sequenced bases in the reference intercept read twice by read 1 and read 2, the number of times the sequence was read twice by read 1 and read 2 in the reference intercept, and the length of the double-sequenced bases in the alternate intercept read twice by read 1 and read 2. It includes one or more of the following selected from the group consisting of: the length of bases; the number of times the sequence was read twice by read 1 and read 2 in the alternative section; a reference sequence (REF_3_BASES) containing one base before and one base after the mutation site; an alternative sequence (ALT_3_BASES) containing one base before and one base after the mutation site; the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in reads in which the reference sequence is observed; and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in reads in which the alternative sequence is observed. In the step of calculating data necessary for determining artificial mutations or somatic mutations based on the received data, the calculated data is: This further includes the cosine similarity between vectors defined as (the number of sequences in the read where the reference sequence is observed where read 1 maps to the forward strand, and the number of sequences in the read where read 2 maps to the reverse strand), or the cosine similarity between a vector defined as (the number of sequences in the read where the reference sequence is observed that maps to the forward strand, and the number of sequences in the read where the reference sequence is observed that maps to the reverse strand) and a vector defined as (the number of sequences in the read where the alternative sequence is observed that maps to the forward strand, and the number of sequences mapped to the reverse strand). A method for providing information regarding the prediction of artificial or somatic mutations.
2. The method for providing information for predicting artificial mutations or somatic mutations according to claim 1, wherein the medical data is NGS (Next Generation Sequencing) data for somatic mutations.
3. The method for providing information for predicting artificial mutations or somatic mutations according to claim 2, wherein the NGS data is one of the following: a WGS (whole gene sequencing) file, a WES (whole exome sequencing) file, an RNA sequencing file, or targeted sequencing performed on a specific region.
4. The step of calculating the data necessary for determining artificial mutations or somatic mutations based on the received data is: A method for providing information for predicting artificial mutations or somatic mutations according to claim 1, comprising the step of vectorizing the received data.
5. The step of calculating the data necessary for determining artificial mutations or somatic mutations based on the received data is: The method for providing information for predicting artificial or somatic mutations according to claim 4, further comprising the step of calculating similarity based on the vectorized data.
6. In the step of determining the prediction of artificial mutations or somatic mutations for a sample obtained from an individual based on the calculated data, The method for providing information for predicting artificial mutations or somatic mutations according to claim 1, characterized in that the sample obtained from the individual is an FFPE sample.
7. In the step of determining the prediction of artificial mutations or somatic mutations for a sample obtained from an individual based on the calculated data, The aforementioned individual is characterized by being an individual that has developed cancer. A method for providing information for predicting artificial mutations or somatic mutations as described in claim 1.
8. The method for providing information for predicting artificial or somatic mutations according to claim 7, wherein the cancer is one of liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma.
9. A communication unit that receives data on somatic cell mutations obtained from medical data or biological samples from a data provision server, and The communication unit includes a processor electrically connected to the communication unit, and the processor is Based on the received data, the data necessary for determining whether an artificial mutation or somatic mutation is used is calculated. Using a neural network model configured to predict artificial or somatic mutations with data necessary for somatic mutation discrimination as input, the calculated data is used to predict artificial or somatic mutations in a sample obtained from an individual. In a device for providing information on the prediction of artificial or somatic mutations, Based on the received data, the data necessary for distinguishing between artificial and somatic mutations is: the number of reads in which the reference sequence is observed / the sequence depth at the relevant position; the number of reads in which the alternative sequence is observed / the sequence depth at the relevant position; the ratio of the difference between the number of sequences where read 1 maps to the forward strand and the number of sequences where read 2 maps to the reverse strand in reads in which the reference sequence is observed; and the ratio of the number of sequences where read 1 maps to the forward strand and the number of sequences where read 2 maps to the reverse strand in reads in which the alternative sequence is observed. The ratio of the difference in the number of sequences mapped to the strand to the reads in which the alternate sequence is observed, the median length of the inserted sequences in the reads in which the alternate sequence is observed divided by the median length of the inserted sequences in the reads in which the reference sequence is observed, the length of the double-sequenced bases in the reference intercept read twice by read 1 and read 2, the number of times the sequence was read twice by read 1 and read 2 in the reference intercept, and the length of the double-sequenced bases in the alternate intercept read twice by read 1 and read 2. The data includes one or more selected from the group consisting of: the length of bases; the number of times the sequence was read twice by read 1 and read 2 in the alternative section; a reference sequence (REF_3_BASES) containing one base before and one base after the mutation site; an alternative sequence (ALT_3_BASES) containing one base before and one base after the mutation site; the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in reads in which the reference sequence is observed; and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand in reads in which the alternative sequence is observed. A device for providing information for predicting artificial or somatic mutations, wherein the data necessary for distinguishing between artificial and somatic mutations based on the received data further includes cosine similarity between vectors defined as (number of sequences where read 1 maps to forward strand and number of sequences where read 2 maps to reverse strand in reads where the reference sequence is observed), or cosine similarity between a vector defined as (number of sequences where forward strand is mapped and number of sequences where reverse strand is mapped in reads where the reference sequence is observed) and a vector defined as (number of sequences where forward strand is mapped and number of sequences where reverse strand is mapped in reads where the alternative sequence is observed).
10. The aforementioned medical data is NGS (Next Generation Sequencing) data for somatic mutations. A device for providing information for predicting artificial or somatic mutations, as described in claim 9.
11. The aforementioned NGS data is one of the following: a WGS (whole gene sequencing) file, a WES (whole exome sequencing) file, an RNA sequencing file, or targeted sequencing, which is a sequence analysis performed on a specific region of the gene. A device for providing information for predicting artificial mutations or somatic mutations according to claim 10.
12. The aforementioned processor, Further configured to vectorize the received data, A device for providing information for predicting artificial or somatic mutations, as described in claim 9.
13. The aforementioned processor, The device for providing information for predicting artificial or somatic mutations according to claim 12, further configured to calculate similarity based on the vectorized data.
14. The device for providing information for predicting artificial mutations or somatic mutations according to claim 9, characterized in that the sample obtained from the individual is an FFPE sample.
15. The aforementioned individual is characterized by being an individual that has developed cancer. A device for providing information for predicting artificial or somatic mutations, as described in claim 9.
16. The device for providing information for predicting artificial or somatic mutations according to claim 15, wherein the cancer is one of liver cancer, colorectal cancer, breast cancer, lung cancer, and fibroma.
Citation Information
Patent Citations
Machine learning system and method for somatic mutation discovery
US20190189242A1