Method and device for distinguishing somatic mutations and artificial mutations in DNA sequencing data generated from formalin-fixed paraffin-embedded samples using deep learning

A deep learning-based neural network model accurately distinguishes somatic and artificial mutations in FFPE samples by analyzing base sequence context, enhancing precision in tumor mutation burden measurements and neoantigen discovery for personalized medicine.

WO2025230098A1PCT designated stage Publication Date: 2025-11-06THERAGEN BIO CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/001704
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2025-02-05
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing methods for analyzing FFPE samples in DNA sequencing face challenges in distinguishing between somatic and artificial mutations, leading to inaccuracies in tumor mutation burden measurements and neoantigen discovery due to high levels of artifactual mutations, which complicates medical applications like immunotherapy response prediction and neoantigen vaccine development.

Method used

A deep learning-based artificial neural network model is employed to classify somatic and artificial mutations by analyzing base sequence context data, including reference sequences before and after the mutation site, without relying on variant callers, using encoder and decoder structures or transformer models to enhance accuracy.

Benefits of technology

The model accurately distinguishes between somatic and artificial mutations in FFPE samples, improving the precision of tumor mutation burden measurements and neoantigen discovery, thereby supporting tailored medical treatments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025001704_06112025_PF_FP_ABST
    Figure KR2025001704_06112025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides: a method for providing information on predictions of artificial mutations or somatic mutations, the method being implemented by a processor; and a device for providing information on predictions of artificial mutations or somatic mutations. The method comprises: a step for receiving medical data or data obtained from a biological sample isolated from an individual; a step for calculating data necessary for a determination of an artificial mutation or somatic mutation on the basis of the received data, wherein the calculated data includes a mutation position sequence of the artificial mutation or somatic mutation and a reference sequence including N bases before and after the mutation position, the sequence of the mutation position is a standard sequence or a mutation sequence, and N is a natural number of 2 or more; and a step for predicting an artificial mutation or somatic mutation for a sample obtained from an individual, the prediction being made by using the calculated data as input by using an artificial neural network model configured to predict artificial mutations or somatic mutations by using the data necessary for a determination of a somatic mutation as input.
Need to check novelty before this filing date? Find Prior Art

Description

Method for distinguishing somatic mutations and artificial mutations using deep learning in DNA sequencing data generated from formalin-fixed, paraffin-embedded samples and device using the same

[0001] The present invention relates to a method for distinguishing somatic mutations and artificial mutations using deep learning in DNA sequencing data generated from formalin-fixed paraffin-embedded samples, and to a device using the same.

[0002] Factors affecting the development of cancer include germline mutations, which include genetic predispositions, and somatic mutations, which occur during the process of somatic cell division limited to specific organs or tissues.

[0003] Somatic mutations include single nucleotide variations, structural variations, and aneuploidy. Specific cancer types are predicted by somatic single nucleotide variations, which arise from specific mutations and malfunctions in repair mechanisms caused by exposure to chemicals, ultraviolet light, and smoking. A broad range of structural variations, including deletions, insertions, inversions, tandem duplications, translocations, and complex rearrangements, can affect the function of normal genes in cancer.

[0004] Each cancer type involves key genes, and identifying mutations in these genes can help determine the stage of cancer progression. Furthermore, even cancers arising in the same organ exhibit distinct combinations of genetic mutations (heterogeneity). These genetic differences can influence cancer prognosis, treatment, and recurrence.

[0005] Since the first human genome was sequenced in 2003, the development and application of genome analysis technology has expanded in various ways. With the development of analytical technologies that can identify multiple genetic mutations in a single test, utilizing next-generation sequencing (NGS), cancer genome diagnosis has begun to play a crucial role in effective cancer treatment. Coupled with the success of developing targeted anticancer treatments targeting mutant proteins generated by somatic mutations in key genes such as EGFR and PIK3CA, clinical demand for cancer genome diagnosis is increasing.

[0006] For genomic analysis, ideally, cancer tissue should be immediately extracted for nucleic acid analysis within 10 minutes of the cessation of vascular supply during surgery, or frozen in liquid nitrogen and stored at cryogenic temperatures below -70℃. When necessary, nucleic acids should be extracted and DNA sequenced. However, it is impossible for hospitals to apply this method to surgical tissues from all patients who are not research subjects, and frozen tissues are not optimal samples for histopathological examination. Therefore, all body-derived tissues should be formalin-fixed and paraffin-embedded (FFPE), and then microdissected for histopathological examination. FFPE blocks should be stored at room temperature for at least 10 years or indefinitely. The need for actual cancer genome analysis is determined after the results of pathological tissue tests using FFPE tissue are available. Because of the size of cancer tissue that can be obtained through surgery or biopsy, there is often no surplus tissue that can be frozen after the tissue required for diagnostic FFPE blocks is secured. In addition, many hospitals do not have the infrastructure for frozen storage, so most clinical genome analyses are conducted using FFPE samples.

[0007] While FFPE samples offer the advantages of convenient storage and optimal use for pathological tissue examination, they also present disadvantages for genomic analysis. Specifically, chemical reactions occurring during formalin fixation can induce artifacts, unrelated to cancer. This can lead to mutations as high as 10-fold higher than the original mutations during DNA sequencing using NGS techniques. By targeting specific hotspot driver mutations in key oncogenes, panels can be developed to detect these mutations with high accuracy. However, distinguishing between true and artifactual mutations is difficult during exome or genome-level analysis. The presence of artifactual mutations leads to inaccuracies in tumor mutation burden measurements and neoantigen discovery, posing a significant obstacle to cutting-edge medical applications such as immunotherapy response prediction and neoantigen vaccine development. Although methods for removing artificial mutations from FFPE NGS DNA sequencing data have been developed, their clinical applicability is limited because they simultaneously remove most real mutations while removing artificial mutations. Therefore, the need for a technology that removes most artificial mutations while preserving real mutations has been highlighted.

[0008] The background technology of the invention has been prepared to facilitate a better understanding of the present invention. It should not be construed as an admission that the matters described in the background technology of the invention constitute prior art.

[0009] The inventors of the present invention recognized that the vast majority of general clinical samples are FFPE tissues, and that NGS analysis is difficult due to the above-mentioned problem in the case of old FFPE tissues or samples in which mutations occurred during the fixation process.

[0010] That is, the inventors of the present invention recognized that when determining somatic mutations using existing FFPE samples, the accuracy is lowered due to the appearance of artificial mutations by various factors, so a method that can accurately predict only somatic mutations is needed, and then noted that the accuracy can be increased when using a deep learning neural network model as in the present invention.

[0011] In particular, the inventors of the present invention were able to analyze paired NGS data generated from frozen tissue and FFPE tissue derived from the same tissue, define mutations found only in FFPE as artificial mutations, and perform deep learning with NGS data to predict somatic mutations compared to artificial mutations with higher accuracy.

[0012] At this time, the inventors of the present invention expected that the artificial neural network model would have higher prediction performance when learning was applied to a sequence containing a mutation position sequence and two or more bases before and after the mutation position sequence.

[0013] Accordingly, the inventors of the present invention sought to construct an artificial neural network model trained to classify artificial mutations and somatic mutations based on data necessary for mutation identification, including the surrounding sequences where the mutation is located.

[0014] Furthermore, the inventors of the present invention sought to construct an artificial neural network model that is independent of a variant caller and can receive base sequence data in various formats.

[0015] As a result, the inventors of the present invention were able to confirm that it is possible to distinguish between artificial mutations and somatic mutations with higher accuracy by applying the constructed artificial neural network model.

[0016] As a result, the inventors of the present invention have developed an information providing system that distinguishes between artificial mutations and somatic mutations generated from FFPE based on the deep learning artificial neural network model of the present invention.

[0017] Accordingly, the inventors of the present invention were able to recognize that the limitations of conventional diagnosis and information provision systems can be supplemented by introducing an information provision system that distinguishes between artificial mutations and somatic mutations generated by FFPE based on the deep learning artificial neural network model of the present invention, and thus, an efficient treatment method can be suggested.

[0018] Therefore, the problem to be solved by the present invention is to provide a method for providing information on somatic mutation prediction through artificial mutation removal and preservation of somatic mutation samples based on deep learning of the present invention, and a device using the same.

[0019] The tasks of the present invention are not limited to the tasks mentioned above, and other tasks not mentioned will be clearly understood by those skilled in the art from the description below.

[0020] In order to solve the above-described problem, a method for providing information on somatic mutation prediction based on deep learning according to one embodiment of the present invention is provided. The method for providing information is implemented by a processor, and includes the steps of: receiving data on mutations obtained from medical data and / or biological samples isolated from an individual; calculating data necessary for determining an artificial mutation or a somatic mutation based on the received data, wherein the calculated data includes a reference sequence including N bases preceding and following a position of an artificial mutation or a somatic mutation, wherein N includes a natural number of 2 or more; and using an artificial neural network model configured to predict an artificial mutation or a somatic mutation by inputting data necessary for determining a somatic mutation, the step of predicting an artificial mutation or a somatic mutation for a sample obtained from an individual by inputting the calculated data.

[0021] At this time, the medical data and / or data obtained from a biological sample isolated from the individual may be data about mutations, specifically data about somatic mutations, and the medical data may be NGS (Next Generation Sequencing) data, specifically data about VCF (Variant Call Format) files, BAM (Binary Alignment Map) files, WGS (Whole genome sequence), WES (whole exome sequencing), Targeted sequencing files, etc., which are base sequences analyzed for the entire genome, and the biological sample isolated from the individual may mean data about mutations observed in the sample, specifically data about somatic mutations observed, but is not limited thereto.

[0022] According to a feature of the present invention, the data produced through the step of producing data necessary for determining artificial mutation or somatic mutation based on the received data may include a reference sequence including two or more bases before and after the mutation position.

[0023] According to another feature of the present invention, the step of generating data necessary for determining artificial mutation or somatic mutation based on the received data may further include the step of vectorizing the received data and the step of calculating similarity based on the vectorized data. In this case, in the step of calculating similarity based on the vectorized data, the similarity may be cosine similarity, but is not limited thereto.

[0024] According to another feature of the present invention, in the step of predicting an artificial mutation or somatic mutation for a sample obtained from an individual by inputting the produced data, the artificial neural network model may be configured to distinguish the type of mutation by adopting an encoder and / or decoder-based structure, but is not limited thereto.

[0025] In a specific example, an artificial neural network model based on an encoder and / or decoder receives as input data including a reference sequence including two or more bases before and after a mutation location, i.e., base sequence context data, extracts key features such as patterns therefrom, and ultimately classifies and outputs FFPE artificial mutations or actual somatic mutations.

[0026] In various embodiments, the encoder unit may be comprised of a self-attention layer capable of calculating the relationship between each base sequence of an input sequence and all other bases, and a feed-forward neural network that processes and transforms information about each base. Furthermore, the decoder unit may be comprised of a self-attention layer that performs self-attention on input values ​​(previous bases) of the decoder unit, an encoder-decoder attention layer that calculates the relationship between the output values ​​of the encoder unit and the input values ​​of the decoder, and a feed-forward neural network that processes and transforms information about each base.

[0027] An artificial neural network model with these structural features can capture relationships between base pairs from multiple perspectives using a multi-head attention mechanism based on sequence context data containing information about mutations and their surrounding base pairs (reference sequences). Optionally, a normalization layer can be placed after each layer to stabilize learning and prevent overfitting.

[0028] That is, the artificial neural network model learns how mutant bases interact with surrounding bases in the input data, and through this, it can distinguish and classify artificial mutations and somatic mutations.

[0029] However, the mutation classification process is not limited to the artificial neural network model based on the encoder and / or decoder structure described above, and may be based on a transformer model based on pattern analysis of an image, such as Transformer, GPT, T5 (Text-To-Text Transfer Transformer), BERT (Bidirectional Encoder Representations from Transformer), ViT (Vision Transformer), PiT (Pooling-based Vision Transformer), CvT (Convolutional Vision Transformer), CrossFormer, CrossViT, NesT, MaxViT, and SepViT (Separable Vision Transformer), or various algorithms, such as CNN (Convolution Neural Network), DNN (Deep Neural Network), DCNN (Deep Convolution Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), SSD (Single Shot Detector), and SVM (Support Vector Machine).

[0030] According to another feature of the present invention, in the step of predicting the artificial mutation or somatic mutation, a step of predicting the artificial mutation or somatic mutation for the sample using the artificial neural network model based on the produced data and the medical data may be performed.

[0031] According to another feature of the present invention, in the step of generating data necessary for determining a mutation or an artificial mutation based on the received data, the generated data may include cosine similarity between vectors defined by the number of sequences mapped to the reverse strand of read 2 or cosine similarity of the number of sequences mapped to the reverse strand, and the number of reads in which the reference sequence is observed / sequence depth at the corresponding position, the number of reads in which the alternate sequence is observed / sequence depth at the corresponding position, the ratio of the difference between the number of sequences mapped to the forward strand of read 1 and the number of sequences mapped to the reverse strand in the read in which the reference sequence is observed in the read, the ratio of the difference between the number of sequences mapped to the forward strand of read 1 and the number of sequences mapped to the reverse strand in the read in which the alternate sequence is observed in the read, and the vector (in which the reference sequence is observed) (the number of sequences in which read 1 is mapped to the forward strand in the read, the number of sequences in which read 2 is mapped to the reverse strand in the read) and (the number of sequences in which read 1 is mapped to the forward strand in the read in which an alternative sequence is observed, the median of the length of inserts in the read in which an alternative sequence is observed divided by the median of the length of inserts in the read in which the reference sequence is observed, the length of double-sequenced bases in the reference fragment, the number of times a sequence is read twice by read 1 and read 2 in the reference fragment,The length of the double-sequenced bases of the sequence read twice by Read 1 and Read 2 in the alternate fragment, the number of times the sequence is read twice by Read 1 and Read 2 in the alternate fragment, the reference sequence (REF_N_BASES) including one base before and after the mutation position, the alternative sequence (ALT_N_BASES) including one base before and after the mutation position, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference sequence is observed in the reads in which the reference sequence is observed, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the alternative sequence is observed in the reads in which the alternative sequence is observed, and the vector (the number of sequences mapped to the forward strand among the reads in which the reference sequence is observed, the number of sequences mapped to the reverse strand) and (the number of sequences mapped to the forward strand among the reads in which the alternative sequence is observed) may further include, but is not limited to, any one or more selected from the group consisting of no.,

[0032] In one embodiment of the present invention, the cosine similarity can be expressed by the following mathematical formula.

[0033] [Mathematical Formula 1]

[0034]

[0035] According to a feature of the present invention, an artificial neural network model configured to predict artificial mutations or somatic mutations by inputting data necessary for somatic mutation determination is used to predict artificial mutations or somatic mutations for a sample obtained from an individual by inputting the produced data. In order to solve the above-described problem, an information providing method for predicting somatic mutations based on deep learning according to an embodiment of the present invention is provided. The present information providing method is implemented by a processor, and includes the steps of: receiving data on a mutation obtained from medical data and / or a biological sample isolated from an individual; producing data necessary for determining an artificial mutation or a somatic mutation based on the received data, wherein the produced data includes a reference sequence including N bases before and after a position of an artificial mutation or a somatic mutation, wherein N includes a natural number of 2 or more; and using an artificial neural network model configured to predict an artificial mutation or a somatic mutation by inputting data necessary for determining a somatic mutation, the step of predicting an artificial mutation or a somatic mutation for a sample obtained from an individual by inputting the produced data.

[0036] At this time, the medical data and / or data obtained from a biological sample isolated from the individual may be data about mutations, specifically data about somatic mutations, and the medical data may be NGS (Next Generation Sequencing) data, specifically data about VCF (Variant Call Format) files, BAM (Binary Alignment Map) files, WGS (Whole genome sequence), WES (whole exome sequencing), Targeted sequencing files, etc., which are base sequences analyzed for the entire genome, and the biological sample isolated from the individual may mean data about mutations observed in the sample, specifically data about somatic mutations observed, but is not limited thereto.

[0037] According to a feature of the present invention, the data produced through the step of producing data necessary for determining artificial mutation or somatic mutation based on the received data may include a reference sequence including two or more bases before and after the mutation position.

[0038] According to another feature of the present invention, the step of generating data necessary for determining artificial mutation or somatic mutation based on the received data may further include the step of vectorizing the received data and the step of calculating similarity based on the vectorized data. In this case, in the step of calculating similarity based on the vectorized data, the similarity may be cosine similarity, but is not limited thereto.

[0039] According to another feature of the present invention, in the step of predicting an artificial mutation or somatic mutation for a sample obtained from an individual using the produced data as input, the artificial neural network model may be configured to distinguish the type of mutation by adopting an encoder and / or decoder-based structure, but is not limited thereto.

[0040] In a specific example, an artificial neural network model based on an encoder and / or decoder receives as input data including a reference sequence including two or more bases before and after a mutation location, i.e., base sequence context data, extracts key features such as patterns therefrom, and ultimately classifies and outputs FFPE artificial mutations or actual somatic mutations.

[0041] In various embodiments, the encoder unit may be comprised of a self-attention layer capable of calculating the relationship between each base sequence of an input sequence and all other bases, and a feed-forward neural network that processes and transforms information about each base. Furthermore, the decoder unit may be comprised of a self-attention layer that performs self-attention on input values ​​(previous bases) of the decoder unit, an encoder-decoder attention layer that calculates the relationship between the output values ​​of the encoder unit and the input values ​​of the decoder, and a feed-forward neural network that processes and transforms information about each base.

[0042] An artificial neural network model with these structural features can capture relationships between base pairs from multiple perspectives using a multi-head attention mechanism based on sequence context data containing information about mutations and their surrounding base pairs (reference sequences). Optionally, a normalization layer can be placed after each layer to stabilize learning and prevent overfitting.

[0043] That is, the artificial neural network model learns how mutant bases interact with surrounding bases in the input data, and through this, it can distinguish and classify artificial mutations and somatic mutations.

[0044] However, the mutation classification process is not limited to the artificial neural network model based on the encoder and / or decoder structure described above, and may be based on a transformer model based on pattern analysis of an image, such as Transformer, GPT, T5 (Text-To-Text Transfer Transformer), BERT (Bidirectional Encoder Representations from Transformer), ViT (Vision Transformer), PiT (Pooling-based Vision Transformer), CvT (Convolutional Vision Transformer), CrossFormer, CrossViT, NesT, MaxViT, and SepViT (Separable Vision Transformer), or various algorithms, such as CNN (Convolution Neural Network), DNN (Deep Neural Network), DCNN (Deep Convolution Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), SSD (Single Shot Detector), and SVM (Support Vector Machine).

[0045] According to another feature of the present invention, in the step of predicting the artificial mutation or somatic mutation, a step of predicting the artificial mutation or somatic mutation for the sample using the artificial neural network model based on the produced data and the medical data may be performed.

[0046] According to another feature of the present invention, in the step of generating data necessary for determining a mutation or an artificial mutation based on the received data, the generated data may include cosine similarity between vectors defined by the number of sequences mapped to the reverse strand of read 2 or cosine similarity of the number of sequences mapped to the reverse strand, and the number of reads in which the reference sequence is observed / sequence depth at the corresponding position, the number of reads in which the alternate sequence is observed / sequence depth at the corresponding position, the ratio of the difference between the number of sequences mapped to the forward strand of read 1 and the number of sequences mapped to the reverse strand in the read in which the reference sequence is observed in the read, the ratio of the difference between the number of sequences mapped to the forward strand of read 1 and the number of sequences mapped to the reverse strand in the read in which the alternate sequence is observed in the read, and the vector (in which the reference sequence is observed) (the number of sequences in which read 1 is mapped to the forward strand in the read, the number of sequences in which read 2 is mapped to the reverse strand in the read) and (the number of sequences in which read 1 is mapped to the forward strand in the read in which an alternative sequence is observed, the median of the length of inserts in the read in which an alternative sequence is observed divided by the median of the length of inserts in the read in which the reference sequence is observed, the length of double-sequenced bases in the reference fragment, the number of times a sequence is read twice by read 1 and read 2 in the reference fragment,The length of the double-sequenced bases of the sequence read twice by Read 1 and Read 2 in the alternate fragment, the number of times the sequence is read twice by Read 1 and Read 2 in the alternate fragment, the reference sequence (REF_N_BASES) including one base before and after the mutation position, the alternative sequence (ALT_N_BASES) including one base before and after the mutation position, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference sequence is observed in the reads in which the reference sequence is observed, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the alternative sequence is observed in the reads in which the alternative sequence is observed, and the vector (the number of sequences mapped to the forward strand among the reads in which the reference sequence is observed, the number of sequences mapped to the reverse strand) and (the number of sequences mapped to the forward strand among the reads in which the alternative sequence is observed) may further include, but is not limited to, any one or more selected from the group consisting of no.,

[0047] In one embodiment of the present invention, the cosine similarity can be expressed by the following mathematical formula. In the step of, the sample obtained from the subject may be an FFPE sample, and the subject may be a subject with cancer. In this case, the cancer may be a subject with any one of liver cancer, colon cancer, breast cancer, lung cancer, and fibroma, but is not limited thereto.

[0048] In order to solve the aforementioned problem, an information providing system including a device for providing information on somatic mutation prediction, a mutation prediction device, and a server for providing medical data according to an embodiment of the present invention is provided.

[0049] Specific details of other embodiments are included in the detailed description and drawings.

[0050] The present invention can provide information on prediction of somatic mutations, enabling tailored personal medicine, by providing information on prediction results for artificial mutations or somatic mutations using a deep learning artificial neural network model.

[0051] More specifically, the present invention can provide an artificial neural network model trained to distinguish somatic mutations or artificial mutations by inputting various formats of base sequence data without relying on a variant caller.

[0052] In particular, the present invention, in the process of generating data related to somatic mutations, further considers the surrounding base sequence (reference sequence) of the mutation location and applies this to an artificial neural network model to determine a prediction result for artificial mutations or somatic mutations, thereby enabling more accurate prediction of artificial mutations or somatic mutations.

[0053] That is, the present invention utilizes information on prediction of artificial mutations or somatic mutations based on a deep-learning artificial neural network model, thereby enabling more precise setting of treatment directions for individuals and contributing to the selection of customized treatment methods.

[0054] The present invention predicts artificial mutations and somatic mutations with high accuracy by applying a deep learning artificial neural network model trained with learning data including cosine similarity for prediction of artificial mutations or somatic mutations, thereby complementing the limitations of using conventional FFPE samples.

[0055] The effects according to the present invention are not limited to those exemplified above, and more diverse effects are included in this specification.

[0056] FIG. 1 illustrates an information providing system for predicting artificial mutations or somatic mutations using a prediction device for artificial mutations or somatic mutations according to one embodiment of the present invention.

[0057] FIG. 2 illustrates an example of a configuration of an artificial mutation or somatic mutation prediction device according to one embodiment of the present invention.

[0058] FIG. 3 is a block diagram of a procedure of a method for providing information on prediction of artificial mutation or somatic mutation according to one embodiment of the present invention.

[0059] FIGS. 4A and 4B illustrate exemplary procedures of a method for providing information on artificial mutations or somatic mutations according to various embodiments of the present invention.

[0060] FIG. 5 is a diagram showing the results of evaluating the prediction performance of an artificial neural network model trained to classify mutations by inputting a reference sequence together with various embodiments of the present invention.

[0061] Figures 6A to 6E and 7A to 7E are diagrams showing the results of evaluating the prediction performance of an artificial neural network model according to the number of reference sequences.

[0062] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, but may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. In connection with the description of the drawings, similar reference numerals may be used for similar components.

[0063] In this document, the expressions "has," "may have," "includes," or "may include" indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), but do not exclude the presence of additional features.

[0064] In this document, the expressions "A or B," "at least one of A and / or B," or "one or more of A or / and B" can include all possible combinations of the listed items. For example, "A or B," "at least one of A and B," or "at least one of A or B" can all refer to cases where (1) at least one A is included, (2) at least one B is included, or (3) at least one A and at least one B are included.

[0065] The terms "first," "second," "first," or "second," as used herein, may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, without limiting the components. For example, a first user device and a second user device may represent different user devices, regardless of order or importance. For example, without departing from the scope of the rights set forth in this document, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.

[0066] When it is said that a component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the component is directly coupled to the other component, or can be connected via another component (e.g., a third component). Conversely, when it is said that a component (e.g., a first component) is "directly coupled to" or "directly connected to" another component (e.g., a second component), it should be understood that no other component (e.g., a third component) exists between the first component and the other component.

[0067] The expression "configured to" as used herein can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something is "specifically designed to" in hardware. Instead, in some contexts, a "device configured to" can mean that the device, together with other devices or components, is "capable of." For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing those operations, or a generic-purpose processor (e.g., a CPU, GPU, or application processor) that can perform those operations by executing one or more software programs stored in a memory device.

[0068] The terms used in this document are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include the plural expression unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this document. Terms defined in general dictionaries among the terms used in this document may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this document. In some cases, even if a term is defined in this document, it cannot be interpreted to exclude the embodiments of this document.

[0069] The individual features of the various embodiments of the present invention can be partially or wholly combined or combined with each other, and as can be fully understood by those skilled in the art, various technical connections and operations are possible, and each embodiment can be implemented independently of each other or can be implemented together in a related relationship.

[0070] For clarity in the interpretation of this specification, the terms used in this specification are defined below.

[0071] As used herein, the term "subject" may refer to any subject for which mutations, specifically somatic mutations, are to be predicted. For example, the subject may be an individual in which a somatic mutation has occurred. In this case, the subject disclosed herein may be any mammal other than a human, but is not limited thereto.

[0072] The term "biological sample" as used herein refers to any biological sample that is isolated from an individual and can be stored in FFPE format, including, but not limited to, urine, tissue, cell lysate, whole blood, plasma, serum, saliva, ocular fluid, cerebrospinal fluid, sweat, milk, ascites fluid, synovial fluid, and peritoneal fluid. Preferably, the biological sample may be a tissue sample isolated from an individual for which somatic cell mutations are to be confirmed, but is not limited thereto. Specifically, the biological sample of the present invention may be an FFPE sample (Formalin Fixed Paraffin Embedded), and an FFPE sample is a standard method used when collecting and storing cancer tissue from a patient, where the tissue is fixed with formalin and solidified in paraffin to make a paraffin block, so that it can be stored at room temperature, and may also mean a sample in which the paraffin block is thinly sliced ​​multiple times and used for various histological examinations.

[0073] The term "somatic variant" used herein means that one of the characteristics of cancer confirmed in cells is that it is accompanied by genetic abnormalities, and thus, by confirming somatic variants, it is possible to distinguish the occurrence and type of cancer. Therefore, in the present invention, data obtained from medical data and / or biological samples isolated from an individual may be data on somatic mutations, and medical data may be NGS (Next Generation Sequencing) data, but is not limited thereto, and may mean that somatic variants are observed in biological samples isolated from an individual, but is not limited thereto.

[0074] As used herein, the term "artifact" may refer to, but is not limited to, an artificial defect that occurs due to a physical or chemical alteration that is not related to cancer.

[0075] The term "NGS (Next Generation Sequencing)" used in this specification refers to next-generation base sequence analysis, which is one of the high-speed analysis methods of base sequences of a genome.

[0076] At this time, NGS can be applied to genome-wide and whole-genome analyses to achieve various purposes, including clinical research. Meanwhile, most NGS analyses can simultaneously analyze multiple target samples.

[0077] As used herein, the term “NGS data” may mean a file containing base sequence data analyzed for a target sample.

[0078] At this time, NGS data may include a VCF (Variant Call Format) file, a BAM (Binary Alignment Map), a WGS (whole genome sequencing) file in which the base sequence is analyzed for the entire genome, a WES (whole exome sequencing) file, an RNA sequencing file, and a targeted sequencing file in which the base sequence is analyzed for a specific region, and is specifically, but not limited to, a VCF file and / or a BAM file.

[0079] The term "reference sequence" as used herein includes a mutation site sequence and its surrounding sequences, and may refer to a sequence including N bases preceding and following an artificial mutation site or somatic mutation site. In this case, the mutation site sequence may include a standard sequence based on the human reference genome or a mutation site sequence in which a mutation has occurred. For example, the reference sequence may be a sequence including N bases preceding and following a standard sequence in which no mutation has occurred, or a sequence including N bases preceding and following a mutation site in which a mutation has occurred at the same site.

[0080] In various embodiments, the variant position sequence may be, but is not limited to, a standard sequence.

[0081] In this specification, reference sequence may be used interchangeably with “base sequence context data.”

[0082] Here, N is a natural number greater than or equal to 2, preferably a natural number greater than or equal to 3, more preferably a natural number greater than or equal to 4, and even more preferably a natural number greater than or equal to 5, but is not limited thereto.

[0083] In a specific example, when N is 5, a total of 11 bases including the 5 bases before and after the mutation position sequence can be used as input data for the artificial neural network model.

[0084] As used herein, the term "artificial neural network model" or "predictive model" may be a model trained to classify and output artificial mutations or somatic mutations by inputting data necessary for artificial mutation or somatic mutation determination, particularly base sequence context data, including a reference sequence.

[0085] In various embodiments, the artificial neural network model may be a model trained to classify mutations using, but is not limited to, a reference sequence of a standard sequence in which no mutation has occurred and a sequence including N bases preceding and following the position thereof as training data.

[0086] In various embodiments, the artificial neural network model may be, but is not limited to, a model with an encoder / decoder. For example, the artificial neural network model may be based on various algorithms, such as a deep neural network (DNN), a direct neural network (DCNN), a recurrent neural network (RNN), a neural network model (RBM), a deep neural network (DBN), a solid-state drive (SSD), and a support vector machine (SVM).

[0087] Hereinafter, with reference to FIGS. 1 and 2, an information providing system for predicting artificial mutations or somatic mutations and a predicting device for artificial mutations or somatic mutations based on a predicting device for artificial mutations or somatic mutations according to one embodiment of the present invention are described.

[0088]

[0089] FIG. 1 illustrates an information providing system for predicting artificial mutations or somatic mutations using a device for predicting artificial mutations or somatic mutations according to one embodiment of the present invention.

[0090] FIG. 2 illustrates an example of a configuration of an artificial mutation or somatic mutation prediction device according to one embodiment of the present invention.

[0091]

[0092] First, referring to FIG. 1, the information providing system (1000) may be a system configured to provide information related to prediction of artificial mutation or somatic mutation by applying a deep-learning artificial neural network model based on data obtained from medical data and / or biological samples. At this time, the information providing system (1000) may be configured with a data providing server / device (100) that provides data obtained from medical data and / or biological samples, a prediction device that predicts artificial mutation or somatic mutation based on the received data, and an information providing device (300) that visually provides information on prediction of artificial mutation or somatic mutation for a sample obtained from an individual based on the received prediction data.

[0093] Here, the information providing device (300) is an electronic device that provides a user interface for displaying information related to prediction of artificial mutation or somatic mutation, and may include at least one of a smartphone, a tablet PC (personal computer), a laptop, and / or a PC.

[0094] A device (300) for providing information on prediction of somatic mutations can receive information related to prediction results of artificial mutations or somatic mutations from a somatic mutation prediction device (200) and display the received results through a display unit to be described later.

[0095] The somatic mutation prediction device (200) may include a general-purpose computer, laptop, and / or data server, etc., which performs various operations to determine information related to prediction of artificial mutation or somatic mutation based on various data (particularly medical data and / or data on somatic mutation) obtainable from a data providing server / device (100). In this case, the data providing server (100) may be a device for accessing a web server providing a web page or a mobile web server providing a mobile website, but is not limited thereto.

[0096] More specifically, the somatic mutation prediction device (200) can receive medical data and / or data on somatic mutations from the server (100) for providing medical data, and predict artificial mutations or somatic mutations based on the received data and provide information related thereto. At this time, the somatic mutation prediction device (200) can perform prediction of artificial mutations or somatic mutations based on medical data received from the server (100) for providing data and / or data on somatic mutations obtained from biological samples by applying a deep-learning artificial neural network model.

[0097] In addition, the somatic mutation prediction device (200) can provide artificial mutation or somatic mutation prediction results to a device (300) for providing information on somatic mutation prediction.

[0098] In this way, information provided from the somatic mutation prediction device (200) may be provided as a web page via a web browser installed in a device (300) for providing information on somatic mutation prediction, or may be provided in the form of an application or program. In various embodiments, such data may be provided in a form included in a platform in a client-server environment.

[0099] Hereinafter, a somatic cell mutation prediction device (200) according to one embodiment of the present invention is described as performing operations by receiving medical data and / or data obtained from a biological sample from a data providing server (100), but the device (300) itself for providing information on somatic cell prediction may perform all operations.

[0100] Next, with reference to FIG. 2, the components of the somatic cell mutation prediction device (200) of the present invention will be described in detail.

[0101] Referring to FIG. 2, the somatic cell mutation prediction device (200) may include a communication interface (210), a memory (220), an I / O interface (230), and a processor (240), and each component may communicate with each other through one or more communication buses or signal lines.

[0102] The communication interface (210) can be connected to a device (300) for providing information on somatic mutation prediction and a server (100) for providing medical data via a wired / wireless communication network to exchange data. For example, the communication interface (210) can receive data on somatic mutation and / or individual data from the server (100) for providing medical data, and can transmit information on a determined artificial mutation or somatic mutation prediction result to the device (300) for providing information on somatic mutation prediction.

[0103] Meanwhile, the communication interface (210) that enables transmission and reception of such data includes a communication port (211) and a wireless circuit (212), wherein the wired communication port (211) may include one or more wired interfaces, for example, Ethernet, Universal Serial Bus (USB), FireWire, etc. In addition, the wireless circuit (212) may transmit and receive data with an external device via an RF signal or an optical signal. In addition, the wireless communication may use at least one of a plurality of communication standards, protocols, and technologies, for example, GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol.

[0104] The memory (220) can store various data used in the somatic mutation prediction device (200). For example, the memory (220) can store data on somatic mutations, or store a deep learning model configured to predict artificial mutations or somatic mutations based thereon.

[0105] In various embodiments, the memory (220) may include a volatile or non-volatile storage medium capable of storing various data, commands, and information. For example, the memory (220) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), RAM, SRAM, ROM, EEPROM, PROM, network storage, cloud, and blockchain database.

[0106] In various embodiments, the memory (220) may store configurations of at least one of an operating system (221), a communication module (222), a user interface module (223), and one or more applications (224).

[0107] An operating system (221) (e.g., an embedded operating system such as LINUX, UNIX, MAC OS, WINDOWS, VxWorks, etc.) may include various software components and drivers to control and manage general system operations (e.g., memory management, storage device control, power management, etc.) and may support communication between various hardware, firmware, and software components.

[0108] The communication module (223) can support communication with other devices through the communication interface (210). The communication module (220) can include various software components for processing data received by the wired communication port (211) or wireless circuit (212) of the communication interface (210).

[0109] The user interface module (223) can receive a user's request or input from a keyboard, touch screen, microphone, etc. through an I / O interface (230) and provide a user interface on the display.

[0110] The application (224) may include a program or module configured to be executed by one or more processors (230). Here, the application for providing information related to somatic mutation prediction may be implemented on a server farm.

[0111] The I / O interface (230) can connect at least one of an input / output device (not shown) of the somatic cell mutation prediction device (200), such as a display, a keyboard, a touch screen, and a microphone, to the user interface module (223). The I / O interface (230) can receive user input (e.g., voice input, keyboard input, touch input, etc.) together with the user interface module (223) and process a command according to the received input.

[0112] The processor (240) is connected to a communication interface (210), a memory (220), and an I / O interface (230) to control the overall operation of the somatic cell mutation prediction device (200), and can perform various commands for providing information through an application or program stored in the memory (220).

[0113] The processor (240) may correspond to a computing device such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an AP (Application Processor). In addition, the processor (240) may be implemented in the form of an integrated chip (Integrated Chip (IC)) such as a SoC (System on Chip) in which various computing devices are integrated. Alternatively, the processor (240) may include a module for calculating an artificial neural network model, such as an NPU (Neural Processing Unit).

[0114] In various embodiments, the processor (240) may be configured to apply a deep learning artificial neural network model to produce data necessary for determining artificial mutations or somatic mutations based on data obtained from medical data and / or biological samples, and to predict and provide artificial mutations or somatic mutations.

[0115] In various embodiments, data required for artificial mutation or somatic mutation determination may be data including a reference sequence including two or more bases before and after the mutation location.

[0116] In more diverse embodiments, the somatic mutation prediction device (200) may receive medical data and / or data obtained from a biological sample isolated from an individual from the medical data providing device (100), and may use the received data to apply a deep-learned artificial neural network model to predict whether there is an artificial mutation or a somatic mutation, and may provide a prediction result for a somatic mutation using the learned model. At this time, the data obtained from the medical data and / or the biological sample isolated from the individual may be data for a somatic mutation, and the medical data may be, but is not limited to, NGS (Next Generation Sequencing) data, and may mean, but is not limited to, data on observation of a somatic mutation in a biological sample isolated from the individual. Specifically, the NGS data may be, but is not limited to, a VCF file and / or a BAM file.

[0117]

[0118] Hereinafter, with reference to FIG. 3 and FIG. 4A to 4C, a method for providing information on prediction of artificial mutation or somatic mutation according to one embodiment of the present invention will be specifically described.

[0119] FIG. 3 is a block diagram of a procedure of a method for providing information on prediction of artificial mutation or somatic mutation according to one embodiment of the present invention.

[0120] FIGS. 4A and 4B illustrate exemplary procedures of a method for providing information on artificial mutations or somatic mutations according to various embodiments of the present invention.

[0121]

[0122] First, referring to FIG. 3, the procedure of a method for providing information on prediction of artificial mutation or somatic mutation according to one embodiment of the present invention is as follows.

[0123] First, data on mutations occurring in medical data and / or biological samples isolated from an individual are received (S310).

[0124] In various embodiments, the sample obtained from the subject may be an FFPE sample, and the subject may be a subject having any one of, but not limited to, liver cancer, colon cancer, breast cancer, lung cancer, and fibroadenoma.

[0125] Next, based on the received data, data required for determining artificial mutation or somatic mutation is generated (S320), and using an artificial neural network model configured to predict artificial mutation or somatic mutation by inputting data required for determining somatic mutation, artificial mutation or somatic mutation is predicted for a sample obtained from an individual by inputting the generated data (S330).

[0126] According to a feature of the present invention, in the step (S310) of receiving data on a mutation that occurred in medical data and / or a biological sample separated from an individual, a WES (whole exome sequencing) file, an RNA sequencing file, or a targeted sequencing file may be received, and specifically, data in the format of a VCF file and / or a BAM file may be received, but is not limited thereto. That is, base sequence data in various formats independent of a mutation caller may be received.

[0127] Next, in a step (S320) where data required for determining artificial mutation or somatic mutation is generated based on the received data, data including a reference sequence including two or more bases before and after the mutation position can be generated.

[0128] According to various embodiments of the present invention, in the step (S320) where data necessary for determining artificial mutation or somatic mutation is generated based on the received data, a step of vectorizing the received data and a step of calculating similarity based on the vectorized data may be further performed.

[0129] At this time, referring to Table 1 together, in the step of producing data necessary for determining a mutation or artificial mutation based on the received data, the produced data includes the number of reads in which the reference sequence is observed / sequence depth at the corresponding position, the number of reads in which the alternate sequence is observed / sequence depth at the corresponding position, the ratio of the difference between the number of sequences in which Read 1 is mapped to the forward strand and the number of sequences in which Read 2 is mapped to the reverse strand in the read in which the reference sequence is observed in the read, the ratio of the difference between the number of sequences in which Read 1 is mapped to the forward strand and the number of sequences in which Read 2 is mapped to the reverse strand in the read in which the alternate sequence is observed in the read, the median of the inserted length in the read in which the alternate sequence is observed divided by the median of the inserted length in the read in which the reference sequence is observed, and the double-sequenced sequence read twice by Read 1 and Read 2 in the reference fragment. The length of the double-sequenced bases in the reference fragment, the number of times the sequence was read twice by Read 1 and Read 2 in the alternate fragment, the length of the double-sequenced bases in the alternate fragment, the number of times the sequence was read twice by Read 1 and Read 2 in the alternate fragment, the reference sequence (REF_3_BASES) including N bases before and after the mutation position (where N is a natural number greater than or equal to 2), the alternate sequence (ALT_3_BASES) including N bases before and after the mutation position, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference sequence is observed in the reads in which the reference sequence is observed,The cosine similarity may further include, but is not limited to, one or more selected from the group consisting of the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the alternative sequence is observed in the reads in which the alternative sequence is observed, and the cosine similarity between vectors defined by (the number of sequences mapped to the forward strand by Read 1 in the reads in which the reference sequence is observed, the number of sequences mapped to the reverse strand by Read 2 in the reads in which the reference sequence is observed) or the cosine similarity between vectors defined by (the number of sequences mapped to the forward strand by Read 1 in the reads in which the reference sequence is observed, the number of sequences mapped to the reverse strand) and (the number of sequences mapped to the forward strand by Read 2 in the reads in which the alternative sequence is observed, the number of sequences mapped to the reverse strand).

[0130] No. Data related to somatic mutation determination 1 Number of reads in which the reference sequence is observed / sequence depth at that position 2 Number of reads in which the alternate sequence is observed / sequence depth at that position 3 Proportion of the difference between fr_ref (the number of sequences in which read 1 is mapped to the forward strand in the read in which the reference sequence is observed) and rf_ref (the number of sequences in which read 2 is mapped to the reverse strand) in the read in which the reference sequence is observed 4 Proportion of the difference between fr_alt (the number of sequences in which read 1 is mapped to the forward strand in the read in which the alternate sequence is observed) and rf_alt (the number of sequences in which read 2 is mapped to the reverse strand) in the read in which the alternate sequence is observed 5 A vector defined by (fr_ref, rf_ref) and (fr_alt,rf_alt) cosine similarity between vectors defined by 6 median insert length in reads where the alternate sequence is observed divided by the median insert length in reads where the reference sequence is observed 7 length of double-sequenced bases in the reference fragment 8 number of times the sequence is read twice by read 1 and read 2 in the reference fragment 9 length of double-sequenced bases in the alternate fragment 10 number of times the sequence is read twice by read 1 and read 2 in the alternate fragment 11 REF_N_BASES: Reference sequence containing N bases before and after the mutation position centered 12 ALT_N_BASES: Alternate sequence containing N bases before and after the mutation position centered 13 fs_ref (forward base of the read where the reference sequence is observed) The proportion of reads in which the reference sequence is observed that is the difference between fs_alt (the number of sequences mapped to the forward strand among reads in which an alternate sequence is observed) and rs_ref (the number of sequences mapped to the reverse strand). The proportion of reads in which an alternate sequence is observed that is the difference between fs_alt (the number of sequences mapped to the forward strand among reads in which an alternate sequence is observed) and rs_alt (the number of sequences mapped to the reverse strand). The cosine similarity between the vector defined by (fs_ref, rs_ref) and the vector defined by (fs_alt, rs_alt).

[0131] As used herein, the term "read" may refer to DNA including an adapter, and the term "fragment" may refer to DNA excluding an adapter. As used herein, the term "read1" refers to a paired end DNA that is read from both ends when entering the NGS equipment and is read in the 5' to 3' direction, and "read2" refers to a paired end DNA that is read in the 3' to 5' direction.

[0132] The term "depth" as used herein may be used interchangeably with the term "read-depth" and means the thickness or depth of a read.

[0133] As used herein, the term “fr_ref” refers to the number of sequences in which Read 1 maps to the forward strand in a read where the reference sequence is observed.

[0134] As used herein, the term “rf_ref” refers to the number of sequences to which read2 is mapped to the reverse strand.

[0135] As used herein, the term “fr_alt” refers to the number of sequences in which read 1 maps to the forward strand in a read in which an alternative sequence is observed.

[0136] As used herein, the term “rf_alt” refers to the number of sequences to which read2 maps to the reverse strand.

[0137] As used herein, the term “fs_ref” refers to the number of sequences that map to the forward strand among the reads in which the reference sequence is observed.

[0138] As used herein, the term “rs_ref” means the number of sequences mapped to the reverse strand.

[0139] As used herein, the term “fs_alt” refers to the number of sequences mapped to the forward strand among reads in which an alternative sequence is observed.

[0140] As used herein, the term “rs_alt” refers to the number of sequences mapped to the reverse strand.

[0141] Referring back to FIGS. 3 and 4A, a step (S330) is performed to predict artificial mutations or somatic mutations for samples obtained from an individual by using an artificial neural network model configured to predict artificial mutations or somatic mutations by inputting data required for somatic mutation determination. At this time, a prediction result for an artificial mutation or somatic mutation for a sample obtained from an individual is determined and provided. For example, the model may be configured to output whether the result value is a real artificial mutation, a fake artificial mutation, a real somatic mutation, or a fake somatic mutation. For example, the output may be expressed as a percentage, and a threshold may be set so that a case where the percentage is greater than the threshold may be determined to be a real artificial mutation, a fake artificial mutation, a real somatic mutation, or a fake somatic mutation.

[0142] In various embodiments of the present invention, an artificial neural network-based model that is input with data on mutations occurring in medical data and / or biological samples isolated from an individual may be an artificial neural network having an encoder / decoder and capable of extracting features from contextual data of a sequence.

[0143] For example, referring to FIG. 4B together, the artificial neural network model (420) may be equipped with an encoder unit (422) and a decoder unit (424). More specifically, when base sequence context data (4122) having a reference sequence including two or more bases before and after the generated mutation position is input, the relationship between the mutation and the surrounding bases is analyzed through the encoder unit. Then, processing such as self-attention is performed and converted through the decoder unit (424), and finally, an artificial mutation or somatic mutation (432) may be classified and output through an output layer (not shown).

[0144] At this time, the encoder unit (422) may be composed of a self-attention layer that can calculate the relationship between each base sequence of the input sequence and all other bases, and a feed-forward neural network that processes and transforms information about each base. Furthermore, the decoder unit (424) may be composed of a self-attention layer that performs self-attention on the input values ​​(previous bases) of the decoder unit (424), an encoder-decoder attention layer that calculates the relationship between the output value of the encoder unit and the input value of the decoder, and a feed-forward neural network that processes and transforms information about each base. However, the present invention is not limited thereto.

[0145] Optionally, NGS data (4124) in the format of a VCF file and / or a BAM file may be input together with an artificial neural network model (420) to perform mutation determination.

[0146] Meanwhile, the structure of the artificial neural network model is not limited to this, and may be based on various algorithms such as a transformer model based on pattern analysis of images, such as Transformer, BERT (Bidirectional Encoder Representations from Transformers), GPT, T5 (Text-To-Text Transfer Transformer), ViT (Vision Transformer), PiT (Pooling-based Vision Transformer), CvT (Convolutional Vision Transformer), CrossFormer, CrossViT, NesT, MaxViT, and SepViT (Separable Vision Transformer), or a CNN (Convolution Neural Network), DNN (Deep Neural Network), DCNN (Deep Convolution Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), SSD (Single Shot Detector), and SVM (Support Vector Machine).

[0147] The information providing procedure for prediction of artificial mutation or somatic mutation is not limited to the above.

[0148] According to the method for providing information on prediction of artificial mutation or somatic mutation, the present invention provides a system for providing information on prediction of artificial mutation or somatic mutation based on an algorithm, thereby enabling easier setting of treatment direction for an individual and contributing to selection of a customized treatment method.

[0149] That is, the present invention can complement the limitations of conventional diagnosis and information provision systems by introducing an information provision system for predicting artificial mutations or somatic mutations based on a neural network model, thereby predicting them with high accuracy, and can more easily set a treatment direction for each individual.

[0150]

[0151] Assessment 1: Prediction of artifactual or somatic mutations in FFPE samples

[0152] To generate a trainable dataset, cancer tissue Formalin-fixed paraffin-embedded (FFPE) samples were collected, and mutations that also appeared in Frozen Fresh (FFPE) samples were predicted from 9,685,786 mutation calls.

[0153] As a result, it was confirmed that 4,826,834 mutations were found in both FF and FFPE samples, and 4,858,952 mutations were found only in FFPE samples, indicating that only about 50% of them were somatic mutations rather than artificial mutations.

[0154] Afterwards, using the above-mentioned acquired samples and deep learning artificial neural network model, features highly correlated with somatic cell prediction were derived, and then the performance for somatic cell mutation prediction was evaluated.

[0155]

[0156] Evaluation 2: Evaluation of the performance of artificial neural network models based on sequence context data in predicting artificial mutations or somatic mutations.

[0157] Hereinafter, with reference to FIG. 5, the performance evaluation results for somatic mutation prediction of an artificial neural network model trained with a reference sequence, i.e., base sequence context data, including a mutation location sequence and its surrounding sequences, according to various embodiments of the present invention will be described.

[0158] FIG. 5 is a diagram showing the results of evaluating the prediction performance of an artificial neural network model trained to classify mutations by inputting base sequence context data together in various embodiments of the present invention.

[0159] Referring to FIG. 5, the evaluation results of an artificial neural network model trained to classify artificial mutations and somatic mutations are shown by inputting a total of 11 sequence data including a mutation location sequence and 5 reference sequences before and after the mutation location sequence.

[0160] Here, precision, recall, F1 score, accuracy, and specificity were calculated as follows.

[0161]

[0162]

[0163]

[0164]

[0165]

[0166] More specifically, in Fig. 5(a), when the base sequence context data including the reference sequence is reflected in the prediction, it shows excellent diagnostic performance with an accuracy, sensitivity, and precision of mutation classification close to 95%. In particular, in Fig. 5(b), when the data with mixed surrounding sequences of the mutation location sequence is reflected in the learning, the result showing lower diagnostic performance may suggest that the surrounding sequences of the mutation location sequence affect the accuracy of mutation classification. In contrast, in Fig. 5(c), FIREVAT, which is a conventional mutation classification platform, shows lower diagnostic performance than the artificial neural network-based model applied to various embodiments of the present invention.

[0167]

[0168] Evaluation 3: Evaluation of the performance of artificial neural network models to predict artificial mutations or somatic mutations according to the number of reference sequences.

[0169] First, in order to evaluate the prediction performance according to the number of reference sequences, in the sequence of TTGCT[G]ATGCA ([G] is the sequence where the mutation is located), the artificial neural network model that used only the mutation location sequence of [G] for learning, the artificial neural network model that used the sequence CT[G]AT with 2 bp each before and after the mutation location sequence for learning, the artificial neural network model that used the sequence GCT[G]ATG with 3 bp each before and after the mutation location sequence for learning, the artificial neural network model that used the sequence TGCT[G]ATGC with 4 bp each before and after the mutation location sequence for learning, and the artificial neural network model that used the sequence TTGCT[G]ATGCA with 5 bp each before and after the mutation location sequence for learning were compared in terms of artificial mutation or somatic mutation prediction performance.

[0170] First, in the case of the artificial neural network model (0 bp context) that uses only the mutation location sequence in Fig. 6A for learning, the accuracy in predicting artificial mutations or somatic mutations is 0.9079, the sensitivity is 0.9383, the specificity is 0.8777, the precision is 0.8840, and the F1-score is 0.9130.

[0171] In the case of an artificial neural network model (2bp context) that uses a total of 5 sequences including the mutation location sequence and the 2bp before / after in Fig. 6B for learning, the accuracy in predicting artificial mutations or somatic mutations is 0.9380, the sensitivity is 0.9281, the specificity is 0.9478, the precision is 0.9464, and the F1-score is 0.9451.

[0172] In the case of an artificial neural network model (3bp context) that uses a total of 7 sequences including the mutation location sequence and the 3bp before / after in Fig. 6C for learning, the accuracy in predicting artificial mutations or somatic mutations is 0.9462, the sensitivity is 0.9440, the specificity is 0.9485, the precision is 0.9479, and the F1-score is 0.9459.

[0173] In the case of an artificial neural network model (4bp context) that uses a total of 9 sequences including the mutation location sequence and the 4bp before / after in Fig. 6D for learning, the accuracy in predicting artificial mutations or somatic mutations is 0.9503, the sensitivity is 0.9438, the specificity is 0.9568, the precision is 0.9559, and the F1-score is 0.9498.

[0174] In particular, in the case of an artificial neural network model (5bp context) that uses a total of 11 sequences of the mutation location sequence and the preceding / following 5bp in Fig. 6E for learning, the accuracy in predicting artificial mutations or somatic mutations is 0.9538, the sensitivity is 0.9472, the specificity is 0.9604, the precision is 0.9596, and the F1-score is 0.9517.

[0175] That is, according to the above results, the prediction performance of an artificial neural network model based on base sequence context data that learns the surrounding sequences, i.e., the reference sequence, together with the mutant position sequence is superior to that of an artificial neural network model that learns only the mutant position sequence.

[0176] In relation to this, with further reference to Figures 7A to 7E, it is shown that as the number of sequences around the location where the mutation occurred increases, the accuracy, sensitivity, specificity, precision and F1-score of the prediction increase.

[0177] Accordingly, in a reference sequence including a sequence including N bases before and after the position of the mutation, N may be a natural number of 2 or more, preferably a natural number of 3 or more, more preferably a natural number of 4 or more, and even more preferably a natural number of 5 or more, but is not limited thereto.

[0178]

[0179] Although the embodiments of the present invention have been described in more detail with reference to the attached drawings, the present invention is not necessarily limited to these embodiments, and various modifications may be implemented without departing from the technical spirit of the present invention. Therefore, the embodiments disclosed in the present invention are not intended to limit the technical spirit of the present invention, but to explain it, and the scope of the technical spirit of the present invention is not limited by these embodiments. Therefore, it should be understood that the embodiments described above are exemplary in all aspects and not restrictive. The protection scope of the present invention should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present invention.

[0180] [Explanation of symbols]

[0181] 100: Server for providing data

[0182] 200: Somatic mutation prediction device

[0183] 210: Communication Interface

[0184] 211: Wired communication port 212: Wireless circuit

[0185] 220: Memory

[0186] 221: Operating System 222: Communication Module

[0187] 223: User Interface Module 224: Application

[0188] 230: I / O interface 240: Processor

[0189] 300: Device for providing information on somatic mutation prediction

Claims

1. A method for providing information on the prediction of artificial mutations or somatic mutations implemented by a processor, A step of receiving data obtained from medical data or a biological sample isolated from an individual; A step of producing data necessary for determining an artificial mutation or somatic mutation based on the received data, wherein the produced data includes a mutation position sequence of the artificial mutation or the somatic mutation and a reference sequence including N bases before and after the mutation position, wherein the mutation position sequence is a standard sequence or a mutation sequence, and wherein N includes a natural number of 2 or more, the producing step, and A method for providing information on prediction of artificial mutations or somatic mutations, comprising a step of predicting artificial mutations or somatic mutations for a sample obtained from an individual by using an artificial neural network model configured to predict artificial mutations or somatic mutations by inputting data necessary for determining somatic mutations.

2. In paragraph 1, The above medical data is NGS (Next Generation Sequencing) data for somatic cell mutations. A method for providing information on the prediction of artificial mutations or somatic mutations.

3. In paragraph 2, The above NGS data is any one of a VCF (Variant Call Format) file, a BAM (Binary Alignment Map) file, a WGS (whole genome sequencing) file in which the base sequence is analyzed for the entire genome, a WES (whole exome sequencing) file, an RNA sequencing file, and a targeted sequencing file in which the base sequence is analyzed for a specific region. A method for providing information on the prediction of artificial mutations or somatic mutations.

4. In paragraph 1, The step of predicting the above artificial mutation or somatic mutation is, Based on the above-mentioned generated data and the above-mentioned medical data, it further includes a step of predicting an artificial mutation or somatic mutation for the sample using the artificial neural network model. A method for providing information on the prediction of artificial mutations or somatic mutations.

5. In paragraph 1, In the step of producing data necessary for mutation or artificial mutation determination based on the received data, the produced data is The number of reads in which the reference sequence is observed / sequence depth at that position, the number of reads in which the alternate sequence is observed / sequence depth at that position, the ratio of the difference between the number of sequences in which Read 1 is mapped to the forward strand and the number of sequences in which Read 2 is mapped to the reverse strand in the reads in which the reference sequence is observed in the reads in which the reference sequence is observed, the ratio of the difference between the number of sequences in which Read 1 is mapped to the forward strand and the number of sequences in which Read 2 is mapped to the reverse strand in the reads in which the alternate sequence is observed in the reads in which the alternate sequence is observed, the median of the inserted length in the reads in which the alternate sequence is observed divided by the median of the inserted length in the reads in which the reference sequence is observed, the length of the double-sequenced bases in the reference fragment, the number of times the sequence is read twice by Read 1 and Read 2 in the reference fragment, and the number of times the sequence is read twice by Read 1 and Read 2 in the alternate fragment. The method further comprises at least one selected from the group consisting of the length of the double-sequenced bases of the sequence read twice by Read 2, the number of times the sequence is read twice by Read 1 and Read 2 in the replacement fragment, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference sequence is observed in the reads in which the reference sequence is observed, and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the replacement sequence is observed in the reads in which the replacement sequence is observed. A method for providing information on the prediction of artificial mutations or somatic mutations.

6. In paragraph 1, The step of producing data necessary for determining mutation or artificial mutation based on the above received data is: A step of vectorizing the received data is included. A method for providing information on the prediction of artificial mutations or somatic mutations.

7. In paragraph 6, The step of producing data necessary for determining mutation or artificial mutation based on the above received data is: Further comprising a step of calculating similarity based on the vectorized data. A method for providing information on the prediction of artificial mutations or somatic mutations.

8. In paragraph 5, In the step of producing data necessary for mutation or artificial mutation determination based on the received data, the produced data is The cosine similarity between vectors defined by (the number of sequences mapped to the forward strand of read 1 in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand of read 2) or the cosine similarity between vectors defined by (the number of sequences mapped to the forward strand of reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and vectors defined by (the number of sequences mapped to the forward strand of reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand), A method for providing information on the prediction of artificial mutations or somatic mutations.

9. In paragraph 1, In the step of determining a prediction for an artificial mutation or somatic mutation for a sample obtained from an individual based on the above-mentioned data, The sample obtained from the above object is characterized in that it is an FFPE sample. A method for providing information on the prediction of artificial mutations or somatic mutations.

10. In paragraph 1, In the step of determining a prediction for an artificial mutation or somatic mutation for a sample obtained from an individual based on the above-mentioned data, The above object is characterized in that it is an object that has developed cancer. A method for providing information on the prediction of artificial mutations or somatic mutations.

11. In paragraph 10, The above cancer is any one of liver cancer, colon cancer, breast cancer, lung cancer, and fibroma. A method for providing information on the prediction of artificial mutations or somatic mutations.

12. A communication unit that receives data on somatic cell mutations obtained from medical data or biological samples from a data provision server, and Including a processor electrically connected to the above communication unit, The above processor, Based on the received data, data required for determining an artificial mutation or somatic mutation is generated, wherein the generated data includes a mutation position sequence of the artificial mutation or the somatic mutation and a reference sequence including N bases before and after the mutation position, wherein the mutation position sequence is a standard sequence or a mutation sequence, and wherein N includes a natural number of 2 or more, An artificial neural network model configured to predict artificial mutations or somatic mutations by inputting data required for somatic mutation determination, and configured to predict artificial mutations or somatic mutations for a sample obtained from an individual by inputting the above-described data. A device for providing information on the prediction of artificial mutations or somatic mutations.

13. In paragraph 12, The above medical data is NGS (Next Generation Sequencing) data for somatic cell mutations. A device for providing information on the prediction of artificial mutations or somatic mutations.

14. In paragraph 12, The above NGS data is any one of a VCF (Variant Call Format) file, a BAM (Binary Alignment Map) file, a WGS (whole genome sequencing) file in which the base sequence is analyzed for the entire genome, a WES (whole exome sequencing) file, an RNA sequencing file, and a targeted sequencing file in which the base sequence is analyzed for a specific region. A device for providing information on the prediction of artificial mutations or somatic mutations.

15. In paragraph 12, The above processor, Based on the above-mentioned generated data and the above-mentioned medical data, the artificial neural network model is further configured to predict an artificial mutation or somatic mutation for the sample. A device for providing information on the prediction of artificial mutations or somatic mutations.

16. In paragraph 12, Based on the data received above, the data required to determine mutation or artificial mutation is: The number of reads in which the reference sequence is observed / sequence depth at that position, the number of reads in which the alternate sequence is observed / sequence depth at that position, the ratio of the difference between the number of sequences mapped to the forward strand by Read 1 and the number of sequences mapped to the reverse strand by Read 2 in the reads in which the reference sequence is observed in the reads in which the reference sequence is observed, the ratio of the difference between the number of sequences mapped to the forward strand by Read 1 and the number of sequences mapped to the reverse strand by Read 2 in the reads in which the alternate sequence is observed in the reads in which the alternate sequence is observed, the median of the inserted length in the reads in which the alternate sequence is observed divided by the median of the inserted length in the reads in which the reference sequence is observed, the length of the double-sequenced bases in the reference fragment, the number of times the sequence is read twice by Read 1 and Read 2 in the reference fragment, and in the alternate fragment. The method further comprises at least one selected from the group consisting of the length of the double-sequenced bases of the sequence read twice by Read 1 and Read 2, the number of times the sequence is read twice by Read 1 and Read 2 in the replacement fragment, the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the reference sequence is observed in the reads in which the reference sequence is observed, and the ratio of the difference between the number of sequences mapped to the forward strand and the number of sequences mapped to the reverse strand among the reads in which the replacement sequence is observed in the reads in which the replacement sequence is observed. A device for providing information on the prediction of artificial mutations or somatic mutations.

17. In paragraph 12, The above processor, Further configured to vectorize the received data, A device for providing information on the prediction of artificial mutations or somatic mutations.

18. In paragraph 17, The above processor, Further configured to calculate similarity based on the above vectorized data, A device for providing information on the prediction of artificial mutations or somatic mutations.

19. In paragraph 16, Based on the data received above, the data required to determine mutation or artificial mutation is: The cosine similarity between vectors defined by (the number of sequences mapped to the forward strand of read 1 in the reads where the reference sequence is observed, the number of sequences mapped to the reverse strand of read 2) or the cosine similarity between vectors defined by (the number of sequences mapped to the forward strand of reads where the reference sequence is observed, the number of sequences mapped to the reverse strand) and vectors defined by (the number of sequences mapped to the forward strand of reads where the alternative sequence is observed, the number of sequences mapped to the reverse strand), A device for providing information on the prediction of artificial mutations or somatic mutations.

20. In paragraph 12, The sample obtained from the above object is characterized in that it is an FFPE sample. A device for providing information on the prediction of artificial mutations or somatic mutations.

21. In paragraph 12, The above object is characterized in that it is an object that has developed cancer. A device for providing information on the prediction of artificial mutations or somatic mutations.

22. In paragraph 21, The above cancer is any one of liver cancer, colon cancer, breast cancer, lung cancer, and fibroma. A device for providing information on the prediction of artificial mutations or somatic mutations.

Citation Information

Patent Citations

  • System for assisting parking vehicle and method for the same

    KR1020210070878A

  • Manufacturing method of functional liquid manure

    KR102376447B1

  • Method of determining kinship using gene sequence variation

    KR102391084B1

  • A system for discriminating zygosity of variant

    KR102470337B1

  • Conductive pattern and display device including the same

    KR102881000B1