Somatic Mutation Detection Device and Method for Reducing Sequencing Platform-Specific Errors
By applying neural networks to detect mutations, the problem of specific false positives when traditional software detects mutations on the second-generation sequencing platform is solved, and the accuracy of mutation detection is improved.
Patent Information
- Application Number
- CN201980101042.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2039-10-25
AI Technical Summary
Traditional software is prone to specific false positives when detecting mutations on second-generation sequencing platforms, resulting in a decrease in the accuracy of mutation detection.
The neural network is used to detect mutations, and the specific false positives occurring on the sequencing platform are corrected by learning data, thereby improving the accuracy of mutation detection.
Effectively prevent the reduction of mutation detection accuracy due to specific false positives of the sequencing platform and improve the accuracy of detection.
Smart Images

Figure CN114467144B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for detecting mutations and an apparatus for implementing the method. More specifically, the present disclosure relates to an apparatus and a method for detecting mutations using a neural network, which can reduce sequencing platform-specific errors through learning. Background Art
[0002] Next-generation sequencing (NGS) refers to a method of dividing DNA into multiple fragments and performing parallel sequencing. Different from traditional Sanger sequencing, next-generation sequencing can analyze multiple DNA fragments simultaneously, and thus is more advantageous in terms of analysis time, analysis cost, and analysis accuracy.
[0003] Figure 1 A graph 100 comparing next-generation sequencing 110 and Sanger sequencing 120 is shown. As shown in graph 100, the performance of next-generation sequencing 110 is superior to that of Sanger sequencing 120. In addition, as shown in the horizontal axis of graph 100, next-generation sequencing 110 can have multiple read lengths.
[0004] Next-generation sequencing can be used for DNA sequencing of cancer patients to detect mutations. A next-generation sequencing method can be adopted, and mutations in cancer cells can be detected through various software for DNA sequencing.
[0005] When detecting mutations using traditional software, especially when performing DNA sequencing through a specific sequencing platform such as short read sequencing, due to the characteristics of the sequencing platform, even if there is actually no mutation, it will be misdetected as a mutation, resulting in false positives. Such specific false positives occurring on the sequencing platform will reduce the accuracy of mutation detection.
[0006] Therefore, in order to prevent specific false positives from occurring on the sequencing platform and reducing the accuracy of mutation detection, it is necessary to improve the mutation detection method. Summary of the Invention
[0007] Problems to be Solved
[0008] The object of the present disclosure is to eliminate the problems that occur in traditional software, solve the problem of reduced mutation detection accuracy due to specific false positives occurring on the sequencing platform, and improve the performance of mutation detection.
[0009] Solutions to the Problems
[0010] As a technical means for solving the above technical problems, the present disclosure provides a mutation detection device on the one hand, which is characterized by including: a memory for storing neural network implementation software; and a processor for running the software to detect mutations, the processor being configured to generate first genomic data extracted from a detection target cell and second genomic data extracted from a normal cell, preprocess the first genomic data and the second genomic data so as to extract image data, and detect mutations of the detection target cell based on the image data through the neural network, and the neural network has been learned to correct specific false positives occurring on a sequencing platform.
[0011] On the other hand, the present disclosure provides a method for running neural network implementation software to detect mutations, which includes the following steps: generating first genomic data extracted from a detection target cell and second genomic data extracted from a normal cell; preprocessing the first genomic data and the second genomic data so as to extract image data; and detecting mutations of the detection target cell based on the image data through the neural network, and the neural network has been learned to correct specific false positives occurring on a sequencing platform.
[0012] Advantages of the Invention
[0013] In the process of detecting mutations, the device and method of the present disclosure can apply a neural network, and the neural network can be pre-learned to correct specific false positives occurring on the sequencing platform, thereby preventing the reduction of the accuracy of mutation detection due to the occurrence of specific false positives on the sequencing platform. In particular, different from traditional statistical methods, it can apply a neural network to detect mutations, and detect mutations based on stronger performance compared to traditional methods. Brief Description of the Drawings
[0014] Figure 1 It is a curve graph comparing the second-generation sequencing method and the traditional sequencing method;
[0015] Figure 2 It shows a neural network of some embodiments;
[0016] Figure 3 It shows a mutation detection process of some embodiments;
[0017] Figure 4 It is a block diagram of components of a mutation detection device in some embodiments;
[0018] Figure 5 It shows the structure and learning method of a neural network in some embodiments;
[0019] Figure 6 shows the generation process of neural network learning data in some embodiments;
[0020] Figure 7 is a flowchart of the steps constituting the mutation detection method in some embodiments. Detailed implementation manners
[0021] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. The following description is only for specifically describing the embodiments and is not intended to limit or define the claims of the present disclosure. It should be construed that the content easily inferred by those of ordinary skill in the technical field to which the present disclosure pertains from the description and embodiments of the present invention should fall within the scope of the claims of the present disclosure.
[0022] The terms used in the present disclosure are common terms widely used in the technical field to which the present disclosure pertains. However, the meanings of the terms in the present disclosure may change according to the intentions of those skilled in the art, the emergence of new technologies, review standards, or precedents. Some terms can be arbitrarily selected by the applicant. In this case, the meanings of the arbitrarily selected terms will be described in detail. It should be construed that the terms used in the present disclosure do not only have the meanings explained in the dictionary, and their meanings reflect the overall idea of the specification.
[0023] It should not be construed that the terms "constitute", "include", etc. used in the present disclosure must include all the components or steps described in the specification, and it should be construed that when some components or steps are not included and when additional components or steps are further included, it also stems from these terms.
[0024] The ordinal terms including "first" or "second" etc. used in the present disclosure can be used to describe various components or steps. However, these components or steps should not be limited by the ordinal numbers. It should be construed that the ordinal terms are only used to distinguish one component or step from another component or step.
[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. The content well-known to those of ordinary skill in the technical field to which the present disclosure pertains is omitted here and will not be described further.
[0026] Figure 2 shows a neural network in some embodiments.
[0027] Figure 2 shows the components constituting the neural network 200. The neural network 200 is an artificially implemented neural network. In addition to the input layer and the output layer, it also has hidden layers and can effectively implement various non-linear functions. The neural network 200 includes multiple hidden layers and can be equivalent to a deep neural network. In addition to Figure 2In addition to the structures illustrated by way of example, neural network 200 can also be implemented as various architectures such as a recurrent neural network (RNN) or a convolutional neural network (CNN).
[0028] Through learning, neural network 200 can become a pattern for adjusting numerical values, which are various parameters constituting neural network 200. When neural network 200 is properly learned according to various machine learning and deep learning methods, it can perform functions based on the learning purpose with high performance. Thus, in addition to fields such as speech recognition, natural language, and image analysis, neural network 200 can also be widely applied to various fields. In particular, as in the present disclosure, in order to solve the technical problems existing in the prior art, neural network 200 can be applied to biological fields such as mutation detection.
[0029] Figure 3 The mutation detection process of some embodiments is shown.
[0030] As Figure 3 shown, a series of processing procedures can be implemented on the first genomic data 310 and the second genomic data 320 within the mutation detection device 300, and a mutation detection result 350 can be generated. As described below, the mutation detection device 300 operates in the same manner as Figure 4 the device 400 shown.
[0031] A series of processing procedures inside the device 300 can be implemented in the form of software or a program. Each step of a series of processing procedures inside the device 300 can be implemented by a module that implements a specific function, such as an image generation module 330 or a mutation detection module 340. For example, the software that implements a series of processing procedures can be presented in the form of a Python script and run in an environment such as LINUX CentOS release 7.6.
[0032] The first genomic data 310 can mean genomic data extracted from a cell to be detected. The cell to be detected is a cell that is the object of mutation detection and can mean a cancer cell. The second genomic data 320 can mean genomic data extracted from a normal cell.
[0033] In order to accurately grasp which genes in the genes of the cell to be detected have mutated, the second genomic data 320 can be considered on the basis of considering the first genomic data 310. At the same time, although Figure 3 not shown, the process of extracting the first genomic data 310 from the cell to be detected and the process of extracting the second genomic data 320 from the normal cell can also be implemented by another module that constitutes the software inside the device 300.
[0034] The device 300 does not simply detect mutations statistically based on the genomic data of cancer patients. Instead, it can detect mutations by extracting the first genomic data 310 and the second genomic data 320 from the actual cancer cells and the corresponding normal cells. Therefore, the different characteristics of each cancer patient and each cancer cell can be reflected one by one in the mutation detection process. As a result, it is possible to accurately detect which genes in the genes of cancer cells have mutated.
[0035] The image generation module 330 can extract image data from the first genomic data 310 and the second genomic data 320. The image data can mean the visualization data of the first genomic data 310 and the second genomic data 320 for the neural network 200 that has been learned / trained to detect mutations.
[0036] The mutation detection module 340 can detect mutations in the target cells to be detected based on the image data. To this end, the neural network 200 can be implemented in the mutation detection module 340. The neural network 200, after being learned, can detect which genes in the genes of the target cells to be detected have mutated. For example, as Figure 5 and Figure 6 shown, the neural network 200 can extract features from the image and be implemented as a convolutional neural network (CNN). The convolutional neural network (CNN), after being learned, performs specific functions based on the features.
[0037] The mutation detection module 340 can further process and handle the output of the neural network 200 to generate a mutation detection result 350. The mutation detection result 350 can be generated in a standard format (VCF). This standard format (VCF) is compared with the reference gene to display information related to the genes considered to have mutated.
[0038] The device 300 can apply the neural network 200 learned for a specific purpose to the detection of mutations. Therefore, the accuracy of mutation detection can be further improved. As described below, the neural network 200, after being learned, can correct the specific false positives that occur on the sequencing platform. Therefore, it is possible to prevent the drawbacks that occur in traditional mutation detection software, that is, the problem that the accuracy is reduced due to false positives.
[0039] In addition, through the device 300, the mutation detected from the target cell to be detected may be a somatic single nucleotide variant (sSNV). As a somatic mutation, the somatic single nucleotide variant may mean that among several bases constituting the base sequence, only a single base has mutated. The somatic single nucleotide variant may be suitable for detection by next-generation sequencing and may also be suitable for detection by the neural network 200, where the neural network 200 has undergone special learning to correct specific false positives occurring on the sequencing platform. However, the device 300 is not limited thereto, and in addition to somatic single nucleotide variants, other types of mutations can also be detected.
[0040] Figure 4 It is a block diagram of the components of the mutation detection device in some embodiments.
[0041] As Figure 4 shown, the mutation detection device 400 may include a memory 410 and a processor 420. However, it is not limited thereto. In addition to Figure 4 the components shown, the device 400 may further include other general components. In addition, Figure 4 the device 400 shown may be an example of implementing Figure 3 the device 300 shown.
[0042] The device 400 may correspond to various devices for detecting mutations. For example, the device 400 may be various computing devices, such as a PC, a server device, a smartphone, a tablet computer, and other mobile devices.
[0043] The memory 410 may store software for implementing the neural network 200. For example, data related to several layers and several nodes constituting the neural network 200, operations performed at several nodes, and several parameters applicable during the operation process may be stored in the memory 410 in the form of at least one instruction, program, or software.
[0044] The memory 410 can be implemented as a non-volatile memory, such as ROM (read only memory), PROM (programmable ROM), EPROM (electrically programmable ROM), EEPROM (electrically erasable and programmable ROM), flash memory, PRAM (phase-change RAM), MRAM (magnetic RAM), RRAM (resistive RAM), FRAM (ferroelectric RAM), etc., or can be implemented as a volatile memory, such as DRAM (dynamic RAM), SRAM (static RAM), SDRAM (synchronous DRAM), PRAM (phase-change RAM), RRAM (resistive RAM), FeRAM (ferroelectric RAM), etc. Also, the memory 410 can be implemented as an HDD (hard disk drive), SSD (solid state drive), SD (secure digital), Micro-SD (micro secure digital), etc.
[0045] The processor 420 can run software stored in the memory 410 to detect mutations. The processor 420 executes a series of processes for detecting mutations and detects mutations of the cells to be detected. The processor 420 can execute to control the overall functions of the device 400 and process various operations inside the device 400.
[0046] The processor 420 can be implemented as an array of multiple logic gates or a general-purpose microprocessor. The processor 420 can be configured as a single processor or multiple processors. The processor 420 may not be independent of the memory 410 storing the software and be integrated with the memory 410. The processor 420 can be at least one of a CPU (central processing unit), GPU (graphics processing unit), and AP (application processor) provided in the device 400. However, this is only illustrative, and the processor 420 can be implemented in many other forms.
[0047] The processor 420 may generate first genomic data extracted from the cells of the detection target and second genomic data extracted from normal cells. The processor 420 may embed the result data set of sequencing the cells of the detection target into the genomic data, extract the first genomic data, embed the result data set of sequencing normal cells into the genomic data, and extract the second genomic data.
[0048] For example, the processor 420 may generate the first genomic data and the second genomic data through, such as, the HCC1143 cell line. Additionally, the first genomic data and the second genomic data may be whole genome data.
[0049] The processor 420 may preprocess the first genomic data and the second genomic data to extract image data. The processor 420 may perform preprocessing to make the first genomic data and the second genomic data have a form suitable for processing by the neural network 200.
[0050] For example, the first genomic data and the second genomic data may be converted into an image form, such as, image data. However, the conversion into an image form is merely illustrative, and depending on how the neural network 200 is implemented, in addition to images, the first genomic data and the second genomic data may be converted into various forms.
[0051] The processor 420 may preprocess the first genomic data and the second genomic data by correcting them based on mapping quality and depth. The processor 420 can, based on the mapping quality, remove low-quality reads and adjust the depth of the first genomic data and the second genomic data. Through the above preprocessing process, the processor 420 may generate image data that has a form suitable for processing by the neural network 200.
[0052] The processor 420 may, through the neural network 200, detect mutations in the cells of the detection target based on the image data. The neural network 200 has been trained to correct specific false positives that occur on the sequencing platform. The processor 420 applies the trained neural network 200 to detect which genes in the cells of the detection target have mutated from the image data.
[0053] A sequencing platform can refer to the specific method of detecting the sequencing of target cells. Depending on the applicable sequencing platform, the sequencing method will also be different. For example, in next-generation sequencing (NGS), the type of sequencing platform is determined based on the size of the DNA fragments being decomposed, that is, based on the read length of the DNA fragments being processed in parallel. For example, sequencing platforms can include long-read sequencing and short-read sequencing, etc. However, not limited to this read-length-based classification, a sequencing platform can refer to various analysis methods for the purpose of sequencing.
[0054] The neural network 200 can be pre-trained, receive image data, and output the mutations of the target cells to be detected. The pre-trained neural network 200 can be stored in the memory 410 in the form of software, and the processor 420 runs the software to detect the mutations of the target cells from the image data. This software is used to implement the learned neural network 200.
[0055] The learning or training of the neural network 200 can be performed by the device 400. To enable the neural network 200 to learn, the device 400 and the processor 420 can complete the learning of the neural network 200 by repeatedly updating the parameter values that make up the neural network 200. Alternatively, the neural network 200 can be learned externally to the device 400 and then implemented as software.
[0056] After learning, the neural network 200 can correct the specific false positives that occur on the sequencing platform. For example, after learning, the neural network 200 can correct the specific false positives that occur in short-read sequencing, and the read length of short-read sequencing can be 100 or less. However, not limited by this specific value, short-read sequencing can be a sequencing method with a shorter read length than long-read sequencing.
[0057] The specific false positives that occur on the sequencing platform mean that either a specific gene mutation is detected by a specific sequencing platform, but in fact, the gene has not mutated. That is, this means that a false positive can be judged by a specific sequencing platform as having mutated, but is judged by other sequencing platforms as not having mutated.
[0058] For example, the specific false positives that occur on a specific sequencing platform can be the specific false positives of short-read sequencing. The specific false positives that occur in short-read sequencing may be detected as normal by long-read sequencing, but are misdetected as having mutated by short-read sequencing. When specific false positives occur in short-read sequencing, it may misjudge a gene that has not actually mutated as having mutated, which will reduce the accuracy of mutation detection.
[0059] After learning, the neural network 200 can correct the specific false positives that occur on the sequencing platform. Therefore, when applying the neural network 200 to detect the mutations of the target cells, the accuracy of mutation detection can be improved. Next, throughFigure 5 and Figure 6 shows the specific content learned by the neural network 200.
[0060] Figure 5 shows the structure and learning method of the neural network in some embodiments.
[0061] Figure 5 shows the structure of the neural network 530 and the process of the neural network 530 learning based on the first learning image data 510 and the second learning image data 520. Figure 5 The shown neural network 530 can be Figures 2 to 4 an example implemented by the shown neural network 200.
[0062] As described above, the neural network 530 can be a convolutional neural network, which extracts features from the image data and calculates the probability of a gene mutation in the detected target cell based on the features.
[0063] The neural network 530 can be implemented as a convolutional neural network (CNN) including a first network 531 and a second network 532. The first network 531 can include a convolutional layer and a pooling layer, and the second network 532 can include a fully connected network. When the neural network 530 completes learning, the first network 531 can extract features representing the features of the input data from the input data, and the second network 532 can perform functions based on the features according to the purpose of the neural network 530.
[0064] As described above, the learning of the neural network 530 can be executed by the device 400. Alternatively, after the learning of the neural network 530 is completed outside the device 400, the device 400 can only execute the inference of the neural network 530.
[0065] The neural network 530 can use the first learning image data 510 and the second learning image data 520 as learning data for learning. Specifically, after learning, the neural network 530 can distinguish actual mutations and false positive mutations based on the first learning image data 510 and the second learning image data 520, where the first learning image data 510 represents the learning data related to the actual mutation, and the second learning image data 520 represents the learning data related to the false positive misdetected as a mutation.
[0066] The first learning image data 510 can display the learning data related to the actual mutation. A certain sequencing platform can judge the actual mutation as a mutation, which also means that other sequencing platforms can also judge it as a mutation. For example, the actual mutation means that both short-read sequencing and long-read sequencing judge it as a mutation.
[0067] The second learning image data 520 may represent relevant learning data misdetected as mutations due to false positives. As described above, a specific sequencing platform may misdetect as mutations those that are actually not mutations. Therefore, through learning, the neural network 530 can correct false positives by applying mutations misdetected due to false positives. For example, for a misdetected mutation, long-read sequencing may determine that no mutation has occurred, while short-read sequencing may determine that a mutation has occurred.
[0068] In order for the neural network 530 to learn, the first learning image data 510 and the second learning image data 520 can be applied together as learning data. Therefore, the learned neural network 530 can correct specific false positives occurring on the sequencing platform. Setting both the first learning image data 510 and the second learning image data 520 as learning data can improve the accuracy of the neural network 530 in detecting mutations.
[0069] Figure 6 Illustrated is the generation process of neural network learning data in some embodiments.
[0070] Figure 6 Illustrated by way of example are different sequencing platforms for generating the first learning image data 510 and the second learning image data 520, and long-read sequencing 610 and short-read sequencing 620 are shown.
[0071] The first learning image data 510 and the second learning image data 520 can be generated based on the results of long-read sequencing 610 and short-read sequencing 620 of the same learning cells. In order to obtain learning data for the neural network 530 to learn, long-read sequencing 610 and short-read sequencing 620 can be performed on the same cancer cells locally including mutant genes, and the results of the two can be compared.
[0072] For example, Pacbio sequencing can be performed through long-read sequencing 610, and Illumina sequencing can be performed through short-read sequencing 620. However, it is not limited thereto. For short reads and long reads, other sequencing methods with appropriate read lengths can be adopted.
[0073] Figure 6 Illustrated by way of example are the implementation results of long-read sequencing 610 and short-read sequencing 620. When the same criteria are applied, there will be local differences between the mapping results of long-read sequencing 610 and the mapping results of short-read sequencing 620. For example, according to the comparison result 630, it is considered that both long-read sequencing 610 and short-read sequencing 620 have mutations. Therefore, the bases corresponding to the comparison result 630 can be set as the actual mutations.
[0074] However, in the comparison result 640, long-read sequencing 610 detected a mutation that had occurred, while short-read sequencing 620 detected no mutation. Therefore, the base corresponding to the comparison result 640 can be set as follows: on short-read sequencing, it was misdetected as a mutation due to specific false positives.
[0075] The actual mutation-related data corresponding to the comparison result 630 can be labeled as the first learning image data 510, and the misdetected mutation-related data corresponding to the comparison result 640 can be labeled as the second learning image data 520. The neural network 530 can learn based on the first learning image data 510 and the second learning image data 520 generated in the above manner. Therefore, after learning, false positives similar to the situation of the comparison result 640 can be corrected.
[0076] In addition, the actual mutation-related data corresponding to the comparison result 630 and the misdetected mutation-related data corresponding to the comparison result 640 can be implemented as virtual cancer cell genome data through HCC1143 cell line, etc., and the first learning image data 510 and the second learning image data 520 can be generated respectively for the actual mutation and the misdetected mutation through the process of obtaining information such as gene sequence, insertion / deletion (indel), and mapping quality from the virtual cancer cell genome data. That is, the first learning image data 510 and the second learning image data 520 can include at least one of gene sequence, insertion / deletion, and mapping quality.
[0077] Figure 7 It is a flowchart of the steps constituting the mutation detection method in some embodiments.
[0078] As Figure 7 shown, the mutation detection method may include step 710 to step 730. However, it is not limited thereto. In addition to Figure 7 the several steps shown, the mutation detection method may further include other several general steps.
[0079] Figure 7 The mutation detection method shown can be composed of the steps processed in sequence by Figures 3 to 6 the device 300 or the device 400 shown. Therefore, even if Figure 7 the content omitted in the mutation detection method shown below, Figures 3 to 6 the above content of the device 300 or the device 400 shown is equally applicable to Figure 7 the mutation detection method shown.
[0080] In step 710, the apparatus 400 may generate first genomic data extracted from the target cells to be detected and second genomic data extracted from normal cells.
[0081] The apparatus 400 may preprocess the first genomic data and the second genomic data based on mapping quality and depth.
[0082] In step 720, the apparatus 400 may preprocess the first genomic data and the second genomic data to extract image data.
[0083] In step 730, the apparatus 400 may detect mutations in the target cells to be detected based on the image data through a neural network, and the neural network has been trained to correct specific false positives that occur on the sequencing platform.
[0084] The neural network has been trained to distinguish actual mutations and misdetected mutations based on first training image data and second training image data, where the first training image data represents training data related to actual mutations, and the second training image data represents training data related to mutations misdetected as false positives.
[0085] The first training image data and the second training image data may be generated based on the results of long read sequencing and short read sequencing of the same training cells.
[0086] The first training image data and the second training image data may include at least one of gene sequence, indel (insertion / deletion), and mapping quality.
[0087] The neural network may be a convolutional neural network (CNN) that extracts features from the image data and calculates the probability of gene mutations in the target cells to be detected based on the features.
[0088] The mutations detected from the target cells to be detected may be somatic single nucleotide variants (sSNVs).
[0089] Figure 7The disclosed mutation detection method can be recorded in a computer-readable recording medium that records at least one program or software including instructions for executing the method.
[0090] The computer-readable recording medium may include, for example, a specially configured hardware device for storing and executing program instructions of magnetic media such as hard disks, floppy disks, magnetic tapes, optical media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and read-only memories (ROMs), random access memories (RAMs), flash memories, etc. The program instructions may include, for example, machine language codes made by compilers and high-level language codes that can be executed by a computer using an interpreter or the like.
[0091] The above details the embodiments of the present disclosure, but the claims of the present disclosure are not limited thereto, and it should be construed that various variations and improvements made by those skilled in the art to which the present disclosure pertains using the basic concepts of the present disclosure described in the appended claims also fall within the scope of the claims of the present disclosure.
Claims
1. An apparatus, which is a mutation detection apparatus, characterized in that, Comprising: A memory for storing neural network implementation software; And A processor for running the software, detecting mutations, The processor is used to generate first genomic data extracted from a target cell to be detected and second genomic data extracted from a normal cell, Preprocess the first genomic data and the second genomic data to extract image data, Through the neural network, based on the image data, detect mutations in the target cell to be detected. This neural network has been trained to correct specific false positives that occur on the sequencing platform; Wherein, the neural network has been trained to distinguish normal mutations and misdetected mutations based on first training image data and second training image data. The first training image data represents training data related to normal mutations detected normally, and the second training image data represents training data related to mutations misdetected as false positives; and Generate the first training image data and the second training image data based on the results of long read sequencing and short read sequencing respectively performed on the same training cells.
2. The apparatus according to claim 1, characterized in that: The first training image data and the second training image data include at least one of gene sequence, indel (insertion / deletion), and mapping quality.
3. The apparatus according to claim 1, characterized in that: The neural network is a convolutional neural network (CNN), which extracts features from the image data and calculates the probability of gene mutations in the target cell to be detected based on the features.
4. The apparatus according to claim 1, characterized in that: The processor corrects the first genomic data and the second genomic data based on mapping quality and depth, and performs the preprocessing.
5. The apparatus according to claim 1, characterized in that: The mutation detected from the target cell to be detected is a somatic single nucleotide variant (sSNV).
6. A method, which is a method of running a neural network implementation software to detect mutations, characterized in that Including the following steps: Generate first genomic data extracted from a target cell to be detected and second genomic data extracted from a normal cell; Preprocess the first genomic data and the second genomic data to extract image data; Through the neural network, based on the image data, detect mutations in the target cell to be detected. This neural network has been trained to correct specific false positives that occur on the sequencing platform; Among them, the neural network is trained to distinguish the normal mutations and the false positive mutations based on the first training image data and the second training image data, where the first training image data represents the training data related to the normal (actual) mutations detected normally, and the second training image data represents the training data related to the mutations misdetected as false positives; and The first training image data and the second training image data are generated based on the results of long read sequencing and short read sequencing respectively performed on the same training cells.
7. The method according to claim 6, characterized in that: The first training image data and the second training image data include at least one of gene sequence, indel (insertion / deletion), and mapping quality.
8. The method according to claim 6, characterized in that: The neural network is a convolutional neural network (CNN) that extracts features from the image data and calculates the probability of gene mutation in the target cell to be detected based on the features.
9. The method according to claim 6, characterized in that: The step of extracting the image data includes the following steps: correcting the first genomic data and the second genomic data based on the mapping quality and depth, and performing the preprocessing.
10. The method according to claim 6, characterized in that: The mutations detected from the target cell to be detected are somatic single nucleotide variants (sSNVs).
Citation Information
Patent Citations
Deep learning analysis pipeline for next generation sequencing
US10354747B1