Gene Mutation Detection Method, Device, Electronic Device and Storage Medium

A hybrid gene variant detection method using long-read and short-read sequencing with neural networks enhances accuracy by leveraging their respective strengths, effectively addressing inaccuracies in complex genomic regions.

CN119832980BActive Publication Date: 2025-07-15BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309305.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-15
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing gene variant detection technology is insufficient in the accuracy of complex genome regions, and it is difficult to accurately compare short-read long-sequencing platforms, and long-read long-sequencing platforms have challenges in accuracy and error rates.

Method used

Combining long read sequencing technology and short read sequencing technology, the data are compared and detected through deep learning methods. The advantages of long read sequencing in complex genome areas and the accuracy of short read sequencing in non-complex regions are used to comprehensively determine the mutation site.

Benefits of technology

Improve the accuracy of genetic variant detection, especially in complex genome areas, and improve the overall accuracy of variant detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832980B_ABST
    Figure CN119832980B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, electronic device, and storage medium for gene mutation detection, belonging to the technical field of gene mutation detection. The method fuses the long-read mutation detection results and short-read mutation detection results of the gene fragment to be detected; for gene loci with inconsistent sequencing results, it first detects whether the gene locus is in a complex genomic region. If it is in a complex genomic region, the mutation detection result of the gene locus is determined based on the long-read mutation detection results; if it is not in a complex genomic region, the mutation detection result of the gene locus is determined based on the short-read mutation detection results. This method can improve the accuracy of genomic mutation detection for the gene fragment to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of gene sequencing technology, and particularly to a gene variant detection method, device, electronic device and storage medium. Background Art

[0002] Gene variant detection is an important area in genomics research, which involves identifying and analyzing variants in an individual's genome, and these variants may be the causes of diseases, affecting gene expression or individual phenotypic differences. Gene variant detection technology has a wide range of applications in many fields, including but not limited to genetic disease diagnosis, tumor gene detection, and drug development.

[0003] Genomic variant detection generally identifies the differences between an individual's genome and a linear reference genome sequence by comparing the sequencing data of the individual's genome with the reference genome sequence, so as to characterize the variant information of the individual's genome.

[0004] In the prior art, the accuracy of gene variant detection for genes in the genome is not high. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to propose a gene variant detection method, device, electronic device and storage medium, aiming to improve the accuracy of gene variant detection.

[0006] To achieve the above object, in the first aspect of the embodiments of the present application, a gene variant detection method is proposed, and the method includes:

[0007] Aligning the long-read sequencing data of the gene fragment to be tested with the reference genome to obtain the first sequencing result alignment data, and aligning the short-read sequencing data of the gene fragment to be tested with the reference genome to obtain the second sequencing result alignment data;

[0008] Determining the second variant detection result corresponding to each gene locus in the gene fragment to be tested based on the first sequencing result alignment data, and determining the first variant detection result corresponding to each gene locus in the gene fragment to be tested based on the second sequencing result alignment data;

[0009] Identifying the first target gene locus where the first variant detection result is different from the second variant detection result, and when the first target gene locus is in the genomic complex region of the gene fragment to be tested, determining the third variant detection result of the first target gene locus according to the second variant detection result and the variant detection parameters of each gene locus;

[0010] When the first target gene locus is not in the genomic complex region, determining the third variant detection result of the first target gene locus according to the first variant detection result and the variant detection parameters;

[0011] Determine multiple first mutation sites with mutations based on the third mutation detection result, and determine the target gene mutation detection result of the gene fragment to be tested according to the first mutation site and the second mutation site, where the second mutation site is a mutation site that exists in the intersection of the first mutation detection result and the second mutation detection result.

[0012] To achieve the above object, a second aspect of the embodiments of the present application proposes a gene mutation detection device, which includes:

[0013] A comparison unit, configured to align the long-read sequencing data of the gene fragment to be tested to a reference genome to obtain first sequencing result alignment data, and align the short-read sequencing data of the gene fragment to be tested to the reference genome to obtain second sequencing result alignment data;

[0014] A first determination unit, configured to determine a first mutation detection result corresponding to each gene site in the gene fragment to be tested based on the first sequencing result alignment data, and determine a second mutation detection result corresponding to each gene site in the gene fragment to be tested based on the second sequencing result alignment data;

[0015] An identification unit, configured to identify a first target gene site where the first mutation detection result is different from the second mutation detection result. When the first target gene site is in the genomic complex region of the gene fragment to be tested, determine the third mutation detection result of the first target gene site according to the second mutation detection result and the mutation detection parameter of each gene site;

[0016] A second determination unit, configured to determine the third mutation detection result of the first target gene site according to the first mutation detection result and the mutation detection parameter when the first target gene site is not in the genomic complex region;

[0017] A third determination unit, configured to determine multiple first mutation sites with mutations based on the third mutation detection result, and determine the target gene mutation detection result of the gene fragment to be tested according to the first mutation site and the second mutation site, where the second mutation site is a mutation site that exists in the intersection of the first mutation detection result and the second mutation detection result.

[0018] In some embodiments, the identification unit includes:

[0019] A first determination subunit, configured to determine the third mutation detection result of the first target gene site according to the second mutation detection result when the first target gene site is in the genomic complex region and the second mutation detection result corresponding to the first target gene site is a mutation site;

[0020] A first acquisition subunit, configured to, when the first target gene locus is in the genomic complex region and the second variant detection result corresponding to the first target gene locus is a non-variant locus, acquire a first variant detection parameter and a first reference variant parameter range corresponding to the first target gene locus;

[0021] A second determination subunit, configured to determine a third variant detection result of the first target gene locus according to the first variant detection parameter and the first reference variant parameter range.

[0022] Optionally, in some embodiments, the second determination subunit includes:

[0023] A first determination module, configured to, when the variant detection parameter is within the reference variant parameter range, determine that the third variant detection result of the first target gene locus is a variant locus;

[0024] A second determination module, configured to, when the variant detection parameter is not within the reference variant parameter range, determine that the third variant detection result of the first target gene locus is a non-variant locus.

[0025] Optionally, in some embodiments, the second determination unit includes:

[0026] A third determination subunit, configured to, when the first target gene locus is not in the genomic complex region and the first variant detection result corresponding to the first target gene locus is a variant locus, determine the third variant detection result of the first target gene locus according to the first variant detection result;

[0027] A second acquisition subunit, configured to, when the first target gene locus is not in the genomic complex region and the first variant detection result corresponding to the first target gene locus is not a variant locus, acquire a second variant detection parameter and a second reference variant parameter range corresponding to the first target gene locus;

[0028] A fourth determination subunit, configured to determine the third variant detection result of the first target gene locus according to the second variant detection parameter and the second reference variant parameter range.

[0029] Optionally, in some embodiments, the first determination unit includes:

[0030] An evaluation subunit, configured to perform quality evaluation on variant loci determined based on the first sequencing result alignment data by using a first preset neural network model, and classify the variant loci into first candidate variant loci with a quality value higher than a first threshold and second candidate variant loci with a quality value not higher than the first threshold according to the obtained quality value;

[0031] A screening subunit, configured to screen the second candidate mutation sites based on a second preset neural network model, and remove the second candidate mutation sites determined by the model as non-mutated;

[0032] A fifth determination subunit, configured to determine a second mutation detection result corresponding to each gene site in the gene fragment to be tested according to the first candidate mutation site and the screened second candidate mutation sites.

[0033] Optionally, in some embodiments, the evaluation subunit includes:

[0034] A third determination module, configured to determine a plurality of fourth candidate mutation sites according to the first sequencing result alignment data;

[0035] A first acquisition module, configured to acquire the sequencing base proportion information of the fourth candidate mutation sites and their associated sites, and generate a first evaluation feature for each of the fourth candidate mutation sites based on the sequencing base proportion information;

[0036] A first invocation module, configured to invoke the first preset neural network model to perform mutation quality evaluation on the first evaluation feature, and determine a quality value corresponding to each of the fourth candidate mutation sites according to the evaluation result.

[0037] Optionally, in some embodiments, the gene mutation detection device further includes:

[0038] A fourth determination module, configured to determine a first threshold based on a preset proportion quantile value of the quality values corresponding to all the fourth candidate mutation sites.

[0039] Optionally, in some embodiments, the screening subunit includes:

[0040] A second acquisition module, configured to, for each of the second candidate mutation sites, acquire mutation evaluation data within a preset evaluation region of each of the second candidate mutation sites, and determine a second evaluation feature for each of the second candidate mutation sites according to the mutation evaluation data; wherein the preset evaluation region includes the second candidate mutation site and its upstream and downstream regions;

[0041] A second invocation module, configured to invoke the second preset neural network model to perform mutation identification on the second evaluation feature, and determine the second candidate mutation sites determined by the model as mutation sites as third candidate mutation sites.

[0042] Optionally, in some embodiments, the second evaluation feature includes mutation evaluation information, base quality information, and gene typing information.

[0043] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the gene mutation detection method described in the first aspect is implemented.

[0044] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the gene mutation detection method described in the first aspect is implemented.

[0045] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the gene mutation detection method described in the first aspect.

[0046] The gene mutation detection method provided by the embodiments of the present application includes: obtaining first sequencing result alignment data by aligning long-read sequencing data of a gene fragment to be detected to a reference genome, and obtaining second sequencing result alignment data by aligning short-read sequencing data of the gene fragment to be detected to the reference genome; determining a second mutation detection result corresponding to each gene locus in the gene fragment to be detected based on the first sequencing result alignment data, and determining a first mutation detection result corresponding to each gene locus in the gene fragment to be detected based on the second sequencing result alignment data; identifying a first target gene locus where the first mutation detection result is different from the second mutation detection result. When the first target gene locus is in a genomic complex region of the gene fragment to be detected, determining a third mutation detection result of the first target gene locus according to the second mutation detection result and the mutation detection parameters of each gene locus; when the first target gene locus is not in the genomic complex region, determining a third mutation detection result of the first target gene locus according to the first mutation detection result and the mutation detection parameters; determining a plurality of first mutation sites with mutations based on the third mutation detection result, and determining a target gene mutation detection result of the gene fragment to be detected according to the first mutation sites and second mutation sites, where the second mutation sites are mutation sites that have an intersection in the first mutation detection result and the second mutation detection result.

[0047] It can be seen that the gene mutation detection method provided by the embodiments of the present application complements the advantages of mutation detection based on short-read sequencing alignment and the advantages of mutation detection based on nanopore sequencing alignment, so that in non-genomic complex regions, the short-read sequencing alignment results are mainly used to obtain accurate mutation detection results, and in regions with poor short-read sequencing alignment, the nanopore sequencing alignment results are mainly used, thereby improving the accuracy of the entire genome mutation detection. Description of the Drawings

[0048] The accompanying drawings are used to provide a further understanding of the technical solution of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application, and do not constitute a limitation to the technical solution of the present application.

[0049] Figure 1 It is a schematic diagram showing inaccurate gene mutation detection caused by inaccurate alignment of short read lengths for sequencing;

[0050] Figure 2 It is a flowchart of the mutation detection method provided by the present application;

[0051] Figure 3 It is Figure 2 a flowchart of step S202 in

[0052] Figure 4 It is Figure 3 a flowchart of step S301 in

[0053] Figure 5 It is Figure 3 another flowchart of step S301 in

[0054] Figure 6 It is Figure 3 a flowchart of step S302 in

[0055] Figure 7 It is Figure 6 a flowchart of step S601 in

[0056] Figure 8 It is Figure 2 a flowchart of step S203 in

[0057] Figure 9 It is Figure 2 a flowchart of step S204 in

[0058] Figure 10 It is another flowchart of gene mutation detection provided by the present application;

[0059] Figure 11 It is a flowchart of the gene mutation detection device provided by the embodiment of the present application;

[0060] Figure 12 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0061] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations:

[0063] Nanopore Sequencing (NS): Also known as third-generation sequencing technology or long-read sequencing technology, it is a single-molecule sequencing method based on nanopore technology. This method directly detects the base sequence of DNA or RNA molecules using electrical signals. Nanopore sequencing is one of the long-read sequencing technologies.

[0064] Short Read Sequencing (SRS): It is a commonly used sequencing technology in genomics research, mainly involving cutting DNA into shorter fragments (usually dozens to hundreds of base pairs), and then performing high-throughput sequencing on these fragments. Among them, second-generation sequencing technology, also known as next-generation sequencing technology (NGS), is a short-read sequencing technology.

[0065] Gene variant detection: A technology used to identify changes in the structure or composition of DNA or RNA sequences, which may be caused by genetic factors, environmental factors, or the interaction of both. Gene variant detection can be divided into single-gene variant detection and genomic variant detection. In the embodiments of the present application, gene variant detection can specifically be genomic variant detection.

[0066] In related technologies, due to the unique variant information existing in individuals and populations, it contains information related to the occurrence and development of diseases, clinical diagnosis, and evolutionary mechanisms. Therefore, whole-genome variant detection is of great significance in exploring species evolution, understanding human healthy growth and disease treatment, etc. The whole-genome variant calling process includes steps such as sequencing, alignment, variant detection, and false positive site filtering. Among them, the sequencing read length, base recognition accuracy, and alignment accuracy to a certain extent determine whether variants can be correctly identified in a specific genomic background. Existing genomic variant detection schemes, whether based on traditional statistical models or emerging neural network-based variant detection tools, are all aimed at greatly solving the background noise or errors brought by sequencing and alignment, so as to improve the accuracy of variant detection.

[0067] In genomic regions where short reads (usually 150 to 250 base pairs) can be accurately aligned, gene variant detection using short-read sequencing platforms can already obtain relatively accurate detection results. However, short-read sequencing platforms still have significant limitations in complex genomic regions (such as genomic repeat regions). Generally speaking, if the sequencing read length is shorter than the genomic repeat region, it is difficult to align the short reads to their true positions in the genome because it is impossible to determine which segment of the repeat they belong to. Therefore, it is very difficult to obtain accurate variant information by performing variant detection from such genomic repeat regions where short reads are difficult to align using short-read sequencing platforms. Such repeat regions include segmental duplication regions, long interspersed nuclear repeats, short tandem repeats, variable number tandem repeats, telomeres, and microsatellite repeats (up to 30 Mbp).

[0068] As Figure 1 shown, it is a schematic diagram of inaccurate gene variant detection caused by inaccurate short-read sequencing alignment. As shown in the figure, the short-read sequencing result XXXY can be aligned to the first fragment 110, the second fragment 120, and the third fragment 130 in the reference gene sequence. If a gene locus in the short read is detected to have a gene variant, it is impossible to determine which fragment in the entire genome the variant locus is in, resulting in inaccurate genomic variant detection results.

[0069] The emergence of long-read sequencing technologies such as nanopore sequencing has overcome the limitation of inaccurate alignment of short-read platforms in genomic repeat regions. Nanopore sequencing technology is a high-throughput sequencing method based on nanopores that uses the electrochemical resistance in the nanopores to capture and analyze DNA molecules. The core of this technology is that during the process of single-stranded DNA passing through the nanopore, the DNA molecule can be captured and read through the change in current. The advantages of nanopore sequencing technology lie in its powerful capabilities of being fast, portable, real-time, and direct sequencing and its long read length (long read sequencing, LR), making it an important sequencing technology in the fields of studying genomic structure, genetic variation, disease diagnosis, etc. However, the current sequencing accuracy of nanopore sequencing still needs to be improved, and its performance in genomic homopolymer (homopolymers, DNA sequences composed of the same nucleotide units) and heteropolymer regions (heteropolymers, DNA sequences composed of different nucleotide units) is still not good.

[0070] In summary, whether it is a short-read sequencing platform or a long-read sequencing platform, there are certain defects that lead to low accuracy in genomic variant detection.

[0071] Based on this, the embodiments of the present application provide a gene mutation detection method, apparatus, electronic device and storage medium, which combine the advantages of long-read sequencing technology and short-read sequencing technology to improve the accuracy of detecting mutations in the genome. Next, the gene mutation detection method provided by the embodiments of the present application will be described.

[0072] Referring to Figure 2 , in some embodiments, the gene mutation detection method provided by the embodiments of the present application includes but is not limited to steps S201 to S205.

[0073] Step S201: Align the long-read sequencing data of the gene fragment to be tested to the reference genome to obtain the first sequencing result alignment data, and align the short-read sequencing data of the gene fragment to be tested to the reference genome to obtain the second sequencing result alignment data;

[0074] Step S202: Determine the second mutation detection result corresponding to each gene locus in the gene fragment to be tested based on the first sequencing result alignment data, and determine the first mutation detection result corresponding to each gene locus in the gene fragment to be tested based on the second sequencing result alignment data;

[0075] Step S203: Identify the first target gene locus where the first mutation detection result is different from the second mutation detection result. When the first target gene locus is in the genomic complex region of the gene fragment to be tested, determine the third mutation detection result of the first target gene locus according to the second mutation detection result and the mutation detection parameters of each gene locus;

[0076] Step S204: When the first target gene locus is not in the genomic complex region, determine the third mutation detection result of the first target gene locus according to the first mutation detection result and the mutation detection parameters;

[0077] Step S205: Determine multiple first mutation sites with mutations based on the third mutation detection result, and determine the target gene mutation detection result of the gene fragment to be tested according to the first mutation site and the second mutation site.

[0078] Steps S201 to S205 illustrated in the embodiments of the present application, when performing gene mutation detection on a gene fragment to be tested, can first perform long-read sequencing and short-read sequencing on the gene fragment to be tested respectively, and obtain corresponding long-read sequencing data and short-read sequencing data. Then, align the sequencing data obtained by the two sequencing methods to the reference genome respectively to obtain long-read sequencing alignment data and short-read sequencing alignment data. Further, perform mutation detection based on the long-read sequencing alignment data and the short-read sequencing alignment data respectively to determine the long-read mutation detection result and the short-read mutation detection result. Since both the long-read mutation detection result and the short-read mutation detection result have inaccurate mutation detection results at some base sites due to their own defects, this solution proposes that after obtaining the long-read mutation detection result and the short-read mutation detection result of the gene fragment to be tested respectively, the base sites determined to be mutated by both the long-read mutation detection result and the short-read mutation detection result are determined as accurate mutation sites. For the base sites where the long-read mutation detection result and the short-read mutation detection result are inconsistent, if the base site is in a genomic complex region, the long-read mutation detection result and the mutation detection parameters of the base site are used to comprehensively determine whether it is a mutation site; if the base site is not in a genomic complex region, the short-read mutation detection result and the mutation detection parameters of the base site are used to comprehensively determine whether it is a mutation site. In this way, in the genomic complex region, the advantage of long-read sequencing can be used to avoid the problem of inaccurate mutation detection results caused by the inability of short-reads to determine the accurate section, and in the non-genomic complex region, the advantage of accurate short-read sequencing can be used to avoid the problem of inaccurate mutation detection results caused by the low accuracy of long-read sequencing results, thereby improving the accuracy of gene detection for the gene fragment to be tested.

[0079] In step S201 of some embodiments, the gene fragment to be tested can be any randomly sampled gene fragment, such as a DNA fragment or an RNA fragment, or some gene fragment standards can also be used to verify the effect of gene mutation detection. For example, when the gene fragment to be tested is a DNA fragment, using the gene mutation detection method provided by the present application to perform gene mutation detection on the DNA fragment, the DNA fragment can be sequenced simultaneously using nanopore sequencing technology and next-generation sequencing technology to obtain the long-read sequencing data and short-read sequencing data of the DNA fragment. Among them, in some other embodiments, other sequencing technologies can also be used to measure the long-read sequence and short-read sequence of the DNA fragment. The long-read sequencing data measured by nanopore sequencing and the short-read sequencing data measured by next-generation sequencing technology here are only examples and do not limit the specific sequencing means for the DNA fragment.

[0080] After obtaining the long-read sequencing data and short-read sequencing data of the DNA fragment, a mapping software can be further used to map the long-read sequencing data and the short-read sequencing data to a pre-set reference genome respectively to obtain their respective corresponding mapping result data, and then the mapping result data can be further converted in format to a mapping result file in Binary Alignment / Map (BAM) format.

[0081] In step S202 of some embodiments, after mapping the nanopore sequencing sequence and the next-generation sequencing sequence to the reference genome respectively to obtain their respective corresponding mapping result files, variant detection can be performed based on their respective corresponding mapping result files using a short-read sequencing platform and a long-read sequencing platform respectively to obtain variant detection results output by different sequencing platforms. Specifically, the long-read sequencing platform can output a second variant detection result, and the short-read sequencing platform can output a first variant detection result. Among them, in the embodiments of the present disclosure, the variant detection result can specifically be whether each base site in the gene sequence is a variant site, and when it is determined that a certain base site is a variant site, the variant detection result can further include the specific type of the variant. The variant types can include Single Nucleotide Polymorphisms (SNP) and Insertion-Deletion (indel).

[0082] Among them, the relatively mature variant detection tool in the short-read platform is the Genome Analysis Toolkit (GATK). It mainly uses a Gaussian mixture model to recalibrate the variant quality score and machine learning models such as Bayesian models to filter candidate variant sites. With the progress of sequencing technology, a large amount of sequencing data and benchmarks available for training and testing have been accumulated, covering millions of variants in a series of genomic environments, making it possible to use deep learning for the detection of small variants in the human genome. Currently, the mainstream architecture of variant detection tools based on deep learning is the convolutional neural network (CNN). Traditional machine learning-based variant detection tools use manually designed sequence features (e.g., whether it is a homopolymer) and read features aligned to the reference genome (e.g., whether the variant has strand bias). To reduce the need for feature engineering, the CNN architecture utilizes the information from reads and the association between the variant site and the surrounding reference genome sequence. Through appropriate hyperparameter training, the CNN model can approximate a complex non-linear function to classify potential variant sites into homozygous variants, heterozygous variants, or non-variant sites. The representative tool for variant detection using the CNN framework is DeepVariant, which has achieved excellent variant detection performance on both short-read and long-read sequencing platforms. In addition, some methods using a framework that combines RNN and CNN for whole-genome small variant detection (such as clair3) have also achieved excellent performance. Thus, it can be seen that deep learning-based methods have gradually replaced traditional statistical methods and become the mainstream methods for genomic variant detection.

[0083] Referring to Figure 3 , in some embodiments, step S202 includes but is not limited to steps S301 to S303.

[0084] Step S301, perform quality assessment on the variant sites determined based on the first sequencing result alignment data using a first preset neural network model, and divide the variant sites into first candidate variant sites with a quality value higher than a first threshold and second candidate variant sites with a quality value not higher than the first threshold according to the obtained quality values;

[0085] Step S302, screen the second candidate variant sites using a second preset neural network model, and remove the second candidate variant sites determined by the model as non-variants;

[0086] Step S303, determine the second variant detection result corresponding to each gene site in the gene fragment to be tested according to the first candidate variant sites and the third candidate variant sites.

[0087] In step S301 of some embodiments, the first preset neural network model may be a preset recurrent neural network model (RNN), specifically a two-layer bidirectional long short-term memory network (biLSTM). The preset recurrent neural network model is used to evaluate the quality of the mutation sites determined from the first sequencing result alignment data. Specifically, first, multiple potential mutation sites in the gene sequence are determined based on the first sequencing result alignment data. Then, for these potential mutation sites, their corresponding RNN network input features can be obtained. Next, the input features corresponding to each potential mutation site are input into the RNN network, so that the genotype information and mutation quality value (mutation quality) corresponding to each potential mutation site can be obtained. Then, based on a threshold, the potential mutation sites corresponding to different quality values are classified. Specifically, the potential mutation sites with quality values higher than the threshold are classified as high-quality first candidate mutation sites, and the potential mutation sites with quality values not higher than the threshold are classified as low-quality second candidate mutation sites. The high-quality first candidate mutation sites can be saved, and the low-quality second candidate mutation sites can be further evaluated.

[0088] In step S302 of some embodiments, after using the RNN to classify the potential mutation sites into high-quality first candidate mutation sites and low-quality second candidate mutation sites according to the mutation quality, a second preset neural network model, such as a preset convolutional neural network model (CNN), can be further used to further screen the determined low-quality second candidate mutation sites. Among them, the CNN model may specifically be a CNN model including a residual network (ResNet). Specifically, for each second candidate mutation site, the mutation evaluation data corresponding to the candidate mutation site is first obtained, and then the CNN model is used to discriminate the mutation evaluation data to further filter the low-quality mutation sites in the second candidate mutation sites, and the true mutation sites are obtained therefrom as the third candidate mutation sites.

[0089] Among them, both the RNN proposed in step S301 and the CNN proposed in step S302 can be models that have been pre-trained using sample data in advance.

[0090] In step S203 of some embodiments, after using the RNN model and the CNN model to perform two-layer screening on the potential mutation sites, the separately screened high-quality mutation sites are uniformly summarized as the mutation sites determined based on the first sequencing result alignment data, and then the first mutation detection result corresponding to the first sequencing result alignment data can be further determined accordingly.

[0091] Steps S301 to S303 use deep learning methods to detect gene variations in sequencing alignment data measured by different sequencing platforms, which can improve the detection efficiency and accuracy of gene variation detection based on sequencing alignment data. For the second sequencing alignment result, similar steps to steps S301 to S303 can also be used to perform variation detection to obtain the second variation detection result.

[0092] Please refer to Figure 4 , in some embodiments, step S301 includes but is not limited to steps S401 to S403.

[0093] Step S401, determining a plurality of fourth candidate variation sites according to the first sequencing result alignment data;

[0094] Step S402, obtaining the sequencing base proportion information of the fourth candidate variation site and its associated sites, and generating a first evaluation feature for each fourth candidate variation site based on the sequencing base proportion information;

[0095] Step S403, calling a first preset neural network model to evaluate the variation quality of the first evaluation feature, and determining the quality value corresponding to each fourth candidate variation site according to the evaluation result.

[0096] Among them, in the embodiments of the present disclosure, the sequencing base proportion information corresponding to each gene locus in the whole genome of the DNA to be tested can be obtained first. The sequencing base proportion information can be the proportion information of each base (A, G, C or T) in a plurality of sequencing result data obtained by sequencing a certain locus. For example, if the locus is sequenced 100 times, 20 times are A, 70 times are G, 5 times are C, and 5 times are T, then the sequencing base proportion information corresponding to this locus is A: 20%, G: 70%, C: 5%, T: 5%. Generally speaking, if a locus is a non-variation locus, then the base composition of this locus is single (that is, the proportion of a certain base in ATGC exceeds 90%), and at the same time the single base is the same as the base at this locus in the reference genome. Otherwise, if the base composition of this locus is not single, or the base composition is single but inconsistent with the base at this locus in the reference genome, then this locus is determined as an abnormal locus, or called a fourth candidate variation site.

[0097] Then, the sequencing base ratio information of each fourth candidate variant site and its associated sites can be obtained. Among them, the associated site can be specifically the site of the abnormal site and its flanking region, wherein the flanking region can be specifically the region including the abnormal site and its upstream and downstream regions, and the length of the flanking region here can be 128bp. After obtaining the sequencing base ratio information of each fourth candidate variant site and its associated sites, the sequencing base ratio information of the fourth candidate variant site and its associated sites can be summarized into a two-dimensional tensor as the first evaluation feature of the RNN model to evaluate the fourth candidate variant site.

[0098] Furthermore, the RNN model may be called to perform variation quality assessment on the first assessment feature corresponding to each fourth candidate variation site, thereby obtaining a quality value corresponding to each fourth candidate variation site.

[0099] Please refer to Figure 5 In some embodiments, step S301 may also include but is not limited to steps S501 to S503.

[0100] Step S501, determining a first threshold based on a preset ratio quantile value of quality values corresponding to all fourth candidate variant sites;

[0101] Step S502, determining the variation site whose quality value among the fourth candidate variation sites is higher than the first threshold as the first candidate variation site;

[0102] Step S503: Determine the variation sites whose quality values among the fourth candidate variation sites are not higher than the first threshold as the second candidate variation sites.

[0103] In the embodiment of the present disclosure, it is not necessary to determine the threshold for classifying the quality of gene variation in advance, because for different gene fragments to be tested, the threshold for classifying the quality of gene variation sites into high-quality variation sites and low-quality variation sites may be different. Therefore, in the embodiment of the present disclosure, after using RNN to evaluate and obtain the quality value corresponding to each fourth candidate variation site, the quality values of all fourth candidate variation sites can be aggregated, and then the first threshold is determined based on the preset proportional quantile value of the aggregated quality value.

[0104] Specifically, for example, the 70th percentile of the quality value can be set as the first threshold, that is, the quality values of the fourth candidate variant sites are sorted from high to low, and then the quality value sorted at the 70th percentile is determined as the first threshold. Further, the fourth candidate variant site with a quality value higher than the first threshold can be determined as the first candidate variant site, and then the fourth candidate variant site with a quality value lower than the fourth candidate variant site can be determined as the second candidate variant site.

[0105] Reference Figure 6, in some embodiments, step S302 includes but is not limited to steps S601 to S602.

[0106] Step S601, obtain the variant evaluation data within the adjacent region of each second candidate variant site, and determine the second evaluation feature of each second candidate variant site according to the variant evaluation data;

[0107] Step S602, call the second preset neural network model to perform variant identification on the second evaluation feature, and determine the second candidate variant site determined by the model as a variant site as the third candidate variant site.

[0108] In steps S601 to S602 of some embodiments, after determining low-quality second candidate variant sites based on RNN for variant detection of sequencing result alignment data, CNN can be further used to screen the low-quality second candidate variant sites to screen out true variant sites with relatively high variant quality. In this process, the variant evaluation data within the adjacent region of each second candidate variant site can be obtained first, and the second evaluation feature of each second candidate variant site can be determined according to the variant evaluation data. Among them, the adjacent region of the second candidate variant site can be understood as the flanking segment of the second candidate variant site, that is, the aforementioned region including the second candidate variant site and the adjacent sites on its read length; the variant evaluation data can include multiple feature data, and these feature data can support the evaluation of the variant quality of the second candidate variant site.

[0109] After obtaining the variant evaluation data corresponding to each second candidate variant site, the second evaluation feature of each second candidate variant site can be further generated based on the variant evaluation data. Specifically, the variant evaluation data can be embedded according to the input specification of the CNN model, so as to convert the variant evaluation data corresponding to the second candidate variant site into its corresponding feature vector, thereby obtaining the second evaluation feature corresponding to each second candidate variant site.

[0110] Furthermore, the preset CNN model can be called to process the second evaluation feature, so as to realize the variant identification of the second candidate variant site, and then determine the second candidate variant site identified as a variant site by the CNN model as the third candidate variant site according to the variant identification result.

[0111] That is, in the embodiments of the present disclosure, the CNN model can be a binary classification model, the input of the model is the second evaluation feature, and the input is yes or no, that is, the output is the variant detection conclusion of the second candidate variant site, and this conclusion indicates that the second candidate variant site is a variant site or not a variant site.

[0112] Please refer to Figure 7In some embodiments, step S601 includes but is not limited to steps S701 to S702.

[0113] Step S701, obtaining mutation assessment information, base quality information, and genotyping information corresponding to a preset number of base pairs upstream and downstream of the corresponding read length for each second candidate variant site;

[0114] Step S702: Generate a second evaluation feature for each second candidate variant site based on mutation evaluation information, base quality information, and genotyping information corresponding to a preset number of base pairs.

[0115] In some embodiments, in step S701 to step S702, the neighboring region of the second candidate variation site may be a flanking segment within a preset sequence length range upstream and downstream of the read where the second candidate variation site is located. Specifically, the flanking segment here may be a base sequence 15 bp upstream and downstream of the second candidate variation site on the read. When the aforementioned CNN model is required to perform variation evaluation on the second candidate variation site, the mutation evaluation information, base quality information, and genotyping information contained in the flanking segment of the second candidate variation site may be obtained to generate a second evaluation feature of the second candidate variation site.

[0116] The mutation assessment information may include the mutation (variation) base sequence contained in the flanking segment, the mutation marker information, etc., and the assessment feature may also include the mutation marker, insertion sequence, and comparison quality information. In detail, the second assessment feature of the second candidate mutation site may be obtained in the following manner:

[0117] First, for the second candidate variant site, its location on the chromosome and the chromosome segment containing 15bp upstream and downstream can be determined:

[0118] A. Extract linear reference genome base sequence based on chromosome segment location

[0119] B. At the same time, obtain the base quality of the read containing the target segment from fastq according to the positioning

[0120] C. Count the base mutations of the reads in this segment

[0121] D. Mark the target mutation (AC, AG, AT...) based on the statistics of reads in this segment

[0122] E. Get the chain direction information from the bam containing the alignment information according to the read name

[0123] F. Get the insertion sequence of the read in this segment from the bam containing the alignment information according to the read name

[0124] G. Obtain the alignment quality of reads per day from the bam containing alignment information according to the read name

[0125] H. Use a genotyping tool to obtain the genotyping information of each read.

[0126] Then, the obtained information can be embedded to obtain the corresponding feature vector, that is, the second evaluation feature. Then, the second evaluation feature can be input into the aforementioned CNN model for evaluation to determine whether the second candidate variant site is a true variant site.

[0127] In step S203 of some embodiments, after performing variant detection on the first sequencing result alignment data and the second sequencing result alignment data based on the above solution to obtain the corresponding first variant detection result and second variant detection result respectively, the embodiments of the present application can be used to fuse the results of the first variant detection result and the second variant detection result. Specifically, when fusing the results, the first target gene loci where the first variant detection result and the second variant detection result are different can be distinguished first. That is, for the loci where the first variant detection result and the second variant detection result are the same, it can be determined that the variant detection result is accurate. At this time, when the variant detection results simultaneously indicate that a certain locus is a variant locus, it can be determined that it is a true variant locus. Since the first variant detection result of the target gene locus is different from the second variant detection result, that is, there is a result in the first variant detection result and the second variant detection result indicating that the target gene locus is a variant locus. At this time, it can be determined whether to use the first variant detection result or the second variant detection result as the standard according to whether the first target gene locus is in the genomic complex region of the gene fragment to be tested.

[0128] Among them, the genomic complex region may specifically include regions such as segmental duplication sequence regions, tandem repeat sequence regions, variable number tandem repeat sequence regions, etc., which are regions where short reads are difficult to align. In these regions, short read alignment is difficult to align to the accurate region, resulting in inaccurate results of short read gene variant detection. Long read sequencing, due to its long read length advantage, can avoid the problem of not being able to align to the accurate region mentioned above, and thus can obtain more accurate variant detection results.

[0129] Therefore, when the first target gene locus is in the genomic complex region of the gene fragment to be tested, the true variant detection result of the first target gene locus can be mainly determined according to the long read variant detection result, that is, the second variant detection result, that is, the third variant detection result of the first target gene locus is determined.

[0130] Refer to Figure 8 , in some embodiments, step S203 includes but is not limited to steps S801 to S803.

[0131] Step S801: When the first target gene locus is in a complex genomic region and the second variant detection result corresponding to the first target gene locus is a variant locus, determine the third variant detection result of the first target gene locus according to the second variant detection result.

[0132] Step S802: When the first target gene locus is in a complex genomic region and the second variant detection result corresponding to the first target gene locus is not a variant locus, obtain the first variant detection parameter corresponding to the first target gene locus and the first reference variant parameter range.

[0133] Step S803: Determine the third variant detection result of the first target gene locus according to the first variant detection parameter and the first reference variant parameter range.

[0134] In steps S801 to S803 of some embodiments, when it is detected that the first target gene locus is in a complex genomic region, it can be detected whether the second variant detection result corresponding to the first target gene locus determines that the first target gene locus is a variant locus. When the second variant detection result indicates that the first target gene locus is a variant locus, the third variant detection result of the first target gene locus can be determined according to the second variant detection result, that is, it can be determined that the first target gene locus is a true variant locus.

[0135] When the second variant detection result indicates that the first target gene locus is not a variant locus, the first variant detection parameter corresponding to the first target gene locus and the corresponding first reference variant parameter range can be further obtained. Then, the third variant detection result of the first target gene locus can be determined according to the first variant detection parameter and the first reference variant parameter range. Specifically, the first variant detection parameter can be compared with the first reference variant parameter range. When the first variant detection parameter is within the first reference variant parameter range, it is determined that the first target gene locus is a true variant locus; conversely, if the first variant detection parameter is not within the first reference variant parameter range, it is determined that the first target gene locus is not a true variant locus.

[0136] Among them, the first variant detection parameter may specifically include the following information of the first target gene locus:

[0137] a. Mutation frequency (number of reads containing a specific mutation / total number of reads containing the locus);

[0138] b. Mutation type (single nucleotide base mutation, insertion, and deletion);

[0139] c. Mutation quality, output according to the model (which can be used as an indicator to evaluate whether the mutation is reliable).

[0140] In step S204 of some embodiments, when it is detected that the first target gene locus is not in a genomic complex region, considering that the accuracy of short-read sequencing is higher than that of long-read sequencing, the true variant detection result of the first target gene locus can be determined mainly based on the short-read variant detection result.

[0141] Referring Figure 9 , in some embodiments, step S204 includes but is not limited to steps S901 to S903.

[0142] Step S901, when the first target gene locus is not in a genomic complex region and the first variant detection result corresponding to the first target gene locus is a variant locus, determine the third variant detection result of the first target gene locus according to the first variant detection result;

[0143] Step S902, when the first target gene locus is not in a genomic complex region and the first variant detection result corresponding to the first target gene locus is not a variant locus, obtain the second variant detection parameter and the second reference variant parameter range corresponding to the first target gene locus;

[0144] Step S903, determine the third variant detection result of the first target gene locus according to the second variant detection parameter and the second reference variant parameter range.

[0145] In steps S901 to S903 of some embodiments, when it is detected that the first target gene locus is not in a genomic complex region, the first variant detection result corresponding to the first target gene locus can be further obtained. If the first variant detection result corresponding to the first target gene locus indicates that the first target gene locus is a variant locus, the third variant detection result of the first target gene locus can be determined according to the first variant detection result, that is, it is determined that the first target gene locus is a true variant locus. If the first variant detection result corresponding to the first target gene locus indicates that the first target gene locus is not a variant locus, the second variant detection parameter and the second reference variant parameter range corresponding to the first target gene locus can be further obtained.

[0146] Then, the third variant detection result of the first target gene locus can be determined according to the second variant detection parameter and the second reference variant parameter range. Specifically, when the second variant detection parameter is within the second reference variant parameter range, it can be determined that the first target gene locus is a true variant locus; conversely, when the second variant detection parameter is not within the second reference variant parameter range, it can be determined that the first target gene locus is not a true variant locus.

[0147] Among them, for different mutation types, for example, when it is determined that the mutation type of the first target gene locus is SNP, the mutation detection result of the first target gene locus can be used to determine the third mutation detection result with the second mutation detection result as the main and the first mutation detection result as the supplement; when it is determined that the mutation type of the first target gene locus is insertion / deletion (indel), the mutation detection result of the first target gene locus can be used to determine its third mutation detection result with the first mutation detection result as the main and the second mutation detection result as the supplement.

[0148] Therefore, based on the existing sequencing platforms, in the genomic regions where short reads can be accurately aligned, using short-read data for variant detection can already obtain very accurate variant detection results. However, regardless of the base accuracy, short reads will hinder the variant detection in relatively complex regions of the genome such as large tandem repeats and highly homologous regions (such as segmental duplications) as well as highly variable medically relevant human leukocyte antigens. In addition, short-read sequencing platforms have blind spots in high-GC regions and cannot provide effective sequencing data for these regions, which brings difficulties to variant detection. For the challenges encountered by short-read platforms, the direct sequencing and long-read characteristics of nanopore sequencing can be well solved. The longer read length makes the alignment of reads in complex genomic regions more accurate. At the same time, due to the advantages of the sequencing principle, nanopore sequencing is not affected in high-GC regions of the genome. Compared with short-read sequencing platforms, the nanopore sequencing platform has excellent variant detection performance in these regions. However, the sequencing error rate of nanopore long-read sequencing still needs to be further optimized. In addition, the problem of small indels (insertions / deletions) commonly existing in nanopore sequencing data makes the nanopore sequencing platform still face challenges in variant detection. And the variant detection method provided by this application can integrate the excellent performance of long-read sequencing platforms in complex genomic regions and the advantage of the sequencing accuracy of short-read sequencing platforms to greatly improve the accuracy of genomic variant detection.

[0149] As Figure 10 shown, it is another flow schematic diagram of the gene variant detection method provided by this application. As shown in the figure, for the gene variant detection method provided by this application, the long-read sequencing platform and the short-read sequencing platform are first used to sequence the gene fragment to be tested respectively to obtain long-read sequencing data and short-read sequencing data, and then the long-read sequencing data and the short-read sequencing data are respectively aligned to the reference genome to obtain long-read sequence alignment data and short-read sequence alignment data. Further, deep learning methods are used to perform variant detection on the long-read sequence alignment data and the short-read sequence alignment data respectively to obtain the corresponding long-read variant detection data and short-read variant detection data.

[0150] When both the long-read variant detection data and the short-read variant detection data of a gene locus indicate that the gene locus is a variant locus, it is determined that the variant locus is a true variant locus; when the long-read variant detection data and the short-read variant detection data of a gene locus do not both indicate that the gene locus is a variant locus, it is possible to further distinguish whether the gene locus is in a complex genomic region. When the gene locus is in a complex genomic region, it is possible to determine whether the gene locus is a true variant locus based on the long-read variant detection result; when the gene locus is not in a complex genomic region, it is possible to determine whether the gene locus is a true variant locus based on the short-read variant detection result.

[0151] The present application also discloses the following specific embodiments for implementing the gene detection method provided by the present application:

[0152] 1) Simultaneously perform nanopore long-read sequencing and second-generation sequencing on the HG002 standard sample to obtain long-read sequence data and corresponding short-read sequencing data. Then, use the alignment software minimap2 (v 2.24) and bwa (0.7.17-r1188) to align the long-read sequence and the short-read sequence to the GRch38 genome respectively, and obtain the corresponding alignment result files in BAM format.

[0153] 2) Obtain the pileup statistical results of the sequencing data from the BAM file containing alignment information obtained above. Then, use the model of the RNN framework containing a two-layer bidirectional long short-term memory network that has been trained with a standard data set to perform preliminary variant detection on the target locus in the long-read sequence, and obtain the potential variant loci in the whole genome of the HG002 sample.

[0154] 3) For the mutation sites output by the RNN model, mark the variant sites with quality scores in the top 70% as high-quality variant sites, and mark the remaining 30% of the variant sites as low-quality variant sites.

[0155] 4) Use the variant sites marked as low-quality and their corresponding sequences, base qualities, base mutation sequences included in the target interval, markers of target mutations, plus / minus strand markers, insertion sequences, alignment qualities, and genotyping information in the region of the genomic flank chain (i.e., the region including the variant site and 128 bp in length upstream and downstream thereof) as inputs, and use the CNN neural network framework containing a residual neural network model to further screen these low-quality sites to obtain the potential true variant sites among them.

[0156] 5) Obtain the variant detection results corresponding to the long-read and short-read sequencing data respectively, and then use an iterator to screen these variant sites: if the variant sites and variant characteristics exist in both data, write them into the final vcf file; conversely, if the variant detection sites only exist in one type of data, first determine whether the site exists in the complex genomic region (regions that are not friendly to short reads, such as the long-read dominant regions obtained through prior analysis like Figure 3 , which are known and provided here). If so, take the variant sites identified by the long-read data as the main ones, and at the same time filter the variant sites existing in the short-read data based on quality and mutation frequency, retaining high-quality variant sites; if not, take the variant sites identified by the short-read data as the main ones, and at the same time filter the variant sites existing in the long-read data based on quality and mutation frequency, and also retain the high-quality variant sites in the long-read. Thus, the final whole-genome variant detection results are obtained.

[0157] Use the benchmark evaluation tool hap.py (v0.3.15) to evaluate the performance of the results, and the evaluation results are shown in Table 1. It can be seen from the analysis results that compared with the existing variant detection solutions, the method proposed in the present invention has a significant improvement in the detection of small variants in the whole genome, whether it is the identification of SNPs or INDELs (F1-score >= 0.9).

[0158] Next, the gene variant detection device provided by the embodiments of the present application will be described.

[0159] As shown in Table 1, it is a comparison table of the effects corresponding to different variant detection schemes. It can be seen from the data in the table that the long-read and short-read combination scheme provided by the present application can obtain better variant detection accuracy.

[0160] Table 1 Comparison table of variant detection effects

[0161]

[0162] Referring to Figure 11 , in some embodiments, the embodiments of the present application also provide a gene variant detection device 1100, and the gene variant detection device 1100 includes:

[0163] An alignment unit 1101, configured to align the long-read sequencing data of the gene fragment to be tested to a reference genome to obtain the first sequencing result alignment data, and align the short-read sequencing data of the gene fragment to be tested to the reference genome to obtain the second sequencing result alignment data;

[0164] The first determination unit 1102 is configured to determine the second variant detection result corresponding to each gene locus in the gene fragment to be tested based on the first sequencing result alignment data, and determine the first variant detection result corresponding to each gene locus in the gene fragment to be tested based on the second sequencing result alignment data;

[0165] The recognition unit 1103 is configured to recognize the first target gene locus where the first variant detection result is different from the second variant detection result. When the first target gene locus is in the genomic complex region of the gene fragment to be tested, determine the third variant detection result of the first target gene locus according to the second variant detection result and the variant detection parameters of each gene locus;

[0166] The second determination unit 1104 is configured to, when the first target gene locus is not in the genomic complex region, determine the third variant detection result of the first target gene locus according to the first variant detection result and the variant detection parameters;

[0167] The third determination unit 1105 is configured to determine a plurality of first variant sites with variants based on the third variant detection result, and determine the target gene variant detection result of the gene fragment to be tested according to the first variant site and the second variant site. The second variant site is a variant site that has an intersection in the first variant detection result and the second variant detection result.

[0168] In some embodiments, the recognition unit includes:

[0169] The first determination subunit is configured to, when the first target gene locus is in the genomic complex region and the second variant detection result corresponding to the first target gene locus is a variant site, determine the third variant detection result of the first target gene locus according to the second variant detection result;

[0170] The first acquisition subunit is configured to, when the first target gene locus is in the genomic complex region and the second variant detection result corresponding to the first target gene locus is not a variant site, acquire the first variant detection parameter and the first reference variant parameter range corresponding to the first target gene locus;

[0171] The second determination subunit is configured to determine the third variant detection result of the first target gene locus according to the first variant detection parameter and the first reference variant parameter range.

[0172] Optionally, in some embodiments, the second determination subunit includes:

[0173] The first determination module is configured to, when the variant detection parameter is within the reference variant parameter range, determine that the third variant detection result of the first target gene locus is a variant site;

[0174] A second determination module, configured to determine that the third mutation detection result of the first target gene locus is not a mutation locus when the mutation detection parameter is not within the reference mutation parameter range.

[0175] Optionally, in some embodiments, the second determination unit includes:

[0176] A third determination subunit, configured to determine the third mutation detection result of the first target gene locus according to the first mutation detection result when the first target gene locus is not in the genomic complex region and the first mutation detection result corresponding to the first target gene locus is a mutation locus;

[0177] A second acquisition subunit, configured to acquire the second mutation detection parameter and the second reference mutation parameter range corresponding to the first target gene locus when the first target gene locus is not in the genomic complex region and the first mutation detection result corresponding to the first target gene locus is not a mutation locus;

[0178] A fourth determination subunit, configured to determine the third mutation detection result of the first target gene locus according to the second mutation detection parameter and the second reference mutation parameter range.

[0179] Optionally, in some embodiments, the first determination unit includes:

[0180] An evaluation subunit, configured to perform quality evaluation on mutation loci determined based on the first sequencing result alignment data by using a first preset neural network model, and classify the mutation loci into first candidate mutation loci with a quality value higher than a first threshold and second candidate mutation loci with a quality value not higher than the first threshold according to the obtained quality value;

[0181] A screening subunit, configured to screen the second candidate mutation loci by using a second preset neural network model to obtain third candidate mutation loci determined by the second preset neural network model as mutation loci;

[0182] A fifth determination subunit, configured to determine the second mutation detection result corresponding to each gene locus in the gene fragment to be detected according to the first candidate mutation loci and the third candidate mutation loci.

[0183] Optionally, in some embodiments, the evaluation subunit includes:

[0184] A third determination module, configured to determine a plurality of fourth candidate mutation loci according to the first sequencing result alignment data;

[0185] A first acquisition module, configured to acquire sequencing base proportion information of the fourth candidate mutation site and its associated sites, and generate a first evaluation feature for each of the fourth candidate mutation sites based on the sequencing base proportion information;

[0186] A first call module, configured to call a first preset neural network model to evaluate the mutation quality of the first evaluation feature, and determine a quality value corresponding to each of the fourth candidate mutation sites according to the evaluation result.

[0187] Optionally, in some embodiments, the evaluation subunit further includes:

[0188] A fourth determination module, configured to determine a first threshold based on a preset proportional quantile value of the quality values corresponding to all the fourth candidate mutation sites;

[0189] A fifth determination module, configured to determine a mutation site with a quality value higher than the first threshold among the fourth candidate mutation sites as a first candidate mutation site;

[0190] A sixth determination module, configured to determine a mutation site with a quality value not higher than the first threshold among the fourth candidate mutation sites as a second candidate mutation site.

[0191] Optionally, in some embodiments, the screening subunit includes:

[0192] A second acquisition module, configured to acquire mutation evaluation data in a neighboring region of each second candidate mutation site, and determine a second evaluation feature for each second candidate mutation site according to the mutation evaluation data;

[0193] A second call module, configured to call a second preset neural network model to identify mutations of the second evaluation feature, and determine a second candidate mutation site determined by the model as a mutation site as a third candidate mutation site.

[0194] Optionally, in some embodiments, the second acquisition module includes:

[0195] An acquisition sub-module, configured to acquire mutation evaluation information, base quality information, and genotyping information corresponding to a preset number of base pairs upstream and downstream of each second candidate mutation site in its respective read length;

[0196] A generation sub-module, configured to generate a second evaluation feature for each second candidate mutation site based on the mutation evaluation information, base quality information, and genotyping information corresponding to the preset number of base pairs.

[0197] Reference Figure 12 , Figure 12 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0198] The processor 1201 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0199] The memory 1202 can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called by the processor 1201 to execute the gene mutation detection method of the embodiments of the present application;

[0200] The input / output interface 1203 is used to implement information input and output;

[0201] The communication interface 1204 is used to implement communication and interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0202] The bus 1205 transmits information between the various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204);

[0203] Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are communicatively connected to each other inside the device through the bus 1205.

[0204] The embodiments of the present application also provide a computer program product, which includes a computer program. The processor of the computer device reads and executes this computer program, so that the computer device executes to implement the above-mentioned gene mutation detection method.

[0205] The terms "first", "second", "third", "fourth", etc. (if any) in the description of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprise" and "include" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0206] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0207] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.

[0208] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0209] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0210] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0211] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0212] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0213] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A method for detecting gene mutations, characterized in that, The method includes: Aligning the long-read sequencing data of the gene fragment to be tested to a reference genome to obtain first sequencing result alignment data, and aligning the short-read sequencing data of the gene fragment to be tested to the reference genome to obtain second sequencing result alignment data; Performing quality assessment on the variant sites determined based on the first sequencing result alignment data by using a preset recurrent neural network model, and determining a second variant detection result corresponding to each gene site in the gene fragment to be tested according to the quality assessment result, and determining a first variant detection result corresponding to each gene site in the gene fragment to be tested based on the second sequencing result alignment data; Identifying a first target gene site where the first variant detection result is different from the second variant detection result. When the first target gene site is in the genomic complex region of the gene fragment to be tested, determining a third variant detection result of the first target gene site according to the second variant detection result and the variant detection parameters of each gene site; When the first target gene site is not in the genomic complex region, determining a third variant detection result of the first target gene site according to the first variant detection result and the variant detection parameters; Determining a plurality of first variant sites with variants based on the third variant detection result, and determining a target gene variant detection result of the gene fragment to be tested according to the first variant sites and second variant sites, where the second variant sites are variant sites that have an intersection in the first variant detection result and the second variant detection result.

2. The method according to claim 1, characterized in that, The step of when the first target gene site is in the genomic complex region of the gene fragment to be tested, determining a third variant detection result of the first target gene site according to the second variant detection result and the variant detection parameters of each gene site includes: When the first target gene site is in the genomic complex region and the second variant detection result corresponding to the first target gene site is a variant site, determining the third variant detection result of the first target gene site according to the second variant detection result; When the first target gene site is in the genomic complex region and the second variant detection result corresponding to the first target gene site is a non-variant site, obtaining the first variant detection parameter and the first reference variant parameter range corresponding to the first target gene site; Determining the third variant detection result of the first target gene site according to the first variant detection parameter and the first reference variant parameter range.

3. The method according to claim 2, characterized in that, The step of determining the third variant detection result of the first target gene site according to the first variant detection parameter and the first reference variant parameter range includes: When the variant detection parameter is within the reference variant parameter range, determining the third variant detection result of the first target gene site as a variant site; When the variant detection parameter is not within the reference variant parameter range, determining the third variant detection result of the first target gene site as a non-variant site.

4. The method according to any one of claims 1 to 3, characterized in that, When the first target gene locus is not in the genomic complex region, determining a third variant detection result of the first target gene locus according to the first variant detection result and the variant detection parameter includes: When the first target gene locus is not in the genomic complex region and the first variant detection result corresponding to the first target gene locus is a variant locus, determining the third variant detection result of the first target gene locus according to the first variant detection result; When the first target gene locus is not in the genomic complex region and the first variant detection result corresponding to the first target gene locus is not a variant locus, obtaining a second variant detection parameter and a second reference variant parameter range corresponding to the first target gene locus; Determining the third variant detection result of the first target gene locus according to the second variant detection parameter and the second reference variant parameter range.

5. The method according to claim 1, characterized in that The determining the second variant detection result corresponding to each gene locus in the gene fragment to be tested according to the quality assessment result includes: Dividing the variant loci into first candidate variant loci with a quality value higher than a first threshold and second candidate variant loci with a quality value not higher than the first threshold according to the evaluated quality value; Screening the second candidate variant loci based on a second preset neural network model, and removing the second candidate variant loci determined by the model as non-variant; Determining the second variant detection result corresponding to each gene locus in the gene fragment to be tested according to the first candidate variant loci and the screened second candidate variant loci.

6. The method according to claim 5, characterized in that, The performing quality assessment on the variant loci determined based on the first sequencing result alignment data based on a preset recurrent neural network model includes: Determining a plurality of fourth candidate variant loci according to the first sequencing result alignment data; Obtaining the sequencing base proportion information of the fourth candidate variant loci and their associated loci, and generating a first evaluation feature for each of the fourth candidate variant loci based on the sequencing base proportion information; Invoking a preset recurrent neural network model to perform variant quality assessment on the first evaluation feature, and determining the quality value corresponding to each of the fourth candidate variant loci according to the assessment result.

7. The method according to claim 6, wherein It further includes: Determining the first threshold based on a preset proportion quantile value of the quality values corresponding to all the fourth candidate variant loci.

8. The method according to claim 5, wherein The screening the second candidate variant loci based on a second preset neural network model, and removing the second candidate variant loci determined by the model as non-variant includes: For each of the second candidate variant loci, obtaining variant assessment data within a preset assessment region of each of the second candidate variant loci, and determining a second evaluation feature for each of the second candidate variant loci according to the variant assessment data; wherein the preset assessment region includes the second candidate variant locus and its upstream and downstream regions; Invoking the second preset neural network model to perform variant identification on the second evaluation feature, and determining the second candidate variant loci determined by the model as variant loci as third candidate variant loci.

9. The method according to claim 8, wherein The second evaluation feature includes mutation assessment information, base quality information, and genotyping information.

10. A gene mutation detection device, characterized in that, The gene mutation detection device includes: A comparison unit, configured to compare the long-read sequencing data of the gene fragment to be detected with a reference genome to obtain first sequencing result comparison data, and compare the short-read sequencing data of the gene fragment to be detected with the reference genome to obtain second sequencing result comparison data; A first determination unit, configured to determine a second mutation detection result corresponding to each gene locus in the gene fragment to be detected based on the first sequencing result comparison data, and determine a first mutation detection result corresponding to each gene locus in the gene fragment to be detected based on the second sequencing result comparison data; An identification unit, configured to identify a first target gene locus where the first mutation detection result is different from the second mutation detection result, and when the first target gene locus is in a genomic complex region of the gene fragment to be detected, determine a third mutation detection result of the first target gene locus according to the second mutation detection result and the mutation detection parameters of each gene locus; A second determination unit, configured to determine a third mutation detection result of the first target gene locus according to the first mutation detection result and the mutation detection parameters when the first target gene locus is not in the genomic complex region; A third determination unit, configured to determine a plurality of first mutation sites with mutations based on the third mutation detection result, and determine a target gene mutation detection result of the gene fragment to be detected according to the first mutation site and the second mutation site, where the second mutation site is a mutation site that intersects in the first mutation detection result and the second mutation detection result.

11. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the gene mutation detection method according to any one of claims 1 to 9.

12. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the gene mutation detection method according to any one of claims 1 to 9.

13. A computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the gene mutation detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Structural genetic variation detection method and device and computer equipment

    CN118782139A

  • SNP and INDEL detection method based on deep learning and long-reading sequencing

    CN119028431A