Genetic variation detection method and apparatus, storage medium, and computer device
By introducing a genotype filling method with high-quality reference haplotype set and linkage imbalance principle, combined with deep learning models, the accuracy of nanopore sequencing technology in gene variation detection, especially the detection accuracy of insertion/deletion sites is solved, and the problem of insufficient accuracy of sequencing results in nanopore sequencing technology is solved.
Patent Information
- Application Number
- PCT/CN2023/143619
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
Nanopore sequencing technology has the problem of insufficient accuracy of sequencing results in gene variant detection, which leads to low accuracy of gene variant detection, especially poor performance in the detection of insertion/deletion sites.
By introducing a high-quality reference haplotype set, the haplotype of the sample is inferred based on high-quality SNP sites, and the insertion or deletion sites of the sample are determined based on the inferred haplotype. The principle of linkage imbalance is used to perform genotype filling of low-quality sequencing sites, and genetic variation detection is carried out in combination with deep learning models.
The accuracy of nanopore sequencing technology in gene variant detection, especially the detection accuracy of insertion/deletion sites, while maintaining detection efficiency.
Smart Images

Figure CN2023143619_03072025_PF_FP_ABST
Abstract
Description
Gene mutation detection method, device, storage medium and computer equipment Technical Field
[0001] The present disclosure relates to the field of gene detection technology, and in particular to a gene variation detection method, apparatus, storage medium, and computer equipment. Background Art
[0002] Genetic variation detection involves sequencing and differential analysis of the genes of a species or population using sequencing technology to obtain a large amount of genetic variation information, including single nucleotide polymorphisms (SNPs), insertions, deletions, structural variants (SVs), and copy number variations (CNVs) within the genome. This genetic variation information is of great significance and utility for understanding inherited diseases, drug responses, gene expression regulation, and genetic differences between individuals. Therefore, genetic variation detection is of great significance to species health and disease treatment.
[0003] Nanopore sequencing (NS) is a novel single-molecule DNA sequencing technology that uses changes in electrical current within a protein nanopore to identify DNA base sequences. Compared to traditional second-generation sequencing technologies, which are limited by read length, nanopore sequencing boasts longer read lengths, addressing the poor performance of second-generation sequencing technologies in detecting variants in complex genomic regions. However, in related technologies, nanopore sequencing suffers from insufficient sequencing accuracy, resulting in low accuracy for detecting genetic variants based on nanopore sequencing results.
[0004] Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to provide a method, apparatus, storage medium and computer equipment for detecting gene mutations, which can improve the accuracy of gene mutation detection.
[0006] To achieve the above objectives, a first aspect of the embodiments of the present application provides a method for detecting gene mutations, the method comprising:
[0007] Obtaining base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested;
[0008] performing gene variation detection on the sequence alignment result data based on a first gene variation detection model to obtain a plurality of single-base variation sites, wherein the first gene variation detection model is a model for performing gene variation detection based on gene stacking;
[0009] Obtaining quality scores corresponding to the multiple single-base variant sites, and determining the single-base variant sites having quality scores greater than a preset threshold as first-category variant sites;
[0010] A target haplotype corresponding to the gene segment to be tested is determined in a reference haplotype set based on the first type of variation site, and an insertion or deletion site of the gene segment to be tested is determined based on the target haplotype.
[0011] A second aspect of the present application provides a gene variation detection device, comprising:
[0012] A first acquisition unit is used to obtain base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested;
[0013] a detection unit, configured to perform gene variation detection on the sequence alignment result data based on a first gene variation detection model to obtain a plurality of single-base variation sites, wherein the first gene variation detection model is a model for performing gene variation detection based on gene stacking;
[0014] a second acquiring unit, configured to acquire quality scores corresponding to the plurality of single-base variant sites, and determine the single-base variant sites having quality scores greater than a preset threshold as first-category variant sites;
[0015] A determination unit is used to determine a target haplotype corresponding to the gene segment to be tested in a reference haplotype set based on the first type of variation site, and to determine an insertion or deletion site of the gene segment to be tested based on the target haplotype.
[0016] In one embodiment, the determining unit includes:
[0017] A first calculation subunit is used to calculate the genotype likelihood values corresponding to the first type of variant sites and multiple alleles;
[0018] The first determining subunit is configured to determine a target haplotype corresponding to the gene segment to be tested in a reference haplotype set based on the genotype likelihood value.
[0019] In one embodiment, the first determining subunit is further configured to:
[0020] The iterative updating step is executed cyclically until the number of cycles reaches a preset number, and the haplotype obtained by the last update is output as the target haplotype of the gene segment to be tested. The iterative updating step comprises the following steps:
[0021] Determining a candidate haplotype subset in a reference haplotype set based on the first type of variant sites;
[0022] Calculating the probability of occurrence of the base corresponding to each first variant site based on the genotype likelihood value and the candidate haplotype subset;
[0023] Inferring the individual genotype of each single-base variation site using a hidden Markov model based on the probability and the distance between the single-base variation sites in the candidate haplotype subset;
[0024] The haplotype of the gene fragment to be tested is updated based on the combination of individual genotypes of each single-base variation site.
[0025] In one embodiment, the second acquiring unit includes:
[0026] a detection subunit, configured to detect potential variant sites on the base sequence alignment result data, and obtain a plurality of potential variant sites and a quality score corresponding to each potential variant site;
[0027] A second determining subunit is configured to determine quality scores corresponding to the multiple single-base variation sites based on the matching relationships between the multiple single-base variation sites and the multiple potential variation sites;
[0028] The third determining subunit is configured to determine the single-base variation site having a quality score greater than a preset threshold as a first-category variation site.
[0029] In one embodiment, the gene variation detection device provided by the present application further includes:
[0030] A receiving subunit, configured to receive an input quality value to control the ratio;
[0031] The second calculation subunit is configured to calculate a preset threshold based on the quality value control ratio and the quality score corresponding to each potential variation site.
[0032] In one embodiment, the gene variation detection device provided by the present application further includes:
[0033] a fourth determining subunit, configured to determine the single-base variant site having a quality score not greater than the preset threshold as a second-category variant site;
[0034] a review subunit, configured to perform single-base variation review on the second-category variant sites based on a second gene variation detection model, and determine sites with confirmed variations as third-category variant sites based on the review results, wherein the second gene variation detection model is a model for gene variation detection based on full alignment;
[0035] The fifth determining subunit is used to determine the target variation detection result of the gene fragment to be tested based on the first type of variation site, the third type of variation site and the insertion or deletion site.
[0036] In one embodiment, the fifth determining subunit includes:
[0037] A screening module, configured to perform quality screening on the first type of variant sites and the third type of variant sites to obtain target single-base variant sites;
[0038] The first determination module is used to determine the target variation detection result of the gene fragment to be tested according to the target single base variation site and the insertion or deletion site.
[0039] In one embodiment, the screening module includes:
[0040] an acquisition submodule, configured to acquire mutation frequencies corresponding to the first type of variation sites and the third type of variation sites;
[0041] The determination submodule is used to determine that the mutation sites of the first type of mutation sites and the third type of mutation sites whose mutation frequencies are greater than a preset frequency threshold are target single-base mutation sites.
[0042] In one embodiment, the review subunit includes:
[0043] An acquisition module, used to obtain the base sequence features corresponding to each second-category variation site;
[0044] an identification module, configured to perform gene variation identification on each of the base sequence features based on a second gene variation detection model to obtain a gene variation identification result;
[0045] The second determining module is configured to determine a plurality of third-category variation sites from the plurality of second-category variation sites according to the gene variation identification result.
[0046] In one embodiment, the first acquiring unit includes:
[0047] An acquisition subunit, used to obtain nanopore sequencing data of the gene fragment to be tested;
[0048] The alignment subunit is used to compare the nanopore sequencing data to the reference genome to obtain base sequence alignment result data.
[0049] A third aspect of an embodiment of the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the gene variation detection method described in the first aspect is implemented.
[0050] To achieve the above-mentioned objectives, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the gene variation detection method described in the first aspect.
[0051] To achieve the above-mentioned objectives, the fifth aspect of the embodiments of the present application proposes a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the gene variation detection method described in the first aspect.
[0052] The embodiments of the present application propose a method, apparatus, storage medium, and computer equipment for detecting gene variation. The method obtains base sequence alignment result data corresponding to sequencing data of a gene fragment to be tested; performs gene variation detection on the sequence alignment result data based on a first gene variation detection model to obtain multiple single-base variation sites, where the first gene variation detection model is a model for performing gene variation detection based on gene stacking; obtains quality scores corresponding to the multiple single-base variation sites, and determines the single-base variation sites with a quality score greater than a preset threshold as a first-class variation site; determines a target haplotype corresponding to the gene fragment to be tested in a reference haplotype set based on the first-class variation site, and determines the insertion or deletion site of the gene fragment to be tested based on the target haplotype.
[0053] In this way, this method determines the accurate target haplotype of the gene fragment to be tested in the reference haplotype set through the high-quality single-base variation sites measured by the model for gene variation detection based on gene stacking, and then determines the insertion / deletion site of the gene fragment to be tested based on the insertion / deletion site of the target haplotype. In this way, the accuracy of insertion / deletion site detection can be improved when gene variation detection is performed based on sequencing results when the sequencing results are inaccurate, thereby improving the accuracy of gene variation detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0055] FIG1 is a flow chart of a method for detecting gene variation provided in an embodiment of the present application;
[0056] FIG2 is a flow chart of step 101 in FIG1 ;
[0057] FIG3 is a flow chart of step 103 in FIG1 ;
[0058] FIG4 is a schematic diagram of a supplementary process of the gene variation detection method provided in an embodiment of the present application;
[0059] FIG5 is a flow chart of step 104 in FIG1 ;
[0060] FIG6 is a flowchart of the burn-in iteration process;
[0061] FIG7 is a schematic diagram of the overall process of haplotype identification of the gene fragment to be tested using high-quality SNP sites and a reference haplotype set;
[0062] FIG8 is another supplementary flow diagram of the gene variation detection method provided in an embodiment of the present application;
[0063] FIG9 is a flow chart of step 803 in FIG8 ;
[0064] FIG10 is a flow chart of step 901 in FIG9 ;
[0065] FIG11 is a flow chart of step 802 in FIG8 ;
[0066] FIG12 is a schematic diagram of the overall process of the gene variation detection method provided by the present disclosure;
[0067] FIG13 is a comparison of the Indel mutation detection effects of the disclosed solution and related solutions;
[0068] FIG14 is a structural diagram of a gene variation detection device provided in an embodiment of the present application;
[0069] FIG15 is a schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0071] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:
[0072] Haplotype: refers to the combination pattern of alleles at multiple loci of an individual, that is, the arrangement of alleles at a series of connected loci.
[0073] Genotype: The collective term for the entire genetic makeup of an organism. It reflects the organism's genetic makeup, the sum of all genes inherited from both parents. In genetics, genotype often refers to the genotype of a particular trait. Two organisms differ in genotype if they have just one. Therefore, genotype refers to all combinations of alleles at all loci for an individual.
[0074] Alleles: Alleles are genes located at the same position on a pair of homologous chromosomes that control different forms of the same trait. If the two alleles of a gene (or locus) on homologous chromosomes are identical, the individual is homozygous for the trait. If the alleles are different, the individual is heterozygous for the trait.
[0075] Read length: Due to limitations in sequencing technology, genome sequencing requires fragmenting the genome into DNA fragments before library construction and sequencing. Read length refers to the base sequence obtained by a single sequencing pass on a sequencer. Different sequencing instruments produce varying read lengths. Sequencing an entire genome can generate hundreds of millions of reads.
[0076] Due to the high error rate of nanopore sequencing technology, traditional statistical methods (such as pileup and freebayes) are no longer sufficient for variant detection using long-read sequencing data. With the rapid development of deep learning (DL) methods, identifying genomic variant sites based on deep learning models has become the mainstream method for variant detection (long-read variant calling, LRVC) using long-read sequencing sequences.
[0077] Currently, there are generally two methods for long-read sequencing variant detection based on deep learning:
[0078] The first approach uses base stacking as the input to the neural network. This approach works by aligning and stacking multiple DNA sequencing reads at specific locations to form a sequence stacking graph. The base at each position is represented by its height and frequency in the stacking graph, capturing the variation and heterozygosity between individuals and identifying variant sites. This method can efficiently and accurately identify single-base variant sites (SNPs), but when the sequencing data is inaccurate, the accuracy of identifying insertion / deletion sites (Indels) is low.
[0079] The second method uses full alignment as neural network input. This method fully aligns DNA sequencing reads with the reference genome to form a global alignment. This allows direct access to the base and variant status at each position, accurately detecting single-base variants. This method has high accuracy for detecting both SNPs and indels, but it is inefficient.
[0080] Based on this, the present application embodiment provides a kind of gene variation detection method, device, storage medium and computer equipment.This method is by introducing high-quality reference haplotype set as the supplement of long read sequencing data, infers the haplotype of sample based on high-quality reference haplotype set and high-quality SNP site, then determines the accurate Indel site of sample according to the inferred haplotype.Below, the gene variation detection method provided by the present application embodiment is described.
[0081] 1 , in some embodiments, the gene variation detection method provided in the embodiments of the present application includes but is not limited to steps 101 to 104 .
[0082] Step 101, obtaining base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested;
[0083] Step 102: performing gene variation detection on the sequence alignment result data based on a first gene variation detection model to obtain multiple single-base variation sites;
[0084] Step 103: Obtain quality scores corresponding to multiple single-base variant sites, and determine single-base variant sites with quality scores greater than a preset threshold as first-category variant sites;
[0085] Step 104 : determining a target haplotype corresponding to the gene segment to be tested in the reference haplotype set based on the first type of variant sites, and determining the insertion or deletion site of the gene segment to be tested based on the target haplotype.
[0086] The step 101 to step 104 shown in the embodiment of the present application, when gene variation detection is carried out to gene fragment to be tested, the comparison result data of the base sequence corresponding to the sequencing data of gene fragment to be tested can be first obtained.Then, the model (for being distinguished with other gene variation detection models, referred to herein as the first gene variation detection model) for carrying out gene variation detection based on gene accumulation is adopted to carry out gene variation detection to the base sequence comparison result data of gene fragment to be tested.This model can measure SNP site and Indel site in gene fragment to be tested, the data of Indel site measured by this model can be discarded, the SNP site measured by this model, the quality score corresponding to each SNP site can be obtained, then according to quality score and preset threshold, determine that the SNP site that quality score is greater than this preset threshold is high-quality SNP site (for being distinguished with other SNP sites, referred to herein as the first class variant site).Further, high-quality reference haplotype set can be obtained, then according to above-mentioned high-quality SNP site and linkage disequilibrium principle, determine the target haplotype corresponding to gene fragment to be tested in reference haplotype set. In this way, the Indel site in the target haplotype can be determined as the Indel site (ie, insertion or deletion site) of the gene fragment to be tested.
[0087] This method can be called a genotype imputation method. Based on known high-quality SNP sites and known high-quality reference genotypes, this method uses the linkage disequilibrium (LD) principle to impute the genotypes of sites with poor sequencing data in the gene fragment to be tested or sites that are inaccurately identified due to background noise, thereby determining the target haplotype corresponding to the gene fragment to be tested, and then determining the Indel site in the gene fragment to be tested based on the target haplotype, thereby greatly improving the accuracy of detecting Indel sites in the gene fragment to be tested.
[0088] In the disclosed embodiments, the gene fragment to be tested can be a complete DNA fragment or RNA fragment, or a sub-fragment extracted from a DNA fragment or RNA fragment. The DNA fragment or RNA fragment herein can be a DNA or RNA fragment of any species. The sequencing data of the gene fragment to be tested can be any type of sequencing data, that is, sequencing data obtained using any sequencing method, for example, first-generation sequencing methods (such as the Sanger method, also known as the dideoxy chain termination method), second-generation sequencing methods (such as the Illumina method, also known as massively parallel sequencing), and third-generation sequencing methods (such as nanopore sequencing or PacBio sequencing). In this embodiment, the first gene variation detection model can specifically be the model described above that uses base stacking as a neural network input. The principle of linkage disequilibrium specifically refers to the non-random association between two or more loci, which may be on the same chromosome or on different chromosomes. As long as the probability of two loci appearing at the same time is greater than the probability of random combination of the population, it means that the two loci are in a state of linkage disequilibrium. Linkage disequilibrium is caused by mutation or recombination. Linkage disequilibrium occurs when a new mutation occurs near a certain SNP on a chromosome. The strength of linkage disequilibrium is related to the distance between the two SNPs. The smaller the distance, the smaller the chance of recombination, and the stronger the linkage disequilibrium, that is, the higher the probability of these two SNPs appearing simultaneously in the genome. Based on this principle, the identified high-quality SNP sites can be used to search for sequences similar to the target genotype in the existing high-quality haplotype collection, and use this sequence to fill the low-quality sequencing segment of the target sequence, and thereby obtain a reliable set of genomic variations. Among them, the above-mentioned reference haplotype collection can be a reference haplotype collection maintained by the 1000 Genomes Project, obtained from the Internet. This reference haplotype collection is a high-quality haplotype collection.
[0089] Referring to FIG. 2 , in some embodiments, step 101 includes but is not limited to steps 201 to 202 .
[0090] Step 201: Obtain nanopore sequencing data of the gene fragment to be tested.
[0091] Step 202 : Compare the nanopore sequencing data to the reference genome to obtain base sequence alignment result data.
[0092] In an embodiment of the present application, the gene fragment to be tested can be a human whole genome DNA standard (such as HG002). When the standard is subjected to variation detection, the standard can first be subjected to three generations of nanopore sequencing to obtain the nanopore sequencing data of the standard, and the nanopore sequencing data can specifically be a DNA base sequence. In three generations of nanopore long read sequencing, the error rate of sequencing is high, resulting in poor performance of gene variation detection, especially the detection of Indel mutation sites when gene variation detection is performed based on nanopore sequencing results. Therefore, if only nanopore sequencing data is relied upon for variation detection, in sections such as homopolymers and heteropolymers with poor quality in nanopore sequencing, due to the large sequencing background noise, it is difficult to obtain good variation detection results in weak areas of nanopore sequencing using existing variation detection tools based on deep learning methods. Therefore, the genotype filling method provided by the present disclosure is applied in the scene of gene detection by nanopore sequencing, and can have a more obvious effect.
[0093] Then, the DNA base sequence obtained by nanopore sequencing can be aligned to the reference genome to obtain a comparison result file in BAM format. Among them, the reference genome is a special genome that is representative and can be used to compare with the genomes of other organisms. The reference genome can be generated by various methods (such as sequencing, assembly, etc.) and covers the entire genome, including all genes and elements. The reference genome is often used to identify and understand the function of genes, and can also be used to compare genomic differences between multiple species. In the disclosed embodiment, when the gene fragment to be tested is a human whole genome DNA standard, the reference genome can specifically be a human reference genome. BAM is called Binary Alignment / Map format, that is, binary alignment map format, which is a binary representation of the sequence alignment map format (Sequence Alignment / Map format, SAM), and is a commonly used file for storing sequencing data and reference genome alignment results.
[0094] Referring to FIG. 3 , in some embodiments, step 103 includes but is not limited to steps 301 to 303 .
[0095] Step 301: Detect potential variant sites on the base sequence alignment result data to obtain multiple potential variant sites and a quality score corresponding to each potential variant site;
[0096] Step 302: determining quality scores corresponding to the multiple single-base variant sites based on the matching relationships between the multiple single-base variant sites and the multiple potential variant sites;
[0097] Step 303: Determine the single-base variant sites with a quality score greater than a preset threshold as first-category variant sites.
[0098] In the disclosed embodiment, the quality value of the first type of variant site can be obtained when the potential variant site detection is performed on the base sequence comparison result data corresponding to the sequencing data. Specifically, after the gene fragment to be tested is sequenced to obtain sequencing data, and the sequencing data is compared to the reference genome to obtain the comparison result data in the BAM format, the pileup function in the bioinformatics tool (such as Samtools) can be used to detect potential variant sites on the base sequence comparison result data. Among them, the Samtools software is a tool software for processing SAM and BAM formats, which can realize binary viewing, format conversion, sorting and merging and other functions. In combination with the flag, tag and other information in the SAM format, it can also complete the statistical summary of the comparison results. Pileup is a basic command in Samtools, and its function is to perform base stacking for each site in the genome, which is used to recall SNP mutations or Indel mutation sites.
[0099] In addition to outputting potential mutation sites, the quality score of each SNP site can also be output by using Samtools to compare the result data for mutation site detection. These potential mutation sites can be used as target sites, combined with the base composition and ratio (composition ratio matrix) of the sites before and after them (for example, the sites 16bp before and after) and the insertion / deletion characteristics, as the input of the above-mentioned first gene variation detection model to further determine whether the target site actually has SNP variation or Indel variation. Among them, in the embodiment of the present disclosure, the first gene variation detection model can specifically be a two-layer long short-term memory network model (Bi-directional Long Short-Term Memory, BiLSTM) that has been trained.
[0100] Then, for each SNP site output by the first gene variation detection model, the corresponding matching site can be found in the SNP sites output by Samtools. In this way, the quality score of each SNP site output by the first gene variation detection model can be obtained from the quality score of each variation site output by Samtools.
[0101] Furthermore, based on the quality score of each SNP site output by the first genetic variation detection model and a preset threshold, high-quality SNP sites, i.e., the aforementioned first-category variant sites, can be identified from the SNP sites output by the first genetic variation detection model. Specifically, SNP sites output by the first genetic variation detection model with a quality score greater than the aforementioned preset threshold can be identified as first-category variant sites.
[0102] Referring to FIG. 4 , in some embodiments, before step 303 , the following steps 401 to 402 may also be included but not limited to.
[0103] Step 401, receiving an input quality value control ratio;
[0104] Step 402 : A preset threshold is calculated based on the quality value control ratio and the quality score corresponding to each potential variant site.
[0105] In the embodiment of the present disclosure, the preset threshold value can be adjusted according to the manually input quality value control ratio. For example, when the SNP sites in the top 70% of the quality value are manually determined as high-quality sites, the quality values of all SNP sites output by the first variation detection model can be counted, and then the quality control line of the top 70% of the quality value can be calculated to obtain the preset threshold value. Alternatively, if the manually input quality value control ratio is 40%, the quality values of all SNP sites output by the first variation detection model can be counted, and then the quality control line of the top 40% of the quality value can be calculated to obtain the preset threshold value. The manually input quality value control ratio here can be obtained based on experience.
[0106] Referring to FIG. 5 , in some embodiments, step 104 includes but is not limited to steps 501 and 502 .
[0107] Step 501, calculating the genotype likelihood values corresponding to the first type of variant sites and multiple alleles;
[0108] Step 502 : determining a target haplotype corresponding to the gene segment to be tested in the reference haplotype set based on the genotype likelihood value.
[0109] In the disclosed embodiments, when determining the target haplotype corresponding to the gene fragment to be tested in the reference haplotype set based on the first type of variant site, the genotype likelihoods corresponding to the first variant site and multiple alleles can be calculated. Specifically, the first type of variant site (i.e., the high-quality SNP site output by the first gene variation detection model) can be used, and based on the sequencing bases of the sequencing read length, the sequencing quality, and the comparison quality, etc., the genotype likelihood values for various alleles can be calculated, and the genotype likelihood values include the haplotype information of the sequence.
[0110] Further, can determine the target haplotype corresponding to gene fragment to be tested based on genotype likelihood value in reference haplotype set.Wherein, because linkage disequilibrium may take place in gene on same chromosome, the allele on a plurality of sites often can exist simultaneously, and row forms specific haplotype.Therefore, can determine the target haplotype corresponding to gene fragment to be tested based on existing high-quality reference haplotype set, in conjunction with the genotype likelihood value of the above-mentioned variable site that calculates.
[0111] In certain embodiments, in reference haplotype set, determine the process of the target haplotype corresponding to gene fragment to be tested based on genotype likelihood value, specifically can be according to existing high-quality reference haplotype set, carry out the Burn-in iterative process of Markov Chain Monte Carlo algorithm in conjunction with the genotype likelihood value of the variable site of above-mentioned calculation.Particularly, can set the number of times of iterative process, after the number of times of iterative process reaches preset number of times, just can export the target haplotype corresponding to gene fragment to be tested.Referring to Fig. 6, above-mentioned iterative process process specifically can comprise step 601 to step 604.
[0112] Step 601, determining a candidate haplotype subset in a reference haplotype set based on the first type of variant sites;
[0113] Step 602, calculating the probability of occurrence of the base corresponding to each first variant site based on the genotype likelihood value and the candidate haplotype subset;
[0114] Step 603, using a hidden Markov model to infer the individual genotype of each single-base variant site based on the probability and the distance between the single-base variant sites in the candidate haplotype subset;
[0115] Step 604 : updating the haplotype of the gene fragment to be tested based on the combination of individual genotypes at each single-base variant site.
[0116] In the disclosed embodiments, the target genotype corresponding to gene fragment to be tested is determined in high-quality reference genotype set by the Burn-in iteration of multiple rounds. Specifically, the haplotype corresponding to gene fragment to be tested can be found out in the reference haplotype set based on high-quality SNP sites, constituting a candidate haplotype subset. Specifically, the PBWT algorithm can be adopted to determine a plurality of haplotypes most similar to the sequencing sequence of gene fragment to be tested based on high-quality SNP sites in the reference genotype set, constituting a candidate haplotype subset. Then, the probability of occurrence of the base corresponding to each first variable site can be further calculated according to genotype likelihood value and candidate haplotype subset, and this probability can be used as the emission probability of a hidden Markov model. Then, just can utilize hidden Markov model, specifically can be Li and Stephens hidden Markov model (Li and Stephens HMM), in conjunction with the distance between the SNP sites in aforementioned emission probability and candidate haplotype subset, infer the individual genotype of each SNP site in gene fragment to be tested. Further, the different combinations of the individual genotypes based on each SNP site can be used to update the haplotype of gene fragment to be tested.
[0117] In this way, after multiple rounds of updating, the target haplotype of the gene fragment to be tested can be obtained, and the accurate genotype of the gene fragment to be tested can also be obtained.
[0118] Please refer to Figure 7 for a schematic diagram of the overall process for haplotype identification of the target gene segment using high-quality SNP sites and a reference haplotype set. Specifically, the genotype likelihood value of the gene variant site is first calculated based on the long-read sequencing data and the existing high-quality reference haplotype set. The occurrence probability of the corresponding base is then calculated based on the genotype likelihood value of the gene variant site. This is then used as the emission probability input into the Lee and Stephens Hidden Markov Model. Multiple rounds of iterations are then used to determine the target haplotype corresponding to the target gene segment based on the reference haplotype set.
[0119] In some embodiments, after obtaining the accurate Indel site of the gene segment to be tested, a gene variation detection method is also provided that can obtain the complete gene variation site (i.e., including SNP sites and Indel sites) of the gene segment to be tested. Specifically, referring to Figure 8, after step 104, steps 801 to 803 may also be included.
[0120] Step 801: Determine a single-base variant site with a quality score not greater than a preset threshold as a second-category variant site;
[0121] Step 802: Perform single-base variation verification on the second-category variant sites based on the second gene variation detection model, and determine the sites with confirmed mutations as third-category variant sites.
[0122] Step 803 : determining the target variation detection result of the gene fragment to be tested based on the first category variation sites, the third category variation sites, and the insertion or deletion sites.
[0123] In the embodiment of the present disclosure, among the SNP sites and Indel sites obtained based on the first gene variation detection model, the Indel sites can be directly eliminated due to their low quality. Then, the genotype is filled in according to the high-quality SNPs in the SNP to obtain the target haplotype, and the accurate Indel sites of the gene fragment to be tested are determined according to the target haplotype. As for the SNP sites output by the first gene variation detection model, the high-quality SNP sites therein (i.e., the first type of variation sites) can be directly used as part of the SNP sites of the gene fragment to be tested due to their high quality. For the low-quality SNP sites (herein referred to as the second type of variation sites) among the SNP sites output by the first gene variation detection model, the embodiment of the present disclosure can use the second gene variation detection model to perform variation review on them. If the second variation detection model confirms that it is a SNP site, it is confirmed to be a third type of variation site, and then it can be used together with the aforementioned high-quality SNP sites as the final SNP sites of the gene fragment to be tested.
[0124] Specifically, single-base variant sites with a quality score no greater than a preset threshold can be identified as second-category variant sites. Then, based on the second genetic variation detection model, single-base variants of the second-category variant sites are reviewed for single-base variants, and sites with confirmed variants are identified as third-category variant sites. Third-category variant sites are low-quality SNP sites output by the first variation detection model that have been further reviewed as high-quality by the second variation detection model. Furthermore, the target variant detection results corresponding to the gene fragment to be tested can be determined based on the first-category variant sites, the third-category variant sites, and the insertion or deletion sites.
[0125] Among them, in the embodiment of the present disclosure, the second gene variation detection model can be the aforementioned gene variation detection model based on full alignment as the neural network input. Specifically, the second gene variation detection model can be a trained residual neural network model (ResNet). When the second type of variation site is subjected to single-base variation review based on the second gene variation detection model, the genotype typing, base quality, alignment quality, and different dimensional features such as positive and negative chains of each second type of variation site and its preceding and following base sequences can be specifically determined first as the input of the above-mentioned residual neural network model, so as to further screen the second type of variation site based on the above-mentioned residual neural network model to determine whether the second type of variation site is a site where a variation actually exists. When the second gene variation detection model detects and confirms that any second type of variation site is a truly credible variation site, the credible variation site can be determined as a third type of variation site. In this way, the union of the first type of variation site and the third type of variation site is the final SNP site of the gene fragment to be tested, and the Indel site determined by the target haplotype determined by genotype filling is the final Indel site of the gene fragment to be tested. Then the final SNP site and the final Indel site together constitute the final target variation detection result of the gene fragment to be tested.
[0126] Using the gene variation detection method provided in the embodiments of the present disclosure to perform gene variation detection on nanopore sequencing data can not only improve the accuracy of detecting Indel variation sites, but also greatly improve the efficiency of overall gene variation detection of the gene fragments to be tested while ensuring the efficiency of gene variation detection.
[0127] Referring to FIG. 9 , in some embodiments, step 803 includes but is not limited to steps 901 to 902 .
[0128] Step 901: Perform quality screening on the first and third category variant sites to obtain target single-base variant sites.
[0129] Step 902 : determining the target variation detection result of the gene fragment to be tested based on the target single base variation site and the insertion or deletion site.
[0130] In the embodiment of the present disclosure, when determining the final SNP site of the gene fragment to be tested based on the high-quality SNP sites (first type of variation sites) output by the first gene variation detection model and the credible SNP sites (third type of variation sites) obtained by screening the low-quality SNP sites output by the first gene variation detection model according to the second gene variation detection model, the first type of variation sites and the third type of variation sites can be quality-screened again to filter out some low-quality SNP sites. Specifically, the first type of variation sites and the third type of variation sites can be quality-screened to obtain the target single-base variation sites. Then, the target variation detection result of the gene fragment to be tested can be determined based on the target single-base variation site and the insertion or deletion site.
[0131] Referring to FIG. 10 , in some embodiments, step 901 includes but is not limited to steps 1001 to 1002 .
[0132] Step 1001, obtaining the mutation frequencies corresponding to the first type of variant sites and the third type of variant sites;
[0133] Step 1002 : Determine the mutation sites in the first category and the third category whose mutation frequencies are greater than a preset frequency threshold as target single-base mutation sites.
[0134] In the disclosed embodiments, the mutation frequencies of variant sites can be used to screen for first-category variant sites and third-category variant sites. Specifically, the mutation frequencies corresponding to the first-category variant sites and the third-category variant sites can be obtained, and then the variant sites with mutation frequencies greater than a preset frequency threshold among the first-category variant sites and the third-category variant sites are designated as target single-base variant sites.
[0135] Among them, mutation frequency or variant frequency (allele frequency, AF) refers to the frequency of a specific allele in a species' mutagenic population, or the ratio of all alleles. This frequency reflects the frequency of occurrence of an allele in a specific population. For high-frequency variant sites, the AF value is relatively high, which means that the frequency of the allele in the population is high. The AF value of a low-frequency variant site is relatively low, indicating that the frequency of the allele in the population is low. In genetic studies, the AF value is usually calculated by comparing the genotype frequencies between different individuals.
[0136] After obtaining the mutation frequencies corresponding to each first-category variant site and third-category variant site, the first-category variant sites and second-category variant sites can be screened based on a preset frequency threshold. Specifically, first-category variant sites and second-category variant sites with mutation frequencies greater than the preset frequency threshold can be identified as target single-base variant sites. For example, the frequency threshold can be set to 0.1, and first-category variant sites and second-category variant sites with mutation frequencies greater than 0.1 can be identified as target single-base variant sites.
[0137] Referring to FIG. 11 , in some embodiments, step 802 may include but is not limited to steps 1101 to 1103 .
[0138] Step 1101, obtaining the base sequence features corresponding to each second-category variant site;
[0139] Step 1102: performing gene variation identification on each base sequence feature based on the second gene variation detection model to obtain a gene variation identification result;
[0140] Step 1103 : determining a plurality of third-category variation sites from the plurality of second-category variation sites based on the gene variation identification results.
[0141] In the disclosed embodiments, the specific process of performing single-base variation verification on the second-class variant sites based on the second gene variation detection model can obtain the base sequence features corresponding to each second-class variant site. The base sequence features include features of different dimensions, such as genotyping, base quality, alignment quality, and positive and negative strands, for the second-class variant site and its surrounding base sequences.
[0142] The base sequence features corresponding to each second-category variant site can then be used as input to a second genetic variation detection model, resulting in a genetic variation identification result output by the second genetic variation detection model. Furthermore, multiple third-category variant sites can be identified from the multiple second-category variant sites based on the genetic variation identification results output by the second genetic variation detection model.
[0143] Please refer to Figure 12, which is a schematic diagram of the overall process of the genetic variation detection method provided by the present disclosure. As shown in the figure, the present embodiment aligns the long-read sequencing sequence obtained by nanopore sequencing of the gene fragment to be tested to a reference genome to obtain a BAM format alignment result file. Then, a stacking-based neural network model is used to perform genetic variation detection on the BAM format alignment result file to obtain genetic variation detection results. The genetic variation detection results output by the stacking-based neural network model are then subjected to variation quality screening, and SNP sites are classified into high-quality variation sites and low-quality variation sites based on the quality value of the SNP sites. Indel sites in the genetic variation detection results output by the stacking-based neural network model can be eliminated. Then, for high-quality variation sites, a genotype imputation method based on haplotype information can be further used to obtain more accurate indel sites; for low-quality variation sites, a full-alignment-based neural network model can be used to screen out high-quality variation sites. Finally, the high-quality variation sites output by the stacking-based neural network model and the full-alignment-based neural network model, as well as the indel sites obtained by genotype imputation, are used as the final variation detection results. This method can obtain more accurate genetic variation detection results based on higher genetic variation detection efficiency.
[0144] Based on a deep learning model, the disclosed embodiments propose a scheme for genotyping low-quality sequencing sites based on the principle of linkage disequilibrium. This method uses SNP sites with good long-read sequence performance to genotype-impaired low-quality sequencing sites or segments in long-read sequencing, inferring the haplotype sequence of the sequenced sample and obtaining potential mutation sites. This method can significantly improve the accuracy and sensitivity of indel identification in long-read sequencing even at low depth. The method is applicable to the analysis of both single samples and population samples.
[0145] The following is a detailed description of the gene variation detection method provided by the present application using a specific embodiment.
[0146] 1. First, perform nanopore long-read sequencing on the HG002 standard to obtain a long-read sequencing sequence. Then, use the alignment software minimap2 (v2.24) to align the long-read sequence to the GRch38 genome and obtain the alignment result file in BAM format.
[0147] 2. A two-layer bidirectional long short-term memory network (BiLSTM) model trained on a standard dataset was used to perform preliminary mutation detection on target sites in long-read sequences to obtain potential mutation sites in the entire genome of the HG002 sample.
[0148] 3. Using the top 80% of homozygous and heterozygous SNP sites by mutation site quality score (qual) output by the BiLSTM model as high-quality SNP sites, their genotypes are confirmed using the selected high-quality sites and long-read sequences. The haplotype sequences of the long-read sequences are then confirmed using different genotype combinations. Next, haplotypes similar to those in the long-read sequences are searched for in the existing high-quality reference haplotype sequence set. Based on these haplotype subsets, the Lee and Stephens hidden Markov model is used to infer the haplotype of the sample and the indel sites present in the obtained haplotypes.
[0149] 4. The 30% of SNPs with the lowest quality scores (qual) among the mutation sites output by the BiLSTM model are used as low-quality mutation sites identified by the stacking-based deep learning network. These low-quality sites are then used as input for further model screening and judgment using the fully aligned residual neural network model (ResNet) to obtain highly reliable SNP sites.
[0150] 5. First, the SNP sites with the quality score (qual) of the mutation sites output by the BiLSTM model in the top 70% and the highly reliable SNP sites screened by the ResNet model are filtered with a mutation frequency (AF) less than 0.1 to remove low-quality SNP sites. Then, the filtered SNP sites are merged with the Indel sites identified based on the inferred haplotypes to output the final results of genomic variation detection.
[0151] Furthermore, the benchmarking tool hap.py (v0.3.15) can be used to perform a performance evaluation on the results. The evaluation results are shown in Figure 13. Figure 13 compares the performance of this solution and related solutions in indel variation detection. As can be seen from the figure, compared with existing traditional deep learning methods, the method proposed in the present invention significantly improves the recognition of indels in long-read sequencing data (F1 score >= 0.9), and the present invention still has a relatively ideal variation detection effect when the sequencing depth is low (below 20X). Among them, the F1 score is a comprehensive indicator for evaluating the recall rate and precision of the model. The higher the F1 score, the better the model effect.
[0152] The following describes the gene variation detection device provided in the embodiments of the present application.
[0153] 14 , in some embodiments, the present application further provides a gene variation detection device, which includes:
[0154] The first acquiring unit 1401 is used to acquire base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested;
[0155] A detection unit 1402 is configured to perform gene variation detection on the sequence alignment result data based on a first gene variation detection model to obtain multiple single-base variation sites, where the first gene variation detection model is a gene variation detection model based on gene stacking.
[0156] The second acquiring unit 1403 is configured to acquire quality scores corresponding to multiple single-base variant sites, and determine a single-base variant site having a quality score greater than a preset threshold as a first-category variant site;
[0157] The determining unit 1404 is configured to determine a target haplotype corresponding to the gene segment to be tested in the reference haplotype set based on the first type of variant sites, and determine an insertion or deletion site of the gene segment to be tested based on the target haplotype.
[0158] In one embodiment, the determining unit includes:
[0159] A first calculation subunit is used to calculate the genotype likelihood values corresponding to the first type of variant sites and multiple alleles;
[0160] The first determining subunit is used to determine the target haplotype corresponding to the gene segment to be tested in the reference haplotype set based on the genotype likelihood value.
[0161] In one embodiment, the first determining subunit is further configured to:
[0162] The iterative updating step is executed cyclically until the number of cycles reaches a preset number, and the haplotype obtained by the last update is output as the target haplotype of the gene segment to be tested. The iterative updating step includes the following steps:
[0163] Determining a candidate haplotype subset in the reference haplotype set based on the first type of variant sites;
[0164] Calculate the probability of occurrence of the base corresponding to each first variant site based on the genotype likelihood value and the candidate haplotype subset;
[0165] Based on the probability and the distance between the single-base variant sites in the candidate haplotype subset, the hidden Markov model is used to infer the individual genotype of each single-base variant site;
[0166] The haplotype of the gene fragment to be tested is updated based on the combination of individual genotypes at each single-base variant site.
[0167] In one embodiment, the second acquiring unit includes:
[0168] The detection subunit is used to detect potential mutation sites in the base sequence alignment result data, and obtain multiple potential mutation sites and the quality score corresponding to each potential mutation site;
[0169] A second determining subunit is configured to determine quality scores corresponding to the plurality of single-base variation sites based on a matching relationship between the plurality of single-base variation sites and the plurality of potential variation sites;
[0170] The third determination subunit is used to determine a single-base variation site with a quality score greater than a preset threshold as a first-category variation site.
[0171] In one embodiment, the gene variation detection device provided by the present application further includes:
[0172] A receiving subunit, configured to receive an input quality value to control the ratio;
[0173] The second calculation subunit is used to calculate a preset threshold based on the quality value control ratio and the quality score corresponding to each potential variation site.
[0174] In one embodiment, the gene variation detection device provided by the present application further includes:
[0175] a fourth determining subunit, configured to determine a single-base variant site having a quality score not greater than a preset threshold as a second-category variant site;
[0176] a review subunit, configured to review single-base variations of the second-category variant sites based on a second gene variation detection model, and to determine sites with confirmed variations as third-category variant sites based on the review results. The second gene variation detection model is a model for gene variation detection based on full alignment.
[0177] The fifth determination subunit is used to determine the target variation detection result of the gene fragment to be tested based on the first type variation site, the third type variation site and the insertion or deletion site.
[0178] In one embodiment, the fifth determining subunit includes:
[0179] A screening module is used to perform quality screening on the first and third category variant sites to obtain the target single-base variant sites;
[0180] The first determination module is used to determine the target variation detection result of the gene fragment to be tested based on the target single base variation site and the insertion or deletion site.
[0181] In one embodiment, the screening module includes:
[0182] An acquisition submodule, used to obtain the mutation frequencies corresponding to the first type of mutation sites and the third type of mutation sites;
[0183] The determination submodule is used to determine the mutation sites with mutation frequencies greater than a preset frequency threshold among the first and third category mutation sites as target single-base mutation sites.
[0184] In one embodiment, the review subunit includes:
[0185] An acquisition module, used to obtain the base sequence features corresponding to each second-category variation site;
[0186] an identification module, configured to identify gene variations for each base sequence feature based on a second gene variation detection model, and obtain a gene variation identification result;
[0187] The second determination module is used to determine multiple third-category variation sites from multiple second-category variation sites based on the gene variation identification results.
[0188] In one embodiment, the first acquiring unit includes:
[0189] An acquisition subunit, used to obtain nanopore sequencing data of the gene fragment to be tested;
[0190] The alignment subunit is used to compare the nanopore sequencing data to the reference genome to obtain base sequence alignment result data.
[0191] It can be seen that the contents of the above-mentioned gene variation detection method embodiment are all applicable to the embodiment of the present gene variation detection device. The functions specifically implemented by the present gene variation detection device embodiment are the same as those of the above-mentioned gene variation detection method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned gene variation detection method embodiment.
[0192] 15 , which illustrates a hardware structure of a computer device according to another embodiment, the computer device includes:
[0193] The processor 1501 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0194] Memory 1502 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). Memory 1502 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 1502 and is called by processor 1501 to execute the genetic variation detection method or genetic variation detection method of the embodiments of this application;
[0195] Input / output interface 1503, used to implement information input and output;
[0196] Communication interface 1504, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0197] Bus 1505 , which transmits information between various components of the device (e.g., processor 1501 , memory 1502 , input / output interface 1503 , and communication interface 1504 );
[0198] The processor 1501 , the memory 1502 , the input / output interface 1503 and the communication interface 1504 are connected to each other in communication within the device via the bus 1505 .
[0199] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device executes the above-mentioned gene variation detection method or gene variation detection method.
[0200] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.
[0201] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0202] It should be understood that in the description of the embodiments of the present application, multiple (or multiple items) means more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.
[0203] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0204] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0205] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0206] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0207] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0208] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A method for detecting gene mutations, characterized in that, The method includes: Obtaining the base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested; Performing gene variant detection on the sequence alignment result data based on a first gene variant detection model to obtain a plurality of single nucleotide variant sites, where the first gene variant detection model is a model for gene variant detection based on gene stacking; Obtaining the quality scores corresponding to the plurality of single nucleotide variant sites, and determining the single nucleotide variant sites with quality scores greater than a preset threshold as the first type of variant sites; Determining a target haplotype corresponding to the gene fragment to be tested in a reference haplotype set based on the first type of variant sites, and determining the insertion or deletion sites of the gene fragment to be tested based on the target haplotype.
2. The method according to claim 1, wherein The determining a target haplotype corresponding to the gene fragment to be tested in a reference haplotype set based on the first type of variant sites includes: Calculating the genotype likelihood values corresponding to the first type of variant sites and multiple alleles; Determining a target haplotype corresponding to the gene fragment to be tested in a reference haplotype set based on the genotype likelihood values.
3. The method according to claim 2, wherein The determining a target haplotype corresponding to the gene fragment to be tested in a reference haplotype set based on the genotype likelihood values includes: Repeatedly executing an iterative update step until the number of loops reaches a preset number, and outputting the haplotype obtained in the last update as the target haplotype of the gene fragment to be tested. The iterative update step includes the following steps: Determining a candidate haplotype subset in a reference haplotype set based on the first type of variant sites; Calculating the probability of occurrence of the base corresponding to each first variant site according to the genotype likelihood values and the candidate haplotype subset; Inferring the individual genotype of each single nucleotide variant site by using a hidden Markov model based on the probability and the distance between single nucleotide variant sites in the candidate haplotype subset; Updating the haplotype of the gene fragment to be tested based on the combination of the individual genotypes of each single nucleotide variant site.
4. The method according to claim 1, wherein The obtaining the quality scores corresponding to the plurality of single nucleotide variant sites, and determining the single nucleotide variant sites with quality scores greater than a preset threshold as the first type of variant sites includes: Performing potential variant site detection on the base sequence alignment result data to obtain a plurality of potential variant sites and the quality scores corresponding to each potential variant site; Determining the quality scores corresponding to the plurality of single nucleotide variant sites according to the matching relationship between the plurality of single nucleotide variant sites and the plurality of potential variant sites; Determining the single nucleotide variant sites with quality scores greater than a preset threshold as the first type of variant sites.
5. The method according to claim 4, wherein Before the determining the single nucleotide variant sites with quality scores greater than a preset threshold as the first type of variant sites, it further includes: Receiving an input quality value control ratio; Calculating a preset threshold based on the quality value control ratio and the quality scores corresponding to each potential variant site.
6. The method according to claim 1, wherein After the determining a target haplotype corresponding to the gene fragment to be tested in a reference haplotype set based on the first type of variant sites, and determining the insertion or deletion sites of the gene fragment to be tested based on the target haplotype, it further includes: Determine the single-base mutation sites with a mass fraction not greater than the preset threshold as the second type of mutation sites; Based on the second gene mutation detection model, perform single-base mutation verification on the second type of mutation sites, and determine the sites with a verification result of confirmed mutation as the third type of mutation sites. The second gene mutation detection model is a model for detecting gene mutations based on global alignment; Determine the target mutation detection result of the gene fragment to be tested according to the first type of mutation sites, the third type of mutation sites, and the insertion or deletion sites.
7. The method according to claim 6, wherein The determining the target mutation detection result of the gene fragment to be tested according to the first type of mutation sites, the third type of mutation sites, and the insertion or deletion sites includes: Perform quality screening on the first type of mutation sites and the third type of mutation sites to obtain target single-base mutation sites; Determine the target mutation detection result of the gene fragment to be tested according to the target single-base mutation sites and the insertion or deletion sites.
8. The method according to claim 7, characterized in that The performing quality screening on the first type of mutation sites and the third type of mutation sites to obtain target single-base mutation sites includes: Obtain the mutation frequencies corresponding to the first type of mutation sites and the third type of mutation sites; Determine the mutation sites with a mutation frequency greater than the preset frequency threshold among the first type of mutation sites and the third type of mutation sites as the target single-base mutation sites.
9. The method according to claim 6, wherein The performing single-base mutation verification on the second type of mutation sites based on the second gene mutation detection model and determining the sites with a verification result of confirmed mutation as the third type of mutation sites includes: Obtain the base sequence features corresponding to each second type of mutation site; Based on the second gene mutation detection model, perform gene mutation recognition on each of the base sequence features to obtain a gene mutation recognition result; Determine multiple third type of mutation sites among the multiple second type of mutation sites according to the gene mutation recognition result.
10. The method according to claim 1, characterized in that, The obtaining the base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested includes: Obtain the nanopore sequencing data of the gene fragment to be tested; Compare the nanopore sequencing data with the reference genome to obtain the base sequence alignment result data.
11. A gene mutation detection device, characterized in that, The device includes: A first acquisition unit for acquiring the base sequence alignment result data corresponding to the sequencing data of the gene fragment to be tested; A detection unit for detecting gene mutations on the sequence alignment result data based on the first gene mutation detection model to obtain multiple single-base mutation sites. The first gene mutation detection model is a model for detecting gene mutations based on gene stacking; A second acquisition unit for obtaining the mass fractions corresponding to the multiple single-base mutation sites and determining the single-base mutation sites with a mass fraction greater than the preset threshold as the first type of mutation sites; A determination unit for determining the target haplotype corresponding to the gene fragment to be tested in the reference haplotype set based on the first type of mutation sites and determining the insertion or deletion sites of the gene fragment to be tested based on the target haplotype. When the processor executes the computer program, it implements the gene mutation detection method according to any one of claims 1 to 10.
12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, 13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the gene mutation detection method according to any one of claims 1 to 10.
14. A computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the gene mutation detection method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Gene variation detection method and device
CN106611106A
Bioinformatics systems, apparatus, and methods for performing secondary and / or tertiary processing
CN109416928A
Genome short variation deep learning detection method and system based on third-generation sequencing
CN116959560A
Allele-specific sequence variation analysis
US20050089904A1
Systems and methods for identifying sequence variation
US20130073214A1