Gene variation detection method and device, electronic equipment and storage medium
Through long read and long sequencing technology and pre-trained models combined with internal reference gene depth correction, the accuracy problem of SMN1 and SMN2 gene variant detection is solved, achieving more efficient and accurate copy number variant detection, and simplifying the operation steps.
Patent Information
- Application Number
- CN202510829915.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the prior art, the variation detection methods of SMN1 gene and SMN2 gene have the problem of low detection accuracy, especially when distinguishing SMN1 and SMN2 genes, it is difficult, and has complex operation, strong subjectivity and limited detection ability.
The sequencing depth was obtained by long read and long sequencing technology. The pre-trained detection model was used to perform copy number variation detection based on the sequencing depth of SMN1 and SMN2 genes, and the sequencing depth of the internal reference gene was used for depth correction and standardization. The hybrid Gaussian model was used to distinguish copy number states.
It improves the accuracy and efficiency of copy number variation detection of SMN1 and SMN2 genes, simplifies the operation process, reduces subjective errors, and enhances the comprehensiveness and stability of the detection.
Smart Images

Figure CN120340601A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of bioinformatics technology, and particularly relates to a gene mutation detection method and device, an electronic device, and a storage medium. Background Art
[0002] Spinal Muscular Atrophy (SMA) is a genetic disease caused by gene defects, mainly caused by copy number variations of the SMN1 gene. Approximately 95% of patients are affected by homozygous deletion of the SMN1 gene. The SMN2 gene plays a compensatory role when SMN1 is absent, and the more copies of the SMN2 gene, the milder the patient's condition. Thus, it is particularly important to detect mutations in the SMN1 gene and the SMN2 gene.
[0003] However, in the related art, there are problems with low detection accuracy in the mutation detection methods for the SMN1 gene and the SMN2 gene. Summary of the Invention
[0004] The main objective of the embodiments of the present application is to propose a gene mutation detection method and device, an electronic device, and a storage medium, aiming to improve the accuracy of gene mutation detection.
[0005] To achieve the above objective, in the first aspect of the embodiments of the present application, a gene mutation detection method is proposed, and the method includes: Obtain the first long-read sequencing data of the sample to be tested, and align the first long-read sequencing data with a reference genome to obtain first alignment data; Determine the region of the target gene in the reference genome, and calculate the first sequencing depth of the target gene according to the determined region and the first alignment data, where the target gene includes a first sub-target gene and a second sub-target gene; Determine the first base distribution ratio between the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth; Call a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested.
[0006] In some embodiments, the step of calling a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested includes: Perform depth correction on the first sub-sequencing depth and the second sub-sequencing depth respectively according to the second sequencing depth of the reference gene; Call the detection model to perform copy number variation detection based on the first sub-sequencing depth after depth correction and the second sub-sequencing depth after depth correction, so as to obtain the target copy number variation type.
[0007] In some embodiments, the training process of the detection model includes the following steps: Obtain the second long-read sequencing data of the training sample, and align the second long-read sequencing data with the reference genome to obtain second alignment data; Extract the read segment data in the regions where the first sub-target gene and the second sub-target gene are located from the second alignment data respectively, and re-align the read segment data with the reference sequence of the first sub-target gene to obtain re-alignment data; Determine the second base distribution ratio of the first sub-target gene and the second sub-target gene according to the re-alignment data, calculate the third sequencing depth of the target gene according to the re-alignment data, and calculate the fourth sequencing depth of the internal reference gene according to the second alignment data; Determine the third sub-sequencing depth of the first sub-target gene and the fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth; Train the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth and the fourth sequencing depth.
[0008] In some embodiments, the calculating the third sequencing depth of the target gene according to the re-alignment data includes: Calculate the preliminary sequencing depth of the target gene according to the re-alignment data; Calculate the sample depth ratio according to the preliminary sequencing depth and the fourth sequencing depth, and perform normalization processing on the sample depth ratio according to the reference depth ratio of the internal reference sample corresponding to the training sample to obtain the third sequencing depth.
[0009] In some embodiments, the training the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth and the fourth sequencing depth includes: Perform depth correction on the third sub-sequencing depth and the fourth sub-sequencing depth respectively according to the fourth sequencing depth; Call the detection model to determine the first copy number state corresponding to the first sub-target gene based on the third sub-sequencing depth after depth correction, and determine the second copy number state corresponding to the second sub-target gene based on the fourth sub-sequencing depth after depth correction, so as to determine the predicted copy number variation type based on the first copy number state and the second copy number state; Train the detection model according to the standard copy number variation type of the training samples and the predicted copy number variation type.
[0010] In some embodiments, obtaining the first long-read sequencing data of the sample to be tested includes: Construct amplification primers for the target gene, where the amplification primers include first sub-amplification primers; Perform amplification processing on the sample to be tested according to the first sub-amplification primers, and perform sequencing processing on the result of the amplification processing to obtain the first long-read sequencing data.
[0011] In some embodiments, obtaining the first long-read sequencing data of the sample to be tested and aligning the first long-read sequencing data with a reference genome to obtain first alignment data includes: Construct amplification primers for the target gene, where the amplification primers include first sub-amplification primers and second sub-amplification primers, and the length of the second sub-amplification primers is greater than the length of the first sub-amplification primers; Perform amplification processing on the sample to be tested according to the first sub-amplification primers and the second sub-amplification primers, and perform sequencing processing on the result of the amplification processing to obtain the first long-read sequencing data; Split out first sub-amplification data corresponding to the first sub-amplification primers from the second long-read sequencing data, and align the first sub-amplification data with the reference genome to obtain the first alignment data; In some embodiments, the method further includes: Split out second sub-amplification data corresponding to the second sub-amplification primers from the first long-read sequencing data, and align the second sub-amplification data with the reference genome to obtain the third alignment data; Perform variant detection on the sample to be tested according to the third alignment data.
[0012] To achieve the above object, a second aspect of the embodiments of the present application provides a gene variant detection device, the device includes: An alignment unit, configured to obtain the first long-read sequencing data of the sample to be tested, and align the first long-read sequencing data with a reference genome to obtain first alignment data; A first sequencing depth calculation unit, configured to determine the region of the target gene in the reference genome, and calculate the first sequencing depth of the target gene according to the determined region and the first alignment data, where the target gene includes a first sub-target gene and a second sub-target gene; A second sequencing depth calculation unit, configured to determine a first base distribution ratio between a first sub-target gene and a second sub-target gene, and calculate a first sub-sequencing depth of the first sub-target gene and a second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and a first sequencing depth; A variation detection unit, configured to call a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth, so as to obtain a target copy number variation type of the sample to be tested.
[0013] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0014] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0015] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer program product, where the computer program product includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the method described in the first aspect.
[0016] The gene variation detection method, device, electronic device, and storage medium provided in the embodiments of the present application can obtain the sequencing depth of a first sub-target gene and the sequencing depth of a second sub-target gene based on the first long-read sequencing data and the first alignment data of the sample to be tested, and can perform copy number variation detection through a pre-trained detection model and the above two sequencing depths. It can be seen that the embodiments of the present application can implement copy number variation detection of the SMN1 gene (i.e., the first sub-target gene) and the SMN2 gene (i.e., the second sub-target gene) based on the detection model. In addition, since the sequencing depth in the embodiments of the present application is obtained based on long-read sequencing data, based on the characteristics of long-read sequencing (long-read sequencing can read longer DNA fragments), the specific sites of the SMN1 gene and the SMN2 gene can be determined based on the long-read sequencing data, so that the SMN1 gene and the SMN2 gene can be distinguished based on the specific sites, and further the accuracy of copy number variation detection can be improved. Description of the Drawings
[0017] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0018] Figure 1It is a flowchart of the gene variant detection method provided by the embodiments of the present application; Figure 2 It is a schematic diagram of 15 types of copy number variations provided by the embodiments of the present application; Figure 3 It is a flowchart of the detection model training process provided by the embodiments of the present application; Figure 4 It is a schematic diagram of the gene variant detection method provided by the embodiments of the present application; Figure 5 It is a schematic diagram of the sequencing quality control reports and clinical phenotype information of 35 SMA samples provided by the embodiments of the present application; Figure 6 It is a schematic diagram of the variant detection results of the long amplicons of the standard NA12878 provided by the embodiments of the present application; Figure 7 It is a schematic diagram of the SMA copy number variation types predicted by the detection model provided by the embodiments of the present application; Figure 8 It is a schematic diagram of the gene variant detection device provided by the embodiments of the present application; Figure 9 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0019] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0020] First, the terms involved in the present application are explained.
[0021] Short-read sequencing: By fragmenting DNA into short fragments of 50-300 bases for high-throughput parallel sequencing, it has the characteristics of high single-base accuracy, low cost and high throughput. Short-read sequencing is suitable for scenarios such as whole-genome sequencing, transcriptome analysis and epigenetics research. However, short-read sequencing is easily interfered by repetitive sequences when assembling complex genomes and is difficult to analyze long-fragment structural variations.
[0022] Long-read sequencing: By capturing DNA fragments with lengths ranging from thousands to tens of thousands of bases for sequencing. The core advantage of long-read is to span complex regions in the genome (such as repetitive sequences, structural variations or highly polymorphic regions), significantly improving the integrity and accuracy of genome assembly. Long-read sequencing is crucial for studying gene structural variations, transcript isoform identification and metagenomic analysis.
[0023] Sequencing depth: Also known as coverage depth, it is a key indicator for measuring the quality of genomic sequencing data, representing the average number of times a specific genomic region is sequenced. After mapping sequencing reads to the reference genome using alignment tools (such as BWA, Bowtie), an analysis tool is used to count the coverage times at each locus, and then the sequencing depth is calculated.
[0024] Normalization: It is a data preprocessing method. The core of normalization is to transform the original data with different dimensions or magnitudes into a unified standard scale through mathematical transformation. The data after normalization follows the standard normal distribution with a mean of 0 and a standard deviation of 1. Normalization can eliminate the differences in units and magnitudes between different features and improve the comparability of data.
[0025] Spinal Muscular Atrophy (SMA) is a common autosomal recessive genetic disorder. SMA is a genetic disorder caused by gene defects, mainly due to copy number variations of the SMN1 gene. Approximately 95% of patients are affected by the homozygous deletion of the SMN1 gene, while heterozygous deletions result in SMA carriers. The SMN2 gene plays a compensatory role when SMN1 is deleted, and the more copies of SMN2, the milder the condition. Due to the high homology between the SMN1 and SMN2 genes, with only 5 different loci, especially the C / T difference in exon 7, it increases the difficulty of detecting variations. The techniques for detecting the copy number of the SMN1 gene mainly include Sanger sequencing, multiplex ligation-dependent probe amplification (MLPA) (the current gold standard), traditional fluorescence PCR technology (real-time quantitative polymerase chain reaction, real-time qPCR), HRM high-resolution melting curve method (high-resolution melting analysis, HRM), and next-generation high-throughput sequencing technology (next-generation sequencing, NGS), etc.
[0026] Sanger sequencing is a method based on the termination of dideoxyribonucleotides (ddNTPs). By incorporating ddNTPs with different color labels into the reaction, DNA extension is randomly terminated to generate nucleic acid fragments of different lengths, and then the nucleotide sequence is read by capillary electrophoresis separation and color labeling. Sanger sequencing is a classic method for DNA sequence analysis, which is used to detect rare and complex mutations of the SMN1 gene and is the "gold standard" for mutation detection. In addition, Sanger sequencing can discover unknown SNP sites and determine the type and location of SNP sites, but it cannot distinguish between normal people and carriers, nor can it quantitatively analyze the number of SMN2 copies. The basic principle of MLPA technology is that the probe hybridizes with the target sequence DNA, and after connection and PCR amplification, the amplification product is separated by capillary electrophoresis, and the analysis software processes the data to obtain the results. The disadvantages of MLPA technology are that it cannot detect point mutations and special SMN12+0 carriers, as well as high cost, long time consumption, cumbersome operation, and susceptibility to contamination. The MGB probe real-time multiplex fluorescence quantitative PCR method determines the absence of the SMN1 gene by calculating the ΔΔCt value through relative quantitative detection of the SMN1 gene and the internal standard gene. The result interpretation of this method relies on the ΔΔCt calculation, which involves the reading of the Ct values of different detection channels, is highly subjective, and is easily affected by the operator. In addition, the SMN2 gene is highly homologous to SMN1, which can easily interfere with the detection, resulting in false negative or false positive results. In addition, the amplification efficiency of the target gene and the internal reference gene is also required to be high, otherwise the accuracy of the results may be affected. HRM (high-resolution melting curve) technology combines PCR amplification and melting curve analysis, and analyzes gene sequence variation by detecting the melting curve of PCR amplification products, which is particularly suitable for distinguishing single-base differences. Based on the differences in the length and GC content of the amplified products, HRM technology detects changes in the melting temperature (Tm) and obtains different melting curves to determine the variation and copy number of the gene. HRM technology has high requirements for data analysis and requires software normalization, which increases the difficulty of result interpretation. In addition, it is sensitive to the quality of the amplified product, and the results depend on the stability of the PCR amplification efficiency. In addition, the detection sensitivity for low-frequency mutations is low and it is not suitable for detecting a small number or rare mutations. NGS technology has the advantages of high throughput and high sensitivity in genetic testing, but complex analysis methods are often required to identify SMN1 gene copy number variations. This is because there are only a few base differences between the SMN1 and SMN2 gene sequences, making it difficult to distinguish between the SMN1 and SMN2 genes in sequencing data. Complex algorithms and analysis steps must be used to identify these subtle differences, but these methods are often difficult to interpret and their accuracy may also be limited.
[0027] In summary, the detection methods in the related art have the following problems: 1. Complex data parsing. Since the SMN1 and SMN2 genes are highly homologous, complex analysis algorithms are required to process the data. This increases the difficulty of parsing and the possibility of errors.
[0028] 2. Limited detection ability. Some related technologies are mainly used to detect copy number variations of SMN1, and are not accurate enough for complex mutation types or quantification of the copy number of the SMN2 gene.
[0029] 3. Strong subjectivity. There is a large degree of subjectivity in setting thresholds and interpreting data during the operation of traditional fluorescence PCR technology and HRM high-resolution melting curve method, which may lead to inaccurate results.
[0030] 4. Complex operation. The experimental steps of related technologies are cumbersome, accurate operation is required to avoid contamination, and high requirements are imposed on equipment and operators.
[0031] Based on this, the embodiments of the present application provide a gene mutation detection method, device, electronic device and storage medium, so as to improve the accuracy of gene mutation detection.
[0032] The gene mutation detection method provided by the embodiments of the present application relates to the field of bioinformatics technology. The gene mutation detection method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing the gene mutation detection method, etc., but is not limited to the above forms.
[0033] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0034] The following describes the gene mutation detection method provided by the embodiments of this application.
[0035] Refer to Figure 1 , in some embodiments, the gene mutation detection method provided by the embodiments of this application includes but is not limited to steps S101 to S104.
[0036] Step S101, obtain the first long-read sequencing data of the sample to be tested, and compare the first long-read sequencing data with the reference genome to obtain the first comparison data; Step S102, determine the region of the target gene in the reference genome, and calculate the first sequencing depth of the target gene according to the determined region and the first comparison data; Step S103, determine the first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth; Step S104, call the pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested.
[0037] Steps S101 to S104 shown in the embodiments of the present application can obtain the sequencing depth of the first sub-target gene and the sequencing depth of the second sub-target gene based on the first long-read sequencing data and the first alignment data of the sample to be tested, and copy number variation detection can be performed through a pre-trained detection model and the above two sequencing depths. It can be seen from this that the embodiments of the present application can implement copy number variation detection of the SMN1 gene (i.e., the first sub-target gene) and the SMN2 gene (i.e., the second sub-target gene) based on the detection model. In addition, since the embodiments of the present application obtain the sequencing depth based on long-read sequencing data, based on the characteristics of long-read sequencing (long-read sequencing can read longer DNA fragments), the specific sites of the SMN1 gene and the SMN2 gene can be determined based on the long-read sequencing data. In this way, the SMN1 gene and the SMN2 gene can be distinguished based on the specific sites, and thus the accuracy of copy number variation detection can be improved.
[0038] In step S101 of some embodiments, the sample to be tested may refer to an SMA sample to be subjected to gene variation detection. The DNA of the sample to be tested is subjected to PCR amplification and long-read sequencing, and data quality control is performed on the sequencing results to obtain the first long-read sequencing data. The reference genome may refer to the human reference genome (such as GRCh38). The first long-read sequencing data is aligned with the reference genome to identify potential variation information in the sample to be tested with the reference genome as a benchmark, and the first alignment data is obtained.
[0039] In some embodiments, "obtaining the first long-read sequencing data of the sample to be tested" in step S101 includes the following sub-steps: Construct amplification primers for the target gene, and the amplification primers include the first sub-amplification primers; Perform amplification processing on the sample to be tested according to the first sub-amplification primers, and perform sequencing processing on the amplification processing results to obtain the first long-read sequencing data.
[0040] In the embodiments of the present application, the amplification primers may refer to a class of short single-stranded DNA fragments for PCR amplification. The amplification primers are used to guide DNA polymerase to start replicating the target DNA sequence from a specific position, so as to achieve specific amplification of the target DNA. The amplification primers of the present application may include short amplification primers (i.e., the first sub-amplification primers), and the short amplification primers may include amplification primers for exons 7 to 8 of the SMN1 gene and the SMN2 gene. In this way, PCR amplification is performed on the sample to be tested based on the short amplification primers, and sequencing and data quality control are performed on the PCR amplification results to obtain the first long-read sequencing data.
[0041] It can be understood that when the amplification primers only include short amplification primers, the amplification products of PCR also only include short amplicons. Thus, the alignment operation can refer to aligning the short amplicons after data quality control with the reference genome.
[0042] In some other embodiments, the amplification primers may further include long amplification primers (i.e., the second sub-amplification primers, and the length of the second sub-amplification primers is greater than that of the first sub-amplification primers). In this case, step S101 may include the following sub-steps: Construct amplification primers for the target gene, where the amplification primers include the first sub-amplification primers and the second sub-amplification primers, and the length of the second sub-amplification primers is greater than that of the first sub-amplification primers; Perform amplification treatment on the sample to be tested according to the first sub-amplification primers and the second sub-amplification primers, and perform sequencing treatment on the amplification treatment result to obtain the first long-read sequencing data; Split out the first sub-amplification data corresponding to the first sub-amplification primers from the first long-read sequencing data, and align the first sub-amplification data with the reference genome to obtain the first alignment data.
[0043] In the embodiments of the present application, the amplification primers may include long amplification primers (i.e., the second sub-amplification primers) and short amplification primers (i.e., the first sub-amplification primers, and the length of the first sub-amplification primers is less than that of the second sub-amplification primers). Among them, the short amplification primers are the same as those described in the above embodiments, and the long amplification primers may include amplification primers for exons 2b to 8 of the SMN1 gene and the SMN2 gene. It can be understood that the amplification primers may further include amplification primers for internal reference genes without copy number (i.e., exons 4 to 6 of the CFTR gene).
[0044] Thus, perform PCR amplification on the sample to be tested based on the long amplification primers and the short amplification primers, and perform sequencing and data quality control on the PCR amplification result to obtain the first long-read sequencing data.
[0045] It can be understood that when the amplification primers include long amplification primers and short amplification primers, the amplification products of PCR include long amplicons corresponding to the long amplification primers and short amplicons corresponding to the short amplification primers (i.e., the first sub-amplification data). Thus, software such as Citrus can be called to split out the corresponding long amplicons and short amplicons from the first long-read sequencing data based on the long amplification primers and the short amplification primers. At this time, the alignment operation can refer to aligning the split short amplicons with the reference genome. The role of the long amplicons will be introduced in the following embodiments.
[0046] In step S102 of some embodiments, the target gene may refer to a gene related to SMA (hereinafter referred to as the SMA gene, which includes the SMN1 gene and the SMN2 gene. Among them, the SMN1 gene can be called the first sub-target gene, and the SMN2 gene can be called the second sub-target gene). According to the relevant knowledge base, the position (i.e., the region where it is located) of the target gene in the reference genome can be obtained. Thus, according to the determined position and the first alignment data, the sequencing depth of the target gene (i.e., the first sequencing depth) can be calculated.
[0047] In step S103 of some embodiments, the SMN1 gene and the SMN2 gene are highly homologous, with only five different sites, especially the C / T difference in exon 7. For example, at the c.840C>T site in exon 7, the SMN1 gene is usually C (cytosine), while the SMN2 gene is usually T (thymine). Therefore, by statistically analyzing the base distribution ratio of the SMN1 gene and the SMN2 gene at the c.840C>T site in exon 7 (i.e., the first base distribution ratio), the copy number ratio of the SMN1 gene and the SMN2 gene can be determined. Since the first sequencing depth calculated in the above step is the sum of the sequencing depths of the SMN1 gene and the SMN2 gene, according to the first base distribution ratio and the first sequencing depth, the sequencing depth of the SMN1 gene (i.e., the first sub-sequencing depth) and the sequencing depth of the SMN2 gene (i.e., the second sub-sequencing depth) can be calculated.
[0048] In step S104 of some embodiments, the detection model may refer to a pre-trained model with the ability to detect copy number variations. Call this detection model and use the first sub-sequencing depth and the second sub-sequencing depth as the input data of this detection model to obtain the target copy number variation type of the sample to be tested. As Figure 2 shown, in the embodiments of the present application, according to the four possible copy number states of the SMN1 gene and the SMN2 gene (i.e., 0 copies, 1 copy, 2 copies, and ≥3 copies), 15 copy number variation types can be combined (i.e., Figure 2 types A to O shown). Thus, the detection model may refer to a Gaussian mixture model (GMM). The GMM model is a probability model that assumes a mixture of multiple Gaussian distributions, and each Gaussian distribution can represent a different copy number state. During prediction, the GMM model assigns different probabilities to different copy number states according to the first sub-sequencing depth and the second sub-sequencing depth, and then obtains the target copy number variation type according to the copy number state with the maximum probability corresponding to the SMN1 gene and the copy number state with the maximum probability corresponding to the SMN2 gene. That is, the target copy number variation type refers to one of these 15 copy number variation types.
[0049] It can be understood that in some embodiments, other analyses can be performed on the sample to be tested based on the split long amplicons. For example, in some embodiments, the gene mutation detection method provided by the embodiments of the present application may further include the following steps: Split the second sub-amplification data corresponding to the second sub-amplification primer from the first long-read sequencing data, and align the second sub-amplification data with the reference genome to obtain the third alignment data; Perform mutation detection on the sample to be tested according to the third alignment data.
[0050] In the embodiments of the present application, mutation detection can be performed on the sample to be tested based on the split long amplicons (i.e., the second sub-amplification data). Specifically, the long amplicons are aligned with the reference genome to obtain the third alignment data (in BAM format), and the clair3 software applicable to long-read sequencing data is called to perform mutation detection on the third alignment data. In this way, the mutation detection result of the sample to be tested can be obtained. It can be understood that the mutation detection result can be used to study other mutation types of the SMA gene.
[0051] It can be understood that the PCR detection method in the related art cannot cover copy number typing and mutation detection, while the present application can simultaneously achieve copy number type detection and mutation detection based on the characteristics of long-read sequencing data. For example, when it is necessary to simultaneously achieve copy number type detection and mutation detection, PCR amplification is performed based on short amplification primers and long amplification primers. At this time, the DNA fragments obtained after PCR amplification include short amplicons and long amplicons. Among them, the short amplicons are used for the detection and analysis of the copy number variation type of the SMA gene, while the long amplicons are used for the mutation detection and analysis of SNPs (single nucleotide polymorphisms) and INDELs (insertions or deletions). In this case, the sequenced data after data quality control needs to be split to perform different detection and analysis based on the split long amplicons and short amplicons. When only copy number variation detection of the SMA gene is required, PCR amplification is performed based on short amplification primers. At this time, the PCR amplification product only contains short amplicons. In this case, the sequenced data after data quality control can be set not to be split, that is, the downstream analysis is directly performed based on the sequenced data after data quality control (i.e., the first long-read sequencing data).
[0052] In some embodiments, step S104 may include but is not limited to the following sub-steps: Perform depth correction on the first sub-sequencing depth and the second sub-sequencing depth respectively according to the second sequencing depth of the reference gene; Call the detection model to perform copy number variation detection based on the depth-corrected first sub-sequencing depth and the depth-corrected second sub-sequencing depth to obtain the target copy number variation type.
[0053] In the embodiments of the present application, an internal reference gene (such as the CFTR gene) is determined according to a sample to be measured, and a second sequencing depth of the internal reference gene is obtained. The first sub-sequencing depth of the SMN1 gene and the second sub-sequencing depth of the SMN2 gene are normalized (i.e., depth correction) according to the second sequencing depth to eliminate the influence of sequencing conditions and depth differences in different samples. The GMM model performs copy number variation detection based on the depth-corrected first sub-sequencing depth and the depth-corrected second sub-sequencing depth to obtain a target copy number variation type.
[0054] In the embodiments of the present application, the sequencing depth of the SMA gene (including the SMN1 gene and the SMN2 gene) is first normalized based on the sequencing depth of the internal reference gene, and then copy number variation detection is performed based on the normalized sequencing depth, so as to improve the accuracy of the sequencing depth, and further improve the accuracy of the copy number variation detection.
[0055] As follows, the training process of the detection model will be described.
[0056] Refer to Figure 3 In some embodiments, the training process of the detection model includes but is not limited to steps S301 to S305.
[0057] Step S301: Obtain the second long-read sequencing data of the training sample, and align the second long-read sequencing data with the reference genome to obtain second alignment data; Step S302: Extract the read segment data in the regions where the first sub-target gene and the second sub-target gene are located from the second alignment data respectively, and re-align the read segment data with the reference sequence of the first sub-target gene to obtain re-alignment data; Step S303: Determine the second base distribution ratio of the first sub-target gene and the second sub-target gene according to the re-alignment data, calculate the third sequencing depth of the target gene according to the re-alignment data, and calculate the fourth sequencing depth of the internal reference gene according to the second alignment data; Step S304: Determine the third sub-sequencing depth of the first sub-target gene and the fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth; Step S305: Train the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth.
[0058] In step S301 of some embodiments, the training samples may refer to SMA samples used for training the detection model. PCR amplification and long-read sequencing are performed on the DNA of the training samples, and data quality control is performed on the sequencing results to obtain the second long-read sequencing data. Among them, data quality control may refer to using relevant tools to perform quality control on the sequencing data to ensure the accuracy and reliability of subsequent analysis based on the sequencing data. Specifically, after the PCR amplification is completed, the CycloneSEQ 24 barcode library preparation reagent kit can be used to prepare the sequencing library. After obtaining the sequencing library, the CycloneSEQ WT sequencing kit, the CycloneSEQ WT sequencing chip, and the nanopore gene sequencer CycloneSEQ-WT02 adapted to this sequencing chip are used to sequence the DNA of the training samples to obtain the raw sequencing data (i.e., the sequencing results). The raw sequencing data (in FASTQ format) is input into the Bamboo software to generate a data quality control report file. The data quality control report file can be used to evaluate the data quality and provide a basis for subsequent analysis. After performing data quality control on the raw sequencing data, the second long-read sequencing data can be obtained.
[0059] The second long-read sequencing data is aligned with a reference genome (such as the human reference genome GRCh38), and the minimap2 software specialized for processing long reads is used to generate the alignment result (i.e., the second alignment data, which is a SAM file). Among them, the second alignment data may include the position of the read segment, the alignment quality, and the CIGAR string. In addition, the samtools software can also be used to convert the alignment result into a more compact BAM format.
[0060] In step S302 of some embodiments, by obtaining the positions of the SMN1 gene (i.e., the first sub-target gene) and the SMN2 gene (i.e., the second sub-target gene) in the reference genome (such as the human reference genome GRCh38) on a relevant knowledge base (such as the UCSC Genome Browser), and using tools such as samtools to extract the read segment data aligned to the SMN1 gene region and the read segment data aligned to the SMN2 gene region from the second alignment data (i.e., the alignment result of the short amplicon and the reference genome) based on the obtained positions. The obtained read segment data is realigned with the reference sequence of the SMN1 gene, that is, the obtained read segment data is realigned to a sequence file containing only the SMN1 gene to obtain the realignment data (which is a BAM file).
[0061] In step S303 of some embodiments, according to the re-alignment data, the base distribution of the c.840C>T locus in exon 7 of the SMN1 gene and the SMN2 gene is statistically analyzed. According to this base distribution, the distribution ratio of C / T bases (i.e., the second base distribution ratio) can be obtained. Based on the second base distribution ratio, the copy number ratio of the SMN1 gene to the SMN2 gene can be determined. It can be understood that since the SMN1 gene and the SMN2 gene are highly homologous and only 5 bases are different, and the c.840C>T locus in exon 7 affects the gene splicing pattern. Thus, in the case where only 5 bases are different, when the read data of the SMN1 gene region and the read data of the SMN2 gene region extracted are aligned to the reference genome, there may be a situation where it is impossible to accurately distinguish the two genes (because the SMN1 gene may be misaligned to the reference region of the SMN2 gene in the reference genome, and the SMN2 gene may be misaligned to the reference region of the SMN1 gene in the reference genome). Therefore, in the embodiments of the present application, the read data of the two genes extracted are re-aligned to the reference sequence of the SMN1 gene, so that the ratio of the SMN1 gene:SMN2 gene under the reference sequence of the SMN1 gene can be determined according to the base difference characteristics of the two genes (i.e., the SMN1 gene and the SMN2 gene) at the c.840C>T locus.
[0062] Calculate the third sequencing depth of the SMA gene according to the re-alignment data, and calculate the fourth sequencing depth of the reference gene (such as the CFTR gene) according to the whole genome alignment file (i.e., the second alignment data). It can be understood that in order to ensure that the sequencing data is not affected by the amplification fragment length, the statistical value is the average depth information when calculating the sequencing depth (i.e., the third sequencing depth and the fourth sequencing depth). In addition, the third sequencing depth calculated based on the re-alignment data is the sum of the sequencing depth of the SMN1 gene and the sequencing depth of the SMN2 gene.
[0063] In some embodiments, "calculating the third sequencing depth of the target gene according to the re-alignment data" in step S303 includes the following sub-steps: Calculate the preliminary sequencing depth of the target gene according to the re-alignment data; Calculate the sample depth ratio according to the preliminary sequencing depth and the fourth sequencing depth, and standardize the sample depth ratio according to the reference depth ratio of the reference sample corresponding to the training sample to obtain the third sequencing depth.
[0064] In the embodiments of the present application, the third sequencing depth may be the sequencing depth obtained after standardization. Thus, the sequencing depth of the SMA gene directly calculated according to the re-alignment data is used as the preliminary sequencing depth . According to the preliminary sequencing depth and the fourth sequencing depth The sample depth ratio is calculated . In addition, in order to eliminate the fluctuations in the sample sequencing results caused by different batches of experiments, that is, in order to correct the batch effect, a reference sample can be introduced. The reference sample and the training samples are from the same batch. And, according to the same method described above, the reference depth ratio of the reference sample can be calculated. In this way, the sample depth ratio can be standardized based on the reference depth ratio, and the third sequencing depth can be obtained according to the standardization result.
[0065] In step S304 of some embodiments, based on the above description, the third sequencing depth calculated based on the re-alignment data is the sum of the SMN1 gene sequencing depth and the SMN2 gene sequencing depth. In this way, the sequencing depth corresponding to the SMN1 gene (i.e., the third sub-sequencing depth) and the sequencing depth corresponding to the SMN2 gene (i.e., the fourth sequencing depth) can be calculated in combination with the second base distribution ratio.
[0066] In step S305 of some embodiments, the copy number detection error of the detection model can be determined based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth. In this way, the parameters of the detection model can be adjusted according to the copy number detection error, so as to realize the training of the detection model.
[0067] In some embodiments, step S305 may include the following sub-steps: Perform depth correction on the third sub-sequencing depth and the fourth sub-sequencing depth respectively according to the fourth sequencing depth; Call the detection model to determine the first copy number status corresponding to the first sub-target gene based on the depth-corrected third sub-sequencing depth, and determine the second copy number status corresponding to the second sub-target gene based on the depth-corrected fourth sub-sequencing depth, so as to determine the predicted copy number variation type based on the first copy number status and the second copy number status; Train the detection model according to the standard copy number variation type and the predicted copy number variation type of the training samples.
[0068] In the embodiments of the present application, in order to eliminate the influence of the sequencing conditions and depth differences in different samples, the sequencing depth of the SMN1 gene (i.e., the third sub-sequencing depth) and the sequencing depth of the SMN2 gene (i.e., the fourth sub-sequencing depth) can be normalized respectively based on the sequencing depth of the reference gene (i.e., the fourth sequencing depth) to achieve depth correction. In this way, the detection model can be called to perform copy number variation detection on the depth-normalized third sub-sequencing depth and the depth-normalized fourth sub-sequencing depth to obtain the predicted copy number variation type. It can be understood that the predicted copy number variation type can be one of the 15 types as Figure 2 shown.
[0069] Specifically, in the embodiments of the present application, the detection model can be a GMM model, and the GMM model can include Gaussian distributions corresponding to different copy number states (including 0 copy, 1 copy, 2 copies, etc.). The detection model can assign probabilities to different copy number states based on the sequencing depth of the SMN1 gene (i.e., the third sub-sequencing depth), and take the copy number state with the highest probability as the first copy number state of the SMN1 gene. Similarly, the second copy number state of the SMN2 gene can be obtained. Thus, as Figure 2 shown, the predicted copy number variation type can be determined based on the first copy number state and the second copy number state.
[0070] Thus, the detection model is supervised trained based on the predicted copy number variation type and the true copy number of the training sample (i.e., the standard copy number variation type). Specifically, the copy number detection error of the detection model can be calculated based on the predicted copy number variation type and the standard copy number variation type. Thus, the parameters of the detection model can be adjusted based on the copy number detection error.
[0071] The gene variation detection method provided by the embodiments of the present application has the following beneficial effects: 1. The detection accuracy is improved. The embodiments of the present application perform relevant detections based on long-read sequencing data, and the long-read sequencing data can provide more comprehensive genomic result information. Specifically, compared with short-read sequencing, long-read sequencing has obvious advantages in detecting complex gene regions, repetitive sequences, etc. Thus, for the copy number variation detection of the SMN1 gene and the SMN2 gene, more accurate discrimination can be achieved based on long-read sequencing data, thereby improving the detection quality and detection accuracy. In addition, the embodiments of the present application also perform normalization processing based on the sequencing depth of the reference gene, so that the stability of the detection can be maintained under different experimental conditions, thereby improving the detection precision.
[0072] 2. The detection efficiency is improved. Since the long-read sequencing technology generates longer sequence fragments, the alignment difficulty in complex genomic regions can be reduced, thereby reducing the data analysis and processing time. In addition, the efficiency of the entire detection process can be made higher by optimizing the alignment algorithm (such as minimap2) and model modeling. Thus, a large number of sample analyses can be completed in a shorter time.
[0073] 3. Operational simplicity. The mixture Gaussian model applied in the embodiments of the present application can not only automatically distinguish samples with different copy numbers, but also has strong adaptability, can process samples with complex genotypes, and reduces the cumbersome manual alignment and result interpretation operations.
[0074] 4. The comprehensiveness of the detection is improved. The embodiments of the present application can perform PCR amplification based on long amplification primers and short amplification primers. Thus, copy number variation detection and other variation detections can be achieved.
[0075] Referring to Figure 4 , in a specific embodiment, the gene mutation detection method provided by the application example may include the following steps: 1. PCR amplification and nanopore sequencing.
[0076] As shown in Table 1 below, long amplification primers, short amplification primers, and internal reference gene amplification primers without copy number variations are designed. Among them, as shown in Table 1 below, the short amplification primers may be amplification primers for exons 7 to 8 of the SMN1 gene and the SMN2 gene, and the amplification primers for the internal reference gene may be amplification primers for exons 4 to 6 of the CFTR gene.
[0077]
[0078] Table 1 Based on the above-designed amplification primers, 15 columns of standard products obtained clinically are subjected to multiple PCR amplifications to obtain the DNA of 35 SMA samples. Among them, the amplification reaction system is 50 μL, including: 25 μL 2× ApexHF CL Buffer, 1 μL primer mixture (0.2 μM), 30 ng template DNA, 1 μL ApexHF HS DNA polymerase-CL, and nuclease-free water is added to 50 μL. The amplification program is 94°C for 1 min; 98°C for 10 s, 60°C for 15 s, 68°C for 4 min, cycling 25 times; 68°C for 2 min; keep at 4°C.
[0079] After the amplification is completed, the CycloneSEQ 24 barcode library preparation reagent kit (CycloneSEQ, H940-000018) can be used to prepare the sequencing library. After obtaining the sequencing library, the CycloneSEQ WT sequencing kit (CycloneSEQ, H940-000016), the CycloneSEQ WT sequencing chip (CycloneSEQ, H930-000001-00), and the nanopore gene sequencer CycloneSEQ-WT02 (CycloneSEQ, H900-000001-00) adapted to the sequencing chip are used to sequence the SMA samples to obtain sequencing data.
[0080] 2. Data quality control, primer splitting, and read filtering.
[0081] First, convert the sequencing data obtained in step 1 into a FASTQ format file. Second, use Bamboo software to perform quality control on the sequencing data, evaluate whether the data volume is sufficient, and identify whether there are biases during the sequencing process to ensure that the data meets the requirements of subsequent analysis. Then, call Citrus software to split the short amplicons and long amplicons in the sequencing data based on the primer sequences used during PCR amplification for subsequent analysis separately. Finally, to further improve the accuracy and analysis efficiency of the sequencing data of these 35 samples, use the seq module in the seqkit toolkit to filter the sequencing data according to the read length (≥4000bp), remove the reads that do not meet the requirements, optimize the quality of the sequencing data, and lay the foundation for subsequent precise analysis. The sequencing quality control reports and clinical phenotype information of the 35 SMA samples are shown as Figure 5 shown below.
[0082] 3. Genome alignment.
[0083] The purpose of step 3 is to align the sequencing data obtained in step 2 with a reference genome (such as the human reference genome GRCh38) to generate an alignment file in BAM format. Specifically, the input data can be the long amplicons and short amplicons obtained after data quality control, as well as the human reference genome GRCh38. Connect the minimap2 and samtools commands using a pipe symbol to process the alignment of long amplicons with the reference genome and the alignment of short amplicons with the reference genome respectively, and finally generate a BAM file containing alignment information (including the BAM file corresponding to long amplicons and the BAM file corresponding to short amplicons).
[0084] 4. Variant detection for long amplicons.
[0085] After completing the genome alignment, the whole-genome alignment result file (BAM file) of long amplicons can be obtained from step 3, and call the clair3 software for long-read sequencing data to perform variant detection on this alignment result file.
[0086] The variant detection results of the long amplicons of the standard NA12878 are shown as Figure 6 shown below. Figure 6The header in it includes the following content: the name of the chromosome or the name of the reference sequence, the position of the variant site on the reference genome, the unique identifier of the variant (represented by "." when empty in the table), the base or sequence on the reference genome, the base or sequence after the variant (i.e., the sequence different from the reference genome), the alignment quality score (indicating the confidence of the corresponding variant), whether the variant passes the quality filter (usually including "PASS" indicating passing the filter), the model used by the clair3 software for variant detection (F indicates that the clair3 software uses the "Full alignment" model for variant detection, P indicates that the clair3 software uses the pileup model for variant detection, and different models will have corresponding identifiers here), the data format fields (GT represents genotype, GQ represents genotype quality, DP represents sequencing depth, AD represents allele depth, AF represents allele frequency), and the genotype information of the sample.
[0087] 5. Extract the reads corresponding to the SMN1 gene and the SMN2 gene.
[0088] First, obtain the specific positions of the SMA genes (including the SMN1 gene and the SMN2 gene) in the GRCh38 reference genome on the UCSC Genome Browser. Then, call the samtools tool to extract the reads aligned to the SMN1 gene region and the reads aligned to the SMN2 gene region from the whole-genome alignment results (the alignment results corresponding to short amplicons, which are BAM files) obtained in step 3 based on the coordinate information of the SMN1 gene and the SMN2 gene.
[0089] 6. Re-align.
[0090] According to the reads obtained in step 5, re-align the reads to the reference sequence file containing only the SMN1 gene to obtain the alignment result file (which is a BAM file). According to this alignment result file, calculate the base distribution ratio of the SMN1 gene and the SMN2 gene at the c.840C>T site of exon 7.
[0091] 7. Calculate the sequencing depth.
[0092] Calculate the sequencing depth of the SMA gene based on the re-alignment result file obtained in step 6, and calculate the sequencing depth of the reference gene based on the alignment result file obtained in step 3. Among them, the sequencing depth of the SMA gene is the sum of the sequencing depths of the SMN1 gene and the SMN2 gene. In addition, to ensure that the sequencing information is not affected by the length of the amplified fragment, the statistical value is the average depth information when calculating the sequencing depth.
[0093] 8. Standardization processing.
[0094] Based on the ratio of the sequencing depth of the internal reference sample to the sequencing depth of the SMA gene (i.e., ), the of the samples in the same batch are standardized.
[0095] 9. Model training.
[0096] Samples No. 1 to No. 21 and No. 35 among the 35 SMA samples are used as the training set, and a detection model is constructed based on this training set. Among them, sample No. 35 is used as the internal reference sample, and the internal reference sample is used as the reference for the standardization of the training set. When constructing the model, the following information of samples No. 1 to No. 21 and No. 35 needs to be used: the sequencing depth of the SMN1 gene, the sequencing depth of the SMN2 gene, and the sequencing depth of the internal reference gene (such as CFTR). Among them, the sequencing depth of the SMN1 gene and the sequencing depth of the SMN2 gene can be obtained based on the base distribution ratio calculated in step 6 and the sequencing depth of the SMA gene calculated in step 7. The sequencing depth of the internal reference gene is used as the reference for normalization, that is, the sequencing depth of the SMN1 gene and the sequencing depth of the SMN2 gene need to be normalized based on the sequencing depth of the internal reference gene.
[0097] The four possible copy number states of the SMN1 gene and the SMN2 gene are combined, so that the SMA copy number variation can be classified into Figure 2 15 types as shown.
[0098] The detection model can detect the copy number variation of these samples based on the normalized sequencing depth (including the sequencing depth of the SMN1 gene and the sequencing depth of the SMN2 gene) of samples No. 1 to No. 21, and obtain the predicted copy number type corresponding to each sample. The detection model is supervised and trained based on the predicted copy number type and the true copy number type corresponding to the sample.
[0099] As Figure 4 shown, when performing copy number variation detection, the detection model can first determine whether the corresponding sample is a zero-copy, and then select the zero-copy module or the non-zero-copy module for detection based on the zero-copy judgment result to determine the copy number variation type. In this way, the efficiency of copy number variation detection can be improved. Among them, the zero-copy module can detect samples with a copy number state of zero-copy, and the non-zero-copy module can detect samples with a copy number state of non-zero-copy (i.e., 1 copy, 2 copies, etc.).
[0100] 10. Model verification.
[0101] Samples numbered 22 to 35 out of the 35 SMA samples were used as the validation set, which was used to evaluate the detection accuracy of the detection model constructed based on Step 9. Among them, Sample No. 35 was a positive control sample. Before inputting the relevant data into the detection model, it was necessary to standardize the relevant sequencing depths of the data in the training set based on the relationship between the positive control sample and the reference sample. When validating the model, the sequencing depth of the reference gene in each sample was used to normalize the sequencing depth of the SMA gene in that sample, and the detection model was called to perform copy number variation detection on the normalized sequencing depths (including the sequencing depths of the SMN1 gene and the SMN2 gene). Finally, the predicted SMA copy number variation types obtained by the detection model were as Figure 7 shown. In addition, the prediction results of the detection model were compared with the "gold standard" provided by the experiment, and the results showed that the prediction accuracy of the detection model constructed in this application reached 100%.
[0102] It can be understood that the above Steps 1 to 10 are the model construction and training processes. As Figure 4 shown, among them, Steps 1 to 2 are the data preprocessing processes, and Steps 3 to 9 are the data analysis processes and result analysis processes. In practical applications, the relevant data of the sample to be tested (i.e., the sequencing depths of the SMN1 gene and the SMN2 gene) can be obtained based on the above steps, so that the trained model can be called to perform copy number variation detection on the sample to be tested based on the relevant data.
[0103] Referring to Figure 8 , this embodiment of the present application also provides a gene variation detection device, which includes: A comparison unit 801, configured to obtain the first long-read sequencing data of the sample to be tested, and compare the first long-read sequencing data with the reference genome to obtain first comparison data; A first sequencing depth calculation unit 802, configured to determine the region of the target gene in the reference genome, and calculate the first sequencing depth of the target gene according to the determined region and the first comparison data, where the target gene includes a first sub-target gene and a second sub-target gene; A second sequencing depth calculation unit 803, configured to determine the first base distribution ratio between the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth; A variation detection unit 804, configured to call a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth, and obtain the target copy number variation type of the sample to be tested.
[0104] It can be seen that the content in the embodiments of the above gene mutation detection method is applicable to the embodiments of this gene mutation detection device. The functions specifically implemented in the embodiments of this gene mutation detection device are the same as those in the embodiments of the above gene mutation detection method, and the beneficial effects achieved are also the same as those in the embodiments of the above gene mutation detection method.
[0105] Referring to Figure 9 , Figure 9 FIG. shows the hardware structure of an electronic device according to another embodiment. The electronic device includes: A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application; A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the gene mutation detection method in the embodiments of this application; An input / output interface 903, which is used to implement information input and output; A communication interface 904, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904); Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.
[0106] The embodiments of this application also provide a computer program product, which includes a computer program. The processor of the computer device reads and executes this computer program, so that the computer device executes to implement the above gene mutation detection method.
[0107] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned gene mutation detection method.
[0108] As a non-transitory computer-readable storage medium, a memory can be used to store a non-transitory software program and a non-transitory computer-executable program. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0109] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0110] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.
[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0112] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0113] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0114] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0115] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0116] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0118] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0119] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A method for detecting gene mutations, characterized in that, The method includes: Obtaining first long-read sequencing data of a sample to be measured, and aligning the first long-read sequencing data with a reference genome to obtain first alignment data; Determining a region of a target gene in the reference genome, and calculating a first sequencing depth of the target gene according to the determined region and the first alignment data, where the target gene includes a first sub-target gene and a second sub-target gene; Determining a first base distribution ratio between the first sub-target gene and the second sub-target gene, and calculating a first sub-sequencing depth of the first sub-target gene and a second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth; Invoking a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain a target copy number variation type of the sample to be measured.
2. The method according to claim 1, wherein The step of invoking the pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be measured includes: Performing depth correction on the first sub-sequencing depth and the second sub-sequencing depth respectively according to a second sequencing depth of a reference gene; Invoking the detection model to perform copy number variation detection based on the depth-corrected first sub-sequencing depth and the depth-corrected second sub-sequencing depth to obtain the target copy number variation type.
3. The method according to claim 1, wherein The training process of the detection model includes the following steps: Obtaining second long-read sequencing data of a training sample, and aligning the second long-read sequencing data with the reference genome to obtain second alignment data; Extracting read segment data in regions where the first sub-target gene and the second sub-target gene are located from the second alignment data respectively, and re-aligning the read segment data with a reference sequence of the first sub-target gene to obtain re-alignment data; Determining a second base distribution ratio between the first sub-target gene and the second sub-target gene according to the re-alignment data, calculating a third sequencing depth of the target gene according to the re-alignment data, and calculating a fourth sequencing depth of a reference gene according to the second alignment data; Determining a third sub-sequencing depth of the first sub-target gene and a fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth; Training the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth.
4. The method according to claim 3, wherein The step of calculating the third sequencing depth of the target gene according to the re-alignment data includes: Calculating a preliminary sequencing depth of the target gene according to the re-alignment data; Calculating a sample depth ratio according to the preliminary sequencing depth and the fourth sequencing depth, and performing normalization processing on the sample depth ratio according to a reference depth ratio of a reference sample corresponding to the training sample to obtain the third sequencing depth.
5. The method according to claim 3, wherein The step of training the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth includes: Perform depth correction on the third sub-sequencing depth and the fourth sub-sequencing depth respectively according to the fourth sequencing depth; Invoke the detection model to determine the first copy number status corresponding to the first sub-target gene based on the depth-corrected third sub-sequencing depth, and determine the second copy number status corresponding to the second sub-target gene based on the depth-corrected fourth sub-sequencing depth, so as to determine the predicted copy number variation type based on the first copy number status and the second copy number status; Train the detection model according to the standard copy number variation type of the training sample and the predicted copy number variation type.
6. The method according to claim 1, characterized in that, The obtaining of the first long-read sequencing data of the sample to be tested includes: Construct amplification primers for the target gene, where the amplification primers include first sub-amplification primers; Perform amplification processing on the sample to be tested according to the first sub-amplification primers, and perform sequencing processing on the amplification processing result to obtain the first long-read sequencing data.
7. The method according to claim 1, wherein The obtaining of the first long-read sequencing data of the sample to be tested and aligning the first long-read sequencing data with a reference genome to obtain first alignment data includes: Construct amplification primers for the target gene, where the amplification primers include first sub-amplification primers and second sub-amplification primers, and the length of the second sub-amplification primers is greater than the length of the first sub-amplification primers; Perform amplification processing on the sample to be tested according to the first sub-amplification primers and the second sub-amplification primers, and perform sequencing processing on the amplification processing result to obtain the first long-read sequencing data; Split out first sub-amplification data corresponding to the first sub-amplification primers from the first long-read sequencing data, and align the first sub-amplification data with the reference genome to obtain the first alignment data.
8. The method according to claim 7, wherein The method further includes: Split out second sub-amplification data corresponding to the second sub-amplification primers from the first long-read sequencing data, and align the second sub-amplification data with the reference genome to obtain third alignment data; Perform variant detection on the sample to be tested according to the third alignment data.
9. A gene mutation detection device, characterized in that, The device includes: An alignment unit, configured to obtain the first long-read sequencing data of the sample to be tested and align the first long-read sequencing data with a reference genome to obtain first alignment data; A first sequencing depth calculation unit, configured to determine the region of the target gene in the reference genome and calculate the first sequencing depth of the target gene according to the determined region and the first alignment data, where the target gene includes a first sub-target gene and a second sub-target gene; A second sequencing depth calculation unit, configured to determine the first base distribution ratio between the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth; A variant detection unit, configured to invoke a pre-trained detection model to perform copy number variant detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variant type of the sample to be tested.
10. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the gene mutation detection method described in any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the gene mutation detection method described in any one of claims 1 to 8.
12. A computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the gene mutation detection method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for detecting homologous sequences on basis of high-throughput sequencing
CN112513292A
Method and device for detecting alpha-globin gene copy number variation
CN114512187A
Method, equipment and medium for detecting SMN gene copy number variation
CN117153249A
Detection method and device suitable for multiple mutation types of multiple disease-related genes, computer equipment, storage medium and computer program product
CN119068981A
Copy number variation detection process method based on sequencing depth
CN119091954A
Cited By
Weighted voting-based cancer-related gene intelligent analysis method and apparatus, and medium
CN121789766A