Gene mutation detection method and device, electronic device and storage medium
By combining long-read sequencing technology and pre-trained models with deep correction of internal reference genes, the accuracy problem of SMN1 and SMN2 gene variation detection was solved, and efficient and reliable copy number variation detection was achieved.
Patent Information
- Application Number
- CN202510829915.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing methods for detecting mutations in the SMN1 and SMN2 genes have the problem of low detection accuracy, especially the difficulty in distinguishing between the SMN1 and SMN2 genes, and the operation is complex and highly subjective.
Long-read sequencing technology is used to obtain the sequencing depth of the samples to be tested, and copy number variation detection is performed based on the sequencing depth of the SMN1 and SMN2 genes through a pre-trained detection model. Combined with the depth correction of the internal reference gene, the Gaussian mixture model (GMM) is used to predict the copy number variation type.
The accuracy of SMN1 and SMN2 gene copy number variation detection is improved, the operation process is simplified, subjective errors are reduced, and the reliability of detection is enhanced.
Smart Images

Figure CN120340601B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of bioinformatics, and in particular to a gene variation detection method and device, an electronic device, and a storage medium. Background Art
[0002] Spinal muscular atrophy (SMA) is a genetic disease caused by a gene defect, primarily due to copy number variations in the SMN1 gene. Approximately 95% of patients develop the disease due to homozygous deletion of the SMN1 gene. The SMN2 gene compensates for SMN1 loss, and the greater the number of SMN2 copies present, the milder the disease. Therefore, mutation testing for both SMN1 and SMN2 is crucial.
[0003] However, the SMN1 gene and SMN2 gene mutation detection methods in related technologies have the problem of low detection accuracy. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to provide a gene variation detection method and device, electronic equipment and storage medium, aiming to improve the accuracy of gene variation detection.
[0005] To achieve the above objectives, the first aspect of the embodiments of the present application provides a method for detecting gene mutations, the method comprising:
[0006] Acquire first long-read sequencing data of the sample to be tested, and align the first long-read sequencing data with the reference genome to obtain first aligned data;
[0007] Determining a region of a target gene in the reference genome, and calculating a first sequencing depth of the target gene based on the determined region and the first alignment data, wherein the target gene includes a first sub-target gene and a second sub-target gene;
[0008] Determine the first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth;
[0009] The pre-trained detection model is called to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested.
[0010] In some embodiments, calling a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain a target copy number variation type of the sample to be tested includes:
[0011] Performing depth correction on the first sub-sequencing depth and the second sub-sequencing depth according to the second sequencing depth of the internal reference gene;
[0012] The detection model is called to perform copy number variation detection based on the first sub-sequencing depth after depth correction and the second sub-sequencing depth after depth correction to obtain the target copy number variation type.
[0013] In some embodiments, the training process of the detection model includes the following steps:
[0014] Acquire second long-read sequencing data of the training sample, and align the second long-read sequencing data with the reference genome to obtain second aligned data;
[0015] Extracting read data of the regions where the first sub-target gene and the second sub-target gene are located from the second aligned data, respectively, and realigning the read data with the reference sequence of the first sub-target gene to obtain realigned data;
[0016] Determining the second base distribution ratio of the first sub-target gene and the second sub-target gene according to the re-alignment data, calculating the third sequencing depth of the target gene according to the re-alignment data, and calculating the fourth sequencing depth of the internal reference gene according to the second alignment data;
[0017] Determining a third sub-sequencing depth of the first sub-target gene and a fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth;
[0018] The detection model is trained based on the third sub-sequencing depth, the fourth sub-sequencing depth and the fourth sequencing depth.
[0019] In some embodiments, calculating the third sequencing depth of the target gene based on the re-aligned data includes:
[0020] Calculating the initial sequencing depth of the target gene based on the re-alignment data;
[0021] A sample depth ratio is calculated based on the preliminary sequencing depth and the fourth sequencing depth, and the sample depth ratio is normalized based on a reference depth ratio of an internal reference sample corresponding to the training sample to obtain the third sequencing depth.
[0022] In some embodiments, training the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth includes:
[0023] Performing depth correction on the third sub-sequencing depth and the fourth sub-sequencing depth respectively according to the fourth sequencing depth;
[0024] Calling the detection model to determine a first copy number state corresponding to the first sub-target gene based on the third sub-sequencing depth after depth correction, and determining a second copy number state corresponding to the second sub-target gene based on the fourth sub-sequencing depth after depth correction, so as to determine the predicted copy number variation type based on the first copy number state and the second copy number state;
[0025] The detection model is trained according to the standard copy number variation type and the predicted copy number variation type of the training sample.
[0026] In some embodiments, obtaining first long-read sequencing data of the sample to be tested includes:
[0027] constructing amplification primers for the target gene, wherein the amplification primers include a first sub-amplification primer;
[0028] The sample to be tested is amplified according to the first sub-amplification primer, and the amplification result is sequenced to obtain the first long-read sequencing data.
[0029] In some embodiments, obtaining first long-read sequencing data of the sample to be tested, and aligning the first long-read sequencing data with a reference genome to obtain first aligned data includes:
[0030] Constructing amplification primers for the target gene, wherein the amplification primers include a first sub-amplification primer and a second sub-amplification primer, wherein the length of the second sub-amplification primer is greater than the length of the first sub-amplification primer;
[0031] Amplifying the sample to be tested according to the first sub-amplification primer and the second sub-amplification primer, and sequencing the amplification result to obtain the first long-read sequencing data;
[0032] Splitting first sub-amplification data corresponding to the first sub-amplification primer from the second long-read sequencing data, and aligning the first sub-amplification data with the reference genome to obtain first aligned data;
[0033] In some embodiments, the method further comprises:
[0034] Splitting second sub-amplification data corresponding to the second sub-amplification primer from the first long-read sequencing data, and aligning the second sub-amplification data with the reference genome to obtain the third aligned data;
[0035] Perform variation detection on the sample to be tested according to the third comparison data.
[0036] To achieve the above objectives, a second aspect of the embodiments of the present application provides a gene mutation detection device, comprising:
[0037] An alignment unit is used to obtain first long-read sequencing data of the sample to be tested, and align the first long-read sequencing data with the reference genome to obtain first alignment data;
[0038] a first sequencing depth calculation unit, configured to determine a region of a target gene in a reference genome, and calculate a first sequencing depth of the target gene based on the determined region and the first alignment data, wherein the target gene includes a first sub-target gene and a second sub-target gene;
[0039] a second sequencing depth calculation unit, configured to determine a first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculate a first sub-sequencing depth of the first sub-target gene and a second sub-sequencing depth of the second sub-target gene based on the first base distribution ratio and the first sequencing depth;
[0040] The variation detection unit is used to call the pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested.
[0041] To achieve the above-mentioned purpose, a third aspect of an embodiment of the present application proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0042] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0043] To achieve the above-mentioned purpose, the fifth aspect of the embodiments of the present application proposes a computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the method described in the first aspect.
[0044] The gene variation detection method and device, electronic device and storage medium proposed in the embodiment of the present application can obtain the sequencing depth of the first sub-target gene and the sequencing depth of the second sub-target gene based on the first long-read sequencing data and the first alignment data of the sample to be tested, and copy number variation detection can be performed through the pre-trained detection model and the above two sequencing depths. It can be seen that the embodiment of the present application can realize copy number variation detection of the SMN1 gene (i.e., the first sub-target gene) and the SMN2 gene (i.e., the second sub-target gene) based on the detection model. In addition, since the embodiment of the present application is based on the sequencing depth obtained by long-read sequencing data, based on the characteristics of long-read sequencing (long-read sequencing can read longer DNA fragments), the specific sites of the SMN1 gene and the SMN2 gene can be determined based on the long-read sequencing data, so that the SMN1 gene and the SMN2 gene can be distinguished based on the specific sites, thereby improving the accuracy of copy number variation detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.
[0046] Figure 1 This is a flow chart of the gene variation detection method provided in the embodiments of the present application;
[0047] Figure 2 Schematic diagram of 15 types of copy number variation provided in the examples of this application;
[0048] Figure 3 This is a flow chart of the detection model training process provided by an embodiment of the present application;
[0049] Figure 4 is a schematic diagram of a gene variation detection method provided in an embodiment of the present application;
[0050] Figure 5 Schematic diagram of sequencing quality control reports and clinical phenotype information of 35 SMA samples provided in the examples of this application;
[0051] Figure 6 Schematic diagram of the variation detection results of the long amplicon of the standard NA12878 provided in the examples of the present application;
[0052] Figure 7 Schematic diagram of SMA copy number variation types predicted by the detection model provided in the examples of the present application;
[0053] Figure 8 is a schematic diagram of a gene variation detection device provided in an embodiment of the present application;
[0054] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.
[0056] First, the terms used in this application are explained.
[0057] Short-read sequencing breaks DNA into short fragments of 50-300 bases for high-throughput parallel sequencing. It offers high single-base accuracy, low cost, and high throughput. Short-read sequencing is suitable for scenarios such as whole-genome sequencing, transcriptome analysis, and epigenetic research. However, short-read sequencing is susceptible to interference from repetitive sequences when assembling complex genomes, making it difficult to resolve structural variations in long fragments.
[0058] Long-read sequencing: This technology captures and sequences DNA fragments ranging from thousands to tens of thousands of bases in length. The core advantage of long reads is that they span complex regions of the genome (such as repetitive sequences, structural variants, or highly polymorphic regions), significantly improving the completeness and accuracy of genome assemblies. Long-read sequencing is crucial for studying gene structural variation, identifying transcript isoforms, and performing metagenomic analysis.
[0059] Sequencing depth, also known as coverage depth, is a key metric for measuring genomic sequencing data quality. It represents the average number of times a specific genomic region has been sequenced. Sequencing depth is calculated by mapping sequencing reads to a reference genome using alignment tools (such as BWA and Bowtie). Analysis tools then count the number of times each site has been covered, which is then used to calculate sequencing depth.
[0060] Standardization is a data preprocessing method. Its core is to convert raw data of varying dimensions or magnitudes into a unified standard scale through mathematical transformation. Standardized data conforms to a standard normal distribution with a mean of 0 and a standard deviation of 1. Standardization eliminates differences in units and magnitudes between features, improving data comparability.
[0061] Spinal muscular atrophy (SMA) is a common autosomal recessive genetic disorder. SMA is a genetic disorder caused by a gene defect, primarily due to copy number variation in the SMN1 gene. Approximately 95% of patients develop the disease due to homozygous deletion of the SMN1 gene, while those with heterozygous deletions are SMA carriers. The SMN2 gene compensates for SMN1 loss, and a higher copy number indicates a milder disease. Because the SMN1 and SMN2 genes are highly homologous, differing only at five sites, particularly the C / T difference in exon 7, variant detection is challenging. SMN1 gene copy number detection technologies are mainly divided into Sanger sequencing, multiplex ligation-dependent probe amplification (MLPA) MPLA (the current gold standard), traditional fluorescence PCR technology (real-time quantitative polymerase chain reaction, real-time qPCR), HRM high-resolution melting curve analysis (HRM) and next-generation high-throughput sequencing technology (next-generation sequencing, NGS).
[0062] Sanger sequencing is a method based on dideoxyribonucleotide (ddNTP) end-termination. By incorporating color-labeled ddNTPs into the reaction, DNA extension is randomly terminated, generating nucleic acid fragments of varying lengths. These fragments are then separated and color-coded using capillary electrophoresis to determine the nucleotide sequence. Sanger sequencing is a classic method for DNA sequence analysis, used to detect rare and complex mutations in the SMN1 gene and remains the gold standard for mutation detection. While Sanger sequencing can identify unknown SNPs and determine their type and location, it cannot distinguish between healthy individuals and carriers, nor can it quantify SMN2 copy number. The basic principle of MLPA involves hybridization of a probe to target DNA. Following ligation and PCR amplification, the amplified products are separated by capillary electrophoresis, and analysis software is used to process the data to generate the results. However, MLPA has drawbacks such as its inability to detect point mutations and the unique SMN12+0 carrier condition, as well as its high cost, time consumption, cumbersome procedures, and susceptibility to contamination. The MGB probe-based real-time multiplex fluorescence quantitative PCR method determines SMN1 gene deletion by relative quantification of the SMN1 gene against an internal reference gene and calculating ΔΔCt values. This method relies on ΔΔCt calculations, which involve reading Ct values across different detection channels. This method is highly subjective and susceptible to operator influence. Furthermore, the SMN2 gene is highly homologous to SMN1, potentially interfering with the assay, leading to false-negative or false-positive results. Furthermore, high amplification efficiencies of both the target gene and the internal reference gene are required, otherwise the accuracy of the results may be compromised. High-resolution melting (HRM) technology combines PCR amplification with melting curve analysis, analyzing gene sequence variation by examining the melting curves of PCR amplification products. It is particularly well-suited for distinguishing single-base differences. HRM technology detects variations in melting temperature (Tm) based on differences in amplified product length and GC content, generating distinct melting curves that are used to determine gene variation and copy number. HRM technology places high demands on data analysis and requires software normalization, complicating result interpretation. Furthermore, HRM is sensitive to amplified product quality, and its results rely on the stability of PCR amplification efficiency. In addition, the detection sensitivity for low-frequency mutations is low and it is not suitable for detecting a small number or rare mutations. NGS technology has the advantages of high throughput and high sensitivity in genetic testing, but it often requires complex analytical methods to identify SMN1 gene copy number variations. This is because there are only a few base differences between the SMN1 and SMN2 gene sequences, making it difficult to distinguish between the SMN1 and SMN2 genes in sequencing data. Complex algorithms and analytical steps must be used to identify these subtle differences, but these methods are often difficult to interpret and their accuracy may also be limited.
[0063] In summary, the detection methods in related technologies have the following problems:
[0064] 1. Data analysis is complex. Because the SMN1 and SMN2 genes are highly homologous, complex analysis algorithms are required to process the data. This increases the difficulty of analysis and the potential for error.
[0065] 2. Limited detection capabilities. Some related technologies are primarily used to detect SMN1 copy number variations and are not accurate enough for complex mutation types or quantification of SMN2 gene copy number.
[0066] 3. High subjectivity. Traditional fluorescent PCR technology and HRM high-resolution melting curve analysis involve significant subjectivity in threshold setting and data interpretation, which can lead to inaccurate results.
[0067] 4. Complex operation. The relevant technical experimental steps are cumbersome and require precise operation to avoid contamination. It also has high requirements for equipment and operators.
[0068] Based on this, the embodiments of the present application provide a gene variation detection method and device, an electronic device, and a storage medium to improve the accuracy of gene variation detection.
[0069] The gene variation detection method provided in the embodiment of the present application relates to the field of bioinformatics. The gene variation detection method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the gene variation detection method, etc., but is not limited to the above forms.
[0070] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0071] The following describes the gene variation detection method provided in the examples of the present application.
[0072] Reference Figure 1 In some embodiments, the gene variation detection method provided in the embodiments of the present application includes but is not limited to steps S101 to S104.
[0073] Step S101, obtaining first long-read sequencing data of a sample to be tested, and aligning the first long-read sequencing data with a reference genome to obtain first aligned data;
[0074] Step S102, determining a region of the target gene in the reference genome, and calculating a first sequencing depth of the target gene based on the determined region and the first alignment data;
[0075] Step S103, determining the first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculating the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene based on the first base distribution ratio and the first sequencing depth;
[0076] Step S104 , calling a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain a target copy number variation type of the sample to be tested.
[0077] In steps S101 to S104 shown in the embodiment of the present application, the sequencing depth of the first sub-target gene and the sequencing depth of the second sub-target gene can be obtained based on the first long-read sequencing data and the first alignment data of the sample to be tested, and copy number variation detection can be performed through the pre-trained detection model and the above two sequencing depths. It can be seen that the embodiment of the present application can realize copy number variation detection of the SMN1 gene (i.e., the first sub-target gene) and the SMN2 gene (i.e., the second sub-target gene) based on the detection model. In addition, since the embodiment of the present application is based on the sequencing depth obtained by long-read sequencing data, based on the characteristics of long-read sequencing (long-read sequencing can read longer DNA fragments), the specific sites of the SMN1 gene and the SMN2 gene can be determined based on the long-read sequencing data, so that the SMN1 gene and the SMN2 gene can be distinguished based on the specific sites, thereby improving the accuracy of copy number variation detection.
[0078] In step S101 of some embodiments, the sample to be tested may be an SMA sample to be tested for genetic variation. PCR amplification and long-read sequencing are performed on the DNA of the sample to be tested, and data quality control is performed on the sequencing results to obtain first long-read sequencing data. The reference genome may be a human reference genome (e.g., GRCh38). The first long-read sequencing data is aligned with the reference genome, using the reference genome as a benchmark to identify potential variation information in the sample to be tested, thereby obtaining first alignment data.
[0079] In some embodiments, the step S101 of "obtaining first long read sequencing data of the sample to be tested" includes the following sub-steps:
[0080] constructing amplification primers for the target gene, wherein the amplification primers include a first sub-amplification primer;
[0081] The sample to be tested is amplified according to the first sub-amplification primer, and the amplification result is sequenced to obtain first long-read sequencing data.
[0082] In the embodiments of the present application, an amplification primer may refer to a short single-stranded DNA fragment amplified by PCR. Amplification primers are used to guide DNA polymerase to replicate a target DNA sequence starting at a specific location, thereby achieving specific amplification of the target DNA. The amplification primers of the present application may include short amplification primers (i.e., first sub-amplification primers), which may include primers for amplification of exons 7 and 8 of the SMN1 and SMN2 genes. Thus, PCR amplification of the test sample is performed based on the short amplification primers, and the PCR amplification results are sequenced and data quality controlled to obtain first long-read sequencing data.
[0083] It is understood that, when the amplification primers only include short amplification primers, the amplification products of the PCR also only include short amplicons. In this way, the comparison operation can refer to comparing the short amplicons after data quality control with the reference genome.
[0084] In other embodiments, the amplification primers may also include long amplification primers (i.e., second sub-amplification primers, the length of which is greater than the length of the first sub-amplification primer). In this case, step S101 may include the following sub-steps:
[0085] Constructing amplification primers for the target gene, the amplification primers comprising a first sub-amplification primer and a second sub-amplification primer, wherein the length of the second sub-amplification primer is greater than the length of the first sub-amplification primer;
[0086] Amplifying the sample to be tested using the first sub-amplification primer and the second sub-amplification primer, and sequencing the amplification result to obtain first long-read sequencing data;
[0087] First sub-amplification data corresponding to the first sub-amplification primer is split from the first long-read sequencing data, and the first sub-amplification data is aligned with the reference genome to obtain first aligned data.
[0088] In this embodiment of the present application, the amplification primers may include a long amplification primer (i.e., the second sub-amplification primer) and a short amplification primer (i.e., the first sub-amplification primer, which is shorter than the second sub-amplification primer). The short amplification primers are the same as those described in the above embodiment, and the long amplification primers may include amplification primers for exons 2b to 8 of the SMN1 and SMN2 genes. It is understood that the amplification primers may also include amplification primers for an internal reference gene with no copy number (i.e., exons 4 to 6 of the CFTR gene).
[0089] In this way, PCR amplification is performed on the sample to be tested based on the long amplification primer and the short amplification primer, and the PCR amplification results are sequenced and data quality controlled to obtain the first long read sequencing data.
[0090] It is understood that when the amplification primers include long amplification primers and short amplification primers, the PCR amplification product includes a long amplicon corresponding to the long amplification primer and a short amplicon corresponding to the short amplification primer (i.e., first sub-amplification data). Thus, software such as Citrus can be used to separate the corresponding long and short amplicons from the first long-read sequencing data based on the long and short amplification primers. In this case, the alignment operation can refer to aligning the separated short amplicons with the reference genome. The role of the long amplicon will be described in the following examples.
[0091] In step S102 of some embodiments, the target gene may refer to a gene related to SMA (hereinafter referred to as the SMA gene, which includes the SMN1 gene and the SMN2 gene, wherein the SMN1 gene may be referred to as the first sub-target gene, and the SMN2 gene may be referred to as the second sub-target gene). The location (i.e., the region) of the target gene in the reference genome may be obtained based on the relevant knowledge base, and the sequencing depth (i.e., the first sequencing depth) of the target gene may be calculated based on the determined location and the first alignment data.
[0092] In step S103 of some embodiments, the SMN1 and SMN2 genes are highly homologous, differing only at five sites, particularly the C / T ratio in exon 7. For example, at the c.840C>T site in exon 7, the SMN1 gene typically has a C (cytosine), while the SMN2 gene typically has a T (thymine). Therefore, by calculating the base distribution ratio (i.e., the first base distribution ratio) at the c.840C>T site in exon 7 in the SMN1 and SMN2 genes, the copy number ratio of the SMN1 and SMN2 genes can be determined. Since the first sequencing depth calculated in the above step is the sum of the SMN1 gene measurement depth and the SMN2 gene sequencing depth, the sequencing depth of the SMN1 gene (i.e., the first sub-sequencing depth) and the sequencing depth of the SMN2 gene (i.e., the second sub-sequencing depth) can be calculated based on the first base distribution ratio and the first sequencing depth.
[0093] In step S104 of some embodiments, the detection model may refer to a pre-trained model with copy number variation detection capabilities. The detection model is called, and the first sub-sequencing depth and the second sub-sequencing depth are used as input data of the detection model to obtain the target copy number variation type of the sample to be tested. Figure 2 As shown in the examples of the present application, 15 types of copy number variations (i.e. Figure 2 Types A to O are shown in the figure). In this case, the detection model can be referred to as a Gaussian mixture model (GMM). A GMM is a probabilistic model that assumes a mixture of multiple Gaussian distributions, each representing a different copy number state. During prediction, the GMM model assigns different probabilities to different copy number states based on the first and second subsequencing depths. The target copy number variation type is then determined based on the most probable copy number state corresponding to the SMN1 gene and the most probable copy number state corresponding to the SMN2 gene. This means that the target copy number variation type refers to one of these 15 types.
[0094] It is understood that in some embodiments, other analyses can be performed on the sample to be tested based on the separated long amplicons. For example, in some embodiments, the gene variation detection method provided in the embodiments of the present application can also include the following steps:
[0095] Splitting second sub-amplification data corresponding to the second sub-amplification primer from the first long-read sequencing data, and aligning the second sub-amplification data with the reference genome to obtain third aligned data;
[0096] The variation detection is performed on the sample to be tested according to the third comparison data.
[0097] In this embodiment of the present application, variant detection can be performed on the sample to be tested based on the separated long amplicon (i.e., the second sub-amplification data). Specifically, the long amplicon is aligned with the reference genome to obtain third alignment data (in BAM format). Clair3 software, which is suitable for long-read sequencing data, is then used to perform variant detection on this third alignment data. In this way, variant detection results for the sample to be tested can be obtained. It is understood that the variant detection results can be used to study other mutation types in the SMA gene.
[0098] It is understandable that PCR detection methods in related art cannot cover both copy number typing and variant detection. However, the present application, leveraging the characteristics of long-read sequencing data, can simultaneously perform copy number typing and variant detection. For example, when both copy number typing and variant detection are required, PCR amplification is performed using short and long amplification primers. In this case, the DNA fragments obtained after PCR amplification include short amplicons and long amplicons. Short amplicons are used for typing and analysis of copy number variations in the SMA gene, while long amplicons are used for variant detection and analysis of SNPs (single nucleotide polymorphisms) and INDELs (insertions or deletions). In this case, the sequencing data after data quality control needs to be split so that different detection and analysis can be performed based on the split long and short amplicons. When only copy number variation detection of the SMA gene is required, PCR amplification is performed using short amplification primers. In this case, the PCR amplification products only contain short amplicons. In this case, parameter settings can be used to avoid splitting the sequencing data after data quality control, allowing downstream analysis to be performed directly on the sequencing data after data quality control (i.e., the first long-read sequencing data).
[0099] In some embodiments, step S104 may include but is not limited to the following sub-steps:
[0100] Perform depth correction on the first sub-sequencing depth and the second sub-sequencing depth according to the second sequencing depth of the internal reference gene;
[0101] The detection model is called to perform copy number variation detection based on the depth-corrected first sub-sequencing depth and the depth-corrected second sub-sequencing depth to obtain the target copy number variation type.
[0102] In this embodiment of the present application, an internal reference gene (e.g., the CFTR gene) is determined based on the sample to be tested, and a second sequencing depth for the internal reference gene is obtained. The first subsequencing depth of the SMN1 gene and the second subsequencing depth of the SMN2 gene are normalized (i.e., depth-corrected) based on the second sequencing depth to eliminate the effects of differences in sequencing conditions and depth across samples. A GMM model is then used to detect copy number variations based on the depth-corrected first and second subsequencing depths to determine the target copy number variation type.
[0103] In the embodiment of the present application, the sequencing depth of the SMA gene (including the SMN1 gene and the SMN2 gene) is first normalized based on the sequencing depth of the internal reference gene, and then copy number variation detection is performed based on the normalized sequencing depth. This can improve the accuracy of the sequencing depth, and thus improve the accuracy of copy number variation detection.
[0104] The following describes the training process of the detection model.
[0105] Reference Figure 3 In some embodiments, the training process of the detection model includes but is not limited to steps S301 to S305.
[0106] Step S301, obtaining second long-read sequencing data of the training sample, and aligning the second long-read sequencing data with the reference genome to obtain second aligned data;
[0107] Step S302: extracting read data of the regions where the first sub-target gene and the second sub-target gene are located from the second aligned data, and realigning the read data with the reference sequence of the first sub-target gene to obtain realigned data;
[0108] Step S303, determining the second base distribution ratio of the first sub-target gene and the second sub-target gene based on the re-alignment data, and calculating the third sequencing depth of the target gene based on the re-alignment data, and calculating the fourth sequencing depth of the internal reference gene based on the second alignment data;
[0109] Step S304: determining a third sub-sequencing depth of the first sub-target gene and a fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth;
[0110] Step S305: training a detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth.
[0111] In step S301 of some embodiments, the training sample may refer to an SMA sample used to train the detection model. The DNA of the training sample is PCR amplified and long-read sequenced, and the sequencing results are subjected to data quality control to obtain second long-read sequencing data. Data quality control may refer to the use of relevant tools to perform quality control on the sequencing data to ensure the accuracy and reliability of subsequent analyses based on the sequencing data. Specifically, after PCR amplification, a sequencing library can be prepared using the CycloneSEQ 24 Barcoded Library Preparation Kit. After obtaining the sequencing library, the DNA of the training sample is sequenced using the CycloneSEQ WT Sequencing Kit, the CycloneSEQ WT Sequencing Chip, and the CycloneSEQ-WT02 nanopore gene sequencer adapted for this sequencing chip to obtain raw sequencing data (i.e., sequencing results). The raw sequencing data (in FASTQ format) is input into Bamboo software to generate a data quality control report file. The data quality control report file can be used to assess data quality and provide a basis for subsequent analysis. After performing data quality control on the raw sequencing data, the second long-read sequencing data can be obtained.
[0112] Align the second long-read sequencing data to a reference genome (e.g., the human reference genome GRCh38) and generate the alignment results (i.e., secondary alignment data, a SAM file) using minimap2, a software specifically designed for processing long reads. This secondary alignment data can include read positions, alignment quality, and CIGAR strings. Additionally, the alignment results can be converted to the more compact BAM format using samtools.
[0113] In some embodiments, step S302 involves obtaining the locations of the SMN1 gene (i.e., the first sub-target gene) and the SMN2 gene (i.e., the second sub-target gene) in a reference genome (e.g., the human reference genome GRCh38) from a relevant knowledge base (e.g., the UCSC Genome Browser). Based on the obtained locations, tools such as samtools are used to extract reads aligned to the SMN1 gene region and reads aligned to the SMN2 gene region from the second alignment data (i.e., the alignment results of the short amplicon with the reference genome). The obtained reads are then realigned with the reference sequence of the SMN1 gene, i.e., the obtained reads are realigned to a sequence file containing only the SMN1 gene, thereby generating realigned data (a BAM file).
[0114] In step S303 of some embodiments, the base distribution of the c.840C>T site in exon 7 of the SMN1 and SMN2 genes is statistically analyzed based on the realignment data. Based on this base distribution, the C / T base distribution ratio (i.e., the second base distribution ratio) can be obtained. Based on this second base distribution ratio, the copy number ratio of the SMN1 and SMN2 genes can be determined. It is understood that because the SMN1 and SMN2 genes are highly homologous, they differ by only five bases, with the c.840C>T site in exon 7 affecting the splicing patterns of the genes. Therefore, when the extracted read data from the SMN1 and SMN2 gene regions are aligned to the reference genome with only five base differences, the two genes may not be accurately distinguished (because the SMN1 gene may be incorrectly aligned to the reference region of the SMN2 gene in the reference genome, and the SMN2 gene may be incorrectly aligned to the reference region of the SMN1 gene in the reference genome). Therefore, in the embodiment of the present application, the extracted read data of the two genes are re-aligned to the reference sequence of the SMN1 gene. In this way, the ratio of SMN1 gene:SMN2 gene under the alignment to the SMN1 gene reference sequence can be determined based on the base difference characteristics of the two genes (i.e., SMN1 gene and SMN2 gene) at the c.840C>T site.
[0115] The third sequencing depth for the SMA gene was calculated based on the realignment data, and the fourth sequencing depth for the internal reference gene (e.g., CFTR) was calculated based on the whole-genome alignment file (i.e., the second alignment data). It should be understood that to ensure that the sequencing data is not affected by the length of the amplified fragments, the statistical value used in calculating the sequencing depth (i.e., the third and fourth sequencing depths) is the average depth information. Furthermore, the third sequencing depth calculated based on the realignment data is the sum of the sequencing depths of the SMN1 and SMN2 genes.
[0116] In some embodiments, the step S303 of “calculating the third sequencing depth of the target gene based on the re-aligned data” includes the following sub-steps:
[0117] The preliminary sequencing depth of the target gene was calculated based on the re-alignment data;
[0118] The sample depth ratio is calculated based on the preliminary sequencing depth and the fourth sequencing depth, and the sample depth ratio is normalized according to the reference depth ratio of the internal reference sample corresponding to the training sample to obtain the third sequencing depth.
[0119] In the present embodiment, the third sequencing depth can be the sequencing depth obtained after normalization. In this way, the sequencing depth of the SMA gene directly calculated based on the re-aligned data is used as the initial sequencing depth. According to the initial sequencing depth and fourth sequencing depth Calculate the sample depth ratio In addition, to eliminate fluctuations in sample sequencing results caused by different batches of experiments, that is, to correct for batch effects, an internal reference sample can be introduced. This internal reference sample is from the same batch as the training samples. Furthermore, the reference depth ratio of this internal reference sample can be calculated using the same method as above. In this way, the sample depth ratio can be normalized based on the reference depth ratio, and a third sequencing depth can be obtained based on the normalization result.
[0120] In step S304 of some embodiments, as described above, the third sequencing depth calculated based on the re-aligned data is the sum of the SMN1 gene sequencing depth and the SMN2 gene sequencing depth. Thus, combined with the second base distribution ratio, the sequencing depth corresponding to the SMN1 gene (i.e., the third sub-sequencing depth) and the sequencing depth corresponding to the SMN2 gene (i.e., the fourth sequencing depth) can be calculated.
[0121] In step S305 of some embodiments, the copy number detection error of the detection model can be determined based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth. In this way, the detection model parameters can be adjusted based on the copy number detection error, thereby achieving training of the detection model.
[0122] In some embodiments, step S305 may include the following sub-steps:
[0123] Perform depth correction on the third sub-sequencing depth and the fourth sub-sequencing depth according to the fourth sequencing depth;
[0124] Calling the detection model to determine a first copy number state corresponding to the first sub-target gene based on the depth-corrected third sub-sequencing depth, and determining a second copy number state corresponding to the second sub-target gene based on the depth-corrected fourth sub-sequencing depth, so as to determine a predicted copy number variation type based on the first copy number state and the second copy number state;
[0125] The detection model is trained based on the standard copy number variation types and predicted copy number variation types of the training samples.
[0126] In the embodiment of the present application, in order to eliminate the influence of sequencing conditions and depth differences in different samples, the sequencing depth of the SMN1 gene (i.e., the third sub-sequencing depth) and the sequencing depth of the SMN2 gene (i.e., the fourth sub-sequencing depth) can be normalized based on the sequencing depth of the internal reference gene (i.e., the fourth sequencing depth) to achieve depth correction. In this way, the detection model can be called to perform copy number variation detection on the normalized third sub-sequencing depth and the normalized fourth sub-sequencing depth to obtain a predicted copy number variation type. It is understandable that the predicted copy number variation type can be as follows Figure 2 One of 15 types shown.
[0127] Specifically, in the embodiment of the present application, the detection model can be a GMM model, which can include Gaussian distributions corresponding to different copy number states (including 0 copies, 1 copy, 2 copies, etc.). The detection model can assign probabilities to different copy number states based on the sequencing depth of the SMN1 gene (i.e., the third subsequencing depth), and use the copy number state with the highest probability as the first copy number state of the SMN1 gene. Similarly, the second copy number state of the SMN2 gene can be obtained. In this way, Figure 2 As shown, a predicted copy number variation type can be determined based on the first copy number state and the second copy number state.
[0128] In this manner, supervised training of the detection model is performed based on the predicted copy number variation type and the actual copy number of the training sample (i.e., the standard copy number variation type). Specifically, the copy number detection error of the detection model can be calculated based on the predicted copy number variation type and the standard copy number variation type. In this manner, parameters of the detection model can be adjusted based on this copy number detection error.
[0129] The gene variation detection method provided in the embodiments of the present application has the following beneficial effects:
[0130] 1. Improved detection accuracy. The embodiments of the present application perform relevant detection based on long-read sequencing data, and long-read sequencing data can provide more comprehensive genomic result information. Specifically, compared with short-read sequencing, long-read sequencing has obvious advantages in detecting complex gene regions, repetitive sequences, etc. In this way, for copy number variation detection of SMN1 gene and SMN2 gene, more accurate resolution can be achieved based on long-read sequencing data, thereby improving the quality and accuracy of detection. In addition, the embodiments of the present application also perform normalization based on the sequencing depth of the internal reference gene, so that the stability of the detection can be maintained under different experimental conditions, thereby improving the detection accuracy.
[0131] 2. Improved detection efficiency. Because long-read sequencing technology generates longer sequence fragments, it can reduce the difficulty of alignment in complex genomic regions, thereby reducing data analysis and processing time. In addition, by optimizing alignment algorithms (such as minimap2) and model building, the entire detection process can be made more efficient. This allows large numbers of samples to be analyzed in a shorter time.
[0132] 3. Ease of operation: The mixed Gaussian model used in the examples of this application can not only automatically distinguish samples with different copy numbers, but also has strong adaptability and can handle samples with complex genotypes, reducing tedious manual comparison and result interpretation operations.
[0133] 4. Improved comprehensiveness of detection. The present embodiment can perform PCR amplification based on long and short amplification primers. This allows for copy number variation detection and other variation detection.
[0134] Reference Figure 4 In a specific embodiment, the gene variation detection method provided in the present application example may include the following steps:
[0135] 1. PCR amplification and nanopore sequencing.
[0136] As shown in Table 1 below, long amplification primers, short amplification primers, and primers for amplifying an internal reference gene without copy number variation were designed. As shown in Table 1 below, the short amplification primers can be primers for amplifying exons 7 and 8 of the SMN1 and SMN2 genes, and the internal reference gene amplification primers can be primers for amplifying exons 4 and 6 of the CFTR gene.
[0137]
[0138] Table 1
[0139] Multiple PCR amplifications were performed using the designed primers, using 15 clinically obtained standards, to obtain DNA from 35 SMA samples. The amplification reaction volume (50 μL) consisted of: 25 μL 2× ApexHF CL Buffer, 1 μL 0.2 μM primer mix, 30 ng template DNA, 1 μL ApexHF HS DNA Polymerase-CL, and nuclease-free water to 50 μL. The amplification protocol was: 94°C for 1 minute; 98°C for 10 seconds, 60°C for 15 seconds, and 68°C for 4 minutes, for 25 cycles; 68°C for 2 minutes; and incubation at 4°C.
[0140] After amplification, sequencing libraries can be prepared using the CycloneSEQ 24 Barcoded Library Preparation Kit (CycloneSEQ, H940-000018). Sequencing data can then be generated from SMA samples using the CycloneSEQ WT Sequencing Kit (CycloneSEQ, H940-000016), the CycloneSEQ WT Sequencing Chip (CycloneSEQ, H930-000001-00), and the CycloneSEQ-WT02 nanopore sequencer (CycloneSEQ, H900-000001-00).
[0141] 2. Data quality control, primer splitting and read filtering.
[0142] First, the sequencing data obtained based on step 1 is converted into a FASTQ format file. Secondly, the Bamboo software is used to perform quality control on the sequencing data to evaluate whether the data volume is sufficient and to identify whether there are deviations in the sequencing process to ensure that the data meets the requirements of subsequent analysis. Then, the Citrus software is called to split the short amplicons and long amplicons in the sequencing data based on the primer sequences used in PCR amplification for subsequent analysis. Finally, in order to further improve the accuracy and analysis efficiency of the sequencing data of these 35 samples, the seq module in the seqkit toolkit is used to filter the sequencing data according to the read segment (ie, reads) length (≥4000bp) to remove reads that do not meet the requirements to optimize the quality of the sequencing data and lay the foundation for subsequent precise analysis. The sequencing quality control report and clinical phenotype information of the 35 SMA samples are shown below. Figure 5 shown.
[0143] 3. Genome comparison.
[0144] The purpose of step 3 is to align the sequencing data obtained in step 2 with a reference genome (such as the human reference genome GRCh38) to generate a BAM-formatted alignment file. Specifically, the input data can be the long and short amplicons obtained after data quality control, as well as the human reference genome GRCh38. Use the pipe character to connect the minimap2 and samtools commands to process the alignment of the long and short amplicons against the reference genome, respectively. Ultimately, a BAM file containing the alignment information (including a BAM file for the long and short amplicons) is generated.
[0145] 4. Perform mutation detection on long amplicons.
[0146] After completing the genome alignment, you can obtain the whole genome alignment result file (BAM file) of the long amplicon from step 3, and call the clair3 software for long-read sequencing data to perform variation detection on the alignment result file.
[0147] The results of the mutation detection of the standard NA12878 long amplicon are as follows Figure 6 shown. Figure 6 The table header includes the following: chromosome name or reference sequence name, position of the variant site on the reference genome, unique identifier of the variant (the table is empty and is indicated by "."), base or sequence on the reference genome, mutated base or sequence (i.e., sequence different from the reference genome), alignment quality score (indicating the confidence of the corresponding variant), whether the variant passes quality filtering (usually includes "PASS" to indicate passing filtering), model used by clair3 software for variant detection (F indicates that the clair3 software variant detection uses the "Full alignment" model, P indicates that the clair3 software variant detection uses the pileup model, different models will correspond to different identifiers here), data format field (GT represents genotype, GQ represents genotype quality, DP represents sequencing depth, AD represents allele depth, and AF represents allele frequency), and sample genotype information.
[0148] 5. Extract the reads corresponding to the SMN1 and SMN2 genes.
[0149] First, the specific locations of SMA genes (including SMN1 and SMN2) in the GRCh38 reference genome were obtained from the UCSC Genome Browser. Then, based on the coordinate information of the SMN1 and SMN2 genes, the samtools tool was used to extract reads aligned to the SMN1 and SMN2 gene regions from the whole-genome alignment results (BAM files corresponding to the short amplicon) obtained in Step 3.
[0150] 6. Re-compare.
[0151] Based on the reads obtained in step 5, re-align them to the reference sequence file containing only the SMN1 gene to generate an alignment result file (a BAM file). Based on this alignment result file, calculate the base distribution ratio of the SMN1 gene and SMN2 gene at the c.840C>T site in exon 7.
[0152] 7. Calculate sequencing depth.
[0153] The sequencing depth of the SMA gene was calculated based on the realigned alignment file obtained in step 6, and the measured depth of the internal reference gene was calculated based on the alignment file obtained in step 3. The sequencing depth of the SMA gene is the sum of the sequencing depths of the SMN1 and SMN2 genes. Furthermore, to ensure that sequencing information is not affected by amplified fragment length, the statistical value used in calculating sequencing depth is the average depth information.
[0154] 8. Standardized processing.
[0155] The ratio of the sequencing depth of the internal reference sample to the sequencing depth of the SMA gene (i.e. ) relationship, for the same batch of samples Perform standardization.
[0156] 9. Model training.
[0157] Samples 1 to 21 and 35 of the 35 SMA samples were used as the training set, and a detection model was constructed based on this training set. Sample 35 was used as the internal reference sample, and the internal reference samples were used as a reference for training set normalization. The following information from samples 1 to 21 and 35 was used when constructing the model: sequencing depth of the SMN1 gene, sequencing depth of the SMN2 gene, and sequencing depth of an internal reference gene (such as CFTR). The sequencing depths of the SMN1 and SMN2 genes were calculated based on the base distribution ratios calculated in step 6 and the SMA gene sequencing depths calculated in step 7. The sequencing depths of the internal reference genes served as a reference for normalization; that is, the sequencing depths of the SMN1 and SMN2 genes were normalized based on the sequencing depths of the internal reference genes.
[0158] The four possible copy number states of SMN1 and SMN2 genes are combined, so that SMA copy number variations can be classified as follows: Figure 2 15 types shown.
[0159] The detection model can perform copy number variation detection on samples 1 to 21 based on the normalized sequencing depth (including the sequencing depth of the SMN1 gene and the SMN2 gene), and obtain the corresponding predicted copy number type for each sample. The detection model is supervised and trained based on the predicted copy number type and the actual copy number type of the sample.
[0160] like Figure 4As shown in the figure, when performing copy number variation detection, the detection model can first determine whether the corresponding sample has zero copy. Based on the zero copy determination result, it then selects either the zero copy module or the non-zero copy module for detection to determine the copy number variation type. This can improve the efficiency of copy number variation detection. The zero copy module can detect samples with a copy number status of zero copy, while the non-zero copy module can detect samples with a copy number status of non-zero copy (i.e., 1 copy, 2 copies, etc.).
[0161] 10. Model verification.
[0162] Samples 22 to 35 of the 35 SMA samples were used as a validation set, which was used to evaluate the detection accuracy of the detection model constructed based on step 9. Among them, sample 35 was a positive control sample. Before the relevant data were input into the detection model, the relevant sequencing depth of the data in the training set needed to be standardized based on the relationship between the positive control sample and the internal reference sample. When validating the model, the sequencing depth of the internal reference gene of each sample was used to normalize the sequencing depth of the SMA gene of the sample, and the detection model was called to perform copy number variation detection on the normalized sequencing depth (including the sequencing depth of the SMN1 gene and the sequencing depth of the SMN2 gene). Finally, the SMA copy number variation type predicted by the detection model was obtained as follows: Figure 7 In addition, the prediction results of the detection model were compared with the "gold standard" provided by the experiment, and the results showed that the prediction accuracy of the detection model constructed in this application reached 100%.
[0163] It is understandable that the above steps 1 to 10 are the model building and training process. Figure 4 As shown in the figure, steps 1 and 2 are the data preprocessing process, and steps 3 to 9 are the data analysis process and result analysis process. In actual applications, the relevant data of the test sample (i.e., the sequencing depth of the SMN1 gene and the sequencing depth of the SMN2 gene) can be obtained based on the above steps. In this way, the trained model can be used to detect copy number variations in the test sample based on the relevant data.
[0164] Reference Figure 8 , the embodiment of the present application also provides a gene variation detection device, the device comprising:
[0165] The alignment unit 801 is configured to obtain first long-read sequencing data of a sample to be tested, and align the first long-read sequencing data with a reference genome to obtain first aligned data;
[0166] A first sequencing depth calculation unit 802 is configured to determine a region of a target gene in the reference genome, and calculate a first sequencing depth of the target gene based on the determined region and the first alignment data, wherein the target gene includes a first sub-target gene and a second sub-target gene;
[0167] A second sequencing depth calculation unit 803 is used to determine the first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth;
[0168] The variation detection unit 804 is used to call the pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested.
[0169] It can be seen that the contents of the above-mentioned gene variation detection method embodiment are all applicable to the embodiment of the present gene variation detection device. The functions specifically implemented by the present gene variation detection device embodiment are the same as those of the above-mentioned gene variation detection method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned gene variation detection method embodiment.
[0170] Reference Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0171] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0172] Memory 902 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). Memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 902 and is called by processor 901 to execute the genetic variation detection method of the embodiments of this application.
[0173] Input / output interface 903, used to implement information input and output;
[0174] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0175] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0176] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0177] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned gene variation detection method.
[0178] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned gene variation detection method.
[0179] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0180] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0181] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0182] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0183] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0184] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0185] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0186] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0187] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0188] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0189] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0190] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for detecting gene variation, characterized in that: The method comprises: Acquire first long-read sequencing data of the sample to be tested, and align the first long-read sequencing data with the reference genome to obtain first aligned data; Determining a region of a target gene in the reference genome, and calculating a first sequencing depth of the target gene based on the determined region and the first alignment data, wherein the target gene includes a first sub-target gene and a second sub-target gene; Determine the first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculate the first sub-sequencing depth of the first sub-target gene and the second sub-sequencing depth of the second sub-target gene according to the first base distribution ratio and the first sequencing depth; Calling a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain a target copy number variation type of the sample to be tested; The training process of the detection model includes the following steps: Acquire second long-read sequencing data of the training sample, and align the second long-read sequencing data with the reference genome to obtain second aligned data; Extracting read data of the regions where the first sub-target gene and the second sub-target gene are located from the second aligned data, respectively, and realigning the read data with the reference sequence of the first sub-target gene to obtain realigned data; Determining a second base distribution ratio of the first sub-target gene to the second sub-target gene based on the re-alignment data, calculating a third sequencing depth of the target gene based on the re-alignment data, and calculating a fourth sequencing depth of the internal reference gene based on the second alignment data; Determining a third sub-sequencing depth of the first sub-target gene and a fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth; The detection model is trained based on the third sub-sequencing depth, the fourth sub-sequencing depth and the fourth sequencing depth.
2. The method according to claim 1, characterized in that The calling of the pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain the target copy number variation type of the sample to be tested includes: Performing depth correction on the first sub-sequencing depth and the second sub-sequencing depth according to the second sequencing depth of the internal reference gene; The detection model is called to perform copy number variation detection based on the first sub-sequencing depth after depth correction and the second sub-sequencing depth after depth correction to obtain the target copy number variation type.
3. The method according to claim 1, characterized in that Calculating the third sequencing depth of the target gene according to the re-aligned data includes: Calculating the initial sequencing depth of the target gene based on the re-alignment data; A sample depth ratio is calculated based on the preliminary sequencing depth and the fourth sequencing depth, and the sample depth ratio is normalized based on a reference depth ratio of an internal reference sample corresponding to the training sample to obtain the third sequencing depth.
4. The method according to claim 1, wherein The step of training the detection model based on the third sub-sequencing depth, the fourth sub-sequencing depth, and the fourth sequencing depth includes: Performing depth correction on the third sub-sequencing depth and the fourth sub-sequencing depth respectively according to the fourth sequencing depth; Calling the detection model to determine a first copy number state corresponding to the first sub-target gene based on the third sub-sequencing depth after depth correction, and determining a second copy number state corresponding to the second sub-target gene based on the fourth sub-sequencing depth after depth correction, so as to determine a predicted copy number variation type based on the first copy number state and the second copy number state; The detection model is trained according to the standard copy number variation type and the predicted copy number variation type of the training sample.
5. The method according to claim 1, wherein The step of obtaining first long read sequencing data of the sample to be tested includes: constructing amplification primers for the target gene, wherein the amplification primers include a first sub-amplification primer; The sample to be tested is amplified according to the first sub-amplification primer, and the amplification result is sequenced to obtain the first long-read sequencing data.
6. The method according to claim 1, characterized in that The step of obtaining first long-read sequencing data of the sample to be tested, and aligning the first long-read sequencing data with a reference genome to obtain first aligned data includes: Constructing amplification primers for the target gene, wherein the amplification primers include a first sub-amplification primer and a second sub-amplification primer, wherein the length of the second sub-amplification primer is greater than the length of the first sub-amplification primer; amplifying the sample to be tested according to the first sub-amplification primer and the second sub-amplification primer, and sequencing the amplification result to obtain the first long-read sequencing data; First sub-amplification data corresponding to the first sub-amplification primer is separated from the first long-read sequencing data, and the first sub-amplification data is aligned with the reference genome to obtain the first aligned data.
7. The method according to claim 6, characterized in that The method further comprises: Splitting second sub-amplification data corresponding to the second sub-amplification primer from the first long-read sequencing data, and aligning the second sub-amplification data with the reference genome to obtain third aligned data; Perform variation detection on the sample to be tested according to the third comparison data.
8. A gene mutation detection device, characterized in that: The device comprises: An alignment unit is used to obtain first long-read sequencing data of the sample to be tested, and align the first long-read sequencing data with the reference genome to obtain first alignment data; a first sequencing depth calculation unit, configured to determine a region of a target gene in a reference genome, and calculate a first sequencing depth of the target gene based on the determined region and the first alignment data, wherein the target gene includes a first sub-target gene and a second sub-target gene; a second sequencing depth calculation unit, configured to determine a first base distribution ratio of the first sub-target gene and the second sub-target gene, and calculate a first sub-sequencing depth of the first sub-target gene and a second sub-sequencing depth of the second sub-target gene based on the first base distribution ratio and the first sequencing depth; A variation detection unit is used to call a pre-trained detection model to perform copy number variation detection based on the first sub-sequencing depth and the second sub-sequencing depth to obtain a target copy number variation type of the sample to be tested; The training process of the detection model includes the following steps: Acquire second long-read sequencing data of the training sample, and align the second long-read sequencing data with the reference genome to obtain second aligned data; Extracting read data of the regions where the first sub-target gene and the second sub-target gene are located from the second aligned data, respectively, and realigning the read data with the reference sequence of the first sub-target gene to obtain realigned data; Determining a second base distribution ratio of the first sub-target gene to the second sub-target gene based on the re-alignment data, calculating a third sequencing depth of the target gene based on the re-alignment data, and calculating a fourth sequencing depth of the internal reference gene based on the second alignment data; Determining a third sub-sequencing depth of the first sub-target gene and a fourth sub-sequencing depth of the second sub-target gene according to the second base distribution ratio and the third sequencing depth; The detection model is trained based on the third sub-sequencing depth, the fourth sub-sequencing depth and the fourth sequencing depth.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the gene variation detection method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the gene variation detection method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the gene variation detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for detecting homologous sequences on basis of high-throughput sequencing
WO2020041946A1