Rapid detection method for gene editing efficiency and gene editing linkage phenomenon of CRISPR-cas9 based on three-generation sequencing technology

By developing a TGS-GEM detection method based on third-generation sequencing technology, the problems of gene editing efficiency and accuracy and cost of target site mutation detection in the existing technology are solved, and efficient and accurate gene editing detection and linkage analysis are achieved.

CN120126564APending Publication Date: 2025-06-10SHANGHAI WEIKE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188798.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing gene editing efficiency and target site mutation detection methods mainly rely on second-generation sequencing technology, which has problems with PCR bias and high cost; due to the high error rate of third-generation sequencing technology, it is difficult to accurately detect gene editing target site mutations, and lacks effective error correction algorithms.

Method used

A TGS-GEM detection method based on third-generation sequencing technology was developed, including sequence preprocessing and mutation detection, construction of sequencing error correction model, model-based correction of error sites and gene editing mutation linkage analysis. This method uses Smith-Waterman alignment, naive Bayesian model and other technologies to identify and correct sequencing errors and improve detection accuracy.

Benefits of technology

It realizes rapid and accurate detection of CRISPR-cas9 gene editing efficiency and gene editing linkage phenomenon, reduces the third-generation sequencing error rate, and improves the reliability and efficiency of gene editing mutation detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126564A_ABST
    Figure CN120126564A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid detection method for CRISPR-Cas9 (clustered regularly interspaced short palindromic repeats-associated 9) gene editing efficiency and a linkage phenomenon based on a three-generation sequencing technology. The method comprises the four steps of sequence pretreatment and mutation detection, sequencing error correction model construction, model-based error site correction and gene editing mutation linkage analysis. The method comprises the following steps: acquiring third-generation sequencing data in an FASTQ format, cleaning the data and identifying mutation types; a naive Bayesian model is constructed to correct sequencing errors, so that the data accuracy is improved; utilizing an error correction model to classify mutation and recover false positive mutation; and finally carrying out single-target or multi-target gene editing mutation linkage analysis. The three-generation sequencing technology is applied to detection of gene editing efficiency and target site mutation for the first time, the detection efficiency and accuracy are improved, the sequencing error rate is remarkably reduced, a brand-new and efficient detection means is provided for the field of gene editing, and the method is suitable for popularization and application. Particularly, the method has important application value in the fields of basic research, gene therapy, genetic improvement and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the direction of the analysis of third-generation sequencing data in the field of bioinformatics, and discloses a rapid detection method for gene editing efficiency and gene editing linkage phenomenon of CRISPR-cas9 based on third-generation sequencing technology. Background Art

[0002] Genome editing nucleases are technologies widely used to modify the genomes of living cells, and have many important applications in the fields of biology and medicine. Among them, the CRISPR-Cas nuclease technology has been widely adopted because it can precisely target gene loci through guide RNA (gRNA) for gene editing. In mammalian cells, DNA double-strand breaks (DSBs) induced by engineered nucleases are usually repaired through competitive endogenous cellular DNA repair pathways, and the repair pathways will introduce indel mutations, resulting in gene disruption. Therefore, for the identification of detailed mutation information after gene editing, since the mutation types generated by different editing tools are different, sensitive and accurate detection means are required to clarify the editing efficiency, so as to improve the safety and reliability of the technology.

[0003] Traditional Sanger sequencing has insufficient analytical capabilities in aspects such as low-frequency mutations, chimeric mutations, and complex mutations. When dealing with the mutation identification of polyploids, multiple samples or multiple loci, it is also limited by the high sequencing cost and analysis difficulty; while second-generation sequencing technology has become an important means for various mutation detection and analysis, but its complex library construction process requirements and the characteristics of short sequencing fragments have restricted its wide application in the detection of gene editing target site mutations to a certain extent.

[0004] Third-generation sequencing technology can provide read lengths of dozens of kb or even 100 kb, without PCR amplification, thus avoiding the biases that may be introduced during the PCR process. In addition, third-generation sequencing technology has the characteristics of high throughput and can complete the sequencing of a large number of samples in a short time; however, there is currently no method for accurately detecting gene editing target site mutations through third-generation sequencing technology. Third-generation sequencing technology reduces the splicing cost through long read lengths, saves memory and computing time, and at the same time improves the quality and integrity of genome assembly.

[0005] However, due to some problems still existing in third-generation sequencing technology, it has a high error rate. Especially for Oxford Nanopore technology, the error rate of its single read length is relatively high, reaching 15%-40%, which is much higher than that of second-generation sequencing technology. Moreover, its sequencing errors occur randomly rather than being systematic biases, and it is difficult to predict and correct them with existing methods. Therefore, the premise of using the advantages of third-generation sequencing for gene editing efficiency and target site mutation detection is to solve the problem of high error rate in third-generation sequencing.

[0006] Specifically, 1) The current methods for gene editing efficiency and target site mutation detection mainly rely on second-generation sequencing technology, which has the following deficiencies: During library construction, second-generation sequencing relies on PCR amplification to enhance fluorescence signals, which will introduce PCR bias and affect the accuracy of gene editing target site mutation detection; When the target fragment is long and there are many target sites, contamination may occur between primers, resulting in data waste. 2) Although the current third-generation sequencing technology has the advantage of long length, it has not been applied to gene editing target site mutation detection because of its high sequencing error, leading to a high false positive rate. Currently, there is no efficient and general algorithm to identify the error rate of third-generation sequencing data and correct it, thus improving the accuracy of third-generation sequencing for gene editing efficiency and target site mutation detection. 3) The current methods for gene editing efficiency and target site mutation detection rely on the splicing of short sequences and cannot obtain the information of simultaneous linked mutations from a long sequence, nor can they perform multi-target gene editing mutation linkage analysis on polyploid organisms. Summary of the Invention

[0007] Aiming at the above problems that there is currently no method for third-generation sequencing technology to detect gene editing efficiency and target site mutations, and there are defects in third-generation sequencing for gene editing detection, and there is currently no detection method for multi-target linkage analysis, the present invention has developed a TGS-GEM (Third-Generation Sequencing for Gene Editing Mutation Detection) detection method, that is, a rapid detection method for gene editing efficiency and gene editing linkage phenomenon of CRISPR-cas9 based on third-generation sequencing technology.

[0008] The present invention includes the following technical solutions:

[0009] A rapid detection method for gene editing efficiency and gene editing linkage phenomenon of CRISPR-cas9 based on third-generation sequencing technology, as Figure 1 shown, includes the following steps:

[0010] S1. Pre - processing of Sequences and Mutation Detection: Obtain third - generation sequencing data in FASTQ format and a given target region sequence; clean the data by removing low - quality reads and adapter sequences; perform local Smith - Waterman alignment of the processed sequences with a known reference genome, and count the number and proportion of reads with the same base; compare the sequencing reads with the reference sequence to identify base substitution, insertion, and deletion mutations.

[0011] S2. Construction of a Sequencing Error Correction Model: Extract mutation events from the sequencing data. Combine with the reference sequence and PCR verification to classify mutations into true positives and false positives, and divide the training set and test set; screen classification features, including adjacent base sequences, read positions, base quality scores, GC content, coverage depth, distance between the mutation position and the cutsite, and mutation frequency; construct a Naive Bayes Model (NBM), train and optimize it; evaluate the model performance using the test set, and measure it by indicators such as accuracy, sensitivity, and specificity.

[0012] S3. Correcting Error Sites Based on the Model: Use the model in S2 to classify the mutations detected in S1, and determine true positives or false positives; locate the positions of false - positive mutations according to the results of S1, including single - base substitutions, insertions, and deletions; perform recovery processing on false - positive mutations, restoring SNPs to the original bases, and removing or restoring insertions and deletions.

[0013] S4. Gene Editing Mutation Linkage Analysis: For single - target samples, directly analyze the mutations at the editing sites; for multi - target samples, construct a mutation matrix and perform clustering to identify mutation linkage "clusters", and analyze the multi - target gene editing mutation linkage effect to understand the simultaneously occurring mutations and their impact on gene editing efficiency or specificity.

[0014] Furthermore, the sequence pre - processing and mutation detection in step S1 specifically include the following steps:

[0015] S1.1 Acquisition of Sequencing Data:

[0016] The third - generation sequencing data is in FASTQ format, and a target region sequence is given.

[0017] S1.2 Filtering Low - quality Reads:

[0018] Perform preliminary cleaning on the sequencing data to remove reads with lower quality to ensure the data accuracy for subsequent steps.

[0019] S1.3 Adapter Trimming:

[0020] Remove the attached adapter sequences from the reads and retain the actual useful sequence data.

[0021] S1.4 Alignment to the Reference Genome:

[0022] Align the processed sequences to the known reference genome based on the Smith-Waterman local alignment algorithm. Reads with the same base at each coordinate position are counted into a group. Count the number and proportion of reads in each group, and give a representative reads sequence for each group.

[0023] S1.5 Mutation Identification:

[0024] Compare the sequencing reads with the target reference genome sequence to identify base substitution (Substitutions), insertion (Insertions), and deletion (Deletions) mutation types. If the base of the sequencing reads is different from the reference sequence, it is marked as a substitution. If the sequencing reads contain a base fragment that is more than the reference sequence, it is defined as an insertion. If the sequencing reads lack a base fragment corresponding to the reference sequence, it is defined as a deletion.

[0025] Furthermore, to solve the problem of false positive mutations caused by sequencing errors, we also developed a sequencing error correction model aimed at improving gene editing efficiency and the accuracy of target site mutation detection. This model systematically identifies and filters false positive mutations through feature extraction, classification algorithms, and data correction. The specific method is as follows:

[0026] S2. Construction of Sequencing Error Correction Model:

[0027] S2.1 Sample Collection, Training Set and Test Set Division:

[0028] Extract mutation events from the raw data of third-generation sequencing, combine with the reference sequence and PCR-verified samples, and divide them into

[0029] two categories:

[0030] True positive mutations: Real variations caused by gene editing design target mutations or off-target effects.

[0031] False positive mutations: Pseudo-mutations caused by sequencing errors.

[0032] And divide the collected dataset into a training set and a test set according to 1:1, which are used for model training and

[0033] performance verification.

[0034] S2.2 Classification Feature Screening:

[0035] 1) Sequence information of adjacent bases (k-mer feature): Considering that sequencing errors are often associated with adjacent bases, for example, some repetitive sequences may be more likely to cause errors. Therefore, we select the 5bp of adjacent base information before and after as a k-mer feature input.

[0036] 2) Reads position information: Record the position of the mutation site in the reads (close to the head / tail or in the middle). Since the position where the mutation occurs affects the error probability, the error rate at the head / tail is usually higher.

[0037] 3) Base Quality Score: Record the base quality score for each mutation site. Usually, low-quality bases may be the source of errors, while high-quality bases are more reliable.

[0038] 4) GC content: Calculate the GC ratio in the region (±10bp) before and after the mutation site. An excessively high or low GC ratio may lead to sequencing bias.

[0039] 5) Coverage Depth: Measure the number of read supports for each mutation. Errors are more likely to occur in low-coverage regions.

[0040] 6) Distance between the mutation position and the cutsite: Considering the distance between the mutation site and the cleavage site (cutsite) of the Cas9 enzyme, this distance may affect the efficiency and accuracy of the mutation.

[0041] 7) Mutation frequency: Record the occurrence frequency of the mutation site in all reads. Mutations with high frequencies are more likely to be real variations, while low-frequency mutations may be sequencing errors.

[0042] S2.3 Model construction, training, and optimization

[0043] Construct a Naive Bayes model, which is a classification technique based on Bayes' theorem and feature conditional independence. Its core lies in applying Bayes' theorem to estimate the posterior probability of a specific data sample belonging to each category, and finally selecting the category with the highest probability as the predicted classification of the sample. The model we constructed is named the sequencing error correction model NBM. Each site in the training set data contains these feature information such as k-mer information, Reads position information, base quality score, GC content, coverage depth, distance between the mutation position and the cutsite, and mutation frequency. Each site has a class label, true positive mutation or false positive mutation. Input the training set data into the NBM model for training and learning of the NBM model, and further optimize the parameter tuning to improve the prediction ability of the NBM model.

[0044] S2.4 Model performance evaluation:

[0045] The performance of the NBM model was evaluated using a test set. The predicted results of the NBM model were compared with the true positive and false positive real labels of the test set to verify the classification performance of the NBM model, and the following metrics were recorded: Accuracy, Sensitivity / Recall, Specificity, Positive Predictive Value (PPV), and Negative Predictive Value (NPV). The performance of the NBM model in identifying true positive sites and false positive sites was measured by the above performance evaluation metrics.

[0046] The specific method for constructing the sequencing error correction model NBM is as Figure 1 shown in the left model construction part in

[0047] Furthermore, step S3 corrects the error sites (false positive sites) based on the model judgment results, including the following specific steps:

[0048] S3.1 Classify the detected mutations in step S1 using the model to obtain whether the mutation site is a true positive mutation or a false positive mutation type.

[0049] S3.2 For each read, restore the mutations classified as false positive mutations by the model to obtain a sequence that theoretically does not carry false positive mutations. First, locate the mutation position: in each read, the position of the false positive mutation has been determined in the sequence mutation detection step of step S1 (for example, single base substitution, insertion, deletion, etc.); then replace the mutation: for each false positive mutation, we can restore the original sequence. For SNP (single nucleotide polymorphism) mutations, if detected as false positive, restore to the original base at that position. For insertion and deletion mutations, if marked as false positive, remove the inserted base or restore the deleted base. Assuming the mutation is at a specific position, we can restore the base (or the inserted / deleted region) at that position to the original sequenced base.

[0050] Furthermore, the specific steps of the single-target / multi-target gene editing mutation linkage analysis in step S4 are as follows: If it is a single-target gene editing sample, we directly analyze the mutations at the editing sites of the sample. If it is a multi-target gene editing sample, we use the following steps for multi-target mutation linkage analysis:

[0051] S4.1 Construct a mutation matrix based on the true positive mutations on each read, and cluster the mutation matrix to obtain "clusters" of mutation linkages. The mutation matrix is a two-dimensional matrix, where the rows represent different read clusters and the columns represent mutation sites (SNPs, insertions or deletions) at each coordinate position. Each element of the matrix represents whether the mutation at that position appears in that read. By performing hierarchical clustering on the mutation matrix, we can identify which mutations often occur together in different reads, thus forming "clusters" of linked mutations.

[0052] S4.2 Discover the multi-target gene editing mutation linkage effect based on the information in the "clusters". When editing multiple target sites in the genome, a mutation linkage effect may occur, that is, some mutations may occur simultaneously under the same genotype, affecting the effect of gene editing. By analyzing the "clusters" formed by mutations, it is possible to identify which mutations occur simultaneously, and these combined mutations may affect the efficiency or specificity of gene editing.

[0053] The present invention also discloses a storage device for executing the above-mentioned rapid detection method for CRISPR-cas9 gene editing efficiency and linkage phenomenon based on third-generation sequencing technology. The storage device stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the above-mentioned method steps are implemented.

[0054] The present invention also discloses a computer device, comprising:

[0055] A processor;

[0056] A memory for storing computer-executable instructions;

[0057] Wherein, the processor is configured to execute the computer-executable instructions stored in the memory to implement the above-mentioned rapid detection method for CRISPR-cas9 gene editing efficiency and linkage phenomenon based on third-generation sequencing technology.

[0058] The present invention also discloses a server, the server comprising:

[0059] A data receiving module for receiving third-generation sequencing data in FASTQ format;

[0060] A data processing module for executing the sequence preprocessing, mutation detection, sequencing error correction model construction, error site correction based on the model, and gene editing mutation linkage analysis steps of the present invention;

[0061] A result output module for outputting the analysis results of gene editing efficiency and linkage phenomenon.

[0062] The present invention also discloses a cloud computing-based platform, which includes a plurality of servers. Each server is as described above and is used for distributed processing and analysis of third-generation sequencing data from different samples to achieve rapid detection of CRISPR-cas9 gene editing efficiency and linkage phenomenon based on third-generation sequencing technology.

[0063] The present invention also discloses a gene editing analysis system, which includes:

[0064] A sequencer for generating third-generation sequencing data in FASTQ format;

[0065] A computer device for processing and analyzing the sequencing data;

[0066] A display device for presenting the analysis results of gene editing efficiency and linkage phenomenon.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] The present invention discloses a rapid detection method for CRISPR-cas9 gene editing efficiency and gene editing linkage phenomenon based on third-generation sequencing technology:

[0069] 1) For the first time, the third-generation sequencing technology is applied to the detection of gene editing efficiency and target site mutations. This method provides a new and efficient detection means for the gene editing field. This method not only improves the detection efficiency but also ensures high accuracy, which is crucial for quickly and accurately identifying mutations after gene editing, especially in the fields of basic research, gene therapy, and genetic improvement.

[0070] 2) The present invention develops a new error correction algorithm specifically for the unique errors of third-generation sequencing to correct the errors of third-generation sequencing. This algorithm can significantly reduce the high error rate in third-generation sequencing and improve the sequencing data from low quality to high quality without affecting real mutations, thereby improving the reliability and accuracy of the data.

[0071] 3) In the presence of multiple target sites, the present invention can perform linkage analysis on the edited mutation sites of multiple target sites. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Flowchart of the detection method for gene editing efficiency and gene linkage analysis based on third-generation sequencing, where the left vertical column is the construction process of the sequencing error correction model, and the right vertical column is the entire process of gene editing efficiency and multi-target gene linkage analysis of third-generation samples;

[0073] Figure 2Distribution map of the distances between true positive sites and false positive sites from the cut site. The x-axis represents the distance from the mutation position to the cut site, and the y-axis represents the mutation frequency of the sites. In the graph, green represents true positive sites and red represents false positive sites. It can be seen that the coordinate positions of true positive sites are closer to the cut site and are concentrated, while the coordinate positions of false positive sites show a random distribution from the cut site;

[0074] Figure 3 ROC curve of the sequencing error correction model NBM. The x-axis is the False Positive Rate (FPR), also known as 1 - Specificity, which represents the proportion of samples that are actually negative but are wrongly predicted as positive. The y-axis is the True Positive Rate (TPR), also known as Sensitivity or Recall, which represents the proportion of samples that are actually positive and are correctly predicted as positive; The blue solid line represents the ROC curve of the model; This curve shows the changes in FPR and TPR of the model at different thresholds; The closer the curve is to the upper left corner, the better the performance of the model. The Area Under the Curve (AUC) is an indicator to measure the overall performance of the model. The value of AUC ranges from 0 to 1, and the closer the value is to 1, the stronger the discrimination ability of the model; In the figure, AUC = 0.96, indicating that the model has a very strong classification ability;

[0075] Figure 4 Original mutation detection (>0.2%) in Example 1. Each position horizontally is labeled with the base and the mutations that occurred, and vertically represents the number of reads and the proportion grouped together. It can be seen that the reads group with the highest proportion only accounts for 18.69%, with only 263 reads, indicating that due to sequencing errors, a large number of reads did not get grouped but formed groups of single reads;

[0076] Figure 5 For Figure 4 Magnified view of the reads proportion;

[0077] Figure 6 Original top 10 reads group distribution (before using the sequencing error correction model). The x-axis is the ranking ID of the proportion of the top 10 reads groups, and the y-axis is the number of reads. It can be seen that the number of reads at the first site is only 263;

[0078] Figure 7TOP10 reads grouped distribution after removing false positive sites by the sequencing error correction model (after using the sequencing error correction model). The x-axis is the proportion ranking ID of the top 10 reads groups, and the y-axis is the number of reads. After the sequencing error correction model and site correction, the number of reads at the first site increased significantly;

[0079] Figure 8 Multi-target linkage analysis site clustering map. The x-axis represents the mutation position, the columns represent the reads clusters, and the color intensity represents the number of reads. It can be seen that the mutation sites at the coordinate positions 30, 50 and 270, 280 and 288, 296 will appear simultaneously and form a mutation "cluster". Detailed implementation manners

[0080] In order to enable those skilled in the art to more fully understand the technical solutions of the present invention, the exemplary embodiments of the present invention will be described more comprehensively and in detail below with reference to the accompanying drawings.

[0081] The following examples are used to illustrate the present invention, but are not used to limit the scope of the present invention. The experimental methods used in the following examples are all conventional methods unless otherwise specified. The materials, reagents, etc. used in the following examples can be obtained from commercial sources unless otherwise specified.

[0082] The specific meanings of the terms used are as follows:

[0083] TGS-GEM: Third-Generation Sequencing for Gene Editing Mutation Detection, the third-generation sequencing technology for gene editing mutation detection.

[0084] FASTQ: FASTQ is a text file format used to store biological sequence data and its corresponding base quality scores.

[0085] reads: Short sequence fragments obtained in sequencing technology, used for subsequent data analysis.

[0086] Adapter: Sequencing adapter, which refers to a known short nucleotide sequence in the sequencing field, used to link unknown target sequencing fragments

[0087] Smith-Waterman: A dynamic programming local alignment algorithm method used to find the best local similarity region between two sequences, especially suitable for the alignment of partially similar sequences.

[0088] Substitutions: Base substitution, that is, one base is replaced by another different base.

[0089] Insertions: Insertions refer to the addition of extra base pairs that occur in a DNA sequence.

[0090] Deletions: Deletions refer to the loss of base pairs that occur in a DNA sequence.

[0091] k-mer: A k-mer refers to a DNA fragment of length k.

[0092] Base Quality Score: The base quality score is an integer mapping that measures the probability of error in identifying sequencing bases.

[0093] GC: GC is the ratio of guanine and cytosine in a DNA molecule. The level of GC content may affect the uniformity of sequencing coverage, resulting in the GC bias phenomenon.

[0094] Coverage Depth: The sequencing depth refers to the ratio of the total number of bases sequenced to the genome size, that is, the average number of times each base in the genome is covered by sequencing reads.

[0095] Cutsite: In gene editing, a cutsite refers to a specific site on a DNA molecule that a restriction enzyme recognizes and cuts.

[0096] Naive Bayes: The Naive Bayes classifier is a probabilistic classifier based on Bayes' theorem.

[0097] Accuracy: Accuracy refers to the proportion of correctly predicted samples in the total number of samples, including true positives and true negatives.

[0098] Sensitivity / Recall: Sensitivity (or recall) refers to the proportion of true positive samples correctly identified by the model among all actual positive samples.

[0099] Specificity: Specificity refers to the proportion of true negative samples correctly identified by the model among all actual negative samples.

[0100] PPV (Positive Predictive Value, Precision): It represents the proportion of actually true positive samples among the samples predicted as true positive.

[0101] NPV (Negative Predictive Value): It represents the proportion of actually false positive samples among the samples predicted as false positive.

[0102] Example 1

[0103] A rapid detection method for gene editing efficiency and gene editing linkage phenomenon of CRISPR-cas9 based on third-generation sequencing technology, such as Figure 1As shown, it includes the following steps:

[0104] S1. Sequence preprocessing and mutation detection: Obtain third-generation sequencing data in FASTQ format and given the target region sequence; clean the data, remove low-quality reads and adapter sequences; perform Smith-Waterman local alignment of the processed sequences with the known reference genome, and count the number and proportion of reads with the same base; compare the sequencing reads and the reference sequence to identify base substitution, insertion, and deletion mutations;

[0105] Specifically, it includes the following steps:

[0106] S1.1 Acquisition of sequencing data:

[0107] The third-generation sequencing data is in FASTQ format and the target region sequence is given.

[0108] S1.2 Filtering of low-quality reads:

[0109] Perform preliminary cleaning on the sequencing data to remove low-quality reads to ensure the data accuracy of subsequent steps.

[0110] S1.3 Adapter trimming:

[0111] Remove the ligated adapter sequences from the reads and retain the actual useful sequence data.

[0112] S1.4 Alignment to the reference genome:

[0113] Align the processed sequences to the known reference genome based on the Smith-Waterman local alignment algorithm. Group the reads with the same base at each coordinate position, count the number and proportion of reads in each group, and give a representative reads sequence for each group.

[0114] S1.5 Mutation identification:

[0115] Compare the sequencing reads and the target reference genome sequence to identify base substitution (Substitutions), insertion (Insertions), and deletion (Deletions) mutation types. If the base of the sequencing reads is different from the reference sequence, it is marked as a substitution. If the sequencing reads contain a base fragment that is more than the reference sequence, it is defined as an insertion. If the sequencing reads lack a base fragment corresponding to the reference sequence, it is defined as a deletion.

[0116] S2. Construction of sequencing error correction model: Extract mutation events from sequencing data, combine with reference sequences and PCR verification, classify mutations into true positives and false positives, and divide the training set and test set; Screen classification features, including adjacent base sequences, reads positions, base quality scores, GC content, coverage depth, distance between mutation position and cutsite, mutation frequency; Construct a Naive Bayes Model (NBM), train and optimize it; Use the test set to evaluate the model performance, measured by indicators such as accuracy, sensitivity, and specificity.

[0117] S3. Correct error sites based on the model: Use the model in S2 to classify the mutations detected in step S1 to determine true positives or false positives; Locate the positions of false positive mutations according to the results of step S1, including single base substitutions, insertions, and deletions; Perform recovery processing on false positive mutations, restore SNPs to the original bases, and remove or restore insertions and deletions.

[0118] S4. Gene editing mutation linkage analysis: For single-target samples, directly analyze the mutations at the editing sites; For multi-target samples, construct a mutation matrix and cluster it to identify mutation linkage "clusters", analyze the multi-target gene editing mutation linkage effect, and understand the simultaneous mutations and their impact on gene editing efficiency or specificity.

[0119] Example 2

[0120] The main purpose of this Example 2 is to construct a sequencing error correction model based on third-generation sequencing data, which can significantly reduce the high error rate in third-generation sequencing, thereby improving the accuracy of analyzing editing site mutations by third-generation sequencing technology.

[0121] Perform third-generation sequencing analysis on 60 samples, select 356 mutation sites (including 122 true positive sites and 234 false positive sites) that have been verified by PCR among them, describe the characteristics of the 356 mutation sites, and find that there are obvious differences in the position distributions of true positive sites and false positive sites. See Figure 2 , it can be seen that the coordinate positions of true positive sites are closer to the cutsite and are concentrated, while the coordinate positions of false positive sites show a random distribution with respect to the distance from the cutsite. Therefore, we construct a Naive Bayes model, select factors including the distance between the coordinate position and the cutsite, as well as the sequence information of adjacent bases (k-mer features), reads position information, base quality score, GC content, and mutation frequency as feature inputs, train the model, and the final results show that the constructed Naive Bayes model can effectively distinguish true positive and false positive sites, which means that the model can accurately identify the true mutation sites, and this is very valuable for subsequent biological research and clinical applications. The application of the model can reduce false positive results, save the time and resources required for subsequent verification, and improve the accuracy of mutation detection at the same time.

[0122] 1) Sample collection, division of training set and test set

[0123] 60 samples were subjected to third-generation sequencing analysis to obtain preliminary mutation detection results. 356 mutation sites were screened out from the mutation detection results, including: 122 true positive sites and 234 false positive sites (verified by PCR). Statistical analysis was performed on the basic information of the mutation sites, including the coordinate positions and the distance distribution from the cutsite as shown in Figure 2 .

[0124] 2) Feature selection: The following features were used for model training, and feature extraction was performed on all mutation sites. The extracted features include:

[0125] Position feature: The distance between the coordinate position of the mutation site and the cutsite;

[0126] Sequence feature: The k-mer feature of the adjacent base sequence;

[0127] Reads information: The position information of the reads corresponding to the mutation site;

[0128] Base quality: The base quality score of the sequencing;

[0129] GC content: The GC content in the neighborhood of the mutation site;

[0130] Mutation frequency: The occurrence in the supporting mutation reads / all reads

[0131] The features were standardized to adapt to subsequent model training.

[0132] 3) Construction, training and optimization of the Naive Bayes model

[0133] First, the dataset of the mutation sites of 60 samples was divided into a training set and a test set, using a 1:1 sample ratio for division:

[0134] Training set: Used for model training; Test set: Used for model performance evaluation. The site statistics of the training set and the test set are shown in Table 1 below.

[0135] Table 1: Site statistics of the training set and the test set

[0136] Data set Sample quantity Number of true positive sites Number of false positive sites Total number of sites Training set 30 64 128 192 Test set 30 58 106 164 Total 60 122 234 356

[0137] Secondly, a Naive Bayes model is constructed. In this embodiment, the constructed model is named NBM. The NBM model is trained using 192 sites in the training set. Each of the 192 sites in the 30 samples of the training set contains feature information such as the coordinate position, distance to the cutsite, k-mer features, read information, base quality score, GC content, and mutation frequency. Each site has a class label, true positive mutation or false positive mutation. The training set data is input into the NBM model, and the NBM model parameters are adjusted to optimize the performance of the NBM model.

[0138] 4) Model performance evaluation

[0139] The performance of the model is evaluated on the test set. The prediction results of the NBM model are compared with the actual true positive and false positive sites to verify the accuracy of the model, and the following metrics are recorded: Accuracy, Sensitivity / Recall, Specificity, Positive Predictive Value (PPV), Negative Predictive Value (NPV), as shown in Table 2 below, where:

[0140] The sensitivity is 94.92%, indicating that the model can detect true positive mutations (i.e., real mutation sites) well;

[0141] The specificity is 96.23%, indicating that the model can effectively distinguish false positives and avoid excessive false alarms;

[0142] The accuracy is 95.76%, indicating that the overall prediction accuracy is very high and most samples are correctly classified;

[0143] The PPV (Positive Predictive Value) is 93.33% and the NPV (Negative Predictive Value) is 97.14%, indicating that the model is highly reliable in predicting true positives and true negatives.

[0144] Finally, the ROC curve of the model is plotted and the AUC value is calculated. The ROC curve is shown in Figure 3 : ROC curve, where the abscissa is 1 - Specificity (FPR) and the ordinate is Sensitivity (TPR). The AUC value of the curve is 0.96, indicating that the discrimination ability of the model is very strong. In summary, the model performs well and has a good discrimination effect between true positives and false positives (high sensitivity and specificity). We name the NBM model: Third-generation sequencing error correction model.

[0145] Table 2: Table of actual and predicted results of the test set

[0146]

[0147] Remarks:

[0148] TP (True Positive) in Table 2: The number of samples predicted as positive and actually positive (56);

[0149] FP (False Positive): The number of samples predicted as positive but actually negative (4);

[0150] FN (False Negative): The number of samples predicted as negative but actually positive (3);

[0151] TN (True Negative): The number of samples predicted as negative and actually negative (102);

[0152]

[0153]

[0154]

[0155]

[0156]

[0157] According to the content of Example 1, combined with Figure 4 and Figure 5 , it shows excellent performance in the detection of gene editing target site mutations by the NBM model, with high sensitivity and specificity, and can effectively distinguish true positive and false positive mutation sites.

[0158] Example 3

[0159] The main purpose of Example 3 is to demonstrate how to use third-generation sequencing technology combined with a sequencing error correction model to accurately analyze the mutation situation of gene editing sites.

[0160] Select a switchgrass sample for third-generation sequencing analysis, use the sequencing error correction model to distinguish true positive and false positive sites, and the comparison results before and after removing false positive sites are shown in Figure 6 : The original TOP10 reads grouped distribution (before using the sequencing error correction model) and Figure 7 : The TOP10 reads grouped distribution after removing false positive sites using the NBM model (after using the sequencing error correction model), which can verify that the model can effectively and truly remove false positive sites.

[0161] 1) Select 1 switchgrass sample and obtain the FASTQ raw data through third-generation sequencing;

[0162] 2) Filter the sequencing data to remove low-quality reads (Q < 20);

[0163] 3) Remove the Adapter sequences from the reads to obtain clean reads;

[0164] 4) Align the clean reads to the reference genome. The sequence information of the amplified fragment is: ATGGTGGCCGTGCCCAAGGTCGCGATGGAGTGGCT CCAAGACCCTCTGAGCTGGGTGT TGCTGGCGTCTCTGGCCTTATTCCTCCTGCAGCTGCGGCGGTGGGGCAAGGCGCCGCTGCCGCCGGGCCCGAAGCCGCTGCCGATCATCGGGAACATGACGATGATGGACCAGCTGACCCACCGC (SEQ ID No.1), where the underlined part is the gene editing guide sequence of this sample, including:

[0165] Target: CCAAGACCCTCTGAGCTGGGTGT (SEQ ID No.2);

[0166] Based on the reads distribution obtained by alignment, according to the alignment results, reads with the same sequence are grouped into the same group. We found that there are a total of 924 types of reads with different sequences. The first 20 pieces of information of the reads grouping are shown in Table 3 below (due to sequence length limitations, we have omitted the sequence information after the second row and use "......" to represent it). The last 20 pieces of information of the reads grouping are shown in Table 4 below (due to sequence length limitations, we have omitted the sequence information after the second row and use "......" to represent it). The first column of the table refers to the aligned reads, the second column refers to the reference sequence, the third column refers to the number of reads in the group, and the fourth column refers to the proportion of the reads. From the number of reads and the proportion in Table 3 and Table 4, it can be seen that the mutations caused by the third-generation sequencing technology at random positions in the sequence affect the grouping judgment of reads of the same category, resulting in the number of reads in the top-ranked group accounting for only 18.69%, and the remaining 80% of the reads will be divided into separate groups. Figure 5 Intuitively shows the original reads grouping distribution of the top 10. It can be seen that the proportion of reads distribution at these sites is relatively small, which further illustrates that the mutations generated at random positions have affected the judgment of reads consistency.

[0167] Table 3: The top 20 reads ranked according to the reads support number in the original reads distribution

[0168]

[0169]

[0170] Table 4: Reads ranked last 20 according to the number of reads supported in the original reads distribution

[0171]

[0172]

[0173] 5) Identify the mutations of each read, including insertions, deletions, and substitutions, and keep them in the same order as the reads clustered in step 4). As shown in Table 5 (Top10), there are a total of 924 actual mutation combinations, among which there are mutations caused by a large number of sequencing errors.

[0174] Table 5: Original Mutation Detection Table Top-10

[0175]

[0176] 6) Extract the read features, input the read features into the third-generation sequencing error correction model NBM, and use NBM to judge the type of mutation, so as to achieve the purpose of identifying true positive mutation sites and false positive mutation sites. According to the judgment results of NBM on each mutation, finally retain the sites identified as true positive mutation sites and filter out false positive sites. The resulting retained sites are only the following 24 site combinations, and the reads of false positive sites have been mutated and restored, as shown in Table 6.

[0177] Table 6: Result Table after Model Correction

[0178]

[0179]

[0180] The above Example 2 shows through the analysis of the validation samples that the third-generation sequencing error correction model NBM can accurately distinguish true positive and false positive mutations, thus indicating that third-generation sequencing can be applied to the detection of gene editing efficiency and target site mutations.

[0181] Example 4

[0182] The main purpose of Example 4 is to verify that the detection method can effectively identify whether there is a linkage effect between gene editing mutations at different target sites when analyzing samples with multi-target gene editing.

[0183] Select an 84K poplar sample that has undergone multi-target gene editing. After removing false positive sites using the third-generation sequencing error correction NBM model, use the true positive mutation sites as input, and analyze the multi-target gene editing mutation sites through hclust hierarchical clustering. It can be found that there is a linkage effect of gene editing mutations between multi-targets.

[0184] The main steps of Example 4 are as follows:

[0185] The amplified sequence of the 84K sample is:

[0186] CAAATCAATGTCTCTTCCAGAAACACAAG GCAAGACCCTCCCAGATGCG TGGGACTATAAGGGCCGGCCTGCCGAGCGGTCCAAAACTGGTGGCTGGACCAGTGCTGCCATGATTCTAGGTTACAGCTTTCTTTCATTTTCTATGATATTTATATCCTACTTATCTGTATATGTTTCTCCATCTGATCTCTCACTCTCCTGTTCTGGGTTTTATCAGGCGGAGAGGCAATGGAGAGACTAACAACACTTGGTATTGCTGTTAATCTGGTGACGTA TCTGACTGGTACTATGCACC TGGGCAATGCTACCTCTGCCAACACTGTCACCAACTTCCTTGGAACATCTTTCATGCTTTGTCTGCTTGGTGGTTTTATCGCTGACACCTTTCTTGGAAGGTTTAACCAGTTTTTTGCACTTTCGGTCTTTATTTATTTTGTGCTATATTTGTTTCTGAATTCAGGATTATTGTCTGCTTTCTGATTTCTTAACCCTTTCCT TTGTTTCATGACAGGTATCT CACCATCGCCATCTTCGCCACCGTGCAAG (SEQ ID No.7), where the underlined part is the gene editing guide sequence of the sample, including:

[0187] Target 1: GCAAGACCCTCCCAGATGCG (SEQ ID No.8);

[0188] Target 2: TCTGACTGGTACTATGCACC (SEQ ID No.9);

[0189] Target 3: TTGTTTCATGACAGGTATCT (SEQ ID No.10);

[0190] 1) Perform third-generation sequencing on the samples, filter the sequencing data to remove low-quality reads (Q < 20), and remove the Adapter sequences in the reads to obtain clean reads. Align the clean reads to the amplified sequences to obtain the reads distribution, and then identify the mutations in each read, including insertions, deletions, and substitutions. Through the sequencing error correction model, identify the true positives and false positives at the mutation sites, retain the finally identified true positive mutation sites, and filter out the false positive sites. The final results are shown in Table 7:

[0191] Table 7: Analysis results table of Populus 84K

[0192]

[0193]

[0194] 2) According to the true positive mutations on each read, construct a mutation matrix, where the mutation matrix is a two-dimensional matrix, with rows representing different read clusters and columns representing the mutation sites (SNP, insertion, or deletion) at each coordinate position.

[0195] 3) Use the hclust hierarchical clustering algorithm to cluster the mutation matrix to obtain the "clusters" of mutation linkages. Each element of the matrix represents whether the mutation at that position appears in that read.

[0196] 4) Visualize the clustering results, and it can be identified that the mutation sites at coordinate positions 30, 50 and 270, 280 and 288, 296 will appear simultaneously and form a "cluster" of mutations, such as Figure 8 , indicating that there will be a linked mutation reaction between the first target and the second target.

[0197] Example 3 proves that in the case of multi-target gene editing, the method of the present invention can effectively identify the linked effects of gene editing target site mutations between different targets.

[0198] Combining the above Examples 1, 2, and 3, the rapid detection method for gene editing efficiency based on third-generation sequencing of the present invention, through the sequencing error correction model constructed in the invention, for the first time applies the third-generation sequencing technology to the field of gene editing efficiency and target site mutation detection. The detection results show high accuracy. In addition, in the presence of multiple targets, this method can also perform linkage analysis on the mutation sites between multiple targets.

[0199] The above are only a limited number of preferred embodiments of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A rapid detection method for the gene editing efficiency of CRISPR-cas9 and the gene editing linkage phenomenon based on the third-generation sequencing technology, characterized in that: The following steps are involved: S1. Sequence pre-processing and mutation detection: Obtain third-generation sequencing data in FASTQ format and give the target region sequence; clean the data and remove low-quality reads and adapter sequences; perform Smith-Waterman local alignment on the processed sequence and the known reference genome, and count the number and proportion of reads with the same base; compare sequencing reads and reference sequences to identify base substitution, insertion, and deletion mutations; S2. Construction of sequencing error correction model: Extract mutation events from sequencing data, combine reference sequences and PCR verification, divide mutations into true positives and false positives, and divide training sets and test sets; screen classification features, including adjacent base sequences, reads positions, base quality scores, GC content, coverage depth, distance between mutation position and cutsite, and mutation frequency; construct a naive Bayes model NBM, train and optimize; use the test set to evaluate model performance, measured by accuracy, sensitivity, specificity and other indicators; S3, model-based correction of erroneous sites: Use the model of S2 to classify the mutations detected in step S1 to determine true positives or false positives; locate the position of false positive mutations according to the results of step S1, including single base substitutions, insertions, and deletions; restore false positive mutations, restore SNPs to original bases, and remove or restore insertions and deletions; S4. Gene editing mutation linkage analysis: Single-target samples directly analyze editing site mutations; multi-target samples construct mutation matrices and cluster them to identify mutation linkage "clusters", analyze the linkage effects of multi-target gene editing mutations, and understand simultaneous mutations and their effects on gene editing efficiency or specificity.

2. According to claim 1, a rapid detection method for the gene editing efficiency of CRISPR-cas9 and the gene editing linkage phenomenon based on the third-generation sequencing technology, characterized in that: The S1, sequence pre-processing and mutation detection, includes the following specific steps: S1.1 Acquisition of sequencing data: The third-generation sequencing data is in FASTQ format, and the target region sequence is given; S1.2 Filter low-quality reads: Perform preliminary cleaning of sequencing data and remove low-quality reads to ensure data accuracy in subsequent steps; S1.3 Adapter pruning: Remove the connected adapter sequence from the reads to retain the actual useful sequence data; S1.4 Alignment to the reference genome: The processed sequences are aligned to the known reference genome based on the Smith-Waterman local alignment algorithm, and the reads with the same base at each coordinate position are counted into a group. The number and proportion of reads in each group are counted, and the representative read sequence is given for each group. S1.5 Mutation Identification: Sequencing reads are compared with the target reference genome sequence to identify base substitutions, insertions, and deletions. If the bases of sequencing reads are different from those of the reference sequence, they are marked as substitutions; if sequencing reads contain more base fragments than the reference sequence, they are defined as insertions; if sequencing reads lack the base fragments corresponding to the reference sequence, they are defined as deletions.

3. According to claim 1, a rapid detection method for the gene editing efficiency of CRISPR-cas9 and the gene editing linkage phenomenon based on the third-generation sequencing technology, characterized in that: The S2, sequencing error correction model construction includes the following specific steps: S2.1 Sample collection, training set and test set division: Mutation events were extracted from the raw data of the third-generation sequencing, combined with reference sequences and PCR verification samples, and analyzed. There are two categories: True positive mutation: a real mutation caused by a target mutation or off-target effect of gene editing; False positive mutation: pseudo mutation caused by sequencing error; The collected data set is divided into a training set and a test set according to a 1:1 ratio, which are used for model training and Performance verification; S2.2 Classification feature screening: 1) k-mer feature of sequence information of adjacent bases: select 5bp of adjacent base information before and after as a k-mer feature input; 2) Reads position information: record the position of the mutation site in the reads; 3) Base Quality Score: records the base quality score of each mutation site; 4) GC content: GC ratio of ±10 bp before and after the mutation site was counted; 5) Coverage Depth: measures the number of reads supporting each mutation; 6) Distance between the mutation site and the cutsite: Measure the distance between the mutation site and the cleavage site of the Cas9 enzyme; 7) Mutation frequency: record the frequency of occurrence of mutation sites in all reads; S2.3 Model construction, training and optimization Construct a Naive Bayes model, apply Bayes' theorem to estimate the posterior probability that a specific data sample belongs to each category, and select the category with the highest probability as the predicted classification of the sample; name the constructed model the sequencing error correction model NBM; each site in the training set data contains characteristic information such as k-mer information, reads position information, base quality score, GC content, coverage depth, distance between the mutation location and cutsite, and mutation frequency. Each site has a category label, true positive mutation or false positive mutation. Input the training set data into the NBM model, train and learn the NBM model, and further optimize the parameters to improve the prediction ability of the NBM model; S2.4 Model performance evaluation: The test set was used to evaluate the performance of the NBM model. The prediction results of the NBM model were compared with the true positive and false positive labels of the test set to verify the classification performance of the NBM model. The following indicators were recorded: accuracy, sensitivity / recall, specificity, positive predictive value (PPV), and negative predictive value (NPV). The above performance evaluation indicators were used to measure the performance of the NBM model in identifying true positive sites and false positive sites.

4. According to claim 1, a rapid detection method for the gene editing efficiency of CRISPR-cas9 and the gene editing linkage phenomenon based on the third-generation sequencing technology, characterized in that: The S3, correcting the error site based on the model, comprises the following specific steps: S3.1 uses the model to classify the detected mutations in step S1 to determine whether the mutation site is a true positive mutation or a false positive mutation type; S3.2 For each read, the mutations classified as false positive mutations by the model are restored to obtain a sequence that theoretically does not carry false positive mutations; First, locate the mutation position: in each read, the position of the false positive mutation has been determined in the sequence mutation detection step in step 1; then replace the mutation: For each false positive mutation, we restore the original sequence. For SNP single nucleotide polymorphism mutations, if the detection is false positive, we restore it to the original base at that position; for insertion and deletion mutations, if it is marked as a false positive, we remove the inserted base or restore the deleted base; Assuming that the mutation is located at a specific position, we restore the bases at that position or the region of indel to the original sequenced bases.

5. According to claim 1, a rapid detection method for the gene editing efficiency of CRISPR-cas9 and the gene editing linkage phenomenon based on the third-generation sequencing technology, characterized in that: The S4, gene editing mutation linkage analysis, includes the following specific steps: If it is a single-target gene editing sample, directly analyze the editing site mutation of the sample. If it is a multi-target gene editing sample, use the following steps to perform multi-target mutation linkage analysis: S4.1 constructs a mutation matrix based on the true positive mutations on each read, and clusters the mutation matrix to obtain mutation linkage "clusters"; The mutation matrix is ​​a two-dimensional matrix, in which rows represent different reads clusters and columns represent mutation sites at each coordinate position; each element of the matrix represents whether the mutation at that position appears in that read; hierarchical clustering of the mutation matrix is ​​performed to identify which mutations often appear together in different reads, thus forming "clusters" of linked mutations; S4.2 Discover the linkage effect of gene editing mutations at multiple targets based on the information in the "clusters". When editing multiple target sites in the genome, a mutation linkage effect may occur, that is, some mutations may appear simultaneously under the same genotype, affecting the effect of gene editing. By analyzing the "clusters" of mutations, it is possible to identify which mutations occur simultaneously, and these combined mutations may affect the efficiency or specificity of gene editing.

6. A storage device for executing the method for rapid detection of CRISPR-cas9 gene editing efficiency and linkage phenomenon based on third-generation sequencing technology as described in claim 1, characterized in that: The storage device stores computer executable instructions, and when the computer executable instructions are executed by the processor, the method steps described in claim 1 are implemented.

7. A computer device, characterized in that: include: processor; A memory for storing computer executable instructions; Wherein, the processor is configured to execute computer executable instructions stored in the memory to implement the method for rapid detection of CRISPR-cas9 gene editing efficiency and linkage phenomenon based on third-generation sequencing technology as described in claim 1.

8. A server, characterized in that: The server comprises: A data receiving module, used to receive third-generation sequencing data in FASTQ format; A data processing module for performing the sequence pre-processing, mutation detection, sequencing error correction model construction, model-based correction of error sites, and gene editing mutation linkage analysis steps described in claim 1; The result output module is used to output the analysis results of gene editing efficiency and linkage phenomenon.

9. A cloud computing-based platform, characterized in that: The platform includes multiple servers, each of which is used for distributed processing and analysis of third-generation sequencing data from different samples as described in claim 8, so as to achieve the rapid detection of CRISPR-cas9 gene editing efficiency and linkage phenomenon based on third-generation sequencing technology as described in claim 1.

10. A gene editing analysis system, characterized in that: The system comprises: Sequencer, used to generate third-generation sequencing data in FASTQ format; A computer device as claimed in claim 7, for processing and analyzing the sequencing data; Display device used to show the analysis results of gene editing efficiency and linkage phenomenon.