A method for quality control of full-length plasmids
By using the Cas9/sgRNA complex targeting the plasmid backbone sequence and nanopore sequencing technology, the problems of short read length, inability to control full length, and contamination identification in existing plasmid sequencing methods have been solved, achieving low-cost and efficient full-length plasmid quality control.
Patent Information
- Application Number
- CN202310557911.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-05-17
AI Technical Summary
Existing plasmid sequencing methods have short read lengths, making it impossible to achieve quality control of full-length plasmids, identify minor contamination of plasmids, and are costly.
The circular plasmid was linearized using a Cas9/sgRNA complex targeting the plasmid backbone sequence, followed by nanopore sequencing. Combined with data preprocessing and analysis, full-length plasmid quality control was achieved.
It enables low-cost, high-efficiency full-length plasmid sequencing, can identify minor plasmid contamination, provides full-length plasmid quality control information, simplifies data analysis, and reduces sequencing costs.
Smart Images

Figure CN117275576B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology and relates to a method for quality control of full-length plasmids. Background Technology
[0002] Plasmids are widely used in the field of biology for gene delivery, mRNA synthesis, or direct production of plasmid DNA vaccines. They are also the basis for packaging viral vectors in gene therapy and cell therapy, such as adenovirus (AdV), adeno-associated virus (AAV), and lentivirus (LV). Therefore, accurate validation and contamination control of plasmid sequences are essential for the production of high-quality viral vectors and other clinical therapeutic products.
[0003] Currently, many commercial services can identify plasmid sequences, with Sanger sequencing being the most common. While this technology offers high accuracy, the read length is relatively short (typically <1kb), making it impossible to verify the correctness of the full-length plasmid sequence and hindering high-throughput sequencing. Furthermore, it's difficult to detect minor plasmid contamination by examining Sanger sequencing peak diagrams. Although deep sequencing technology using the Illumina platform can identify minute contamination, piecing together mixed sequencing fragments from multiple plasmids remains challenging, especially for plasmid vectors containing repetitive sequences and strong secondary structures. Moreover, this method is time-consuming and costly. Additionally, both first- and second-generation sequencing technologies rely on DNA amplification, often failing to read DNA fragments with strong secondary structures, such as inverted terminal repeats (ITRs) in AAV vectors.
[0004] With the rapid development of gene therapy and cell therapy in recent years, there is an urgent need to develop more precise and cost-effective methods for quality control of full-length plasmids. Since Oxford Nanopore Technologies (ONT) released the first nanopore sequencer, MinION, in 2014, nanopore sequencing technology has played an increasingly important role in basic and applied research. This technology relies on nanoscale protein pores (nanopores), which are essentially channels formed on a membrane. The nanopores are embedded in a synthetic membrane and immersed in an electrophysiological solution. Applying a constant voltage allows an ionic current to pass through the nanopore. When molecules such as DNA or RNA pass through the nanopore, they interfere with the current. The change in ionic current corresponds to the nucleotide sequence present in the sensing region. Decoding this through computer algorithms enables real-time analysis of individual molecules to determine the base sequence of the DNA or RNA strand passing through the nanopore. Nanopore sequencing technology can sequence ultra-long sequences at the single-molecule level, making full-length plasmid sequencing possible.
[0005] In summary, currently used plasmid sequencing methods suffer from problems such as short read lengths, inability to achieve quality control of full-length plasmids, and inability to detect minor plasmid contamination. Developing a low-cost, high-efficiency full-length plasmid quality control method based on nanopore sequencing technology holds promise for providing numerous gene therapy companies with a powerful full-length plasmid sequencing platform, thus ensuring quality in the development of gene therapy vectors. Summary of the Invention
[0006] To address the shortcomings of existing technologies and practical needs, this invention provides a method for full-length plasmid quality control, which solves the problems of short read lengths, inability to achieve quality control of full-length plasmids, and inability to identify minor contamination in traditional plasmid sequencing methods. It enables low-cost and high-efficiency full-length sequencing of multiple plasmid sequences in the same batch, while also being able to identify minor plasmid contamination.
[0007] To achieve this objective, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a method for quality control of full-length plasmids, the method comprising: linearizing a circular plasmid using a Cas9 / sgRNA complex targeting the plasmid backbone sequence, then performing nanopore sequencing to obtain raw data, and performing data analysis on the obtained raw data after preprocessing to obtain quality control information of the full-length plasmid.
[0009] The method of this invention enables low-cost and high-efficiency full-length plasmid sequencing without PCR dependence. It can directly sequence ultra-long plasmid sequences at the single-molecule level and can also identify minor plasmid contamination. Furthermore, it allows for mixed sequencing of multiple plasmids from the same batch of samples. A schematic diagram of the full-length plasmid quality control method of this invention is shown below. Figure 1 As shown.
[0010] Preferably, the nucleic acid sequence of the sgRNA in the Cas9 / sgRNA complex targeting the plasmid backbone sequence includes the sequences shown in SEQ ID NO.1-SEQ ID NO.15.
[0011] SEQ ID NO. 1: AATAAACCAGCCAGCCGGAA.
[0012] SEQ ID NO. 2: TGGAACGAAAACTCACGTTA.
[0013] SEQ ID NO. 3: TTTTGCTCACCCAGAAACGC.
[0014] SEQ ID NO. 4: GACAACGATCGGAGGACCGA.
[0015] SEQ ID NO. 5: CTACAGAGTTCTTGAAGTGG.
[0016] SEQ ID NO. 6: CAGCGGTCGGGCTGAACGGG.
[0017] SEQ ID NO. 7: TGTGATGCTCGTCAGGGGGG.
[0018] SEQ ID NO. 8: GGCCGATTCATTAATGCAGC.
[0019] SEQ ID NO.9: TAAAGTTCTGCTATGTGGCG.
[0020] SEQ ID NO. 10: CACCACGATGCCTGTAGCAA.
[0021] SEQ ID NO. 11: TCGTAGTTATCTACACGACG.
[0022] SEQ ID NO. 12: TGCTTCAATAATATTGAAAA.
[0023] SEQ ID NO. 13: AAAACCACCGCTACCAGCGG.
[0024] SEQ ID NO. 14: GCTTCCCGAAGGGAGAAAGG.
[0025] SEQ ID NO. 15: GGTATCAGCTCACTCAAAGG.
[0026] Preferably, the data analysis process includes preprocessed data splitting and visualization analysis, sequencing integrity analysis, data early warning analysis, detection and analysis of insertions, deletions and mutations in plasmids, and contamination identification analysis.
[0027] Preferably, the preprocessed data splitting and visualization analysis includes:
[0028] Six 17nt grep sequences were generated based on the 45-50bp upstream and downstream sequences of the Cas9 / sgRNA cleavage site. The original data was split, and the original data and corresponding reference sequences were sorted. Visual images were obtained using the IGV visualization tool.
[0029] The specific point values in the range of 45-50 can be selected as 45, 46, 47, 48, 49, 50, etc.
[0030] Preferably, the sequencing integrity analysis includes:
[0031] Read the data from each plasmid. Based on the length of the vector reference sequence, the sequencing sequence length is considered complete if it is within 90-110% of the reference sequence length. The proportion of complete reads to the total number of reads of the plasmid is considered the completeness rate.
[0032] Preferably, the data early warning analysis includes:
[0033] Sequencing data is compared with a reference sequence, and a difference threshold is set. Data exceeding the difference threshold is identified as warning data, while data below the threshold is identified as normal data. The difference threshold is defined as 90% similarity to the full length of the reference sequence.
[0034] Preferably, the insertion, deletion, and mutation detection analysis includes:
[0035] A consensus sequence is extracted from the normal data after the early warning process. The consensus sequence is compared with the reference sequence, and anomalies are automatically identified. Regions with mutations, insertions, or missing values are recorded in the form of a report.
[0036] Preferably, the pollution identification and analysis includes:
[0037] Analyze the early warning data to identify contaminations that differ significantly from the original plasmid sequence and minor contaminations with a difference of 20 nt or more. The contamination identification includes the contamination category, the amount of contamination data, the contamination rate, the contamination data consensus sequence, and the alignment results between the contamination data consensus sequence and the reference sequence.
[0038] Preferably, the plasmid includes any one or a combination of at least two of lentiviral vector plasmids, adeno-associated virus plasmids, adenovirus plasmids, or mRNA in vitro transcription plasmids.
[0039] Secondly, the present invention provides a software package for analyzing the quality control information of full-length plasmids, the software package comprising:
[0040] The system includes modules for preprocessed data splitting and visualization analysis, sequencing integrity analysis, data early warning analysis, plasmid insertion, deletion and mutation detection and analysis, and contamination identification analysis.
[0041] The preprocessed data splitting and visualization analysis module is used to perform the following:
[0042] Six 17nt grep capture sequences were generated based on the 45-50bp upstream and downstream sequences of the Cas9 / sgRNA cleavage site. The original data was split, and the original data and the corresponding reference sequences were sorted. Visual images were obtained using the IGV visualization tool.
[0043] The sequencing integrity analysis module is used to perform the following:
[0044] Read the data from each plasmid. Based on the length of the vector reference sequence, a sequence length within 90-110% of the reference sequence length is considered complete. The proportion of complete reads to the total number of reads of the plasmid is considered the completeness rate.
[0045] The data early warning and analysis module is used to perform the following:
[0046] Sequencing data is compared with a reference sequence, and a difference threshold is set. Data exceeding the difference threshold is identified as warning data, while data below the threshold is identified as normal data. The difference threshold is defined as 90% similarity to the full length of the reference sequence.
[0047] The insertion, deletion, and mutation detection and analysis module is used to perform the following:
[0048] A consensus sequence is extracted from the normal data after the early warning process. The consensus sequence is compared with the reference sequence, and anomalies are automatically identified. Regions with mutations, insertions, or missing values are recorded in the form of a report.
[0049] The pollution identification and analysis module is used to perform the following:
[0050] Analyze the early warning data to identify contaminations that differ significantly from the original plasmid sequence and minor contaminations with a difference of 20 nt or more. The contamination identification includes the contamination category, the amount of contamination data, the contamination rate, the contamination data consensus sequence, and the alignment results between the contamination data consensus sequence and the reference sequence.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] (1) The method of the present invention can efficiently achieve full-length plasmid sequencing without relying on PCR. It can directly sequence ultra-long plasmid sequences at the single-molecule level, and can also identify a small amount of plasmid contamination. The cycle for obtaining full-length plasmid quality control information is short and the cost is low.
[0053] (2) In order to simplify data analysis, this invention developed an automated analysis process based on Python. It uses 45-50bp sequences upstream and downstream of the plasmid cutting site to split the data and automatically outputs the consensus sequence of the sequenced plasmid, alignment results, integrity analysis, contamination rate and visualization results to a single file. It can complete the full-length sequencing analysis of various types of plasmids, such as lentiviral vectors of more than 10kb, adenovirus vectors of 30-40kb, AAV vectors with ITR structure, and even full-length sequencing of mRNA in vitro transcription (IVT) plasmids with 90 consecutive adenine bases (A90). It can simultaneously sequence multiple plasmids in the same batch, which greatly saves time and sequencing costs.
[0054] (3) This invention extracts consensus sequences from multiple sequencing reads to compensate for the deficiencies of systematic errors. Without providing a reference sequence, more than 100 reads can be used to obtain a consensus sequence with an accuracy of over 99.5%, and can effectively identify plasmid contamination of >0.1%. It is expected to provide a powerful full-length plasmid sequencing platform for many gene therapy companies and provide quality assurance for the development of gene therapy vectors. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of the full-length plasmid quality control method of the present invention;
[0056] Figure 2 This is a visualization of the IGV sequenced from the representative plasmid (8kb) obtained in Specific Example 3;
[0057] Figure 3 This is a visualization of the IGV sequenced from the representative plasmid (18kb) obtained in Specific Example 3;
[0058] Figure 4 This is a visualization of the IGV sequenced from the representative plasmid (35kb) obtained in Specific Example 3;
[0059] Figure 5 This is a visualization of IGV sequencing of the AAV vector plasmid obtained in specific embodiment 3;
[0060] Figure 6 This is a visualization of the A90 fragment in the in vitro transcription (IVT) plasmid of mRNA obtained in Specific Example 3;
[0061] Figure 7 A correlation diagram showing the relationship between the Sanger sequencing ployA length and the nanopore sequencing ployA length obtained in specific Example 3;
[0062] Figure 8 This is a graph showing the relationship between sequencing integrity and plasmid length obtained through specific embodiment 3;
[0063] Figure 9 The graph shows the results of plasmid sequencing integrity in different batches obtained by library construction using the method in Specific Example 2. Detailed Implementation
[0064] To further illustrate the technical means and effects of this invention, the following description, in conjunction with embodiments and accompanying drawings, provides a further explanation of the invention. It is understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it.
[0065] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased through legitimate channels.
[0066] Example 1
[0067] Linearization of lentiviral vector plasmids, adeno-associated virus plasmids, adenovirus plasmids, and mRNA external transcription plasmids was achieved using the Cas9 / sgRNA complex.
[0068] In a 200 μL PCR eight-tube container, 2 μg of Cas9 protein (12 pmol, purchased from IDT) was mixed with 0.8 μL of annealed sgRNA (SEQ ID NO.1-SEQ ID NO.15, purchased from IDT) (24 pmol). The mixture was incubated at 25°C for 10 min to form a Cas9 / sgRNA complex. 500 ng of plasmid and 1 μL of 10×NEB3.1 buffer were added, and double-distilled water was added to a total volume of 10 μL. The mixture was incubated at 37°C for 30 min. Then, 0.1 μL of proteinase K was added to the reaction mixture, and the reaction was terminated at 56°C for 15 min. The mixture was purified using 1.2 volumes of Megabeads magnetic beads, and finally eluted with 10 μL of double-distilled water to obtain the purified linearized plasmid.
[0069] Example 2
[0070] Plasmid sequencing was performed using the MinION sequencing platform.
[0071] (1) Linearized plasmid DNA library construction
[0072] (a) High-quality sample quality control:
[0073] First, check whether the DNA sample obtained in Example 1 has any abnormal appearance. Use agarose gel electrophoresis to detect whether the sample has been degraded, use Nanodrop to detect DNA purity, and use a Qubit 4.0 instrument to detect DNA concentration.
[0074] (b) After the sample passes quality inspection, the sample volume is concentrated to 49 μL using 1.2 times the volume of Megabeads magnetic beads;
[0075] (c) Damage and end repair of the target DNA fragment, and purification of the DNA using 1 volume of Megabeads magnetic beads after the reaction;
[0076] (d) The DNA product obtained in the previous step was ligated into sequencing adapters using the SQK-LSK110 kit. After the reaction, the DNA was purified using 0.4 times the volume of Megabeads magnetic beads, and the constructed DNA library was quantified using Qubit4.0.
[0077] (2) Sequencing library loading
[0078] (a) Thaw sequencing buffer II (SQBII), loading particles II (LBII), flush tether (FLT) and one flush buffer tube (FB) at 25°C;
[0079] (b) Pretreatment solution for preparing sequencing chips: Add 30 μL of thawed and mixed wash tether (FLT) to a whole tube of thawed and mixed wash buffer (FB), and vortex mix at 25°C.
[0080] (c) Open the lid of the MinION Mk1C sequencer, insert the R9.4.1 sequencing chip into the metal clamp, and perform a chip quality check to see the number of usable nanopores;
[0081] (d) Rotate the pretreatment well cap clockwise to expose the pretreatment well, check for small air bubbles around the well, and pretreatment the sequencing chip with pretreatment solution.
[0082] (e) Prepare the sample library, which includes 37.5 μL sequencing buffer II (SBII), 25.5 μL sample loading particles II (LBII), and 12 μL DNA library. (7) Add 75 μL of sample dropwise to the chip through the SpotON sample well, ensuring that the droplet flows into the well before adding the next drop.
[0083] (f) Gently close the SpotON sample loading well cap, rotate the pretreatment well cap counterclockwise to close the pretreatment well, and finally close the top cover of the MinION Mk1C sequencer to start sequencing.
[0084] Figure 9 The results show that the content of this invention has good reproducibility. The plasmid sequencing integrity rate in different batches obtained by library construction using the method of Specific Example 2 is as follows.
[0085] Example 3
[0086] Plasmid sequencing data analysis
[0087] Data preprocessing:
[0088] The raw data obtained from nanopore sequencing consists of several *.fastq.gz files. First, the data is merged. This is done using the Linux command "cat input1.fastq.gz input2.fastq.gz>output.fastq.gz", which is then encapsulated in a self-made Python script (cat_fq_gz.py). Running this script merges all the raw data into a single NP.fastq.gz file. Then, in the corresponding folder, open the WSL environment with Porechop installed and run the Porechop command (porechop -i NP.fastq.gz -o NPchop.fastq.gz) to obtain the adapter-deselected NPchop.fastq.gz file.
[0089] Data Analysis:
[0090] This invention is mainly divided into five modules: data splitting and visualization, sequencing integrity analysis, data early warning, insertion / deletion and mutation detection, and contamination identification. A Python code package named NPlasmid-seq was written and was developed. The package takes the preprocessed sequencing file NPchop.fastq.gz, the original vector sequence, and the sgRNA recognition sequence as input. The required environment variables are configured in WSL, and the NPlasmid-seq command can be run to automatically obtain a data analysis report with one click.
[0091] (1) Data splitting and visualization
[0092] Since the same batch of data contains multiple plasmid vectors, NPlasmid-seq first generates grep sequences based on the 50bp upstream and downstream sequences of the Cas9 / sgRNA cleavage site. Then, it uses a Python script to split the raw data, with each plasmid corresponding to an independent .fastq file. Next, based on the provided vector raw files, it sequentially calls the Minimap2 and SAMtools toolkits to sort the sequencing data with the corresponding reference sequences and obtain .bam files. Finally, it imports the .fasta and corresponding .bam files using the IGV visualization tool to obtain visualization images.
[0093] The code is as follows:
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102] (2) Sequencing integrity analysis
[0103] NPlasmid-seq automatically reads the data from each plasmid. Based on the vector reference sequence length, sequences within 90%-110% of the reference sequence length are considered complete. The percentage of complete reads to the total number of reads in the plasmid is considered the integrity rate, and the sequencing integrity of all plasmids is output. In this invention, we have analyzed the sequencing integrity of plasmids with lengths ranging from 2.5 to 40 kb.
[0104] The code is as follows:
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112] (3) Data early warning
[0113] NPlasmid-seq automatically processes .fastq files in batches, comparing sequencing data with reference sequences. We set a difference threshold; data exceeding this threshold are considered significantly different from theoretical values and are thus flagged as warning data, while data within the threshold are considered normal. All data is ultimately split into two files: "NormalData" containing normal and relatively complete sequencing data, and "FlaggedData" containing other abnormal data. Abnormal data may be caused by sequencing errors or contamination from other vectors.
[0114] The code is as follows:
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129] (4) Insertion, deletion and mutation detection
[0130] NPlasmid-seq can call the longread_umi package to extract the consensus sequence from the "NormalData" after the warning processing, compare the consensus sequence with the reference sequence, and automatically identify outliers. Regions with mutations, insertions, and missing values are recorded in the form of reports.
[0131] The code is as follows:
[0132]
[0133]
[0134]
[0135] (5) Pollution identification
[0136] NPlasmid-seq automatically analyzes "FlaggedData" to detect the presence of contamination data originating from other plasmids. After the contamination detection is completed, all information is recorded in the form of a report. The report includes the report time, program author, vector name, reference sequence, amount of correct data, amount of warning data, and contamination detection results (contamination category and corresponding amount of contamination data, contamination rate, consensus sequence of contamination data and comparison results with the reference sequence).
[0137] The code is as follows
[0138]
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145] Quality control results: Through the above analysis process, a quality control report will be automatically generated, including IGV visualization, consensus sequence and alignment results of complete plasmid sequencing data, amount of contaminated data and consensus sequence. Figure 2 The image shows a visualization of the IGV sequence of the representative plasmid (8kb), demonstrating that nanopore sequencing can achieve full-length sequencing of 8kb long plasmids. Figure 3 The image shows a visualization of the IGV sequence of the representative plasmid (18kb), demonstrating that nanopore sequencing can achieve full-length sequencing of 18kb long plasmids. Figure 4 The IGV visualization of the obtained representative plasmid (35kb) shows that nanopore sequencing can achieve full-length sequencing of 35kb long plasmids. Figure 5 The image shows a visualization of the IGV sequence of the AAV vector plasmid. In the image, BB represents the plasmid backbone, BDDF8 represents the human coagulation factor VIII gene expression box with the B domain deleted, and the area circled in the box is the ITR (inverted terminal repeat). The results show that NPlasmid-seq can successfully sequence the ITR sequence with strong secondary structure in the AAV vector. Figure 6 The image shows a visualization of the A90 fragment in the obtained mRNA in vitro transcription (IVT) plasmid. After magnifying the A90 region, it can be observed that the number of A bases in all reads is not equal, indicating that nanopore sequencing cannot accurately determine 90 consecutive A bases. Figure 7The graph shows the correlation between the length of Sanger sequencing PloyA and the length of nanopore sequencing PloyA obtained in specific example 3. The correlation was calculated using Pearson linear regression analysis, and the length of nanopore sequencing PloyA can be corrected to the true length using the formula Y = 0.74X + 2.82. Figure 8 The graph shows the relationship between sequencing integrity and plasmid length. The results indicate that the longer the plasmid is, the worse the sequencing integrity of the nanopore sequencing.
[0146] Example 4
[0147] The practicality of the NPlasmid-seq method of this invention was tested by artificially creating plasmid contamination.
[0148] We used two methods to test whether NPlasmid-seq could effectively identify plasmid contamination as low as 0.1%: ① Two sgRNA plasmids with only 20 nt difference in their original sequences were digested using the same Cas9 / sgRNA complex, with contamination rates set at 0.5%, 0.2%, and 0.1%, respectively. The plasmids were then mixed and simultaneously digested, purified, and library constructed. The final analysis showed that with a 0.5% contamination rate, the total number of reads was 6417, with 112 reads triggering warnings. After alignment with the contaminated plasmid library sequence, the contamination was determined... The actual number of contaminated sgRNA reads was determined to be 35, with a contamination rate of 0.545%. With a contamination rate of 0.2%, the total number of reads was 21,148, with 283 warning reads. After alignment with the contaminated plasmid library sequence, the actual number of contaminated sgRNA reads was determined to be 49, with a contamination rate of 0.232%. With a contamination rate of 0.1%, the total number of reads was 6,901, with 86 warning reads. After alignment with the contaminated plasmid library sequence, the actual number of contaminated sgRNA reads was determined to be 8, with a contamination rate of 0.116%. ② Based on the read count, sequencing FastQ files cut at the same site but originating from two different plasmids were merged. In this example, the original plasmid sequencing FastQ reads were 4409, and the contaminated plasmid FastQ reads were 88, 44, 22, and 5, respectively, meaning the contamination rates were determined to be 2%, 1%, 0.5%, and 0.1%. NPlasmid-seq analysis showed that the detected contamination rates were 1.86%, 0.94%, 0.48%, and 0.09%, respectively. The results indicate that the NPlasmid-seq method of this invention can effectively identify contamination rates as low as 0.1%, whether the contamination is at the plasmid source or in subsequent sequencing data.
[0149] In summary, the method of the present invention can sequence ultra-long plasmid sequences directly at the single-molecule level at low cost and high efficiency, without relying on PCR, and can also identify minor contamination of plasmids and obtain a full-length plasmid quality control report.
[0150] The applicant declares that the detailed method of the present invention is illustrated by the above embodiments, but the present invention is not limited to the above detailed method, that is, it does not mean that the present invention must rely on the above detailed method to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions of the raw materials of the product of the present invention, addition of auxiliary components, selection of specific methods, etc., all fall within the protection scope and disclosure scope of the present invention.
Claims
1. A method for quality control of full-length plasmids, characterized by, The method comprises linearizing a circular plasmid by using a Cas9 / sgRNA complex targeting a plasmid backbone sequence, then performing nanopore sequencing to obtain raw data, and performing data analysis on the raw data after preprocessing to obtain quality control information of the full-length plasmid. The nucleic acid sequence of the sgRNA in the Cas9 / sgRNA complex targeting the plasmid backbone sequence comprises the sequence shown in SEQ ID NO. 1-SEQ ID NO.
15. The data analysis process comprises preprocessing data shunting and visualization analysis, sequencing integrity analysis, data warning analysis, insertion, deletion and mutation detection analysis existing in the plasmid, and contamination identification analysis. The preprocessing data shunting and visualization analysis comprises: generating 6 17 nt grep capture sequences according to the 45-50 bp sequences upstream and downstream of the Cas9 / sgRNA cleavage site, shunting the raw data, sorting the raw data with the corresponding reference sequence, and obtaining a visualization picture by using an IGV visualization tool. The sequencing integrity analysis comprises: reading the data shunted from each plasmid, taking the length of the vector reference sequence as the standard, and regarding the sequencing sequence length within the range of 90-110% of the reference sequence length as complete, and regarding the proportion of complete reads in the total reads of the plasmid as the completion rate. The data warning analysis comprises: comparing the sequencing data with the reference sequence, setting a difference threshold, and determining the data exceeding the difference threshold as warning data, and otherwise determining as normal data; the difference threshold is 90% similarity with the full-length reference sequence. The insertion, deletion and mutation detection analysis comprises: extracting a consensus sequence from the normal data after warning processing, comparing the consensus sequence with the reference sequence, and automatically identifying abnormal points to record the regions with mutations, insertions and deletions. The contamination identification analysis comprises: analyzing the warning data, identifying the contamination with a larger difference from the original plasmid sequence and the micro-contamination with a difference of 20 nt or more, and the content of the contamination identification comprises the contamination category, the contamination data amount, the contamination rate, the contamination data consensus sequence and the comparison result of the contamination data consensus sequence with the reference sequence.
2. The method of full-length plasmid quality control according to claim 1, wherein, The plasmid comprises any one or a combination of at least two of a lentivirus vector plasmid, an adeno-associated virus plasmid, an adenovirus plasmid or an mRNA in vitro transcription plasmid.
3. A software package for analyzing quality control information of full-length plasmids, characterized in that, The software package comprises: a preprocessing data shunting and visualization analysis module, a sequencing integrity analysis module, a data warning analysis module, an insertion, deletion and mutation detection analysis module existing in the plasmid, and a contamination identification analysis module; The preprocessing data shunting and visualization analysis module is used to perform: generating 6 17 nt grep capture sequences according to the 45-50 bp sequences upstream and downstream of the Cas9 / sgRNA cleavage site, shunting the raw data, sorting the raw data with the corresponding reference sequence, and obtaining a visualization picture by using an IGV visualization tool; The sequencing integrity analysis module is used to perform: Read the data of each plasmid shunt, according to the length of the vector reference sequence, the length of the sequencing sequence is in the range of 90-110% of the reference sequence length is considered complete, the proportion of complete reads in the total reads of the plasmid is considered as the complete rate; The data early warning analysis module is used to execute and include: The sequencing data is compared with the reference sequence, a difference threshold is set, and the data exceeding the difference threshold is determined as early warning data, otherwise it is determined as normal data; the difference threshold is 90% similarity with the full length of the reference sequence; The insertion, deletion and mutation detection analysis module is used to execute and include: The consensus sequence is extracted from the normal data after early warning processing, the consensus sequence is compared with the reference sequence, and the abnormal points are automatically identified, and the regions with mutations, insertions and deletions are recorded in the form of reports; The pollution identification analysis module is used to execute and include: The early warning data is analyzed, and the pollution with large difference from the original plasmid sequence and the micro pollution with difference greater than or equal to 20 nt are distinguished, and the pollution identification content includes pollution category, pollution data volume, pollution rate, pollution data consensus sequence and comparison result of pollution data consensus sequence and reference sequence.
Citation Information
Patent Citations
DNA next generation sequencing whole-process quality analysis method
CN109637581A
Method for detecting intracellular gene editing efficiency and off-target
CN113005185A