Method and device for analyzing sequencing data of low-frequency Indel mutation generated by CRISPR (clustered regularly interspaced short palindromic repeats) gene editing
Through the dual-ended sequencing data acquisition and high-throughput sequencing platform combined with the analysis of CRISPR software, the problem of detecting low-frequency Indel mutations in CRISPR gene editing in the prior art is solved, and efficient and accurate gene editing frequency and pattern analysis is achieved, reducing false positive rates and detection costs.
Patent Information
- Application Number
- CN202510618485.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
AI Technical Summary
Existing gene editing efficiency detection methods have problems such as low detection throughput, time-consuming, expensive and inaccurate results when detecting low-frequency Indel mutations in animal tissues or adults. In particular, existing software has high false positive rates when detecting large-scale Indel mutations, making it difficult to accurately evaluate the frequency and pattern of CRISPR gene editing.
A sequencing data analysis method for low-frequency Indel mutations generated by CRISPR gene editing is adopted, including double-ended sequencing data acquisition, demultiplexing, removal of linker sequences, forward and reverse data splicing and alignment analysis. The data that does not meet the conditions are discarded, data splicing and parameter setting are performed through the designed python script, and high-throughput sequencing is performed by combining the Illumina sequencing platform.
It realizes rapid and accurate analysis of low-frequency and large-variance Indel mutations generated by CRISPR gene editing, significantly improves detection efficiency, reduces false positive rates, and can directly output results without setting up control groups, providing higher data throughput and more accurate gene editing efficiency and pattern information.
Smart Images

Figure CN120544683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sequencing data analysis method and device for low-frequency Indel mutations generated by CRISPR gene editing, and belongs to the field of bioinformatics. Background Art
[0002] Gene editing technology has been widely used and has shown great potential in gene research, gene therapy and genetic improvement.
[0003] The methods and tools for gene editing currently reported in the existing technologies mainly include ZFN (zinc finger nuclease), TALEN (transcription activator-like effector nuclease), and CRISPR; among them, CRISPR technology is widely used in animal breeding and improvement due to its advantages of high efficiency and low cost.
[0004] After gene editing, the target DNA sequence needs to be detected and analyzed to determine the efficiency of gene editing.
[0005] In the past, traditional methods for detecting gene editing efficiency mainly used the T7 nuclease digestion method and the TA cloning Sanger sequencing method. However, the detection throughput of these two methods is very low, the detection process is time-consuming, labor-intensive, and expensive, and the identification results are not very accurate.
[0006] Currently, the most advanced methods for detecting gene editing efficiency all use high-throughput sequencing platforms based on NGS sequencing technology. Among them, the Illumina sequencing platform technology is currently the most widely used and mature high-throughput sequencing platform, with high throughput and high sensitivity.
[0007] In animal breeding, gene editing tools can be used to develop animal breeds with specific traits. However, the efficiency of direct gene editing in animal tissues or adult animals using CRISPR technology is extremely low, typically less than 1%. Therefore, if high-throughput sequencing is performed using the Illumina sequencing platform (with an error tolerance of ±0.1% to 1%), it is difficult to analyze the gene editing efficiency results from the high-throughput sequencing data it outputs.
[0008] Bioinformatics analysis software for high-throughput sequencing data, such as UMItools, smCounter, and Conner, are all suitable for analyzing data from the Illumina platform. However, these software perform well in detecting single-nucleotide variants and are unable to detect indel mutations, especially large-scale indel mutations.
[0009] The above-mentioned existing analysis software and methods have a high false positive rate (eg, about 1%). In particular, the operations of splitting and pruning data usually result in a high error splicing rate (eg, greater than 2%) in the subsequent data splicing.
[0010] In addition, when CRISPR tools are used to edit genes in animal tissues or adult animals, the induced Indel mutations have a highly dynamic range, ranging from 1bp to 50bp, and in extreme cases may be close to 100bp.
[0011] Therefore, for gene editing of animal tissues or adult animals, it is very difficult to accurately evaluate the frequency and pattern of Indel mutations produced by CRISPR gene editing. How to quickly and accurately analyze the gene editing of low-frequency and wide-variety Indel mutations produced by CRISPR gene editing from the massive high-throughput sequencing data obtained from the Illumina sequencing platform is a difficult problem faced by technicians in this field. Summary of the Invention
[0012] In order to solve the above technical problems, the present invention provides a sequencing data analysis method for low-frequency Indel mutations generated by CRISPR gene editing, wherein:
[0013] The method comprises the following steps:
[0014] Step 1) Obtaining double-end sequencing data: constructing a double-end sequencing DNA library for the sample obtained by CRISPR gene editing, and then performing high-throughput sequencing based on the Illumina sequencing platform to obtain R1 fastq data package and R2 fastq data package;
[0015] Step 2a), demultiplexing R1 data: according to the designed forward primer sequence, demultiplex the R1 fastq data packet obtained in step 1), split the R1 forward fastq data file, and obtain the corresponding R1 reverse fastq data file, and store them according to different samples;
[0016] Step 2b), demultiplexing R2 data: According to the designed reverse primer sequence, the R2 fastq data packet obtained in step 1) is demultiplexed to split the R2 reverse fastq data file and obtain the corresponding R2 forward fastq data file, and store them according to different samples;
[0017] The steps 2a) and 2b) are not ordered in any particular order;
[0018] Step 3) Merge data: merge the R1 forward fastq data file with the R2 forward fastq data file, and store them according to different samples to obtain a forward fastq data packet; and merge the R2 reverse fastq data file with the R1 reverse fastq data file, and store them according to different samples to obtain a reverse fastq data packet;
[0019] Step 4), removing the connector sequence: processing the forward fastq data packet and the reverse fastq data packet, and removing the connector sequence of each fastq data;
[0020] Step 5) Splicing forward and reverse data files: Splice the forward fastq data file and the reverse fastq data file of each identical sample, output the full-length fastq data file, and store it according to different samples;
[0021] Step 6), alignment analysis: calling Crispresso2 software, the full-length fastq data file obtained in step 4) is aligned with the target region reference sequence of the CRISPR gene editing and the sgRNA sequence used for analysis and data statistics.
[0022] In a preferred embodiment of the present application, during the comparison analysis in step 6), if any one of the following a) and b) conditions is not met, the full-length fastq data file is discarded; a) the portion of either end sequence that completely overlaps with the target region reference sequence is more than 20 bp; b) the Phred33 quality score within any consecutive 4 bp detection window is greater than 10.
[0023] In a preferred embodiment of the present application, during the comparison analysis in step 6), after ignoring the case of single base substitution, the comparison analysis results with a mutation rate of more than 0.01% are output;
[0024] The mutation rate refers to the ratio of the number of mutated bases to the total number of bases in the reference sequence of the target region.
[0025] In a preferred embodiment of the present application, in step 2a) and step 2b), the AdapterRemoval software is respectively called to execute a demultiplexing-only command without deleting, adding or modifying the metadata.
[0026] In a preferred embodiment of the present application, in step 5), forward and reverse data splicing is performed by a designed python script, and by setting parameters, the forward fastq data file and the reverse fastq data file of the same sample with an overlapping interval greater than or equal to 10bp are spliced; data that does not meet the parameters are discarded.
[0027] In a preferred embodiment of the present application, the CRISPR gene editing uses the CRISPR / Cas9 gene editing tool; preferably, the sample obtained by the CRISPR gene editing is a sample obtained by editing an adult or embryo of an animal using the CRISPR / Cas9 gene editing tool; preferably, the animal is a bird (Aves); more preferably, the bird is a Galliformes, Anseriformes, Otidiformes, Columbiformes or Struthioniformes; more preferably, the bird is a Galliformes and is selected from the family Numididae or the family Phasianidae; more preferably, the bird belongs to the order Galliformes, the family Phasianidae, or the genus Gallus.
[0028] In a preferred embodiment of the present application, in step 1), the forward primer and reverse primer designed for double-end sequencing are respectively located 150 to 250 bp upstream and downstream of the target region of the CRISPR gene editing; preferably, they are respectively located 200 bp upstream and downstream of the target region of the CRISPR gene editing.
[0029] In a preferred embodiment of the present application, in said step 1),
[0030] PE sequencing is performed using the Illumina sequencing platform; preferably, PE200 sequencing is performed using the Illumina sequencing platform.
[0031] On the other hand, the present invention provides a sequencing data analysis device for low-frequency Indel mutations generated by CRISPR gene editing, wherein the sequencing data analysis device is used to perform the steps of the above-mentioned sequencing data analysis method.
[0032] In another aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program / instructions, wherein the computer program / instructions are used to execute the steps of the above-mentioned sequencing data analysis method.
[0033] The present invention provides a sequencing data analysis method for low-frequency Indel mutations generated by CRISPR gene editing. The method of the present invention quickly and accurately analyzes the gene editing status of low-frequency Indel mutations with a wide variation range generated by CRISPR gene editing. In particular, it can accurately assess the frequency and pattern of Indel mutations generated by CRISPR gene editing in animal tissues or adult animals. It can effectively control background noise and directly output results without setting any control group or requiring result correction, significantly improving the efficiency of statistical analysis. In addition, it can directly read the efficiency and pattern information of gene editing of multiple samples at the same sequencing depth. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 The method of Example 1 of the present application is used to analyze the efficiency and editing pattern of cell TYR gene editing; wherein Figure 1 A is the gene editing efficiency diagram, Figure 1 B is the distribution of gene editing break positions, Figure 1 C is the distribution diagram of gene editing insertions and deletions.
[0035] Figure 2 To compare the results of experiment 1 (T7E1 detection) to analyze the TYR gene editing effect in cells (agarose electrophoresis results);
[0036] Figure 3 To compare the results of test 2 (topological cloning Sanger sequencing) to analyze the effect of cell TYR gene editing;
[0037] Figure 4 Schematic diagram of gene editing operations for chicken adults (route a) and chicken embryos (route b);
[0038] Figure 5 This is a graph showing the gene editing efficiency of TYR gene knockout in adult chickens (knockout was achieved by directly injecting adenoviral CRISPR vectors into the jugular vein of chicks);
[0039] Figure 6 This is a graph showing the gene editing efficiency of chicken embryo TYR gene knockout (knockout by directly injecting the adenoviral CRISPR vector into the dorsal aorta of chicken embryos);
[0040] Figure 7 For T5 chicks ( Figure 6 Figure 3. Gene editing frequency pattern of gonads in T5 chicks. DETAILED DESCRIPTION
[0041] The present invention is further described below with reference to examples, but the present invention is not limited to these specific embodiments.
[0042] Example 1 of the present application provides a method for analyzing sequencing data of low-frequency Indel mutations generated by CRISPR gene editing.
[0043] In a specific embodiment of the present application, step 1) obtains double-end sequencing data: constructs a double-end sequencing DNA library for the sample obtained by CRISPR gene editing, and then performs high-throughput sequencing based on the Illumina sequencing platform to obtain R1 fastq data package and R2 fastq data package.
[0044] Regarding the construction method of the double-end sequencing library, the primers are designed first.
[0045] In a specific embodiment of the present application, in step 1), the forward primer and reverse primer designed for double-end sequencing are respectively located 150 to 250 bp upstream and downstream of the target region of the CRISPR gene editing.
[0046] Specifically in this embodiment, the forward primer and reverse primer designed for double-end sequencing are located 200 bp upstream and downstream of the target region of CRISPR gene editing (TYR gene exon 1), respectively.
[0047] Specifically in this embodiment, different barcodes of 6 bases in length containing DNA coding information are added to the forward primer and the reverse primer, and then the test region is amplified using PCR technology to obtain the test gene fragments with different barcode labels.
[0048] Specifically in this embodiment, the PCR amplification products were quality checked on LabChip. After gradient dilution of each qualified sequencing library, they were mixed in equal proportions according to the required sequencing amount.
[0049] In a specific embodiment of the present application, step 2a), demultiplexing R1 data: according to the designed forward primer sequence, the R1 fastq data packet obtained in step 1) is demultiplexed to split the R1 forward fastq data file and obtain the corresponding R1 reverse fastq data file, and store them according to different samples;
[0050] Step 2b), demultiplexing R2 data: According to the designed reverse primer sequence, the R2 fastq data packet obtained in step 1) is demultiplexed to split the R2 reverse fastq data file and obtain the corresponding R2 forward fastq data file, and store them according to different samples;
[0051] The steps 2a) and 2b) are not ordered in any particular order;
[0052] Step 3) Merge data: merge the R1 forward fastq data file and the R2 forward fastq data file, and store the forward fastq data packets according to different samples; and
[0053] The R2 reverse fastq data file is merged with the R1 reverse fastq data file, and reverse fastq data packets are obtained according to different sample storage.
[0054] In a specific embodiment of the present application, the operations of steps 2a) and 2b) above, on the one hand, execute a demultiplexing-only command without deleting, adding, or modifying the metadata; this can retain the original data information to the greatest extent possible and input it into subsequent processes with high fidelity; on the other hand, the original R1 and R2 data packets obtained from the sequencing platform are demultiplexed separately, and the original data are fully split in both directions, and then the data are merged in step 3), which can achieve higher data throughput and, more importantly, significantly reduce the error splicing rate caused by the subsequent splicing step.
[0055] Specifically in this embodiment, in step 2a) and step 2b), the AdapterRemoval software is respectively called to execute a demultiplexing-only command without deleting, adding or modifying the metadata.
[0056] In a specific embodiment of the present application, step 4) removes the adapter sequence: the forward fastq data packet and the reverse fastq data packet are processed to remove the adapter sequence of each fastq data.
[0057] Regarding the removal of adapter sequences, you can call the AdapterRemoval software and execute the command twice to remove the adapter sequences of the data in the forward fastq data packet and the reverse fastq data packet respectively.
[0058] Specifically in this embodiment, a self-designed script is used to execute a command once, and the connector sequence of the data in the forward fastq data packet and the reverse fastq data packet is removed in one operation.
[0059] In a specific embodiment of the present application, step 5), forward and reverse data splicing: splicing the forward fastq data file and the reverse fastq data file of each identical sample, outputting a full-length fastq data file, and storing it according to different samples.
[0060] In step 5), forward and reverse data splicing is performed by a designed python script, and by setting parameters, the forward fastq data file and the reverse fastq data file of the same sample with an overlapping interval greater than or equal to 10bp are spliced; data that does not meet the parameters are discarded.
[0061] Based on the operations of steps 2a) and 2b) and step 3), a higher data throughput is achieved. Therefore, in step 5) of this embodiment, strict parameters can be set to splice forward and reverse fastq data files with an overlap interval greater than or equal to 10bp, and data that does not meet the parameter requirements are directly discarded; the amount of data after discarding can still meet the statistical requirements of subsequent procedures and can significantly reduce the error splicing rate generated by subsequent splicing steps.
[0062] In a specific embodiment of the present application, step 6), alignment analysis: calling Crispresso2 software, the full-length fastq data file obtained in step 4) is aligned with the target region reference sequence of the CRISPR gene editing and the sgRNA sequence used for analysis and data statistics.
[0063] Specifically in this embodiment, during the alignment analysis in step 6), if any of the following a) and b) conditions are not met, the full-length fastq data file is discarded: a) the portion of either end sequence that completely overlaps with the target region reference sequence is more than 20 bp; b) the Phred33 quality score within any consecutive 4 bp detection window is greater than 10.
[0064] Specifically in this embodiment, during the comparison analysis in step 6), after ignoring single-base substitutions, the comparison analysis results with a mutation rate of more than 0.01% are output; the mutation rate refers to the ratio of the number of mutated bases to the total number of bases in the reference sequence of the target region.
[0065] Effect verification
[0066] In order to verify the feasibility of the method in Example 1 above, the inventors of the present application used the CRISPR / Cas9 gene editing method:
[0067] 1. Knockout of the TYR gene in DF-1 cells (chicken embryo fibroblast cell line);
[0068] 2. Perform TYR gene knockout on chicks after birth;
[0069] 3. Knockout of the TYR gene in chicken embryos;
[0070] The gene editing efficiency was detected and analyzed using the method in Example 1 above.
[0071] The TYR gene is a member of the tyrosinase-related protein family. Tyrosinase is the rate-limiting enzyme in the synthesis of melanin. The expression level and activity of TYR directly affect the expression of eumelanin and pheomelanin, thereby affecting the animal's coat color.
[0072] In the genetic breeding research of chickens, skin color and coat color have always been one of the important economic traits in genetic breeding, and the TYR gene is an important candidate gene.
[0073] The experimental method is as follows:
[0074] 1) Plasmid construction
[0075] Guide RNA (gRNA) targeting the chicken TYR gene (NCBI Acc: 373971, GRCg6a) was designed using CHOPCHOP v313. The gRNA and synthesized oligonucleotides are shown in Table 1. PSpCas9(BB)-2A-Puro(PX459) V2.0 (Addgene #62988) was a gift from Feng Zhang. The annealed gRNA oligo was cloned into BbsI-cleaved PX459 to generate the Cas9 / gRNA plasmid.
[0076] Table 1
[0077]
[0078] 2) Production of adenoviral vectors
[0079] The adenoviral Cas9 vector pAV[CRISPR]-hCas9:P2A:EGFP-U6>gRNA3 (Ad5-Cas9 / gRNA3) co-expressing SpCas9, gRNA-TYR3 and eGFP and the control pAV[Exp]-CBh>hCas9(ns):P2A:EGFP (Ad5-Cas9 / ns) of eGFP were produced, purified and concentrated by VectorBuilder, USA.
[0080] 3) Adenoviral vector transfection
[0081] For DF-1 cells, adenovirus with a final concentration of 1E8 VP / ml was directly used for transfection.
[0082] For in vivo transfection in chickens, 100 μL of Ad5-Cas9 / gRNA3 was injected directly into the jugular vein of postnatal chicks using a microsyringe.
[0083] For in ovo transfection of chicken embryos: Fertilized white Leghorn chicken eggs are carefully inoculated at the blunt end of the eggshell. Adenovirus is injected directly into the dorsal aorta of the embryo using a microsyringe. The eggs are then placed in a rotating tray for 16 days before being transferred to a setter tray until hatching.
[0084] 4) Preparation of genomic DNA
[0085] Cells were harvested by scraping directly after a week of drug screening. Chickens injected with the jugular vein were sacrificed and dissected one month after injection, and the heart, liver, spleen, lungs, kidneys, small intestine, thymus, blood, and gonads were collected. For embryonic transfection, the liver, thymus, and gonads were collected from deceased chickens, while blood and feather medulla were collected from live chickens. Genomic DNA was prepared using a tissue DNA kit.
[0086] 5) Barcode library preparation.
[0087] High-fidelity amplification of the target region from genomic DNA was performed using 2× Phanta Max Master Mix. Sample-specific barcodes were fused to the 5' ends of both the forward and reverse primers. PCR conditions were 95°C for 3 minutes; 36× (95°C for 15 seconds, 58°C for 15 seconds, 72°C for 20 seconds); and 72°C for 5 minutes. All amplicons were mixed and purified using the Cycle Pure Kit. Amplicon libraries were prepared using DNAPCR-Free Library Prep according to the manufacturer's instructions. Amplicon libraries were quantified by agarose gel electrophoresis and qPCR.
[0088] 6) Amplicon sequencing.
[0089] Amplicon libraries were sequenced on a NovaSeq 6000 system from Novogene, Beijing, China, in PE250 mode with a sequencing depth of 1 million reads.
[0090] 7) Using the method in Example 1 above, their gene editing efficiency was detected and analyzed.
[0091] Comparative detection experiment 1: mismatch cleavage detection T7E1 detection.
[0092] Briefly, the target region was amplified using 2× Phanta Max Master Mix. The primers were TYR-F: agcagagtcagtggtgaagc and TYR-E1-R: tcttgccctccccttacctt. PCR conditions were 95°C for 3 minutes; 36× (95°C for 15 seconds, 58°C for 20 seconds, 72°C for 20 seconds); and 72°C for 5 minutes. PCR products were purified using a Cycle pure kit (Omegabio-tek, USA). The purified PCR products were then annealed at 95°C for 5 minutes; then annealed at -2°C / min to 85°C, and then -0.1°C / min to 25°C. PCR product fragments were analyzed by 2% agarose gel electrophoresis, and gene editing efficiency was estimated using ImageJ 1.8.0.
[0093] Comparative detection experiment 2: topological cloning Sanger sequencing.
[0094] The PCR product from the T7E1 assay was cloned into the pCE2-Topo vector. The plasmid was heat-shocked into DH5α and then plated on lysis broth plates containing 100 mg / L ampicillin. Individual colonies (n = 10) from the plates were individually Sanger sequenced. Sequences were aligned using the MUSCLE program in MEGA 11.0.10, and gene editing efficiencies were calculated manually.
[0095] Effect data 1: Efficiency testing of gene editing in cells
[0096] In order to verify the feasibility of the method in Example 1 above, the inventors of the present application used the CRISPR / Cas9 gene editing method to knock out the TYR gene in DF-1 cells (chicken embryo fibroblast cell line) and used the method in Example 1 above to detect and analyze the gene editing efficiency of DF-1 cells.
[0097] The Barcode forward and reverse primer sequences designed for the four gRNAs targeting the TYR gene of DF-1 cells are shown in Table 2 below.
[0098] Table 2
[0099] gRNAs Barcode forward primer sequence Barcode reverse primer sequence gRNA1 AGACTCTGATTTTGCCCATGAAGCCCC TTTGCCTGTCCTTACCTGCC gRNA2 AGTCACAGCTCCAACCCCATGTTCAGA ACACAGTCCTCTGCATCTCG gRNA3 TCAGAGAGATTTTGCCCATGAAGCCCC TTTGCCTGTCCTTACCTGCC gRNA4 TGACTGAGCTCCAACCCCATGTTCAGA ACACAGTCCTCTGCATCTCG
[0100] The results of the gene editing efficiency test of DF-1 cells using the method of Example 1 of this application are shown in Figure 1 A- Figure 1 C.
[0101] from Figure 1 As can be seen from the results of A, the gene editing efficiencies of the four gRNAs obtained by the method of Example 1 of this application are 61.27%, 56.25%, 77.09% and 29.75% respectively. Figure 1 As can be seen from the results in B, further pattern analysis showed that most indels occurred at the predicted positions of gRNA. Figure 1 According to the C results, the indel sizes are mainly between -5 and 1 bp.
[0102] Figure 1 The AC results show that the method of Example 1 of the present application can well detect and characterize CRISPR-induced indel mutations in the cell genome.
[0103] The results of comparative detection experiment 1 (T7E1 detection and agarose electrophoresis) are shown in Figure 2 T7E1 detection and agarose gel electrophoresis showed that all four gRNAs underwent gene editing. Grayscale analysis of each lane was performed using ImageJ to estimate the gene editing efficiency, which was 27.2%, 42.5%, 33.7%, and 25.8%, respectively.
[0104] The results of comparative detection test 2 (topological cloning Sanger sequencing) are shown in Figure 3 Sanger sequencing was performed on the Topo clones to analyze the gene editing efficiency and frame shift efficiency; the results showed that the gene editing efficiency was 87.5%, 100%, 100%, and 20.0%, respectively, and the frame shift efficiency was 75.0%, 50.0%, 77.8%, and 0, respectively.
[0105] In comparison, the T7E1 detection method easily overlooks indel mutations, and the topological cloning Sanger sequencing method has low throughput, so it cannot quickly, sensitively and accurately determine the gene editing efficiency.
[0106] In contrast, the method of Example 1 of the present application can display gene editing patterns, indel sites and other information of genes, and the results are highly sensitive, accurate and reliable.
[0107] Effect data 2: Efficiency testing of gene editing in adult chickens
[0108] See also Figure 4 (Route a), the inventors of the present application selected 6 newborn chicks, three of which were directly injected with adenovirus CRISPR / Cas9 vectors through the jugular vein, and then tested the gene editing efficiency of these chicks one month after the injection.
[0109] Barcode primers used to construct the TYR gene knockout chicken amplicon library. See Table 3 below.
[0110] Table 3
[0111] name Sequence(5'-3') Barcoded forward TYRe1NovaA ATCACGTGCTCAGATGAACAACGGCT TYRe1NovaB CGATGTTGCTCAGATGAACAACGGCT TYRe1NovaC TTAGGCTGCTCAGATGAACAACGGCT TYRe1NovaD TGACCATGCTCAGATGAACAACGGCT TYRe1NovaE ACAGTGTGCTCAGATGAACAACGGCT TYRe1NovaF GCCAATTGCTCAGATGAACAACGGCT Barcoded reverse TYRe1Nova1 CAGATCTGCCCTCCCCTTACCTTCAT TYRe1Nova2 ACTTGATGCCCTCCCCTTACCTTCAT TYRe1Nova3 GATCAGTGCCCTCCCCTTACCTTCAT TYRe1Nova4 TAGCTTTGCCCTCCCCTTACCTTCAT TYRe1Nova5 GGCTACTGCCCTCCCCTTACCTTCAT TYRe1Nova6 CTTGTATGCCCTCCCCTTACCTTCAT
[0112] Figure 5This is a graph showing the gene editing efficiency of TYR gene knockout in adult chickens (knockout was achieved by directly injecting adenoviral CRISPR vectors into the jugular vein of chicks); Figure 5 The horizontal axis represents various organs, the vertical axis T1 to 3 represents three chickens that have been gene-edited, and CT1 to 3 represents three chickens that have not been treated. The value represents the gene editing efficiency, in %. Figure 7 It can be seen that 1) the gene editing efficiency of untreated chickens is below 0.03%, which can almost be considered to be 0, indicating that false positives are very well controlled; 2) for gene-edited chickens, the editing efficiency of liver, spleen, breast muscle and blood is relatively higher than that of other tissues.
[0113] The method of Example 1 of the present application successfully detected CRISPR-induced indel mutations in chickens with high sensitivity and accuracy. It can also easily detect inefficient indels in chickens, a feat previously difficult to achieve. Furthermore, due to its ultra-low false positive rate, it effectively controls background noise and directly outputs results without setting up any control group (directly and reasonably omitting indel detection in the control group; no significance testing or result correction is required). This significantly improves the efficiency of statistical analysis and allows for more accurate direct reading of gene editing efficiency and patterns across multiple samples at the same sequencing depth.
[0114] Effect data 3: Efficiency testing of gene editing in chicken embryos
[0115] In order to verify the feasibility of the method in Example 1, the inventors of the present application used CRISPR / Cas9 tools to knock out the TYR gene in chicken embryos. The adenovirus CRISPR / Cas9 vector was directly injected into the embryonic dorsal aorta (see Figure 4 Following route b), 8 chickens (T4-T11 chicks) were hatched, and the livers, sternums, and gonads of dead chickens, as well as the blood and feather medulla of live chickens, were collected. Gene sequencing and gene editing efficiency were analyzed using the method of Example 1 of the present application.
[0116] Results see Figure 6 , a graph of the gene editing efficiency of TYR gene knockout chickens obtained by directly injecting adenovirus CRISPR vectors into the dorsal aorta of chicken embryos; the horizontal axis represents different organs, and the vertical axis represents individual chickens with different gene editing. In general, the gene editing efficiency of T5, T8, and T10 chickens is relatively high; specifically, the gene editing efficiency of the liver of T5 chickens reached 3.99%, the gene editing efficiency of the pectoral muscle reached 3.30%, and the gene editing efficiency of the gonads reached 6.58%; the gene editing efficiency of the blood of T10 chickens reached 5.45%, and the gene editing efficiency of the hair medulla also reached 1.55%; the gene editing efficiency of T7 chickens, which had poor gene editing effects, was no higher than 0.03%.
[0117] Figure 7 This is a pattern diagram of the gene editing frequency of the gonads of T5 chickens. The figure shows the indel pattern with a gene editing efficiency higher than 0.07%, including deletions, additions and changes. It can be seen that there are 12,326 unedited sequences, accounting for 90.75% of all sequences, and among the edited sequences, the most are CC deletions, reaching 211, accounting for 1.55%. Using the method of Example 1 of the present application, the effect diagram is automatically generated as the analysis is performed, which can well characterize and prove the gene editing effect in animal tissues, which facilitates the analysis and judgment of the gene editing effects and patterns of animals, provides more effective information, and improves the credibility of the data.
[0118] The method of the present application can not only detect the efficiency of CRISPR gene editing in cells, but also the efficiency of CRISPR gene editing in chicken embryos and adult chickens.
[0119] In the prior art, the frequency of Indel mutations in samples obtained through CRISPR gene editing in chicken embryos and adult chickens is extremely low, and the variation range of Indel mutations is highly dynamic (ranging from 1bp to 50bp), making it difficult to obtain accurate detection results. Specifically, the inventors of the present application found in experimental research that the use of existing analysis software and methods (statistical analysis of high-throughput sequencing data obtained from the Illumina sequencing platform) is usually unable to obtain a large range of Indel mutations, and it is difficult to analyze the presence of false positives in the results.
[0120] However, the method of the present application can effectively exclude false positive results.
[0121] Specifically, in the bidirectional demultiplexing of steps 2a and 2b), on the one hand, a demultiplex-only command is executed to retain the original data information to the greatest extent possible and input it into the subsequent process with high fidelity; on the other hand, the original data is fully split in both directions and then merged in step 3), which can obtain higher data throughput and significantly reduce the error splicing rate generated by the subsequent splicing step. Furthermore, based on the obtained higher throughput data, combined with the call of Crispresso2 software, unmatched data is discarded, and any data with a Phred33 quality score below 10 within a continuous 4-bp detection window is discarded. After ignoring single-base substitutions, the output of the comparison analysis results with a mutation rate above 0.01% can effectively control the background noise without setting any control group (no significance test and result correction are required), and the results are directly output, which significantly improves the efficiency of statistical analysis and can more accurately and directly read the efficiency and pattern information of gene editing of multiple samples at the same sequencing depth.
[0122] Gene editing in chickens is of great significance in both scientific research and agricultural production. Viruses have long been used as vectors for chicken gene editing, making significant contributions to avian functional genomic research. Since 2006, PGCs-mediated avian gene editing has been developed and has gradually become a trend. However, it still cannot overcome shortcomings such as low efficiency, complex procedures, and harsh conditions. Recently, with the development of viral delivery systems and improvements in transfection efficiency and biosafety, adenoviral CRISPR vector-based avian genome editing methods are being developed. This approach is considered promising for widespread application in poultry research. One advantage of this method is that it can determine the efficiency of gene editing by testing G0 chickens. Traditional methods have difficulty processing such low-frequency samples and are unreliable in estimating gene editing efficiency.
[0123] In the present study, the inventors designed an amplicon sequencing-based method to detect CRISPR-induced indel mutations in cells, chicken embryos, and adult chickens.
[0124] The method of the present application not only maintains the ultra-sensitive and ultra-accurate advantages of NGS, but more importantly, can quantify the indel mutation frequency with high precision, with a sensitivity as low as 0.1% and a false positive rate as low as negligible.
[0125] Furthermore, the method in this application is low-cost and has a simple workflow. The method, which utilizes barcoded primers for multiplexed sequencing library construction, reduces costs to unprecedented levels. It is estimated that the sequencing cost per 10,000 sequences is less than $1, with the potential for further reduction. 12 With more advanced sequencing and data analysis tools, there is no need for expensive servers and / or high-performance processors. This method enables a standard office computer to fully mine 5Mb of data in 30 seconds.
[0126] Another major advantage of the detection method described in this application is its ability to directly and in detail understand the nature and diversity of mutations. This method can simultaneously measure gene editing efficiency and plot gene editing patterns and related statistical charts. The output is intuitive and reliable, providing a multi-dimensional reference for gene-edited chickens.
[0127] Furthermore, the broad feasibility of this method is foreseeable. This method not only meets the needs of avian genome editing experiments, but is also foreseeably applicable to off-target, base editing, plasmid editing, point mutations, and more. The use of this application's detection method will greatly promote the widespread implementation of CRISPR-based precision genetic manipulation.
[0128] It should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each implementation method can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0129] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. Any equivalent implementation methods or changes that do not deviate from the technical spirit of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for analyzing sequencing data of low-frequency Indel mutations generated by CRISPR gene editing, characterized by: The method comprises the following steps: Step 1) Obtaining double-end sequencing data: constructing a double-end sequencing DNA library for the sample obtained by CRISPR gene editing, and then performing high-throughput sequencing based on the Illumina sequencing platform to obtain R1 fastq data package and R2 fastq data package; Step 2a), demultiplexing R1 data: according to the designed forward primer sequence, demultiplex the R1 fastq data packet obtained in step 1), split the R1 forward fastq data file, and obtain the corresponding R1 reverse fastq data file, and store them according to different samples; Step 2b), demultiplexing R2 data: According to the designed reverse primer sequence, the R2 fastq data packet obtained in step 1) is demultiplexed to split the R2 reverse fastq data file and obtain the corresponding R2 forward fastq data file, and store them according to different samples; The steps 2a) and 2b) are not ordered in any particular order; Step 3) Merge data: merge the R1 forward fastq data file and the R2 forward fastq data file, and store the forward fastq data packets according to different samples; and Merging the R2 reverse fastq data file with the R1 reverse fastq data file, and obtaining reverse fastq data packets according to different sample storage; Step 4), removing the connector sequence: processing the forward fastq data packet and the reverse fastq data packet, and removing the connector sequence of each fastq data; Step 5), forward and reverse data splicing: splice the forward fastq data file and the reverse fastq data file of each identical sample, output the full-length fastq data file, and store it according to different samples; Step 6), alignment analysis: calling Crispresso2 software, the full-length fastq data file obtained in step 4) is aligned with the target region reference sequence of the CRISPR gene editing and the sgRNA sequence used for analysis and data statistics.
2. The sequencing data analysis method according to claim 1, wherein: During the alignment analysis in step 6), the full-length fastq data file was discarded if any of the following conditions a) and b) were not met: a) the portion of either end sequence that completely overlapped with the target region reference sequence was more than 20 bp; b) the Phred33 quality score within any consecutive 4 bp detection window was greater than 10.
3. The sequencing data analysis method according to claim 1, wherein: During the comparison analysis in step 6), ignoring single-base substitutions, outputting the comparison analysis results with a mutation rate of more than 0.01%; The mutation rate refers to the ratio of the number of mutated bases to the total number of bases in the reference sequence of the target region.
4. The sequencing data analysis method according to any one of claims 1 to 3, wherein: In step 2a) and step 2b), the AdapterRemoval software is respectively called to execute a demultiplexing-only command without deleting, adding or modifying the metadata.
5. The sequencing data analysis method according to claim 4, wherein: In step 5), forward and reverse data splicing is performed by a designed python script, and by setting parameters, the forward fastq data file and the reverse fastq data file of the same sample with an overlapping interval greater than or equal to 10bp are spliced; data that does not meet the parameters are discarded.
6. The sequencing data analysis method according to any one of claims 1 to 5, wherein: The CRISPR gene editing uses the CRISPR / Cas9 gene editing tool; preferably, the sample obtained by the CRISPR gene editing is a sample obtained by editing an adult or embryo of an animal using the CRISPR / Cas9 gene editing tool; preferably, the animal is an avian (Aves) animal; more preferably, the avian animal is a Galliformes, Anseriformes, Otidiformes, Columbiformes or Struthioniformes animal; more preferably, the avian animal is a Galliformes animal and is selected from the family Numididae or the family Phasianidae; more preferably, the avian animal belongs to the order Galliformes, the family Phasianidae, or the genus Gallus.
7. The sequencing data analysis method according to claim 6, wherein: In step 1), the forward primer and reverse primer designed for double-end sequencing are respectively located 150 to 250 bp upstream and downstream of the target region of the CRISPR gene editing; preferably, they are respectively located 200 bp upstream and downstream of the target region of the CRISPR gene editing.
8. The sequencing data analysis method according to claim 7, wherein: In the step 1), PE sequencing is performed using an Illumina sequencing platform; preferably, PE200 sequencing is performed using an Illumina sequencing platform.
9. A sequencing data analysis device for low-frequency Indel mutations generated by CRISPR gene editing, characterized by: The sequencing data analysis device is used to perform the steps of the sequencing data analysis method according to any one of claims 1 to 9.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program / instruction, wherein: The computer program / instructions are used to execute the steps of the sequencing data analysis method according to any one of claims 1 to 9.