A microbial gene editing scheme automatic design method, system and storage medium
By using automated design methods and optimizing the design of sgRNA and homologous arm primers with CRISPOR software and NCBI database, the problem of cumbersome and time-consuming design of microbial gene editing protocols has been solved, and the success rate and efficiency of gene editing have been improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU UBIGENE BIOSCIENCES CO LTD
- Filing Date
- 2024-12-31
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies for microbial gene editing involve cumbersome and time-consuming design steps, rely heavily on the experience of laboratory technicians, and result in low success rates and high costs.
An automated design method was employed, using CRISPOR software to screen sgRNA sequences and homologous arm primers that met specific conditions. Gene information was obtained from the NCBI database, and different strategies were designed for bacteria and fungi. The gRNA and primer design regions were optimized to ensure specificity and cleavage efficiency and avoid off-target effects.
It improves the success rate of gene editing, reduces the time and cost of manual design, provides an efficient gene editing protocol, and is applicable to experiments with a variety of strains.
Smart Images

Figure CN119905140B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing, and more particularly to an automated design method, system, and storage medium for microbial gene editing schemes. Background Technology
[0002] Gene editing refers to the process of modifying specific targets in the genome of an organism using gene editing technology. It efficiently and precisely inserts, deletes, or replaces genes, thereby altering their genetic information and phenotypic characteristics. However, currently, it is mainly done through manually designed editing protocols, but the design process is cumbersome and time-consuming, leading to increased design costs. Furthermore, it relies heavily on the experience of experimental technicians, and inaccurate results may occur after the design is completed, resulting in a lower success rate of gene editing. Summary of the Invention
[0003] The purpose of this invention is to propose an automated design method for microbial gene editing schemes, which provides scheme design for sgRNA sequences and homologous arm primer sequences for gene editing experiments of various strains. The sgRNA sequence design method comprehensively considers factors such as design region and specificity, resulting in a higher success rate of gene editing. It can replace complicated and time-consuming manual schemes, saving a lot of time for scientific research users.
[0004] The present invention also proposes an automated design system for microbial gene editing schemes, which is used to execute the above-described automated design method for microbial gene editing schemes.
[0005] To achieve this objective, the present invention adopts the following technical solution:
[0006] An automated design method for microbial gene editing protocols includes the following steps:
[0007] (1) Select the design object for different types of strains, which can be bacteria or fungi; after determining the target gene, obtain the target gene information from the NCBI database;
[0008] (2) If the design object of step (1) is selected as bacteria, only step (2-1) is executed; if the design object of step (1) is selected as fungi, step (2-2) is executed.
[0009] (2-1) The bacterial design scheme includes the following steps:
[0010] (2-1-1) Determine the gRNA design region: When the gene is ≤1000bp, the two gRNAs are designed in the region between 1 / 4 and 3 / 4 of the gene, respectively; when the gene is greater than 1000bp, the two gRNAs are designed at both ends of the gene, with the first gRNA designed in the range of 100-400bp on one side and the second gRNA designed in the range of 100-400bp on the other side.
[0011] (2-1-2) Selecting the optimal gRNA sequence: Use CRISPOR software to screen gRNA sequences in the gRNA design region. All sequences must meet the following three conditions:
[0012] ① The specificity score is 100;
[0013] ② The off-target predictions are all 0;
[0014] ③Doench'16 score ≥ 40;
[0015] (2-1-3) Sequence complexity analysis: Perform lattice analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000bp, and identify the location of potential hairpin structures;
[0016] (2-2) The fungal design scheme uses a frameshift scheme, which includes the following steps:
[0017] (2-2-1) The frameshift scheme's gRNA is located within the gene at any position that does not overlap with other genes after the start codon;
[0018] (2-2-2) Use CRISPOR software to screen gRNA sequences in the gRNA design region, and select the optimal gRNA sequence according to the following rules:
[0019] ① Specificity Score ≥ 80, Off-target predictions are all 0, Lindel score ≥ 80, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions;
[0020] (2-2-3) Sequence complexity analysis: Perform lattice analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000bp, and identify the location of potential hairpin structures;
[0021] (3) Design two pairs of four primer sequences for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0022] Upstream homologous arm - reverse primer up-R: Primers can be designed from 200 bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene.
[0023] Upstream homologous arm - forward primer up-F: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp from the up-R primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp from the up-R primer.
[0024] Down-F homologous arm: Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first downstream coding gene, with preference given to those closest to the start or stop codon of the gene.
[0025] Down-R homologous arm: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp from the down-F primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp from the down-F primer.
[0026] (4) Output the design results of sgRNA sequence and homologous arm primer sequence.
[0027] Alternatively, in step (1), the target gene can be determined based on three identifiers: gene name, gene ID, or locus tag.
[0028] Optimally, in step (2-1-2), among the gRNA sequences that simultaneously meet all three conditions, the gRNA with the highest Lindel score is selected.
[0029] Ideally, in step (3), the homologous arm primers should meet the following criteria:
[0030] ①The primer Tm value is between 58 and 62℃;
[0031] ②The GC content of the primers is between 35% and 65%;
[0032] ③The last base at the 3' end is G or C;
[0033] ④ Primer pairs cannot have consecutive 6bp complementary pairings;
[0034] ⑤ Primer sequence length 17–30 bp.
[0035] Alternatively, in step (2-2-2), if there is no sequence that meets rule ①, then rule ② can be used for further filtering.
[0036] ② Specificity Score ≥ 80, Lindel ≥ 70, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions; if there is no sequence that meets rule ②, continue to use rule ③ for screening;
[0037] ③ Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the frameshift knockout scheme cannot be designed for this gene.
[0038] Alternatively, in step (2-2), the fungal design scheme may use a whole-genome knockout scheme instead of a frameshift scheme.
[0039] (2-2-1) gRNA is located within the gene, between 100 and 400 bp on both sides, avoiding areas that overlap with other genes;
[0040] (2-2-2) Use CRISPOR software to screen gRNA sequences in the gRNA design region, and select the optimal gRNA sequence according to the following rules:
[0041] ① Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If there is no sequence that meets this rule, then select according to rule ②.
[0042] ② Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets the above conditions, the gene knockout scheme cannot be designed.
[0043] (2-2-3) Sequence complexity analysis: Perform lattice analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000bp, and identify the location of potential hairpin structures.
[0044] An automated design system for microbial gene editing protocols includes: a gene acquisition module, a bacterial design strategy module, a fungal design strategy module, and a homology arm identification primer design module;
[0045] The gene acquisition module is used to select design objects for different types of strains, namely bacteria or fungi; after determining the target gene, it retrieves the target gene information from the NCBI database; and according to whether the design object is selected as bacteria or fungi, it calls the bacterial design strategy module or the fungal design strategy module.
[0046] The bacterial design strategy module is used to determine the gRNA design region: when the gene is ≤1000bp, two gRNAs are designed in the region between 1 / 4 and 3 / 4 of the gene; when the gene is greater than 1000bp, two gRNAs are designed at both ends of the gene, with the first gRNA designed in the range of 100-400bp on one side and the second gRNA designed in the range of 100-400bp on the other side; and the optimal gRNA sequence is selected: gRNA sequences are screened in the gRNA design region using CRISPOR software, and all sequences must meet the following three conditions:
[0047] ① The specificity score is 100;
[0048] ② The off-target predictions are all 0;
[0049] ③Doench'16 score ≥ 40;
[0050] The bacterial design strategy module is also used for sequence complexity analysis: after performing matrix analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000bp, the location of potential hairpin structures is identified.
[0051] The fungal design strategy module is used to execute a frameshift scheme: it places the gRNA at any non-overlapping position within the gene after the start codon; and uses CRISPOR software to screen gRNA sequences in the gRNA design region, selecting the optimal gRNA sequence according to the following rules:
[0052] ① Specificity Score ≥ 80, Off-target predictions are all 0, Lindel score ≥ 80, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions;
[0053] The fungal design strategy module is also used to perform sequence complexity analysis: perform matrix analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000bp, and identify the location of potential hairpin structures.
[0054] The homologous arm identification primer design module is used to design two pairs of four primer sequences for homologous arm identification. The selection criteria for the primer design region are as follows:
[0055] ① Upstream homologous arm - reverse primer up-R: Primers can be designed from 200bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene;
[0056] ② Upstream homologous arm - forward primer up-F: If the target gene length is ≤2000bp, design primers in the region between 450-600bp from the up-R primer; if the target gene length is greater than 2000bp, design primers in the region between 950-1100bp from the up-R primer.
[0057] ③ Down-F (down homologous arm): Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first coding gene downstream. Primers closer to the start or stop codon of the gene are preferred.
[0058] ④ Down-R homologous arm: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp away from the down-F primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp away from the down-F primer.
[0059] Alternatively, if no sequence matches rule ①, then rule ② can be used for further filtering.
[0060] ② Specificity Score ≥ 80, Lindel ≥ 70: Select the gRNA with the highest Lindel score from the gRNAs that meet the above criteria. If no sequence meets these criteria, then proceed to rule ③ for selection.
[0061] ③ Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the frameshift knockout scheme cannot be designed for this gene.
[0062] Optimally, the fungal design strategy module is also used to execute a whole-gene knockout scheme, placing the gRNA within a range of 100-400 bp flanking the gene, avoiding regions overlapping with other genes; using CRISPOR software to screen gRNA sequences in the gRNA design region, and selecting the optimal gRNA sequence according to the following rules:
[0063] ① Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If there is no sequence that meets this rule, then select according to rule ②.
[0064] ② Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the gene cannot be designed into a whole-gene knockout scheme.
[0065] A storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the above-described automatic design method for a microbial gene editing scheme.
[0066] Compared with the prior art, one of the above technical solutions has the following beneficial effects:
[0067] This protocol provides design schemes for sgRNA sequences and homologous arm primer sequences for gene editing experiments on various strains. The sgRNA sequence design method comprehensively considers factors such as design region and specificity, resulting in a higher gene editing success rate. It can replace complicated and time-consuming manual methods, saving researchers a lot of time and solving the problems of high costs faced in manually designing the optimal site for gene knockout and designing sgRNA sequence schemes. Attached Figure Description
[0068] Figure 1 This is a GC content analysis chart from Example 1;
[0069] Figure 2 A dot matrix analysis diagram of the sequence from Example 1;
[0070] Figure 3 This is a GC content analysis chart from Example 2;
[0071] Figure 4 This is a dot matrix analysis plot of the sequence from Example 2;
[0072] Figure 5 This is a GC content analysis chart from Example 3;
[0073] Figure 6 This is a dot matrix analysis diagram of the sequence from Example 3. Detailed Implementation
[0074] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0075] An automated design method for microbial gene editing protocols includes the following steps:
[0076] (1) Select the design object for different types of strains, which can be bacteria or fungi; after determining the target gene, obtain the target gene information from the NCBI database;
[0077] After selecting the target strain and identifying the target gene, gene information such as its location and sequence can be obtained from the NCBI database. Since bacterial and fungal gene structures differ, the protocol design strategy varies depending on the type of strain. Furthermore, the obtained gene information is used to analyze the gRNA sequence, primarily to assess its potential off-target sites in the genome. Off-target effects can lead to non-specific editing, affecting the accuracy and reliability of experimental results.
[0078] (2) If the design object of step (1) is selected as bacteria, only step (2-1) is executed; if the design object of step (1) is selected as fungi, step (2-2) is executed.
[0079] For bacteria such as Escherichia coli, Salmonella, and Shigella, the specific procedure is (2-1).
[0080] For fungi such as yeast, follow step (2-2);
[0081] (2-1) The bacterial design scheme includes the following steps:
[0082] (2-1-1) Determine the gRNA design region: When the gene is ≤1000bp, the two gRNAs are designed in the region between 1 / 4 and 3 / 4 of the gene, respectively; when the gene is greater than 1000bp, the two gRNAs are designed at both ends of the gene, with the first gRNA designed in the range of 100-400bp on one side and the second gRNA designed in the range of 100-400bp on the other side. According to this gRNA design region, the location selection has a significant effect on improving the success rate of gene editing.
[0083] (2-1-2) Selecting the optimal gRNA sequence: Use CRISPOR software to screen gRNA sequences in the gRNA design region. All sequences must meet the following three conditions:
[0084] ① The specificity score is 100;
[0085] The gRNA specificity score ranges from 0 to 100. The higher the score, the higher the targeting specificity of the gRNA sequence.
[0086] ② Off-target prediction is all 0; when off-target prediction is all 0, it means there is no risk of off-target. Off-target means that the target is hit in a region other than the target region, causing the gene editing of the target gene to fail.
[0087] ③Doench'16 score ≥ 40;
[0088] Doench'16 is a prediction of cutting efficiency. When the Doench'16 score is greater than 40, the cutting effect is around 50%. Therefore, 40 is the minimum value set by this scheme to ensure the best cutting efficiency.
[0089] http: / / crispor.tefor.net / can be used as one embodiment of the CRISPOR software mentioned above. The three conditions in step (2-1-2) fully consider factors such as gRNA specificity, cleavage efficiency, and off-target effects.
[0090] (2-1-3) Sequence complexity analysis: Perform lattice analysis and GC content analysis on the gene and its upstream and downstream 1000bp sequences, and identify the location of potential hairpin structures; after identifying the location of potential hairpin structures, it is easier to avoid the location of hairpin structures when designing primers in the future, as hairpin structures will cause difficulties in primer amplification sequences.
[0091] The system analyzes the complexity and GC content of gene sequences within a certain range before and after the mutation site to ensure the feasibility and accuracy of the editing protocol. If the two primer pairs are not structurally complementary, a single primer cannot form a hairpin structure. If the two primer pairs are structurally complementary, a single primer can form a hairpin structure. A hairpin structure generally refers to a DNA molecule folding back, where some bases come close to each other. The bases within the folded region pair complementarily, and the folded portion forms a hairpin structure. In other words, a single-stranded DNA molecule folds back itself, allowing complementary base pairs to meet and form hydrogen bonds.
[0092] When designing primers, hairpin structures with loops less than 17 bp should be avoided. If a sequence is the reverse complementary sequence of another sequence, the DNA single-strand molecule will fold back to form a hairpin structure. Take 4 bp segments of the sequence sequentially (record the index of the last base) and perform reverse complementary transformations (A to T, T to A, G to C, C to G, then reverse the sequence). Use this as a pattern to search for identical sequences in the entire sequence and obtain the index of the sequence that is identical to the pattern. If the index differs from the index of the pattern by less than 17, it is a hairpin structure with a loop less than 17 bp. If such a structure appears twice, the iteration is terminated and the region is excluded.
[0093] (2-2) The fungal design scheme uses a frameshift scheme, which includes the following steps:
[0094] (2-2-1) The frameshift scheme's gRNA is located within the gene at any position that does not overlap with other genes after the start codon;
[0095] The frameshifting gRNA can be located inside the gene, mainly to avoid regions that overlap with other genes.
[0096] (2-2-2) Use CRISPOR software to screen gRNA sequences in the gRNA design region, and select the optimal gRNA sequence according to the following rules:
[0097] ① Specificity Score ≥ 80, Off-target predictions are all 0, Lindel score ≥ 80, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions;
[0098] (2-2-3) Sequence complexity analysis: Perform lattice analysis and GC content analysis on the gene and its upstream and downstream 1000bp sequences, and identify the location of potential hairpin structures;
[0099] (3) Design two pairs of four primer sequences for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0100] Upstream homologous arm - reverse primer up-R: Primers can be designed from 200 bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene.
[0101] Upstream homologous arm - forward primer up-F: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp from the up-R primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp from the up-R primer.
[0102] Down-F homologous arm: Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first downstream coding gene, with preference given to those closest to the start or stop codon of the gene.
[0103] Down-R homologous arm: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp from the down-F primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp from the down-F primer.
[0104] Primers are used for PCR amplification to identify the genotype after gene editing and confirm whether the gene editing was successful. This protocol designs two pairs of four primer sequences according to the above conditions. Firstly, it requires a certain primer length for PCR amplification (the longer the primer, the more difficult it is to amplify). Secondly, it considers methods for identifying genotypes. Reasonable primer design allows for easier determination of gene editing success by comparing the genotypes before and after PCR amplification.
[0105] (4) Output the design results of sgRNA sequence and homologous arm primer sequence.
[0106] This protocol provides design schemes for sgRNA sequences and homologous arm primer sequences for gene editing experiments on various strains. The sgRNA sequence design method comprehensively considers factors such as design region and specificity, resulting in a higher gene editing success rate. It can replace complicated and time-consuming manual methods, saving researchers a lot of time and solving the problems of high costs faced in manually designing the optimal site for gene knockout and designing sgRNA sequence schemes.
[0107] Alternatively, in step (1), the target gene can be determined based on three identifiers: gene name, gene ID, or locus tag.
[0108] This solution only requires selecting a specific strain and target gene, and it will automatically generate an optimal solution based on steps (2) and (3), saving researchers a lot of time.
[0109] A gene name is a description of a gene's function and is usually composed of symbols that reflect the gene's function or characteristics.
[0110] A gene ID is a unique identifier for a gene in a database;
[0111] A locus tag is a unique identifier for a gene or gene sequence within a genome. It typically consists of two parts: a prefix and a suffix, separated by an underscore. The prefix is the genome identifier, indicating that the gene belongs to a specific genome; the suffix is a number indicating the gene's order or uniqueness within the genome.
[0112] Optimally, in step (2-1-2), among the gRNA sequences that simultaneously meet all three conditions, the gRNA with the highest Lindel score is selected.
[0113] The Lindel score predicts the probability that a given gRNA sequence will cause an insertion or deletion in a DNA sequence. A higher score indicates a higher success rate in causing a change in the target gene.
[0114] Ideally, in step (3), the homologous arm primers should meet the following criteria:
[0115] ①The primer Tm value is between 58 and 62℃;
[0116] ②The GC content of the primers is between 35% and 65%;
[0117] ③The last base at the 3' end is G or C;
[0118] ④ Primer pairs cannot have consecutive 6bp complementary pairings;
[0119] ⑤ Primer sequence length 17–30 bp.
[0120] Alternatively, in step (2-2-2), if there is no sequence that meets rule ①, then rule ② can be used for further filtering.
[0121] ② Specificity Score ≥ 80, Lindel ≥ 70, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions; if there is no sequence that meets rule ②, continue to use rule ③ for screening;
[0122] ③ Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the frameshift knockout scheme cannot be designed for this gene.
[0123] Alternatively, in step (2-2), the fungal design scheme may use a whole-genome knockout scheme instead of a frameshift scheme.
[0124] (2-2-1) gRNA is located within the gene, between 100 and 400 bp on both sides, avoiding areas that overlap with other genes;
[0125] (2-2-2) Use CRISPOR software to screen gRNA sequences in the gRNA design region, and select the optimal gRNA sequence according to the following rules:
[0126] ① Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If there is no sequence that meets this rule, then select according to rule ②.
[0127] ② Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets the above conditions, the gene knockout scheme cannot be designed.
[0128] (2-2-3) Sequence complexity analysis: The gene and its upstream and downstream 1000bp sequences were subjected to lattice analysis and GC content analysis to identify the location of potential hairpin structures.
[0129] This approach allows for the selection of either whole-genome knockout or frameshift as needed. Whole-genome knockout is more difficult, but if successful, it results in a more thorough gene removal. Frameshifting may not affect gene function. Both have their advantages and are feasible.
[0130] An automated design system for microbial gene editing protocols includes: a gene acquisition module, a bacterial design strategy module, a fungal design strategy module, and a homology arm identification primer design module;
[0131] The gene acquisition module is used to select design objects for different types of strains, namely bacteria or fungi; after determining the target gene, it retrieves the target gene information from the NCBI database; and according to whether the design object is selected as bacteria or fungi, it calls the bacterial design strategy module or the fungal design strategy module.
[0132] The bacterial design strategy module is used to determine the gRNA design region: when the gene is ≤1000bp, two gRNAs are designed in the region between 1 / 4 and 3 / 4 of the gene; when the gene is greater than 1000bp, two gRNAs are designed at both ends of the gene, with the first gRNA designed in the range of 100-400bp on one side and the second gRNA designed in the range of 100-400bp on the other side; and the optimal gRNA sequence is selected: gRNA sequences are screened in the gRNA design region using CRISPOR software, and all sequences must meet the following three conditions:
[0133] ① The specificity score is 100;
[0134] ② The off-target predictions are all 0;
[0135] ③Doench'16 score ≥ 40;
[0136] The bacterial design strategy module is also used for sequence complexity analysis: after performing matrix analysis and GC content analysis on the gene and its upstream and downstream 1000bp sequences, the location of potential hairpin structures is identified.
[0137] The fungal design strategy module is used to execute a frameshift scheme: it places the gRNA at any non-overlapping position after the start codon within the gene; and uses CRISPOR software to screen gRNA sequences in the gRNA design region, selecting the optimal gRNA sequence according to the following rules:
[0138] ① The gRNA with the highest Lindel score is selected from the gRNAs that meet the above conditions, with a Specificity Score ≥ 80, Off-target predictions all being 0, and a Lindel score ≥ 80. The fungal design strategy module is also used for sequence complexity analysis: the gene and its upstream and downstream 1000bp sequences are subjected to lattice analysis and GC content analysis, and the location of potential hairpin structures is identified.
[0139] The homologous arm identification primer design module is used to design two pairs of four primer sequences for homologous arm identification. The selection criteria for the primer design region are as follows:
[0140] ① Upstream homologous arm - reverse primer up-R: Primers can be designed from 200bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene;
[0141] ② Upstream homologous arm - forward primer up-F: If the target gene length is ≤2000bp, design primers in the region between 450-600bp from the up-R primer; if the target gene length is greater than 2000bp, design primers in the region between 950-1100bp from the up-R primer.
[0142] ③ Down-F (down homologous arm): Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first coding gene downstream. Primers closer to the start or stop codon of the gene are preferred.
[0143] ④ Down-R homologous arm: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp away from the down-F primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp away from the down-F primer.
[0144] Alternatively, if no sequence matches rule ①, then rule ② can be used for further filtering.
[0145] ② Specificity Score ≥ 80, Lindel ≥ 70: Select the gRNA with the highest Lindel score from the gRNAs that meet the above criteria. If no sequence meets these criteria, then proceed to rule ③ for selection.
[0146] ③ Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the frameshift knockout scheme cannot be designed for this gene.
[0147] Optimally, the fungal design strategy module is also used to execute a whole-gene knockout scheme, placing the gRNA within a range of 100-400 bp flanking the gene, avoiding regions overlapping with other genes; using CRISPOR software to screen gRNA sequences in the gRNA design region, and selecting the optimal gRNA sequence according to the following rules:
[0148] ① Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If there is no sequence that meets this rule, then select according to rule ②.
[0149] ② Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the gene cannot be designed into a whole-gene knockout scheme.
[0150] A storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the above-described automatic design method for a microbial gene editing scheme.
[0151] Example 1:
[0152] An automated design method for microbial gene editing protocols includes the following steps:
[0153] (1) Select the design object for different types of strains. The design object is Escherichia coli, the strain is K-12substr.MG1655, the gene ID is 944749, the gene name is yaaA, and the Locus tag is b0006. After determining the target gene, obtain the target gene information from the NCBI database.
[0154] (2) Select Escherichia coli as the design object and execute step (2-1);
[0155] (2-1) The bacterial design scheme includes the following steps:
[0156] (2-1-1) Determine the gRNA design region: When the gene is 776bp or ≤1000bp, the two gRNAs are designed in the region between 1 / 4 and 3 / 4 of the gene, respectively.
[0157] (2-1-2) Selecting the optimal gRNA sequence: Use CRISPOR software to screen gRNA sequences in the gRNA design region. All sequences must meet the following three conditions:
[0158] ① The specificity score is 100;
[0159] ② The off-target predictions are all 0;
[0160] ③Doench'16 score ≥ 40;
[0161] Among gRNA sequences that simultaneously meet all three conditions, the gRNA with the highest Lindel score is selected.
[0162] The final gRNA sequence screening results are as follows:
[0163] ① gRNA1(rev) is CATCACCAACAAGCTGAACG AGG(100,85); the specificity score is 100 and the cleavage efficiency prediction score is 85.
[0164] ② gRNA2(rev) is TCCGTCTTGAGAATGCCCGA GGG(100,82); where the specificity score is 100 and the cleavage efficiency prediction score is 82.
[0165] (2-1-3) Sequence complexity analysis: Perform raster analysis on the gene and its upstream and downstream 1000bp sequences (e.g., Figure 2 (and Table 1) and GC content analysis (e.g.) Figure 1 ), and identify the locations of potential hairpin structures, as shown in Table 2;
[0166] Table 1 - Data from dot matrix analysis of gene and its upstream and downstream 1000bp sequences.
[0167] End 30 41 190 252 341 384 502 518 651 Start 677 777 869 977 1188 1201 1505 1538 1603 End 686 786 878 990 1198 1211 1515 1548 1613 Start 1701 1796 1894 2013 2163 2311 2396 2421 2498 End 1710 1806 1904 2022 2173 2321 2410 2437 2508 Start 2521 End 2530
[0168] The ±1000bp fragments of the target region are aligned with itself to determine the presence of complex sequences. For Start As shown in Table 1, the average GC content in the target region of the ±1000bp fragment is 50.41%. This region is suitable for PCR screening or sequencing analysis to identify the location of potential hairpin structures, as shown in Table 2.
[0169] Table 2 - Location of Hairpin Structures
[0170] End 40 75 88 131 221 256 318 349 389 Start 390 431 472 498 559 609 628 696 729 End 416 441 496 538 569 622 685 722 739 Start 740 753 804 855 923 997 1027 1050 1080 End 752 776 825 905 987 1025 1044 1072 1105 Start 1114 1146 1184 1209 1234 1259 1330 1395 1420 End 1135 1164 1202 1228 1254 1316 1376 1415 1432 Start 1446 1474 1532 1570 1593 1618 1728 1784 1796 End 1461 1525 1563 1580 1610 1700 1764 1794 1820 Figure 4 1844 1897 1919 1969 1998 2059 2177 2218 2267 Figure 3 1865 1917 1953 1980 2058 2163 2185 2237 2309 End 2310 2388 2407 2442 2476 2538 2591 2616 2660 Start 2336 2406 2439 2471 2498 2557 2612 2636 2686 End 2717 Figure 4 2738
[0171] (3) Upstream homologous arm - reverse primer up-R: Primers can be designed from 200 bp to the front 1 / 3 of the target gene from the right end of the first coding gene upstream. Priority should be given to primers that are close to the start codon or stop codon of the gene.
[0172] Upstream homologous arm - forward primer up-F: Primers are designed in the region between 450 and 600 bp from the up-R primer;
[0173] Down-F homologous arm: Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first downstream coding gene, with preference given to those closest to the start or stop codon of the gene.
[0174] Downstream homologous arm down-R: Primers are designed in the region between 450 and 600 bp from the down-F primer;
[0175] The homologous arm primers meet the following criteria:
[0176] ①The primer Tm value is between 58 and 62℃;
[0177] ②The GC content of the primers is between 35% and 65%;
[0178] ③The last base at the 3' end is G or C;
[0179] ④ Primer pairs cannot have consecutive 6bp complementary pairings;
[0180] ⑤ Primer sequence length 17–30 bp.
[0181] Thus, two pairs of four primer sequences were designed for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0182] Upstream homologous arm - reverse primer up-R: GTTTGCAGTCAATGCCGGATG;
[0183] Upstream homologous arm - forward primer up-F: CAGCCCGGCTTTTTTATGAAG;
[0184] Downstream homologous arm down-F: CAGGTGAAATAAGAATCAGCATATC;
[0185] Downstream homologous arm down-R: AATGGGTTCCTGGGGTGC.
[0186] (4) Output the design results.
[0187] Example 2:
[0188] An automated design method for microbial gene editing protocols includes the following steps:
[0189] (1) Select the design object for different types of strains. The design object is yeast, the strain is Saccharomyces cerevisiae S288c, the gene ID is 851839, the gene name is BTT1, the gene length is 449bp, and the Locus tag is YDR252W. After determining the target gene, obtain the target gene information from the NCBI database.
[0190] (2) Select yeast as the design object and execute step (2-2);
[0191] (2-2) The fungal design scheme uses a frameshift scheme, which includes the following steps:
[0192] (2-2-1) The frameshift scheme's gRNA is located within the gene at any position that does not overlap with other genes after the start codon;
[0193] (2-2-2) Use CRISPOR software to screen gRNA sequences in the gRNA design region, and select the optimal gRNA sequence according to the following rules:
[0194] ① Specificity Score ≥ 80, Off-target predictions are all 0, Lindel score ≥ 80, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions;
[0195] ② Specificity Score ≥ 80, Lindel ≥ 70, select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions; if there is no sequence that meets rule ②, continue to use rule ③ for screening;
[0196] ③ Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the frameshift knockout scheme cannot be designed for this gene.
[0197] gRNA1(rev) was selected as ATGAAGTTCAGCTTGTAACTTGG(100,90); the specificity score was 100 and the cleavage efficiency prediction score was 90.
[0198] (2-2-3) Sequence complexity analysis: Perform raster analysis on the gene and its upstream and downstream 500bp sequences (e.g., End (and Table 3) and GC content analysis (e.g. Start ), and identify the locations of potential hairpin structures, as shown in Table 4;
[0199] Table 3 - Data from dot matrix analysis of gene and its upstream and downstream 500bp sequences.
[0200] End 22 158 190 266 283 402 438 563 1059 Start 1311 1437 End 1320 1450
[0201] The ±500bp fragments of the target region are aligned with itself to determine the presence of complex sequences. For Figure 6 As shown in Table 3, the average GC content in the target region of the ±500bp fragment is 36.62%. This region is suitable for PCR screening or sequencing analysis to identify the location of potential hairpin structures, as shown in Table 4.
[0202] Table 4 - Location of Hairpin Structures
[0203] Figure 5 60 98 214 241 272 286 311 376 416 End 420 466 489 522 560 577 627 739 823 Start 433 487 510 550 570 607 737 821 883 End 899 939 1008 1035 1094 1204 1235 1291 1315 Start 920 992 1026 1078 1104 1226 1286 1314 1385
[0204] (3) Design two pairs of four primer sequences for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0205] Upstream homologous arm - reverse primer up-R: Primers can be designed from 200 bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene.
[0206] Upstream homologous arm - forward primer up-F: Primers are designed in the region between 450 and 600 bp from the up-R primer;
[0207] Down-F homologous arm: Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first downstream coding gene, with preference given to those closest to the start or stop codon of the gene.
[0208] Downstream homologous arm down-R: Primers are designed in the region between 450 and 600 bp from the down-F primer;
[0209] The homologous arm primers meet the following criteria:
[0210] ①The primer Tm value is between 58 and 62℃;
[0211] ②The GC content of the primers is between 35% and 65%;
[0212] ③The last base at the 3' end is G or C;
[0213] ④ Primer pairs cannot have consecutive 6bp complementary pairings;
[0214] ⑤ Primer sequence length 17–30 bp.
[0215] Thus, two pairs of four primer sequences were designed for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0216] Upstream homologous arm - reverse primer up-R: GGCATTTGATGTTGTAGGATTATG;
[0217] Upstream homologous arm - forward primer up-F: TGCGGATTCTGCTCCGTG;
[0218] Downstream homologous arm down-F: CAACAAGTGATGAATAGCTGAC;
[0219] Downstream homologous arm down-R: CACCGAAGTCCAGCAAAAGC.
[0220] (4) Output the design results.
[0221] Example 3:
[0222] An automated design method for microbial gene editing protocols includes the following steps:
[0223] (1) Select the design object for different types of strains. The design object is yeast, the strain is Saccharomyces cerevisiae S288c, the gene ID is 851839, the gene name is BTT1, the gene length is 449bp, and the Locus tag is YDR252W. After determining the target gene, obtain the target gene information from the NCBI database.
[0224] (2) Select yeast as the design object and execute step (2-2);
[0225] (2-2) The fungal design scheme uses a whole-genome knockout protocol; the whole-genome knockout protocol includes the following steps:
[0226] (2-2-1) gRNA is located within the gene, between 100 and 400 bp on both sides, avoiding areas that overlap with other genes;
[0227] (2-2-2) Use CRISPOR software to screen gRNA sequences in the gRNA design region, and select the optimal gRNA sequence according to the following rules:
[0228] ① Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If there is no sequence that meets this rule, then select according to rule ②.
[0229] ② Specificity Score ≥ 70, Lindel ≥ 60. Select the gRNA with the highest Lindel score from the gRNAs that meet the above conditions. If there is no sequence that meets these conditions, the gene cannot be designed into a whole-gene knockout scheme.
[0230] gRNA1(fw) was selected as TTATCAGCAGCAAACAAAGTTGG(100,76); the specificity score was 100 and the cleavage efficiency prediction score was 76.
[0231] gRNA2(fw) was selected as CAAGAACTTGAGTATTTGACAGG(100,79); the specificity score was 100 and the cleavage efficiency prediction score was 79.
[0232] (2-2-3) Sequence complexity analysis: Perform raster analysis on the gene and its upstream and downstream 1000bp sequences (e.g., End (and Table 5) and GC content analysis (e.g. Figure 6 ), and identify the locations of potential hairpin structures, as shown in Table 6;
[0233] Table 5 - Data from dot matrix analysis of gene and its upstream and downstream 1000bp sequences.
[0234] End 272 522 557 658 690 766 783 902 938 Start 1054 1550 1636 1766 1811 1937 2092 2171 2184 End 1063 1559 1648 1776 1820 1950 2101 2180 2196 Start 2357 End 2381
[0235] The ±1000bp fragments of the target region are aligned with itself to determine the presence of complex sequences. For Start As shown in Table 5, the average GC content in the target region within the ±1000 bp fragment is 36.12%. This region is suitable for PCR screening or sequencing analysis to identify the location of potential hairpin structures, as shown in Table 6.
[0236] Table 6 - Location of Hairpin Structures
[0237] End 26 60 128 161 200 252 321 365 420 Start 426 536 565 653 725 747 775 797 816 End 448 560 598 714 741 772 786 811 876 Start 890 920 966 989 1022 1060 1077 1127 1239 End 916 933 987 1010 1050 1070 1107 1237 1321 1323 1399 1439 1508 1535 1594 1704 1735 1791 1383 1420 1492 1526 1578 1604 1726 1786 1814 1815 1947 1966 2009 2131 2231 2258 2352 2406 1885 1960 1982 2119 2172 2244 2351 2396 2430 2433 2448
[0238] (3) Design two pairs of four primer sequences for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0239] Upstream homologous arm - reverse primer up-R: Primers can be designed from 200 bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene.
[0240] Upstream homologous arm - forward primer up-F: Primers are designed in the region between 450 and 600 bp from the up-R primer;
[0241] Down-F homologous arm: Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first downstream coding gene, with preference given to those closest to the start or stop codon of the gene.
[0242] Downstream homologous arm down-R: Primers are designed in the region between 450 and 600 bp from the down-F primer;
[0243] The homologous arm primers meet the following criteria:
[0244] ①The primer Tm value is between 58 and 62℃;
[0245] ②The GC content of the primers is between 35% and 65%;
[0246] ③The last base at the 3' end is G or C;
[0247] ④ Primer pairs cannot have consecutive 6bp complementary pairings;
[0248] ⑤ Primer sequence length 17–30 bp.
[0249] Thus, two pairs of four primer sequences were designed for the identification of homologous arms. The selection criteria for the primer design region are as follows:
[0250] Upstream homologous arm - reverse primer up-R: GGCATTTGATGTTGTAGGATTATG;
[0251] Upstream homologous arm - forward primer up-F: TGCGGATTCTGCTCCGTG;
[0252] Downstream homologous arm down-F: CAACAAGTGATGAATAGCTGAC;
[0253] Downstream homologous arm down-R: CACCGAAGTCCAGCAAAAGC.
[0254] (4) Output the design results.
[0255] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. An automated design method for microbial gene editing protocols, characterized in that, Includes the following steps: (1) Select design objects for different types of strains, with fungi as the design object; After identifying the target gene, obtain the target gene information from the NCBI database; (2) When the design object of step (1) is selected as fungi, step (2-2) is executed. (2-2) The fungal design scheme uses a whole-genome knockout protocol, which includes the following steps: (2-2-1) The sgRNA is located within the gene, between 100 and 400 bp on both sides, avoiding areas that overlap with other genes; (2-2-2) Use CRISPOR software to screen sgRNA sequences in the sgRNA design region, and select the optimal sgRNA sequence according to the following rules: ① Select the sgRNA with the highest Lindel score from those sgRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If no sequence meets this rule, then select according to rule ②. ② Specificity Score ≥ 70, Lindel ≥ 60. Select the sgRNA with the highest Lindel score from the sgRNAs that meet the above conditions. If there is no sequence that meets the conditions, the gene cannot be designed into a whole gene knockout scheme. (2-2-3) Sequence complexity analysis: Perform lattice analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000bp, and identify the location of potential hairpin structures; (3) Design two pairs of four primer sequences for the identification of homologous arms. The selection criteria for the primer design region are as follows: Upstream homologous arm - reverse primer up-R: Primers can be designed from 200 bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene. Upstream homologous arm - forward primer up-F: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp from the up-R primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp from the up-R primer. Down-F homologous arm: Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first downstream coding gene, with preference given to those closest to the start or stop codon of the gene. Down-R homologous arm: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp from the down-F primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp from the down-F primer. (4) Output the design results of the sgRNA sequence and homologous arm primer sequences; In step (3), the homologous arm primers meet the following criteria: ①The primer Tm value is between 58 and 62℃; ②The GC content of the primers is between 35% and 65%; ③The last base at the 3' end is G or C; ④ Primer pairs cannot have consecutive 6bp complementary pairings; ⑤ Primer sequence length 17–30 bp.
2. The method for automatically designing a microbial gene editing scheme according to claim 1, characterized in that, In step (1), the target gene is determined based on three characteristic identifiers: gene name, gene ID, or locus tag.
3. An automated design system for microbial gene editing protocols, characterized in that, An automated design method for executing a microbial gene editing scheme according to any one of claims 1-2, comprising: a gene acquisition module, a fungal design strategy module, and a homology arm identification primer design module; The gene acquisition module is used to select design targets for different types of strains, with fungi as the design target; after determining the target gene, it retrieves the target gene information from the NCBI database; and based on the selection of fungi as the design target, it calls the fungal design strategy module. The fungal design strategy module is used to execute a whole-gene knockout protocol, placing the sgRNA within a range of 100-400 bp flanking the gene, avoiding regions overlapping with other genes; using CRISPOR software to screen sgRNA sequences within the sgRNA design region, and selecting the optimal sgRNA sequence according to the following rules: ① Select the sgRNA with the highest Lindel score from those sgRNAs that meet the above conditions, provided that the Specificity Score is ≥80, the Lindel score is ≥80, and the Off-targets score is 0. If no sequence meets this rule, then select according to rule ②. ② Specificity Score ≥ 70, Lindel ≥ 60. Select the sgRNA with the highest Lindel score from the sgRNAs that meet the above conditions. If there is no sequence that meets the conditions, the gene cannot be designed as a whole gene knockout scheme. The fungal design strategy module is used to perform sequence complexity analysis: perform matrix analysis and GC content analysis on the gene and its upstream and downstream sequences within 1000 bp, and find the location of potential hairpin structures. The homologous arm identification primer design module is used to design two pairs of four primer sequences for homologous arm identification. The selection criteria for the primer design region are as follows: ① Upstream homologous arm - reverse primer up-R: Primers can be designed from 200bp to the right end of the first coding gene upstream to the first 1 / 3 of the target gene, with priority given to those closer to the start or stop codon of the gene; ② Upstream homologous arm - forward primer up-F: If the target gene length is ≤2000bp, design primers in the region between 450-600bp from the up-R primer; if the target gene length is greater than 2000bp, design primers in the region between 950-1100bp from the up-R primer. ③ Down-F (down homologous arm): Primers can be designed from the last 1 / 3 of the target gene to within 200 bp from the left end of the first coding gene downstream. Primers closer to the start or stop codon of the gene are preferred. ④ Down-R homologous arm: If the target gene length is ≤2000bp, design primers in the region between 450 and 600bp away from the down-F primer; if the target gene length is greater than 2000bp, design primers in the region between 950 and 1100bp away from the down-F primer.
4. A storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to perform as claimed in claim 1.
2. An automated design method for a microbial gene editing scheme as described in any one of the above.