A traditional gene editing plasmid and primer design system

The GUI-encapsulated offline gene editing plasmid and primer design system solves the problem of low efficiency in existing systems, enables batch design of multiple targets and high-throughput operation, provides high-quality primer and plasmid maps, improves research efficiency and reduces costs.

CN116994657BActive Publication Date: 2026-07-21SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2023-08-23
Publication Date
2026-07-21

Smart Images

  • Figure CN116994657B_ABST
    Figure CN116994657B_ABST
Patent Text Reader

Abstract

The application discloses a traditional gene editing plasmid and primer design system and belongs to the technical field of bioinformatics. The open source code based on the primer design software PRIMER3 is used, and automatic operation is realized by using a python script in a Linux or windows operating system, so that plasmid mapping, primer design and evaluation are completed without additional operation. The primer design system also introduces a border primer design strategy, allows longer primer sequence design, and obtains high-quality primers with more specific 3' ends. The DNA editing plasmid and primer design system can realize automatic and continuous design of multiple targets, and contains a GUI program encapsulated by Pyinstaller, so that convenient use is realized without dependence on a network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a traditional gene editing plasmid and primer design system, belonging to the field of bioinformatics technology. Background Technology

[0002] Gene editing, including insertion, deletion, and substitution, is a common method in basic and applied research, such as studying gene function and genetically modifying cells. Traditional gene editing is based on the use of intracellular recombinases to perform homologous recombination with the cellular genome through single or double crossovers mediated by homologous arms. Plasmid-mediated gene editing is another common method. Plasmids used for gene editing consist of a vector backbone, a left homologous arm, a knock-in fragment such as a selectable marker gene (optional), and a right homologous arm. Currently, constructing a complete gene-editing plasmid requires, first, determining the editing target and type; then, selecting a vector backbone based on host specificity and designing primers to amplify the backbone, choosing a suitable homologous arm fragment and designing primers to amplify it, selecting a suitable selectable marker gene (optional, as a knock-in fragment) and amplification primers; and since the most efficient method for constructing gene-editing plasmids is Gibson assembly, which relies on a short homologous overlap, typically designed directly onto primers and synthesized together, the overlap is introduced into the DNA fragment via PCR amplification; finally, outputting a complete DNA-editing plasmid map is also necessary for sequence alignment and other analyses. On the other hand, in recent years, researchers have attempted to develop technologies and equipment through multidisciplinary collaboration to automate and achieve high-throughput completion of all stages of design, construction, testing, and learning (DBTL) in synthetic biology, building large-scale "biofactories."

[0003] Achieving laboratory automation and higher levels of reproducibility in bio-foundries has become one of the long-term goals of synthetic biology development. Software capable of automating and performing high-throughput design of plasmids, primers, etc., is also essential for adapting to automated and high-throughput downstream experiments.

[0004] Currently, there is still a lack of dedicated software for the design of plasmids and primers for traditional gene editing. Existing plasmid and primer design systems have four main problems: First, they usually only support the design of a single plasmid map, requiring a large number of manual clicks, which is cumbersome and inefficient. There is a lack of one-stop software specifically for the design of plasmids and primers for traditional gene editing. Second, they cannot automatically provide primer quality prediction and plasmid map information, requiring users to complete subsequent work through other means or methods, which is inefficient, time-consuming, and labor-intensive, and not conducive to subsequent experimental construction. Third, they can only design for one target at a time, making it difficult to design primers and plasmid maps in batches for multiple targets, such as designing knockout plasmids and primers for multiple genes at the genome scale. Fourth, existing software is mainly online and relies on the network.

[0005] Therefore, developing tools that can automate and perform high-throughput design of plasmids and primers can save researchers significant time and money while improving work efficiency. Currently available software often only designs primers, or can only design a single plasmid at a time. It cannot perform batch gene editing plasmid mapping and primer design for multiple targets in a single step, still requiring users to spend considerable time on subsequent processing, resulting in low efficiency and unsuitability for subsequent automation and high-throughput operations. Summary of the Invention

[0006] To address the aforementioned deficiencies in existing technologies, this invention provides a traditional gene-editing plasmid and primer design system. This system allows customers to design gene-editing plasmids in batches for multiple target sites, given a genome and a vector. The output provides primers for amplifying vector backbones, homologous arms, and knock-in fragments (optional), as well as quality assessment primers for colony PCR or DNA sequencing to determine the success of plasmid construction, and plasmid maps for sequence alignment and other analyses. Furthermore, the system is packaged as an offline version with a GUI, making it convenient for users to use anytime, anywhere.

[0007] This invention provides a method for designing gene-editing plasmids and primers, comprising:

[0008] Obtain the information required for gene editing operations, including but not limited to genomic information, vector backbone information, and gene editing target information;

[0009] Primer design and primer design results are performed on the fragments required for gene editing operations; the required fragments include upstream and downstream homologous arms and plasmid backbone; optionally, gene fragments are also included.

[0010] Output the processing results; the processing results include a map of the gene editing plasmid and the primers required for the gene editing operation.

[0011] In one embodiment, the primers required for gene editing include one or more of the following: primers for amplifying upstream and downstream homologous arms of the target site, primers for amplifying the knock-in gene, primers for amplifying the vector, and identification primers.

[0012] In one embodiment, the method is applied to an electronic device; the electronic device performs the plasmid and primer design method based on the open-source code of PRIMER3 and a Python script.

[0013] In one embodiment, the primer design includes upstream and downstream homologous arm primer design, knock-in fragment primer design, vector backbone primer design, plasmid identification primer design, and gene editing identification primer design.

[0014] In one embodiment, the upstream and downstream homologous arm primer design, knock-in fragment primer design, and vector backbone primer design methods may or may not employ a border primer design strategy.

[0015] In one implementation, the border primer design strategy is:

[0016] Forward primer: Design a primer for amplifying the target polynucleotide in a wide range on one side of the 5' end of the target polynucleotide on the template DNA to obtain an efficient primer. Then extend the 5' end of the efficient primer to the 5' end of the target polynucleotide to obtain a complete forward primer.

[0017] Reverse primer: Design a primer for amplifying the target polynucleotide in a wide range on one side of the 3' end of the target polynucleotide on the template DNA to obtain an efficient primer. Then extend the 5' end of the efficient primer to the 3' end of the target polynucleotide to obtain a complete reverse primer.

[0018] The wider range refers to the region that is >0 bp from the 5' or 3' end of the polynucleotide.

[0019] In one implementation, when using a Border primer design strategy, the primer sequence consists of "Primer_Border+

[0020] The primer sequence is "Efficient_Primer"; when the Border primer design strategy is not used, the primer sequence is "Efficient_Primer".

[0021] In one implementation, when using a border primer design strategy, the primer sequence for amplifying the DNA fragment used for homologous recombination also contains overlapping primers, i.e., the complete primer consists of “Primer_Overlap+Primer_Border+Efficient_Primer”.

[0022] In one embodiment, when the gene editing type is gene insertion or replacement, the primer design method for the side of the left and right homologous arms away from the editing target site is as follows: A primer with the structure "Primer_Overlap+Efficient_Primer" is designed; where "Primer_Overlap" is a sequence complementary to the terminal sequence of the vector backbone; and "Efficient_Primer" is an effective primer designed within the extensional region. For the primer design for the side of the left and right homologous arms closer to the editing target site, a primer sequence with the structure "Primer_Overlap+Primer_Border+Efficient_Primer" is designed; where "Efficient_Primer" is an effective primer for amplifying the homologous arm; "Primer_Overlap" is a sequence complementary to the knock-in fragment; and "Primer_Border" is a sequence on the homologous arm where the 5' end of the effective primer extends to the target boundary sequence.

[0023] In one embodiment, when the gene editing type is gene deletion, the primer design method for the side of the left and right homologous arms away from the editing target site is as follows: design primers with the structure shown in "Primer_Overlap+Efficient_Primer"; wherein "Primer_Overlap" is a sequence complementary to the terminal sequence of the vector backbone; and "Efficient_Primer" is an effective primer designed in the extensional region; the primer design method for the side of the left and right homologous arms close to the editing target site is as follows: one primer has the structure "Primer_Overlap+Primer_Border+Efficient_Primer(5'-3')" and the other primer has the structure "Primer_Border+Efficient_Primer(5'-3')".

[0024] In one implementation, the evaluation is a calculation and assessment of information such as primer length, GC content, Tm value, and PCR product length; the evaluation is performed using PRIMER3.

[0025] In one embodiment, the gene-compiling plasmid map simulation is a complete plasmid formed by linking homologous arms, knock-in fragments (if any), and vector backbone fragments.

[0026] In one implementation, the genome file allows multiple genomes to be merged into a single genome file for processing.

[0027] In one implementation, the gene editing includes gene editing of one or more target sites.

[0028] In one embodiment, the information obtained also includes the length of homologous arms, the length of overlap in primers, the range of border length, the length of effective primers, GC content, Tm value and its calculation method, etc.

[0029] The present invention also provides a gene editing plasmid and primer design device, the device comprising:

[0030] The acquisition unit is used to acquire information required for gene editing operations, including but not limited to genomic information, vector backbone information, and gene editing target information.

[0031] The primer processing unit designs primers for the fragments required for gene editing operations and evaluates the primer design results; the required fragments include upstream and downstream homologous arms and plasmid backbone; optionally, it also includes gene fragments;

[0032] The output and display unit is used to output the processing results and display the plasmid map after the homologous arms, knock-in fragments (if any), and vector backbone fragments are ligated; the processing results include the map of the gene editing plasmid and the primers required for the gene editing operation.

[0033] In one embodiment, the device includes an online version or an offline version.

[0034] In one embodiment, the device uses a GUI encapsulation to implement visual operation.

[0035] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the gene editing plasmid and primer design method.

[0036] Beneficial effects:

[0037] (1) This invention proposes a border primer design strategy for the first time for primer design of gene fragments with poor 3' end specificity and many repetitive bases. By allowing longer primer sequence design, repetitive sequences are avoided to a certain extent from appearing at the 3' end of the primer, which is conducive to obtaining high-quality primers with 3' end specificity. In addition, the 5' end of the designed primer is fixed to ensure that the amplification obtains the complete target DNA sequence.

[0038] (2) This invention utilizes the open-source code of the primer design software PRIMER3. Under Linux or Windows operating systems, it employs Python scripts to design multiple gene-editing plasmids and primers for a given genome and vector, targeting a single gene, or designing gene-editing plasmids and primers for multiple targets or the entire genome in batches. Target localization allows for unique locus tags or base positions to meet diverse design needs. This invention's gene-editing plasmid and primer design system provides quality-assessed primer and plasmid maps, achieving, for the first time, the integration of primer design and plasmid map creation. The system uses PyInstaller to encapsulate the program as a GUI, enabling convenient use without relying on a network. It allows users to conduct gene-editing (e.g., deletion, insertion, replacement) research on a wider range of species and provides rapid, one-stop access to primer and plasmid maps needed for downstream wet experiments, effectively reducing the workload of researchers and saving time and economic costs.

[0039] (3) The gene editing plasmid and primer design system of the present invention has been tested in actual experiments. It can accurately design multiple targets, correctly output primer and plasmid maps, and complete the design of a target within 30 seconds. The output data information has been verified by downstream wet experiments. It can be used efficiently and conveniently for gene editing plasmid construction and analysis, which fully proves the feasibility of the invention. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of a traditional gene editing plasmid and primer design system.

[0041] Figure 2 This is a diagram of the interface of a traditional gene editing plasmid and primer design system.

[0042] Figure 3 This is a schematic diagram illustrating the primer design principle of a traditional gene editing plasmid and primer design system.

[0043] Figure 4 This is an application example of a traditional gene editing plasmid and primer design system.

[0044] Figure 5 PCR amplification results for primers designed for different methods. Detailed Implementation

[0045] Technical terms:

[0046] Overlapping primer: In this invention, “primer overlap region”, “overlap sequence”, “primer overlap sequence” and “overlapping primer” can be used interchangeably, all referring to a primer segment with a complementary sequence of a given length to the adjacent segment.

[0047] The primer prediction range (Border_Range) refers to the region greater than 0 bp within the polynucleotide used for amplification, preferably a region 40 bp inside the 5' and / or 3' ends of the polynucleotide. In some embodiments of the present invention, effective primers are designed within the Border_Range, specifically by starting with the first base at the 5' or 3' end of the polynucleotide and moving inwards to design effective primers.

[0048] Primer boundary: refers to the sequence between the 5' end of the effective primer and the 5' or 3' end of the fragment used for amplification within the Border_Range region.

[0049] Efficient Primer: A short single-stranded DNA fragment designed based on the target DNA to be amplified and given parameters. When using a Border primer design strategy, the complete primer for amplifying the target DNA consists of the efficient primer and primer boundaries, allowing the complete primer to bind to the nucleic acid strand of the template DNA, complementing it and serving as the initiation site for nucleotide polymerization to synthesize a new nucleic acid strand identical to the 5' end of the target DNA. Optionally, when the fragment amplified by the primer is also used for homologous recombination, the complete primer may also include overlapping primers. When not using a Border primer design strategy, the efficient primer can be directly used as the primer for amplifying the target DNA.

[0050] Extended region (Extend_size): A region greater than 0 bp extending outward from the 5' end of the left homologous arm or the 3' end of the right homologous arm, which is used by primer3.0 to design effective primers starting from the first base of the extended region in the direction of the homologous arm.

[0051] Left homologous arm (LHA): The sequence upstream of the target site to be edited is completely identical to the genome sequence. It is used to identify and allow recombination to occur. It is also known as the "upstream homologous arm".

[0052] Right homologous arm (RHA): The region to the right of the target site to be edited that is completely identical to the genome sequence. It is used to identify and allow recombination to occur, and is also known as the "downstream homologous arm".

[0053] Knock-in DNA fragment: A DNA fragment to be inserted into the target site (the target site to be edited), or a DNA fragment to replace the gene at the target site.

[0054] Example 1: An offline system for designing traditional gene editing plasmids and primers

[0055] The design system described in this embodiment includes a primer design module and a spectrum design module. The program is packaged into a GUI program using PyInstaller to enable convenient use without relying on the network.

[0056] (1) Primer design module

[0057] The primer design module includes sub-modules for upstream and downstream homologous arm primer design, knock-in fragment primer design, vector backbone primer design, plasmid identification primer design, and gene editing identification primer design. These modules are used to design primers for the following purposes: primers for PCR amplification of homologous arms that undergo homologous recombination with the target genome; primers for PCR amplification of knock-in DNA fragments; primers for PCR amplification of vector backbones; identification primers for colony PCR or DNA sequencing to determine plasmid correctness; and identification primers for colony PCR or DNA sequencing to determine the correctness of gene editing in mutant strains. The module also calculates and evaluates primer length, GC content, Tm value, and PCR product length. Optionally, this evaluation is performed using PRIMER3. Please refer to [link to relevant documentation] for primers designed and amplified fragments designed in the primer design module. Figure 3 .

[0058] (a) Upstream and downstream homologous arm primer design module

[0059] This module can design primers for PCR amplification of homologous arms that undergo homologous recombination with the target genome. That is, after determining the specific sequence of the homologous arm based on information such as the editing target site and the length of the homologous arm specified by the user, the homologous arm primers are designed.

[0060] When the gene editing type is gene insertion or replacement, the design strategy (see reference) Figure 3 a) including:

[0061] For the design of primers on the side of the left and right homologous arms furthest from the editing target, i.e., the forward primer of the left homologous arm (Left_Homologous_Arm, LHA) and the reverse primer of the right homologous arm (Right_Homologous_Arm, RHA): the primer sequence consists of "Primer_Overlap + Efficient_Primer". The corresponding length of the sequence is extracted from the end sequence of the vector backbone according to the user-specified length (Overlap_Length) as "Primer_Overlap". According to the user-given broad extension region (Extend_size), the effective primer (Efficient_Primer) is designed within this region.

[0062] For primer design on the side of the left and right homologous arms closest to the editing target site, i.e., the reverse primer for the left homologous arm (LHA) and the forward primer for the right homologous arm (RHA): The user selects whether to use a Border primer design strategy. If the Border primer design strategy is used, the primer sequence consists of "Primer_Overlap + Primer_Border + Efficient_Primer". If the Border primer design strategy is not used, the primer sequence consists of "Primer_Overlap + Efficient_Primer".

[0063] When using a border primer design strategy, an effective primer (Efficient_Primer) is designed within a given sequence region containing a relatively broad primer prediction range (Border_Range). Then, the 5' end of this effective primer is padded to the target boundary (the reverse primer for the left homologous arm is padded to the 5' end of the editing target, and the forward primer for the right homologous arm is padded to the 3' end of the editing target). The padded sequence is called the primer boundary (Primer_Border). This boundary is then combined with the overlapping primer (Primer_Overlap) used for subsequent DNA fragment assembly to form a complete primer, i.e., Primer_Overlap+.

[0064] Primer_Border+Efficient_Primer(5'-3'). This primer design method ensures high specificity at the 3' end, thereby improving the success rate of downstream PCR amplification, while also ensuring that the expected left and right homologous arms remain unchanged on the target side. The default value for Border_Range is 40 nt, which is a recommended value that takes into account primer synthesis cost and effective overlap length.

[0065] If the Border primer design strategy is not used, and the user does not specify Border_Range, then the Border primer design strategy will not be used. First, the end of the homologous arm is used as the 5' end of the effective primer to generate the effective primer "Efficient_Primer". Then, an overlapping primer (Primer_Overlap) is spliced ​​to the 5' end of this effective primer to form the complete primer, i.e., Primer_Overlap+.

[0066] Efficient_Primer(5'-3').

[0067] When the gene editing type is gene deletion, the design strategy (see reference) Figure 3 b) Includes:

[0068] For the design of primers on the side of the left and right homologous arms furthest from the editing target site, i.e., the forward primer for the left homologous arm (Left_Homologous_Arm, LHA) and the reverse primer for the right homologous arm (Right_Homologous_Arm, RHA), the design strategy is the same as that for gene editing types of gene insertion or replacement.

[0069] For primer design on the side of the left and right homologous arms closest to the editing target site, i.e., the reverse primer for the left homologous arm (LHA) and the forward primer for the right homologous arm (RHA): since the left and right homologous arms are directly connected during assembly, only one primer needs to be added to the "Primer_Overlap" primer during primer design. That is, if a Border primer design strategy is used, one of the reverse primer for the left homologous arm (LHA) or the forward primer for the right homologous arm (RHA) should be selected using "Primer_Overlap + Primer_Border +

[0070] The design pattern is "Efficient_Primer(5'-3')", and the other primer uses the design pattern "Primer_Border + Efficient_Primer(5'-3')". For example,

[0071] Reverse primer design for the left homologous arm: Design an effective primer (Efficient_Primer) within a given sequence region containing a relatively broad primer prediction range (Border_Range). Then, pad the 5' end of the effective primer to the target boundary (edit the 5' end of the target site) to obtain the primer boundary (Primer_Border). Next, select a region from the 5' end of the right homologous arm to be assembled as an overlapping primer (Primer_Overlap), and splice them together to form a primer containing Primer_Overlap + Primer_Border + ...

[0072] Complete primers for Efficient_Primer(5'-3'); Design of forward primers for the right homologous arm: Design effective primers within the given Border_Range, then fill in the 5' end of the effective primers to the 5' end of the right homologous arm to obtain the primer boundary (Primer_Border), assemble the effective primer (Efficient_Primer) and the primer boundary (Primer_Border) to obtain a complete primer fragment containing "Primer_Border+Efficient_Primer(5'-3')".

[0073] (b) Knock-in fragment primer design module

[0074] This module can be used to design primers for PCR amplification of knock-in DNA fragments.

[0075] like Figure 3 As shown in Figure a, using the Border primer design strategy, effective primers (Efficient_Primer) are designed in a given sequence region containing a relatively broad primer prediction range (Border_Range). Then, the 5' end of the effective primer is padded to the target boundary (the reverse primer of the left homologous arm is padded to the 5' end of the editing target, and the forward primer of the right homologous arm is padded to the 3' end of the editing target) to obtain the primer boundary (Primer_Border). The effective primer (Efficient_Primer) and the primer boundary (Primer_Border) are assembled to obtain a complete primer fragment containing "Primer_Border + Efficient_Primer(5'-3')".

[0076] It is important to note that the primers for the knock-in fragment do not contain an overlap region. The "Primer_Overlap" required for the assembly process between the knock-in fragment and its two homologous arms must be designed on the primers of the left and right homologous arms (LHA and RHA) of the knock-in fragment.

[0077] (c) Vector backbone primer design module:

[0078] This module designs primers for obtaining linearized vector backbones for PCR amplification. Using Python scripts and Primer 3.0, it designs effective primers (Efficient_Primer) within a given sequence region containing a relatively broad primer prediction range (Border_Range). The 5' end of the effective primer is then padded to the vector endpoint to obtain the primer boundary (Primer_Border). This, along with "Primer_Overlap" for subsequent DNA fragment assembly, forms the complete primer set: "Primer_Border + Efficient_Primer(5'-3')". Alternatively, the module selects the DNA sequence for primer design using a Python script and the user-defined primer prediction range (Border_Range), and then uses Primer 3.0 to design effective primers.

[0079] (Efficient_Primer), and provide the Tm value, GC content, and length of the primer. Then, use a Python script to splice the effective primer (Efficient_Primer) with the sequence from the 5' end of the effective primer (Efficient_Primer) to the end of the vector to form the complete primer "Primer_Border+Efficient_Primer" (5'-3').

[0080] It is important to note that the primers for the vector backbone do not contain overlap regions. The overlap sequence required for the assembly of the vector backbone and the knock-in fragment must be designed into the primers of the knock-in fragment. Therefore, the "Primer_Overlap" no longer needs to be added to the vector primers. A schematic diagram of the primers and fragments designed in the above process is shown below. Figure 3 As shown in a, 3b.

[0081] (d) Plasmid identification primer design module

[0082] This module can design primers to determine the correctness of constructed gene-editing plasmids. The designed primers can be used for colony PCR, DNA sequencing, etc. Primer 3.0 is used to design identification primers. Based on given parameters, primer design is completed, and information such as primer sequence, length, GC content, Tm value, primer type, template, and PCR product length is output. The primer design method is as follows:

[0083] For target editing types of insertion or replacement, primers are designed for a range of 200–300 bp on both sides of the following junctions: the junction between the vector fragment and the left homologous arm, the junction between the vector fragment and the right homologous arm, the junction between the knock-in fragment and the left homologous arm, and the junction between the knock-in fragment and the right homologous arm. A schematic diagram of the primers and fragments designed in the above process is shown below. Figure 3 As shown in c;

[0084] For target editing types of deletion, primers are designed for a range of 200–300 bp on both sides of the following junctions: the junction between the vector fragment and the left homologous arm, the junction between the vector fragment and the right homologous arm, and the junction between the left and right homologous arms. A schematic diagram of the primers and fragments designed in the above process is shown below. Figure 3 As shown in d.

[0085] (e) Primer design module for gene editing identification

[0086] Primers are used to design primers to determine the success of gene editing in a target host, and can be used for colony PCR, DNA sequencing, etc. Primer identification is achieved using Primer 3.0. Based on given parameters, primer design is completed, and information such as primer sequence, length, GC content, Tm value, primer type, template, and PCR product length is output. The primer design method is as follows:

[0087] When the gene editing type is insertion or replacement of the target site, primer pairs are designed in the following regions: 200 bp upstream of the left homologous arm, 200 bp inside the knock-in fragment near the 5' end, 200 bp inside the knock-in fragment near the 3' end, and 200 bp downstream of the 3' end of the right homologous arm. A schematic diagram of the primers and fragments designed in the above process is shown below. Figure 3 As shown in e;

[0088] When the gene editing type is the deletion of the target gene, primer pairs are designed in the following regions: 200 bp upstream of the left homologous arm, 200 bp inside the right homologous arm sequence near the 5' end, 200 bp inside the left homologous arm sequence near the 3' end, and 200 bp downstream of the right homologous arm. A schematic diagram of the primers and fragments designed in the above process is shown below. Figure 3 As shown in f.

[0089] (2) Atlas Design Module

[0090] The map generation module is used to simulate the complete plasmid after homologous arms, knock-in fragments (if any), and vector backbone are connected, and to draw a complete plasmid map. The plasmid map is annotated with information such as homologous arms, knock-in fragments, and primers.

[0091] Example 2: Method for offline gene editing plasmid and primer design using the system of Example 2

[0092] The system's operation comprises three phases: data input and parameter setting, data processing and analysis, and data output (e.g., ...). Figure 1 As shown in the figure, the system operation process is as follows.

[0093] (I) Data Input and Parameter Setting Stage:

[0094] S001 User-defined input file, setting parameters;

[0095] Custom input files include:

[0096] (a) Target genome file (Genome), which refers to the genome sequence containing the editing target; allows users to upload genome files locally or enter the NCBI Assembly Accession ID;

[0097] (b) Linearized Vector Backbone File (Vector): The main interface allows users to upload linearized vector backbone files, which are sequences that load homologous arms and insert sequences to form complete gene editing plasmids; uploaded linearized vector backbone files are stored on the server to form a linearized vector backbone database; users are allowed to select target vector backbone files from the linearized vector backbone database;

[0098] (c) Target file (Configure) based on the target gene allows users to upload target files. These target files contain information about the target, the type of target editing, and the knock-in fragment. Target information includes the name of the genome containing the target (ChrID), the target name (GeneID), the base position of the target on the genome, and the positive and negative strands. The target editing type includes deletion, insertion, and replacement. The knock-in fragment refers to the DNA sequence introduced into the target site by directly inserting or replacing the existing sequence when the editing type is insertion or replacement. Upload methods include uploading multiple target files at once. Uploaded target files are stored on the server to form a target database; users can select target files from the database. Target files are table-delimited text files and can be uniquely locus_tags. Figure 4 a), or it can be a specific base position ( Figure 4 b) Furthermore, target design can be done in batches, meaning that information about multiple editing targets can be entered into the same file at the same time, and the software will automatically and continuously analyze and process multiple targets.

[0099] Parameter settings include:

[0100] (I) Parameter settings for the design of the same source arm

[0101] The length of the left and right homologous arms (Left_Flank / Right_Flank): This parameter is used to determine the position of the target site on the genome based on the length of the left and right homologous arms.

[0102] Extend_size: The purpose of this parameter is to determine the sequence of the extended primers for the left and right homologous arms based on Extend_size and the information in step (1).

[0103] (II) Parameter settings for primer design

[0104] Users can set the following parameters: minimum effective primer length (Min_Len), maximum effective primer length (Max_Len), optimal effective primer length (Opt_Len), minimum GC content (Min_GC), maximum GC content (Max_GC), optimal GC content (Opt_GC), monovalent ion concentration in the PCR reaction system (Mv_Conc, mM), divalent ion concentration in the PCR reaction system (Dv_Conc, mM), dNTP concentration in the PCR reaction system (dNTP_Conc, mM), and DNA concentration in the PCR reaction system (DNA_Conc). c, nM), Tm value calculation method (Tm_Method), minimum Tm value (Min_Tm), maximum Tm value (Max_Tm), optimal Tm value (Opt_Tm), maximum Tm difference between paired primers (Max_Diff_Tm), maximum allowed number of repeating bases at the 3' end of the primer (Max_PolyX), maximum number of primer pairs returned (Prim_Num); length of the primer overlap region between adjacent fragments (Primer_Overlap); range of predicted effective primers in the Border primer design strategy (Border_Range).

[0105] Select the genome file in step (a) of S001 and the linearized vector backbone file in step (b). Based on the target gene target file in step (c) of S001 and the parameters set in steps (I) and (II), fill in the output compressed file name and click RUN to run and proceed to the next step of data processing and analysis.

[0106] (II) Data Processing and Analysis Stage

[0107] The primer processing module and plasmid map processing module receive the target genome file, linearized vector backbone file, and target file information processed by the user. They perform processes such as homologous arm sequence determination, primer border region determination, effective primer design, primer boundary sequence ligation, primer overlapping sequence ligation, and primer design identification. This yields the homologous arms, insertion sequences, amplification primers and plasmid identification primers required for constructing the gene-editing plasmid, as well as a complete gene-editing plasmid map. Quality assessment is performed by determining whether the primers meet the designed parameters such as length, Tm value, and GC content. Specific steps include:

[0108] The design of upstream and downstream homogeneous arms for S002 was implemented using Python scripts.

[0109] The homologous arm sequences are determined based on information such as the genome, editing target and method (Configure), and user-specified homologous arm lengths (Left_Flank / Right_Flank), and then homologous arm primers are designed. It is important to note that when designing forward primers for LHA and reverse primers for RHA, a broad extension region (Extend_size) is given for primer design. The actual final left and right homologous arms will extend into the region of the extension primers (i.e., including the initially defined sequence and the sequence extending outwards into the extension primers).

[0110] Primers are used to design homologous arms for PCR amplification of the target genome where homologous recombination occurs. Specifically, after determining the specific sequence of the homologous arm based on the editing target site and the user-specified length, primers for amplifying the homologous arm are designed. For primer design on the side of the left and right homologous arms furthest from the editing target site (i.e., the forward primer (LHA) for the left homologous arm and the reverse primer for the right homologous arm (RHA): Efficient primers (Efficient_Primer) are designed based on the user-specified broad extension region (Extend_size) to obtain high-quality primers. The vector backbone end sequence is obtained as Primer_Overlap based on the user-specified length (Overlap_Length); a Python script is used to assemble the primer sequences according to the structure "Primer_Overlap + Efficient_Primer" to obtain the forward primer for the left homologous arm and the reverse primer for the right homologous arm.

[0111] For primer design on the side of the left and right homologous arms closest to the editing target site, the use of the Border primer design strategy is determined by whether the user fills in "Border_Range". If this strategy is adopted, an effective primer (Efficient_Primer) is designed using Primer3 based on the relatively broad primer prediction range (Border_Range) provided by the user. Then, the 5' end of the effective primer is padded to the target site boundary, and this padded sequence is used as the primer boundary (Primer_Border). This boundary is then combined with the overlapping primer (Primer_Overlap) used for subsequent DNA fragment assembly to form the complete primer "Primer_Overlap+Primer_Border+Efficient_Primer" (5'-3'). If Border_Range is not filled in, it means that the Border primer design strategy is not used. In this case, the 5' end of the primer is fixed at the end of the homologous arm, and an effective primer (Efficient_Primer) is designed and generated. Then, an overlapping primer (Primer_Overlap) for constructing the recombinant plasmid is added to the 5' end of the effective primer to form the complete primer "Primer_Overlap+Efficient_Primer" (5'-3').

[0112] For LHA reverse primers and RHA forward primers, the overlap is a user-specified length (Overlap_Length) of a knock-in fragment sequence (target editing type is insert or replace) or the end of another homologous arm (editing type is delete; in this case, there is no knock-in fragment, so the left and right homologous arms are directly connected).

[0113] The design of S003 knock-in primers was carried out using Python scripts and Primer3. The design method included:

[0114] Using a primer design strategy that incorporates a border, the complete primer Primer_Border+Efficient_Primer(5'-3') was constructed. A Python script was used to identify the sequence region of the Border_Range, design an effective primer (Efficient_Primer) using Primer3, confirm the Primer_Border, and generate the complete primer Primer_Border+Efficient_Primer(5'-3'). It is important to note that the primer for the knock-in fragment does not have an overlap region; the overlap required for assembly with the homologous arms on both sides is added to the primers of the left and right homologous arms.

[0115] The design of primers for the S004 vector was carried out using Python scripts and PRIMER3. The design methods included:

[0116] Primers for designing PCR amplification to obtain linearized vector backbones were designed using Python scripts and Primer 3.0, incorporating a Border primer design strategy. The Python script selected the DNA sequence for the primers based on the user-defined Border design region (Border_Range) parameter. Primer 3.0 was then used to design effective primers, specifying their Tm value, GC content, and length. Finally, the Python script used the sequence between the 5' end of the effective primer and the end of the vector as the primer boundary, forming the complete primer "Primer_Border+Efficient_Primer" (5'-3').

[0117] The design of identification primers for plasmid S005 was carried out using Primer 3.0.

[0118] For cases where the target editing type is deletion, primers are designed for a range of 200-300 bp on both sides of the following connection points: the connection point between the vector fragment and the left homologous arm, the connection point between the vector fragment and the right homologous arm, and the connection point between the left and right homologous arms.

[0119] For target editing types of insertion or replacement, primers are designed for a range of 200-300 bp on both sides of the following connection points: the connection point between the vector fragment and the left homologous arm, the connection point between the vector fragment and the right homologous arm, the connection point between the knock-in fragment and the left homologous arm, and the connection point between the knock-in fragment and the right homologous arm.

[0120] Output information such as primer sequence, length, GC content, Tm value, primer type, template, and PCR product length.

[0121] The design of primers for S006 gene editing identification was performed using Primer 3.0, which outputs information such as primer sequence, length, GC content, Tm value, primer type, template, and PCR product length. The primer design method is as follows:

[0122] For cases where the target editing type is deletion, primers are designed for a range of 200-300 bp on both sides of the following connection points: the connection point between the vector fragment and the left homologous arm, the connection point between the vector fragment and the right homologous arm, and the connection point between the left and right homologous arms.

[0123] For target editing types of insertion or replacement, primers are designed for a range of 200–300 bp on both sides of the following connection points: the connection point between the vector fragment and the left homologous arm, the connection point between the vector fragment and the right homologous arm, the connection point between the knock-in fragment and the left homologous arm, and the connection point between the knock-in fragment and the right homologous arm.

[0124] The S007 plasmid map was assembled using a Python script to create a complete plasmid map.

[0125] For plasmid maps of gene editing types such as insertion or replacement, the map includes: simulated left and right homologous arms, complete plasmids after the knock-in fragment is connected to the vector backbone, and their annotation information;

[0126] For plasmid maps of gene editing type deletion, the complete plasmid and its annotation information are included: simulated left and right homologous arms connected to the vector backbone;

[0127] The annotation information includes homologous arms, knock-in fragments, primer names, locations, and other information.

[0128] (3) Data output stage

[0129] During the data output stage, the processing results such as homologous arm sequences, primers, and plasmid maps can be directly output to a specified local path. The output path can be set through "Output".

[0130] The output mainly includes:

[0131] (1) KO_HRA_Fragment is the sequence information of the homologous arm.

[0132] (2) KO_Primer is the final primer information result, which is the result summarized in KO_Final_PrimerSeq.xlsx.

[0133] (3) Primer3_out is the primer prediction result;

[0134] (4) input is the original input data.

[0135] (5) pKO_plasmid is the vector map.

[0136] Example 3: Method for designing multi-target gene editing plasmids and primers using the system of Example 1.

[0137] This example demonstrates the effectiveness of the software by designing plasmids to knock out the target gene of Streptomyces venezuelae ISP5230. Plasmids were designed for target sites SVEN_0673, SVEN_0674, SVEN_0675, ..., SVEN_0692, with the editing type being Replace, meaning the CDS region of the target gene was replaced with a resistance gene (see...). Figure 4 a).

[0138] During the parameter setting phase, upload the Streptomyces venezuelae ISP5230 genome file (SVEN_ISP5230.gb) to "Genome", the vector file (pKC1132_batchA_linear.gb) to "Vector", and the target information file (SVEN_0673_SVEN_0692.txt) to "Configure". Use the default parameters for parameter settings (e.g., ...). Figure 2 As shown), select the output path `outdir` and click `RUN` to run the software. The software runs successfully and outputs the results (see...). Figure 4 b) The output includes homologous arm sequence information (stored in the KO_HRA_Fragment folder) and plasmid map (stored in the pKO_plasmid folder, see...). Figure 4 c) Primer information for plasmid construction (stored in the KO_Primer folder, with the results summarized in KO_Final_PrimerSeq.xlsx, see...) Figure 4 d) Primer information for plasmid identification (stored in the Check_Primer folder), effective primer prediction results (stored in the Primer3_out folder), and raw input data (stored in the input folder). Designing 20 targets can be completed in about 4 minutes.

[0139] Example 4 uses the system from Example 1 to verify the effectiveness of the primer design strategy involving the addition of borders.

[0140] This example uses a plasmid designed to knock out the target gene of Streptomyces venezuelae ISP5230 to test the software's effectiveness. Plasmids were designed for targets SVEN_0276 and SVEN_0279, with the edit type being Replace, meaning the CDS region of the target gene was replaced with an anti-resistance gene. During parameter settings, the Streptomyces venezuelae ISP5230 genome file (SVEN_ISP5230.gb) was uploaded to "Genome," the vector file (pKC1132_batchA_linear.gb) was uploaded to "Vector," and the target information files (SVEN_without_border.txt and SVEN_with_border.txt) were uploaded to "Configure." Border_Range was set to 0 (indicating no Border strategy) and 40 respectively. The remaining parameters used default settings (e.g., ...). Figure 2 (As shown), select the output path `outdir` and click `RUN` to run the software. The software runs successfully and outputs results. From the two output folders, select the primers for amplifying the left homologous arm:

[0141] Primers designed without using the Border design strategy:

[0142] Forward primers for the left homologous arms of plasmids targeting SVEN_0276 and SVEN_0279:

[0143] SIAT-717: CGTATTCAGAGTATTTGTCGGCCCTGGCGGCGGGGCGC (SEQ ID NO. 1); SIAT-718: CGTATTCAGAGTATTTGTCGCCCCTCCAGGCCCCGCGC (SEQ ID NO. 2);

[0144] Reverse primers for the left homologous arms of plasmids targeting SVEN_0276 and SVEN_0279:

[0145] SIAT-722: AGAAGTAGTATGAGATCGACCCGGCGTCGACGACGCAC (SEQ ID NO. 3);

[0146] SIAT-723: GCATGATCTTCTCTTCGAATTCACGGCTTTTACCGGACATG (SEQ ID NO. 4).

[0147] Primers designed using the Border design strategy:

[0148] Forward primers for the left homologous arms of plasmids targeting SVEN_0276 and SVEN_0279:

[0149] SIAT-727: CGTATTCAGAGTATTTGTCGGAACGTGACCTGACCCCGGCTCCTC (SEQ ID NO. 5);

[0150] SIAT-728: CGTATTCAGAGTATTTGTCGACCCCGTCCCACACCCCTTCGCGAAG (SEQ ID NO. 6);

[0151] Reverse primers for the left homologous arms of plasmids targeting SVEN_0276 and SVEN_0279:

[0152] SIAT-732: AGAAGTAGTATGAGATCGACCCGGCGTCGACGACGCACCAGGC (SEQ ID NO. 7);

[0153] SIAT-733: GCATGATCTTCTCTTCGAATTCACGGCTTTACCGGACATGCGGCTCTCCT (SEQ IDNO.8);

[0154] The above primers were submitted to the company for synthesis (Qingke Biotechnology). After synthesis, they were used for PCR amplification to verify the effect. Primer pairs SIAT-717 and SIAT-722 amplified fragment F1, with a length of 2140bp; primer pairs SIAT-718 and SIAT-723 amplified fragment F2, with a length of 2140bp; primer pairs SIAT-727 and SIAT-732 amplified fragment F3, with a length of 2098bp; and primer pairs SIAT-728 and SIAT-733 amplified fragment F4, with a length of 2114bp.

[0155] The PCR reaction system is as follows:

[0156] ddH2O: 20.5 μL;

[0157] Upstream primer / downstream primer (10 μM): 2 μL;

[0158] Template (S. venezuelae ISP5230 genome, 102 ng / μL): 0.5 μL;

[0159] 2x Phanta mix: 25μL;

[0160] Total volume: 50 μL.

[0161] The reaction conditions were: 95℃, 3 min; [95℃, 15 s; 60℃, 15 s; 72℃, 1 min] × 32 cycles; 72℃, 5 min; 4℃, maintenance.

[0162] After the reaction was completed, 5 μL of the reaction solution from each sample was taken for agarose gel electrophoresis analysis, such as... Figure 5 As shown in the results, the band brightness and specificity of fragments F3 and F4 amplified by primers designed using the Border design strategy are significantly stronger than those of F1 and F2 amplified by primers not designed using the Border design strategy. This result fully demonstrates that the Border design strategy proposed in this invention can significantly improve primer quality and is beneficial for obtaining high-concentration, specific PCR products.

[0163] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.

Claims

1. A traditional gene-editing plasmid and primer design method, characterized in that, The method is applied to an electronic device that uses the open-source code of PRIMER3 and Python scripts to execute the traditional gene-editing plasmid and primer design method; including: Obtain the information required for gene editing operations, including genomic information, linearized vector backbone information, and gene editing target information; the gene editing type is gene deletion, insertion, or replacement; Primer design and primer design results are performed for the fragments required for gene editing operations; when deleting a gene, the required fragments include upstream and downstream homologous arms and plasmid backbone; when inserting or replacing a gene, gene fragments are also required. Output processing results; the processing results include a map of the gene editing plasmid and primers required for gene editing operations; the primers required for gene editing operations include one or more of the following: primers for amplifying upstream and downstream homologous arms of the target site, primers for amplifying the knock-in gene, primers for amplifying the vector, and identification primers; the map of the gene editing plasmid simulates the complete plasmid after ligating the homologous arms, the knock-in fragment, and the vector backbone fragment, or simulates the complete plasmid after ligating the homologous arms and the vector backbone fragment; The primer designs include upstream and downstream homologous arm primer designs, knock-in fragment primer designs, vector backbone primer designs, plasmid identification primer designs, and gene editing identification primer designs. The upstream and downstream homologous arm primer design, knock-in fragment primer design, and vector backbone primer design may or may not employ the Border primer design strategy; The Border primer design strategy is as follows: Forward primers: Design primers for amplifying the target polynucleotide within a relatively wide range on one side of the 5' end of the target polynucleotide on the template DNA to obtain effective primers. Then extend the 5' end of the effective primers to the 5' end of the target polynucleotide to obtain complete forward primers. Reverse primer: Design a primer for amplifying the target polynucleotide in a wide range on one side of the 3' end of the target polynucleotide on the template DNA to obtain an effective primer. Then extend the 5' end of the effective primer to the 3' end of the target polynucleotide to obtain a complete reverse primer. The wider range refers to the region that is >0 bp from the 5' or 3' end of the polynucleotide.

2. The method according to claim 1, characterized in that, The evaluation assesses primer length, GC content, Tm value, and PCR product length, and is performed using Primer3.

3. The method according to claim 1 or 2, characterized in that, The gene editing includes gene editing targeting one or more sites.

4. A gene-editing plasmid and primer design device, used to implement the traditional gene-editing plasmid and primer design method as described in any one of claims 1 to 3, characterized in that, Applied to electronic devices, the device includes: The acquisition unit is used to acquire the information required for gene editing operations, including genome information, linearized vector backbone information, and gene editing target information; The primer processing unit designs primers for the fragments required for gene editing operations and evaluates the primer design results; the required fragments include upstream and downstream homologous arms and plasmid backbone; The output and display unit is used to output the processing results and display the plasmid map; the processing results include the map of the gene editing plasmid and the primers required for the gene editing operation; the plasmid map simulates the complete plasmid after the homologous arm, knock-in fragment and vector backbone fragment are connected, or simulates the complete plasmid after the homologous arm and vector backbone fragment are connected.

5. The apparatus according to claim 4, characterized in that, The device can be operated online or offline.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 3.