Online tool for automated design of genome editing based on crispr / cas technology

By developing an online tool for automated design of genome editing using CRISPR/Cas technology, we have solved the problems of low integration and insufficient off-target assessment of existing tools, realizing automated and high-throughput design of genome editing, and improving the success rate of editing and support capabilities for multiple scenarios.

CN118918950BActive Publication Date: 2025-12-19TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410851706.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-12-19
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Existing genome editing design tools have low integration, resulting in low automation, low efficiency, high error rate, and a lack of support for multiple experimental scenarios and off-target assessment capabilities for edited sequences.

Method used

Develop an online tool for automated genome editing design based on CRISPR/Cas technology, which includes an sgRNA design unit, a genome editing design unit, a task management and visualization system, and an error handling unit. By integrating multiple genome editing task workflows, it achieves automated design and high-throughput editing, and performs homologous arm and primer off-target risk assessment and optimization.

Benefits of technology

It improves the automation of genome editing, supports multiple experimental scenarios, reduces the error rate, and enhances the success rate of editing through off-target risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918950B_ABST
    Figure CN118918950B_ABST
Patent Text Reader

Abstract

The application relates to an online tool for CRISPR / Cas technology-based genome editing automation design, which comprises four functional modules, an sgRNA design unit for designing sgRNA, a genome editing design unit for realizing genome editing design, a task management and visualization system for monitoring and managing the progress and state of the gene editing design task, and an error processing unit for automatically identifying and solving errors occurring in the genome editing design process. The online tool is used for CRISPR / Cas technology-based genome editing automation design, the designed sgRNA, primers and homologous arms can realize full-process support for genome editing experiments, various genome editing tasks and processes are integrated to realize support for various genome editing experiment scenes, and the homologous arms and primers are integrated for off-target risk assessment and optimization, which helps to improve the success rate of genome editing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of genome editing, in particular to an online tool for automated design of genome editing based on CRISPR / Cas technology. BACKGROUND

[0002] The existing genome editing design tools have low integration of single-function editing sequence design tools, resulting in low automation degree when dealing with multi-step complex genome editing tasks, low efficiency, high error rate and limited design throughput. Genome editing technology has diversified technical variants and experimental scenarios, but current genome editing design tools lack support functions for multiple experimental scenarios. In addition, the genome contains different types and proportions of repetitive sequences, which may cause off-target of editing sequences and further cause incorrect positioning of genome editing, but the existing genome editing design tools have defects in off-target evaluation of editing sequences.

[0003] The online tool for automated design of genome editing based on CRISPR / Cas technology aims to achieve automation and high-throughput design of genome editing by building an automated editing sequence design process, support multiple genome editing experimental scenarios by integrating design processes of multiple genome editing tasks, and improve genome editing success rate by integrating homologous arm and primer off-target risk evaluation and optimization. SUMMARY

[0004] The present application aims to provide an online tool for automated design of genome editing based on CRISPR / Cas technology to solve the problems of low automation degree of genome editing design, lack of support for multiple experimental scenarios and defects in off-target evaluation of editing sequences in the background art. The main functional modules of the tool include:

[0005] An sgRNA design unit for designing specific sgRNA, facing CRISPR technology, using the high specificity binding ability of sgRNA to achieve precise editing of specific sites in the genome;

[0006] A genome editing design unit for implementing genome editing, facing CRISPR / Cas-mediated homologous recombination technology, achieving precise knockout, insertion and replacement of sequences;

[0007] A task management and visualization system for monitoring and managing the progress and status of genome editing design tasks, using cloud computing and real-time data processing technology to achieve real-time monitoring of tasks, dynamic visualization of results and user interaction;

[0008] Error processing unit, the error processing unit is used for automatic identification and solution of error appearing in genome editing design process, realizes quick positioning and detailed problem explanation to error.

[0009] As a further improvement of the technical solution, the sgRNA design unit comprises a data preprocessing module and an sgRNA generator, and the specific steps involved in sgRNA design are as follows:

[0010] S1.1, the user first needs to specify the reference genome by selecting and uploading the genome sequence, including sgRNA design input file and target modification file containing genome editing design task information, and the data preprocessing module performs format processing on these files.

[0011] S1.2, the user sets the parameters of sgRNA design, including CRISPR / Cas system type, guide sequence (N20) length, upper limit of mismatch base number of potential off-target sites and genome sequence in off-target analysis, and algorithm of sgRNA cutting efficiency, namely Efficiency score, and sgRNA generator generates and optimizes sgRNA according to the related parameters set by the user;

[0012] S1.3, after the user completes the selection / upload of input file and parameter setting, clicks the task submission button;

[0013] S1.4, after the design process is completed, it will automatically jump to the "JOB MANAGER" page, and the "JOB MANAGER" page shows the result list of sgRNA design, including sgRNA sequence, position on target sequence, GC content, number of potential off-target sites containing different mismatch base numbers, Efficiency score and sgRNA ranking;

[0014] As a further improvement of the technical solution, the data preprocessing module is used for receiving the user uploaded genome sequence file, and performing format processing on these files, and the specific steps involved are as follows:

[0015] S1.1.1, receiving the user uploaded genome sequence file, and checking whether the format of the file is correct;

[0016] S1.1.2, using sequence verification technology to check whether there is non-standard base character in the uploaded genome sequence, removing illegal characters and blank lines, converting all bases to uppercase, and unifying the representation of sequences;

[0017] S1.1.3, using genome coordinate system to standardize input file;

[0018] S1.1.4, confirm the consistency of the input sequence with the reference genome of the target species using sequence alignment technology;

[0019] The sgRNA generator is responsible for the actual sgRNA design, including the generation and optimization of sgRNA sequences. The specific steps involved in the generation and optimization of sgRNA by the sgRNA generator are as follows:

[0020] S1.2.1, input the reference genome, target sequence and sgRNA design parameters, use BLAST 2.9.0 alignment tool to align the target sequence with the reference genome sequence, determine the exact position of the target sequence in the reference genome, and get the coordinates of the target sequence on the reference genome;

[0021] S1.2.2, based on the confirmed target sequence position, expand the length of the guide sequence and PAM sequence to both ends, and obtain the sgRNA search region;

[0022] S1.2.3, in the sgRNA search region, identify all candidate sgRNAs adjacent to the PAM sequence;

[0023] S1.2.4, use Bowtie1.3.1 to build a genome index, and align the candidate sgRNAs with the entire reference genome to search for potential off-target sites with high similarity but different positions from the candidate sgRNAs, set an upper limit for the number of base mismatches, and evaluate the potential off-target risk of each candidate sgRNA;

[0024] S1.2.5, use CHOPCHOP v3 software to score the efficiency of each candidate sgRNA according to the preset or user-selected scoring algorithm, including the cutting efficiency and off-target risk of the sgRNA;

[0025] S1.2.6, generate a design result file containing detailed information of all candidate sgRNAs, including sequence, position, efficiency score, off-target risk and ranking. The result file is used for front-end display, and users can view and compare the performance of different sgRNAs online and download the result file.

[0026] As a further improvement of the technical solution, the genome editing design unit designs specific sgRNA and PCR primers for genome editing experiments, uses homologous arms, sgRNA fragments and plasmid template PCR primers to obtain corresponding fragments by PCR and assemble them into a recombinant plasmid, and expresses CRISPR / Cas system after the recombinant plasmid is transformed into cells, and uses the homologous recombination mechanism of cells to accurately repair the DNA double-strand break mediated by CRISPR / Cas system, so as to realize the site-directed editing of target sequence on the genome. The genome editing design unit includes a data preprocessing module, a plasmid map editing module, a PCR primer design module, a homologous arm off-target detection module, a sequencing verification primer design module and a genome sequencing verification primer off-target optimization module.

[0027] The data preprocessing module realizes the formatting, correction and position information extraction of the genomic data and plasmid information.

[0028] The plasmid map editing module realizes user-defined modification and adjustment of plasmid structure, and adapts to different editing strategies.

[0029] The PCR primer design module is used for designing PCR primers in the process of genome editing.

[0030] The homologous arm off-target detection module is used for analyzing the possible off-target positions of the homologous arms and evaluating the off-target risk.

[0031] The sequencing verification primer design module is used for designing sequencing primers for verifying the editing results.

[0032] The genome sequencing verification primer off-target optimization module is used for optimizing the genome sequencing verification primers, reducing the off-target effect and improving the sequencing accuracy.

[0033] As a further improvement of the technical solution, the genome editing design unit involves the following specific steps in the process of genome editing design:

[0034] S2.1, the data preprocessing module formats the user uploaded genome and plasmid data, and parses and confirms the user input parameters, including sgRNA design parameters, editing types, wherein the editing types include knockout, insertion and replacement;

[0035] S2.2, the PCR primer design module designs the PCR primers of homologous arms and sgRNA based on the user input parameters, calculates the optimal length, temperature and GC content of the primers, generates a primer list including sequence, expected product size and expected amplification efficiency;

[0036] S2.3, the sequencing verification primer design module and the genome sequencing verification primer off-target optimization module analyze potential off-target sites, select the best PCR primer position and sequence, optimize the PCR primer conditions, and obtain the sequencing primer design results, including primer position, sequence and predicted off-target analysis results.

[0037] As a further improvement of the technical solution, the PCR primer design module designs PCR primers based on DNA fragment assembly parameters and PCR primer design parameter settings (primer length, Tm value, GC content). It includes functions such as determining the design type, generating the primer design template, primer search and evaluation, primer modification, and generating the recombinant plasmid map;

[0038] The PCR primer design module designs homology arms, sgRNA fragments and plasmid template PCR primers based on parameter input, and the specific steps are as follows:

[0039] S2.2.1, first identify the editing type and target sequence of the design task in the input file, where the editing type includes knockout, insertion and replacement, and determine the design type of each task according to the user-specified plasmid system and additional primer design requirements;

[0040] S2.2.2, according to the requirements of the design type, select the PCR primer design template of the homology arm, sgRNA fragment and plasmid backbone on the reference genome and the starting plasmid, determine the primer search region on the template according to the input DNA fragment assembly parameters, and generate the parameter input of Primer3 according to the primer design parameter input;

[0041] S2.2.3, according to the parameter input of primer design, generate all candidate primers from the primer search region, and generate primer pairs based on them, select the best primer pair according to the scoring rules and algorithms of Primer3, and finally perform off-target evaluation of the homology arm and generate the related result file;

[0042] S2.2.4, according to the design type of the task, determine the overlapping sequence between adjacent recombinant plasmid fragments, add the overlapping sequence and the spacer sequence, the IIS type restriction enzyme recognition site and the protection sequence to the 3' end of the corresponding primer, and use the primer with the added adapter sequence and its key features, including Tm, GC content, to generate the primer design result file;

[0043] S2.2.5, based on the homology arm off-target evaluation strategy, evaluate the off-target risk of the upstream and downstream homology arms, and generate the homology arm off-target risk evaluation report;

[0044] S2.2.6, based on the primer design results and their PCR product sequences, generate a recombinant plasmid map file with recombinant plasmid fragment labels and corresponding PCR primers.

[0045] As a further improvement of the technical solution, the sequencing verification primer design module is used to design sequencing verification primers for recombined plasmids and genome editing, and the sequencing verification primer design module includes determining sequencing targets, primer searching and evaluation, and design result visualization; the specific steps are as follows:

[0046] S2.3.1, determining the target region to be sequenced of the recombined plasmid and the edited genome and the search region of the sequencing verification primer according to the PCR primer design result, and generating the sequencing verification primer design parameter input of Primer3 according to the primer design parameter input (primer length, GC content, Tm);

[0047] S2.3.2, generating all candidate primers from the primer search region according to the primer design parameter input, and generating a primer pair set based thereon, screening the best primer pair according to the scoring rules and algorithms of Primer3, and detecting and optimizing the two primers for PCR of the target sequencing region based on off-target analysis;

[0048] S2.3.3, based on the sequencing primer design result and the recombined vector and the edited genome sequence, generating a recombined plasmid and an edited genome map file containing the tag (homologous arm, sgRNA fragment, inserted fragment, etc.) of the target sequencing fragment and the corresponding sequencing primer tag, for visualization based on the OVE package.

[0049] As a further improvement of the technical solution, the task management and visualization system includes a task submission module, a result display module and a user interaction module;

[0050] The task submission module is used by the user to upload files and input parameters required for editing tasks, including genome sequence, genome editing design requirements, and sgRNA design, PCR primer design and sequencing verification primer design parameters;

[0051] The result display module is used to display the results of gene editing tasks, including sgRNA, PCR primer, sequencing verification primer design results, homologous arm off-target risk evaluation results, and edited recombined plasmid and genome map files;

[0052] The user interaction module is used for the user to navigate to different task management interfaces, modify parameter settings, resubmit modified tasks, and provide user feedback functions.

[0053] As a further improvement of the technical solution, the error processing unit includes an error diagnosis module and a solution module;

[0054] The error diagnosis module is responsible for monitoring the entire genome editing design process, and detects various logical, input and execution errors that may occur in real time, including input file errors, parameter setting errors and primer design failures.

[0055] The solution module generates corresponding solutions and correction suggestions according to the diagnosed error types and specific circumstances.

[0056] As a further improvement of the technical solution, the system architecture of the online tool includes a front-end presentation layer, a logical calculation layer and a data storage layer.

[0057] The front-end presentation layer mainly involves task management and visualization systems, including components that directly interact with users, uses AWS S3 and Cloud Front to provide fast access and delivery of static content, integrates OVE package tools to provide interactive DNA segment visualization, helps users add plasmid tags, specify primers and check design results.

[0058] The logical calculation layer mainly involves sgRNA design units, genome editing design units and error handling units, including managing HTTP requests and providing core computing service functions, and is built using AWS Lambda, AWS API Gateway and AWS Step Functions services.

[0059] The data storage layer mainly involves the persistent storage of the website, including task records (task input, parameter setting, intermediate calculation results) and design results (sgRNA design results, genome editing design results, failed task analysis, homologous arm off-target risk analysis, primer synthesis order), and is built based on AWS DynamoDB and AWS S3 services to manage the persistent storage of the website.

[0060] Compared with the prior art, the present application has the following advantages:

[0061] 1. The genome editing technology based on CRISPR / Cas-mediated homologous recombination realizes automation and high-throughput design of genome editing by building an automated editing sequence design process.

[0062] 2. The design process of multiple genome editing tasks is integrated to support multiple genome editing experimental scenarios.

[0063] 3. The integration of homologous arm and primer off-target risk assessment and optimization helps to improve the success rate of genome editing. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 is the overall architecture of the present application.

[0065] Figure 2 This diagram illustrates the primer design process for sequencing validation. The process comprises three stages: sequencing target identification (①, ②), primer search and evaluation (③), and visualization of design results (④, ⑤). UHA is the upstream homologous arm; DHA is the downstream homologous arm; sgRNA... P sgRNA promoter; sgRNA T sgRNA terminator; GS N , new guide sequence; P, sequencing primer; Test-primer-G, genome sequencing primer; Test-primer-P, recombinant plasmid sequencing primer.

[0066] Figure 3 This is a schematic diagram illustrating the primer optimization strategy for genome editing sequencing validation based on off-target analysis. T represents the target sequence, the region to be sequenced; S represents the spacer sequence, the additional sequences sequenced at both ends of the region to be sequenced to ensure sequencing accuracy; L represents the spacer sequence. U PCR primer search region upstream of the target sequence; L D , downstream PCR primer search region of the target sequence; P, sequencing primer.

[0067] In the diagram: 1. sgRNA design unit; 2. Genome editing design unit; 3. Task management and visualization system; 4. Error handling unit. Detailed Implementation

[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0069] Example

[0070] Please see Figure 1 As shown, an online tool for automated design of genome editing based on CRISPR / Cas technology is provided, including:

[0071] sgRNA design unit 1 is used to design specific sgRNAs. Utilizing CRISPR technology and the highly specific binding ability of sgRNAs, precise editing of specific sites in the genome is achieved. sgRNA design unit 1 includes a data preprocessing module and an sgRNA generator. The specific steps involved in sgRNA design are as follows:

[0072] S1.1, the user first needs to specify the reference genome by selecting or uploading a genomic sequence, including sgRNA design input file and target modification file containing genome editing design task information. If the user's needs are limited to designing sgRNA, the sgRNA design input file needs to be uploaded, and if the needs involve more complex genome editing, the target modification file containing genome editing design task information needs to be uploaded.

[0073] S1.2, the user sets the parameters of sgRNA design, including the type of CRISPR / Cas system, the length of guide sequence (N20), the upper limit of the number of mismatched bases of potential off-target sites and genomic sequence in off-target analysis, and the algorithm of sgRNA cutting efficiency, i.e. Efficiency score, and the sgRNA generator generates and optimizes sgRNA according to the relevant parameters set by the user;

[0074] S1.3, after the user completes the selection / upload of the input file and the parameter setting, click the task submission button;

[0075] S1.4, after the design process is completed, it will automatically jump to the "JOB MANAGER" page, which shows the result list of sgRNA design, including sgRNA sequence, position on target sequence, GC content, number of potential off-target sites containing different number of mismatched bases, Efficiency score and sgRNA ranking. Among them, the sgRNA ranking is usually based on the default algorithm of CHOPCHOP;

[0076] The data preprocessing module is used to receive the user uploaded genomic sequence file, and format the file to provide necessary basic information for sgRNA design. The specific steps involved are as follows:

[0077] S1.1.1, receive the user uploaded genomic sequence file, and check whether the format of the file is correct;

[0078] S1.1.2, check whether the uploaded genomic sequence contains non-standard base characters, remove illegal characters and blank lines, convert all bases to uppercase, and unify the representation of the sequence;

[0079] S1.1.3, use the genomic coordinate system to standardize the input file;

[0080] S1.1.4, use sequence alignment technology to confirm the consistency of the input sequence with the reference genome of the target species;

[0081] The sgRNA generator is responsible for the actual sgRNA design, including the generation and optimization of sgRNA sequences. The specific steps involved in the generation and optimization of sgRNA by the sgRNA generator are as follows:

[0082] S1.2.1, input the reference genome, target sequence and sgRNA design parameters, use the BLAST 2.9.0 alignment tool to align the target sequence with the reference genome sequence, determine the exact position of the target sequence in the reference genome, and obtain the coordinates of the target sequence on the reference genome;

[0083] S1.2.2, based on the confirmed target sequence position, expand the length of the guide sequence and PAM sequence to both ends, and obtain the sgRNA search region;

[0084] S1.2.3, in the sgRNA search region, identify all candidate sgRNAs adjacent to the PAM sequence, and ensure that the candidate sgRNAs meet the specific requirements of the CRISPR / Cas system for the PAM sequence;

[0085] S1.2.4, use Bowtie1.3.1 to build a genome index, and align the candidate sgRNAs with the entire reference genome to search for potential off-target sites with high similarity but different positions from the candidate sgRNAs, set an upper limit for the number of base mismatches, and evaluate the potential off-target risk of each candidate sgRNA;

[0086] S1.2.5, use CHOPCHOP v3 software to score the efficiency of each candidate sgRNA according to the preset or user-selected scoring algorithm, including the cutting efficiency and off-target risk of the sgRNA;

[0087] S1.2.6, generate a design result file containing detailed information of all candidate sgRNAs, including sequence, position, efficiency score, off-target risk and ranking. The result file is used for front-end display, and users can view and compare the performance of different sgRNAs online and download the result file.

[0088] A genome editing design unit 2 is designed for CRISPR / Cas-mediated homologous recombination technology to achieve genome editing design. This unit designs specific sgRNA and PCR primers for genome editing experiments, uses homologous arms, sgRNA fragments, and plasmid template PCR primers to obtain corresponding fragments by PCR and assemble them into a recombination plasmid, and expresses the CRISPR / Cas system after the recombination plasmid is transformed into cells. The homologous recombination mechanism of the cell is used to accurately repair the DNA double-strand break mediated by the CRISPR / Cas system, so as to achieve the site-directed editing of the target sequence on the genome. The genome editing design unit 2 includes a data preprocessing module, a plasmid map editing module, a PCR primer design module, a homologous arm off-target detection module, a sequencing verification primer design module, and a genome sequencing verification primer off-target optimization module; wherein the data preprocessing module realizes the formatting, correction, and position information extraction of the genomic data and plasmid information;

[0089] The plasmid map editing module realizes user-defined modification and adjustment of the plasmid structure, and adapts to different editing strategies;

[0090] The PCR primer design module is used to design PCR primers in the genome editing process;

[0091] The homologous arm off-target detection module is used to analyze the possible off-target positions of the homologous arms and evaluate the off-target risk;

[0092] The sequencing verification primer design module is used to design sequencing primers for verifying the editing results;

[0093] The genome sequencing verification primer off-target optimization module is used to optimize the genome sequencing verification primers, reduce the off-target effect, and improve the sequencing accuracy.

[0094] The genome editing design unit 2 involves the following specific steps in the process of genome editing design:

[0095] S2.1, the data preprocessing module formats the user-uploaded genomic and plasmid data, and parses and confirms the user-input parameters, including sgRNA design parameters, editing types, wherein the editing types include knockout, insertion, and replacement;

[0096] S2.2, the PCR primer design module designs the PCR primers of the homologous arms and sgRNA based on the user-input parameters, calculates the optimal length, temperature, and GC content of the primers, generates a primer list including the sequence, expected product size, and predicted amplification efficiency;

[0097] S2.3, the sequencing verification primer design module and the genome sequencing verification primer off-target optimization module analyze potential off-target sites, select the best PCR primer position and sequence, optimize the PCR primer conditions, and obtain the sequencing primer design results, including primer position, sequence and predicted off-target analysis results.

[0098] The PCR primer design module includes determining the design type, generating the primer design template, primer search and evaluation, primer modification and generating the recombinant plasmid map, designing PCR primers based on DNA fragment assembly parameters and PCR primer design parameter settings (primer length, Tm value, GC content).

[0099] The PCR primer design module designs PCR primers for homologous arms and sgRNAs based on user input parameters, and the specific process steps are as follows:

[0100] S2.2.1, first identify the editing type and target sequence of the design task in the input file, wherein the editing type includes knockout, insertion and replacement, and determine the design type of each task according to the user-specified plasmid system (divided into single-plasmid system and double-plasmid system according to whether sgRNA and homologous arm are on the same plasmid) and additional primer design requirements (no additional primer design, user-specified additional primer sequence, user-specified additional primer design region);

[0101] S2.2.2, according to the requirements of the design type, select the PCR primer design template of the homologous arm, sgRNA fragment and plasmid backbone on the reference genome and the starting plasmid, determine the primer search region on the template according to the input DNA fragment assembly parameters, and generate the parameter input of Primer3 according to the primer design parameter input (such as primer length, GC content, Tm, etc.);

[0102] S2.2.3, according to the parameter input of primer design, generate all candidate primers from the primer search region, and generate primer pairs based thereon, select the best primer pair according to the scoring rules and algorithms of Primer3, and finally perform off-target evaluation of the homologous arm and generate the related result file;

[0103] S2.2.4, determine the overlapping sequence between adjacent recombinant plasmid fragments according to the design type of the task, add the overlapping sequence and the spacer sequence (default "A"), the recognition site of the IIS type restriction enzyme and the protection sequence (default "CCA") to the 3' end of the corresponding primer, and use the primer with the added linker sequence and its key features, including Tm, GC content, to generate the result file of primer design;

[0104] S2.2.5, based on the homologous arm off-target evaluation strategy, evaluate the off-target risk of the upstream and downstream homologous arms, and generate the homologous arm off-target risk evaluation report;

[0105] S2.2.6, based on the primer design results and their PCR product sequences, generate a recombinant plasmid map file with the recombinant plasmid fragment tags (homologous arms, sgRNA promoters and terminators, guide sequences, etc.) and the corresponding PCR primers.

[0106] The sequencing verification primer design module is used for primers for sequencing verification of recombinant plasmids and genome editing, and the sequencing verification primer design module includes determining sequencing targets, primer searching and evaluation, and design result visualization;

[0107] The specific steps are as follows:

[0108] S2.3.1, determine the sequencing target regions of the recombinant plasmid and the edited genome and the search regions of the sequencing verification primers according to the PCR primer design results, and generate the sequencing verification primer design parameter input of Primer3 according to the primer design parameter input (primer length, GC content, Tm);

[0109] The search rules are as follows: Since the length that can be accurately determined by Sanger sequencing is limited (generally 600-700 bp), in order to ensure the sequencing accuracy of each part of the measured fragment, a sequencing primer is designed every 600 bp. For a single recombinant plasmid system, if the distance between the homologous arm and the sgRNA fragment is greater than 600 bp, then the sequencing verification primers are designed respectively, otherwise they are considered as the entire sequence to design the sequencing primer. For a double recombinant plasmid system, the homologous arm and the sgRNA fragment are designed with sequencing primers. The design rules of the recombinant plasmid sequencing primers are as follows: Let the length of the sequencing fragment S be L, if L≤600, then a sequencing primer is designed in the upstream (80-120 bp, the same below) region of S; if 600≤L≤1200, a sequencing primer is designed in the upstream and downstream of S respectively. If 1200≤L, in addition to the above two primers, let S' be the sequence obtained by removing 600 bp from both ends of S, divide S' into 600 bp fragments (the most downstream fragment can be less than 600 bp), and design additional sequencing primers in the upstream of these fragments. The design of genome sequencing primers is basically the same as that of recombinant plasmids, except that at least one pair of primers is needed for PCR and sequencing. Figure 2 The parameter input of Primer3 is the same as that of PCR primer design.

[0110] S2.3.2, generate all candidate primers from the primer search region according to the parameter input of primer design, and generate primer pair sets based on them, screen the best primer pairs according to the scoring rules and algorithms of Primer3, and detect and optimize the two primers for PCR of the sequencing region based on off-target analysis;

[0111] S2.3.3. Based on the sequencing primer design results and the recombinant vector and edited genome sequence, generate a recombinant plasmid containing the tag of the fragment to be sequenced (homologous arm, sgRNA fragment, insert fragment, etc.) and the corresponding sequencing primer tag, and an edited genome map file for OVE package-based visualization;

[0112] wherein the off-target analysis in S2.3.2 is used to ensure that the designed homologous arm only undergoes homologous recombination with the predetermined position of the target gene sequence, and does not undergo unintended editing events at other similar sequences in the genome, and the application of the off-target analysis in homologous arm and primer design is as follows:

[0113] For homologous arm off-target analysis, first determine the upstream and downstream homologous arms of the sequence to be operated according to the parameters input in the front end, and obtain the 20 kbp sequence upstream and downstream thereof as the reference sequence for alignment, and then use BLAST2.9.0

[25] for sequence alignment. If the homologous arm is aligned to multiple positions on the reference sequence, it is considered that the homologous arm needs to be subjected to homologous arm off-target risk assessment. The off-target risk assessment criteria are as follows: (1) identity>90% and coverage>90% are defined as high off-target risk; (2) identity>90% and coverage>70% are defined as medium off-target risk; (3) identity>90% and alignlength>100 are defined as low off-target risk. The homologous arms with medium / high off-target risk levels are defined as off-target homologous arms.

[0114] For primer design, firstly, for recombination vector fragment PCR primers, since their search region is usually short, if the entire search region is off-target, the primer must also be off-target, so there is no need to analyze its off-target. For recombination vector sequencing primers, since there is usually no homologous arm and sgRNA fragment homologous sequence on the plasmid backbone, and there is also usually no repeat sequence on the plasmid backbone, so there is also no need to analyze its off-target. In contrast, genomic sequencing verification mainly includes two steps: (1) obtaining the sequence to be sequenced (a pair of PCR primers are needed); (2) sequencing the sequence (a set of primers including the PCR primer pair are needed). Since the PCR primer pair uses the reference genome as the template sequence, off-target effects can occur once the target sequence has a repeat site. Therefore, we added genome editing sequencing primer optimization based on off-target analysis. This optimization is achieved by optimizing the search region of the PCR primer pair used to obtain the fragment to be sequenced. First, determine the search region of the PCR primer pair according to the front-end input parameters. Then use Bowtie1.3.1

[26] to perform sequence alignment between the search region and the reference genome. Then evaluate the off-target risk according to the alignment results, and sequences with less than 4bp mismatches are defined as potential off-target sites. The off-target risk evaluation criteria are as follows: (1) if both the upstream and downstream primers have potential off-target sites, since the number of real primers available for PCR of the target fragment is affected by off-target, the upstream and downstream primers need to be optimized; (2) if one of the upstream and downstream primers has a potential off-target site, and 0.8*target product length≤off-target primer product length≤1.2*target product length, since the target product and off-target primer product cannot be effectively separated by subsequent methods such as gel electrophoresis, the upstream / downstream primer with potential off-target site needs to be optimized. The optimization strategy is as follows: Figure 3 ): Set the primer search region length L (default 40bp), move the search region of the upstream and downstream primers to their upstream and downstream respectively by L each time, then perform off-target analysis on the new search region, until a sequencing primer needs to be added due to the increase in the length of the sequencing fragment. If all new search regions are off-target, the optimization fails, otherwise it is successful. Finally, for the primer search region that has been successfully optimized, design the sequencing primer set according to the new template determined by the new search region.

[0115] Also included is a task management and visualization system 3 for monitoring and managing the progress and status of gene editing design tasks, using cloud computing and real-time data processing technology to achieve real-time monitoring of tasks, dynamic visualization of results, and user interaction, the task management and visualization system 3 includes a task submission module, a result display module, and a user interaction module;

[0116] The task submission module is used for users to upload files and input parameters required for editing tasks, including genome sequence, genome editing design requirements, and sgRNA design, PCR primer design, and sequencing verification primer design parameters;

[0117] The result display module is used for displaying the results of genome editing tasks, including sgRNA, PCR primer, sequencing verification primer design results, homologous arm off-target risk assessment results, and edited recombinant plasmid and genome map files;

[0118] The user interaction module is used for user navigation to different task management interfaces, modification of parameter settings, resubmission of modified tasks, and provision of user feedback functions.

[0119] The error processing unit 4 is also included, which is used for automatic identification and solution of errors occurring in the process of genome editing design, realizes rapid positioning and detailed problem explanation of errors, and includes an error diagnosis module and a solution module;

[0120] The error diagnosis module is responsible for monitoring the entire genome editing design process, and real-time detection of various logical, input, and execution errors that may occur, including input file errors, parameter setting errors, and primer design failures;

[0121] The solution module generates corresponding solutions and correction suggestions according to the error type and specific circumstances diagnosed.

[0122] The system architecture of the online tool is based on Amazon Web Services, which helps to ensure flexible invocation of computing resources and stability of the system. The architecture mainly includes a front-end presentation layer, a logical computing layer, and a data storage layer;

[0123] The front-end presentation layer mainly involves task management and visualization systems, including components that directly interact with users, uses AWS S3 and Cloud Front to provide fast access and delivery of static content, integrates OVE package tools to provide interactive DNA segment visualization, helps users to add plasmid tags, specify primers, and check design results;

[0124] The logic calculation layer mainly involves sgRNA design unit, genome editing design unit and error processing unit, including management of HTTP request and provision of core calculation service function, and is constructed using AWS Lambda, AWS API Gateway and AWS Step Functions service. The AWS Lambda carries the core calculation program of the genome editing design workflow. The API Gateway processes HTTP request and routes it to the backend. The AWS Step Functions asynchronously calls the AWS Lambda to orchestrate the serverless workflow by processing the message from the API Gateway;

[0125] The data storage layer mainly involves the persistence storage of the website, including task record (task input, parameter setting, intermediate calculation result) and design result (sgRNA design result, genome editing design result, failed task analysis, homologous arm off-target risk analysis, primer synthesis order), and is constructed based on AWS DynamoDB and AWS S3 service, for management of the persistence storage of the website. The AWS DynamoDB provides efficient database service for storage of task state and user configuration and the like dynamic data, and the AWS S3 is used for storage of large-scale data object, such as input file and design result, to ensure long-term reliable storage of data. The above shows and describes the basic principle, main features and advantages of the present application. It should be understood by the person skilled in the art that the present application is not limited by the above embodiments, and the above embodiments and description in the specification are only preferred examples of the present application, and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the present application. The scope of protection of the present application is defined by the appended claims and their equivalents.

Claims

1. An online device for automated design of genome editing based on CRISPR / Cas technology, characterized in that, The application relates to a system for designing and managing genome editing tasks, which comprises the following units: An sgRNA design unit (1) for designing specific sgRNAs, aiming at CRISPR technology, using the highly specific binding ability of sgRNAs to achieve precise editing of specific sites in the genome; A genome editing design unit (2) for achieving editing of the genome, aiming at CRISPR / Cas-mediated homologous recombination technology to achieve precise knockout, insertion and replacement of sequences; A task management and visualization system (3) for monitoring and managing the progress and status of gene editing design tasks, using cloud computing and real-time data processing technology to achieve real-time monitoring of tasks, dynamic visualization of results and user interaction; An error processing unit (4) for automatically identifying and solving errors occurring in the process of genome editing design, achieving rapid positioning of errors and detailed problem explanation; The sgRNA design unit (1) comprises a data preprocessing module and an sgRNA generator, and the specific steps involved in sgRNA design are as follows: S1.1, a user first needs to specify a reference genome by selecting and uploading a genome sequence, including an sgRNA design input file and a target modification file containing genome editing design task information, and a data preprocessing module performs format processing on these files; S1.2, the user sets the parameters of sgRNA design, including the type of CRISPR / Cas system, the length of guide sequence, the upper limit of the number of mismatched bases of potential off-target sites and genome sequence in off-target analysis, and the algorithm of the cutting efficiency of sgRNA, namely Efficiency score, and the sgRNA generator generates and optimizes sgRNA according to the related parameters set by the user; S1.3, after the user completes the selection / uploading of input files and parameter setting, the task submission button is clicked; S1.4, after the design process is completed, the "JOB MANAGER" page will be automatically jumped to, and the "JOB MANAGER" page shows the result list of sgRNA design, including sgRNA sequence, position on target sequence, GC content, number of potential off-target sites containing different numbers of mismatched bases, Efficiency score and sgRNA ranking; The data preprocessing module is used for receiving user-uploaded genome sequence files, and performing format processing on the files, and the specific steps involved are as follows: S1.1.1, receiving user-uploaded genome sequence files and checking whether the format of the files is correct; S1.1.2, using sequence verification technology to check whether the uploaded genome sequence contains non-standard base characters, removing illegal characters and blank lines, converting all bases to uppercase, and unifying the representation of sequences; S1.1.3, using a genome coordinate system to standardize the input file; S1.1.4, using sequence alignment technology to confirm the consistency of the input sequence with the reference genome of the target species; The sgRNA generator is responsible for the actual sgRNA design, including the generation and optimization of sgRNA sequences, and the specific steps involved in the generation and optimization of sgRNA by the sgRNA generator are as follows: S1.2.1, input the reference genome, the target sequence and the parameters of the sgRNA design, use the BLAST 2.9.0 alignment tool to align the target sequence with the reference genome sequence, determine the exact position of the target sequence in the reference genome, and obtain the coordinates of the target sequence on the reference genome; S1.2.2, based on the confirmed target sequence position, expand the length of the guide sequence and PAM sequence to both ends to obtain the sgRNA search region; S1.2.3, in the sgRNA search region, identify all candidate sgRNAs adjacent to the PAM sequence, and ensure that the candidate sgRNAs meet the specific requirements of the CRISPR / Cas system for the PAM sequence; S1.2.4, use Bowtie1.3.1 to build a genome index, and align the candidate sgRNAs with the entire reference genome to search for potential off-target sites with high similarity but different positions to the candidate sgRNAs, set an upper limit for the number of base mismatches, and evaluate the potential off-target risk of each candidate sgRNA; S1.2.5, use CHOPCHOPv3 software to score the efficiency of each candidate sgRNA according to the preset or user-selected scoring algorithm, including the cutting efficiency and off-target risk of the sgRNA; S1.2.6, generate a design result file containing detailed information of all candidate sgRNAs, including sequence, position, efficiency score, off-target risk and ranking, the result file is used for front-end display, users can view and compare the performance of different sgRNAs online, and download the result file; The genome editing design unit (2) designs specific sgRNAs and PCR primers, and provides homologous arm templates, uses the homologous recombination mechanism of cells to accurately repair the DNA double-strand break mediated by the CRISPR / Cas system, thereby achieving site-directed editing of the target sequence on the genome, and the genome editing design unit (2) includes a data preprocessing module, a plasmid map editing module, a PCR primer design module, a homologous arm off-target detection module, a sequencing verification primer design module, and a genome sequencing verification primer off-target optimization module; The data preprocessing module formats, corrects and extracts position information from the genome data and plasmid information; The plasmid map editing module allows users to customize and adjust the plasmid structure to adapt to different editing strategies; The PCR primer design module is used to design PCR primers for the genome editing process; The homologous arm off-target detection module is used to analyze the possible off-target positions of the homologous arm and evaluate the off-target risk; The sequencing verification primer design module is used to design sequencing primers for verifying the editing results; The genome sequencing verification primer off-target optimization module is used to optimize the genome sequencing verification primers to reduce off-target effects and improve sequencing accuracy; The genomic editing design unit (2) involves the following specific steps in the process of genomic editing design: S2.1, the data preprocessing module formats the user uploaded genome and plasmid data, and parses and confirms the user input parameters, including sgRNA design parameters, editing type, wherein the editing type includes knockout, insertion and replacement; S2.2, the PCR primer design module designs the PCR primers of the homologous arm and sgRNA based on the user input parameters, calculates the optimal length, temperature and GC content of the primers, generates a primer list including sequence, expected product size and predicted amplification efficiency; S2.3, the sequencing verification primer design module and the genome sequencing verification primer off-target optimization module analyze potential off-target sites, select the best PCR primer position and sequence, optimize the PCR primer conditions, and obtain the sequencing primer design results, including primer position, sequence and predicted off-target analysis results; The PCR primer design module designs PCR primers based on DNA fragment assembly parameters and PCR primer design parameters, including determining the design type, generating primer design templates, primer search and evaluation, primer modification and generating recombinant plasmid map functions; The PCR primer design module designs PCR primers based on user input parameters, and the specific process steps are as follows: S2.2.1, first identify the editing type and target sequence of the design task in the input file, wherein the editing type includes knockout, insertion and replacement, and determine the design type of each task according to the user specified plasmid system and additional primer design requirements; S2.2.2, according to the requirements of the design type, select the PCR primer design template of the homologous arm, sgRNA fragment and plasmid backbone on the reference genome and the starting plasmid, determine the primer search area on the template according to the input DNA fragment assembly parameters, and generate the parameter input of Primer3 according to the input of the primer design parameters; S2.2.3, according to the parameter input of primer design, generate all candidate primers from the primer search area, and generate primer pairs based on them, select the best primer pair according to the scoring rules and algorithm of Primer3, and finally perform off-target evaluation of the homologous arm and generate the related result file; S2.2.4, determine the overlapping sequence between adjacent recombinant plasmid fragments according to the design type of the task, add the overlapping sequence and spacer sequence, IIS type restriction enzyme recognition site and protection sequence to the 3' end of the corresponding primer, and use the primer with added linker sequence and its key features, including Tm, GC content to generate the result file of primer design; S2.2.5, based on the homologous arm off-target evaluation strategy, evaluate the off-target risk of the upstream and downstream homologous arms, and generate the homologous arm off-target risk evaluation report; S2.2.6, based on the primer design results and their PCR product sequences, generate a recombinant plasmid map file with recombinant plasmid fragment tags and corresponding PCR primers.

2. The online device for automated design of CRISPR / Cas technology based genome editing oriented application according to claim 1, characterized in that: The sequencing verification primer design module is used for primers for sequencing verification of recombination plasmids and genome editing, and the sequencing verification primer design module comprises determining sequencing targets, primer searching and evaluation, and design result visualization; The specific steps are as follows: S2.3.1, determining the target region to be sequenced of the recombination plasmid and the edited genome and the search region of the sequencing verification primer according to the PCR primer design result, and generating the sequencing verification primer design parameter input of Primer3 according to the primer design parameter input; S2.3.2, generating all candidate primers from the primer search region according to the primer design parameter input, and generating a primer pair set based thereon, screening the best primer pair according to the scoring rules and algorithm of Primer3, and detecting and optimizing the two primers for PCR of the target sequencing region based on off-target analysis; S2.3.3, based on the sequencing primer design result and the recombination vector and edited genome sequence, generating a recombination plasmid and edited genome map file containing the tag of the fragment to be sequenced and the corresponding sequencing primer tag, for visualization based on the OVE package.

3. The online CRISPR / Cas technology-based genome editing automation design oriented device according to claim 1, characterized in that: The task management and visualization system (3) comprises a task submission module, a result display module and a user interaction module; Wherein, the task submission module is used for users to upload files and input parameters required for editing tasks, including genome sequence, genome editing design requirements and sgRNA design, PCR primer design and sequencing verification primer design parameters; Wherein, the result display module is used to display the results of gene editing tasks, including sgRNA, PCR primer, sequencing verification primer design result, homologous arm off-target risk evaluation result and edited recombination plasmid and genome map file; Wherein, the user interaction module is used for users to navigate to different task management interfaces, modify parameter settings, resubmit modified tasks, and provide user feedback functions.

4. The online CRISPR / Cas technology-based genome editing automation design oriented device according to claim 1, characterized in that: The error processing unit (4) comprises an error diagnosis module and a solution module; Wherein, the error diagnosis module is responsible for monitoring the entire genome editing design process, and detecting various logical, input and execution errors that may occur in real time, including input file errors, parameter setting errors and primer design failures; Wherein, the solution module generates corresponding solutions and correction suggestions according to the error type and specific situation diagnosed.

5. The online CRISPR / Cas technology-based genome editing automation design oriented device according to claim 1, characterized in that: The architecture of the device comprises a front-end representation layer, a logical calculation layer and a data storage layer; Wherein, the front-end representation layer is mainly related to the task management and visualization system, including components that directly interact with users, using AWS S3 and Cloud Front to provide fast access and delivery of static content, integrating OVE package tools to provide interactive DNA fragment visualization, helping users to add plasmid tags, specify primers and check design results; Among them, the logic calculation layer is mainly related to sgRNA design unit, genome editing design unit and error processing unit, including managing HTTP requests and providing core computing service functions, using AWS Lambda, AWS API Gateway and AWSStep Functions services to build; Among them, the data storage layer is mainly related to the persistence storage of the website, including task records and design results based on AWS DynamoDB and AWS S3 services.

Citation Information

Patent Citations

  • System for multi-round editing of fungal genome by CRISPR system and method thereof

    CN110205334A

  • Accurate and efficient editing method of upland cotton genome

    CN110283840A