A method and system for designing a target sequence
By inserting multiple base sequences and restriction endonuclease sites into the target sequence and combining them with target sequence scoring criteria to optimize the target sequence design, the low efficiency and accuracy problems of the CRISPR-Cas system in multi-gene editing have been solved, achieving efficient multi-gene editing and mutant acquisition.
Patent Information
- Application Number
- CN202310802050.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-07-03
AI Technical Summary
In existing technologies, the CRISPR-Cas system is inefficient when editing multiple genes or multiple sites, has inaccurate target sequence design, and lacks design and scoring tools for multiple target sequences, resulting in editing failures and low conversion efficiency.
A method and system were designed to construct a 50bp target sequence by inserting multiple base sequences into the target sequence, combining restriction endonuclease recognition sites and scoring criteria to optimize the target sequence to improve editing efficiency and accuracy, and using the high-scoring target sequence to link with donor DNA to achieve multi-gene editing.
It improves the multi-gene editing efficiency of the CRISPR-Cas system, reduces design errors, enhances the functional diversity and editing success rate of target sequences, simplifies the operation process, and improves the efficiency of obtaining multiple specific mutants.
Smart Images

Figure CN116884493B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, in particular to a method and system for designing target sequences. BACKGROUND
[0002] Each target sequence (gRNA) can only guide the CRISPR-Cas system to edit the sequence corresponding to it. However, when studying a certain organism, researchers often need to modify multiple genes or sites. Modifying multiple genes or sites requires the design and expression of multiple target sequences, and being able to design and express multiple target sequences at the same time can save a lot of time. If each expression cassette or plasmid expresses one target sequence, when modifying multiple genes or sites, multiple expression cassettes or plasmids containing different target sequences need to be transformed at the same time, and the more the number is, the lower the transformation efficiency is. If homologous recombination repair is needed, one or more donor DNAs also need to be transformed, and the transformation efficiency is even lower. When editing, the probability of obtaining a mutant with multiple target mutations is very low. Connecting multiple target sequences, expressing or releasing multiple target sequences at the same time through one transformation; connecting one or more target sequences with one or more donor DNAs, expressing or releasing these target sequences and donor DNAs at the same time after being transformed into an organism, can greatly improve the efficiency of multi-gene editing and research progress. In addition, if the expression cassette or plasmid connected with the target sequence or donor DNA can be changed at will, it will be helpful to quickly edit other genes or sites, or quickly obtain a mutant with another mutation. This requires the design of precise connection fragments or restriction endonuclease recognition or cleavage sites, and one base error in design will also lead to research failure.
[0003] The environment and epigenetic features in which the CRISPR-Cas system is placed significantly affect the efficacy of CRISPR-Cas. Methylation and chromatin features significantly affect mutagenesis efficiency, with a maximum of 250-fold. Specific chromatin features can have substantial and lasting effects on CRISPR-Cas mutation efficiency and DNA double-strand break repair outcomes. Reducing DNA methylation at CRISPR target sites helps improve editing efficiency.
[0004] The current tools for designing target sequences can only design target sequences for specified reference genomes, which are not necessarily the same as the genome of the target organism, resulting in inaccurate target sequences or target sequences that do not exist in the target organism, thereby causing gene editing to fail. For example, when the CRISPR gRNA Design tool designs a target sequence for the SNF1 gene of wild-type industrial Saccharomyces cerevisiae F01, it can only select a specific reference genome, and the nucleotide sequence of the designed target sequence is GATATGTGCCCCATCCGCTAAGG, but the nucleotide sequence corresponding to the target sequence in Saccharomyces cerevisiae F01 is GATATGTGCACCATCCGCTAAGG, that is, the 10th base in the nucleotide sequence of the target sequence is C, and the 10th base in the nucleotide sequence of Saccharomyces cerevisiae F01 is A, which differs by one base. The current tools for designing target sequences can only design for the genome of a specific species and cannot design target sequences for some organisms or some genes, such as resistance genes. In addition, the length of the target sequence is generally 17-24 base pairs, the target sequence is short, the function of the target sequence design tool is relatively single, and when the target sequence is connected to other sequences or multiple target sequences are co-expressed, it is necessary to artificially design or add enzyme cutting sites or connection fragments, which not only consumes time but also is prone to errors; there is a lack of tools for designing pairs or multiple gRNAs; there is a lack of tools for designing target sequences and scoring target sequences, and it is difficult for researchers to quickly and accurately select the best target sequence from multiple target sequences. This hinders the application of the CRISPR-Cas system.
[0005] The information disclosed in this Background section is only for the purpose of increasing the understanding of the general background of the application and should not be taken as an acknowledgement or any form of suggestion that this information forms prior art with regard to the natural person skilled in the art. SUMMARY
[0006] The purpose of the present application is to provide a method for designing target sequences, which can design target sequences with a length of up to 50 bp, and can insert various base sequences during the design of the target sequence to increase the functionality of the target sequence according to the functional requirements of the target sequence.
[0007] The present application also provides a system for designing target sequences.
[0008] To achieve the above-mentioned purpose, the present application provides a method for designing target sequences, which comprises the following steps:
[0009] When there is a GG (or AG) base on the target sequence, the base and the sequence before the base are read, and the length of the sequence is mbp, wherein m≤50;
[0010] edit the target sequence model to construct a target sequence, wherein the editing comprises cleavage, input sequence information, or reading sequence files;
[0011] According to the target sequence scoring standard, the sequence characteristics of the target sequence are analyzed and judged, and corresponding scores are given, and the target sequence and the score value are output.
[0012] Preferably, in the above technical solution, when it is judged that there is a GG (or AG) base on the target sequence, the sequence before the base is read, and the length of the sequence is mbp, and the specific steps include:
[0013] reading the target sequence;
[0014] determining whether the target sequence contains a GG (or AG) base;
[0015] When the target sequence contains a GG (or AG) base, the position of the GG (or AG) base is marked.
[0016] Taking the base as a base point, the sequence before the base along the target sequence is read, and the length of the sequence is mbp, wherein m≤50.
[0017] Preferably, in the above technical solution, the target sequence includes an arbitrarily input reference genome or DNA sequence, or a built-in genome or DNA sequence.
[0018] Preferably, in the above technical solution, the sequence is a target sequence model, and the target sequence model is edited to construct a target sequence, and the steps specifically include:
[0019] Taking the sequence as a target sequence model, inputting sequence information of a restriction enzyme or reading a sequence file of a restriction enzyme, and obtaining a first preset target sequence;
[0020] Taking the first preset target sequence as a target sequence model, inputting sequence information of a connection vector or an expression cassette or reading a sequence file of a connection vector or a sequence file of an expression cassette, and obtaining a second preset target sequence, wherein the expression cassette is composed of a promoter, a gene or sequence to be expressed, and a terminator;
[0021] Taking the second preset target sequence as a target sequence model, inputting sequence information for multi-fragment connection or reading a sequence file for multi-fragment connection, and obtaining a third preset target sequence, wherein the sequence for multi-fragment connection includes but is not limited to a recognition or cleavage sequence of a type II endonuclease, or a sequence that can be recognized and cleaved by an exogenous or endogenous system;
[0022] When the editing is completed, the constructed target sequence is obtained.
[0023] The target sequence is created by taking the sequence output by the target sequence as a target sequence template. The advantage of the target sequence is that it can solve the problem of design error or poor specificity of the designed target sequence caused by the difference between the reference genome and the genome of the target organism, thereby causing targeting failure or off-targeting. The sequence information of the restriction enzyme is inserted into the target sequence template. First, the restriction enzyme site is added to the target sequence, which facilitates the connection between the preset target sequences or the connection between the preset target sequences and the vector or other sequences, and lays a foundation for seamless cloning in the later stage. Second, it is convenient for later verification. By designing a target sequence containing an endonuclease recognition or cleavage site, the sequence containing the site can be amplified after gene editing, and the cleaved fragments can be analyzed after enzyme digestion to verify whether the editing is successful (for example, the DNA is not cut when the editing is not successful). The second preset target sequence is input or read to obtain a third preset target sequence. Since the third preset target sequence contains a sequence for multi-fragment connection, it has the function of connecting multiple preset target sequences in multiple ways, thereby obtaining multiple gRNAs with different expression modes. After scoring the obtained gRNA, the gRNA with a high score is connected with the donor DNA to obtain a new sequence. The donor DNA contains a base fragment of the corresponding mutation site (to direct the mutation of the target DNA) and a homologous arm on both sides of the base fragment (to perform homologous recombination between the donor DNA and the target DNA). The sequence can produce or release gRNA and donor DNA in the organism at the same time. When the gRNA guides the Cas9 protein to cut the target DNA, the organism repairs the cut DNA, and the donor DNA is used as a repair template for the target DNA, which directs the mutation of the target DNA. The advantages of connecting the gRNA with a high score with the donor DNA and transferring it into the organism are as follows: first, when editing the target DNA, multiple gRNAs guide the Cas9 protein to cut multiple target DNAs or multiple target DNAs at the same time. The cut target DNA uses the donor DNA as a repair template and finally obtains a target DNA with multiple specific mutations or multiple target DNAs with target mutations through homologous recombination. Second, although gRNA and donor DNA can be separately transferred into the organism to achieve the purpose of directing the mutation of the target DNA, multiple gRNAs and donor DNAs need to be transferred into the organism when editing multiple sites. However, the more gRNAs and / or donor DNAs are transferred into the organism at a time, the lower the efficiency of obtaining a mutant with multiple specific mutations. By connecting multiple gRNAs with high scores with multiple donor DNAs and transferring them into the organism, the effect of transferring multiple gRNAs and multiple donor DNAs at a time is achieved, which makes the efficiency of obtaining a mutant with multiple specific mutations higher.
[0024] Preferably, in the above technical solution, the step of analyzing and judging the sequence characteristics of the target sequence according to the target sequence scoring standard and assigning a corresponding score, and then outputting the target sequence and its score value specifically comprises:
[0025] Reading the target sequence scoring standard;
[0026] Reading the target sequence in the created target sequence;
[0027] According to the target sequence scoring standard, the sequence characteristics of the target sequence are analyzed and judged, and a corresponding score is assigned, and then the target sequence and its score value are output.
[0028] According to the above technical solution, the target sequence scoring in the method of the application is specifically (1) setting the characteristics according to the position of the target gene where the target sequence is located, such as the protein coding region, non-coding element or repeated region, and assigning different scores to the target sequence; (2) assigning different scores to the target sequence located in the important region, such as the sequence after the start codon ATG (gene ATG), the protein coding sequence (gene ORF), the regulatory region (such as the promoter, the enhancer region, etc.) and the like; (3) setting the distance between the gene editing region (target point) and the start codon ATG of the gene to assign a corresponding score to the target sequence, the closer the distance, the higher the score; (4) setting the number of different bases on the target sequence to assign a corresponding score to the target sequence, for example: the number of guanine (G) and adenine (A) is large, and a corresponding score is assigned to the target sequence; (5) setting the random and uniform distribution of bases in the target sequence to assign a corresponding score to the target sequence; (6) setting the number of consecutive occurrences of a certain base among the bases A\T\G\C in the target sequence to assign a corresponding negative score; (7) setting the target sequence of the target point located in the essential region of the genome to assign a corresponding score to the target sequence; (8) setting the target sequence containing off-target sites or different off-target site numbers to assign a corresponding negative score to the target sequence; (9) setting the off-target situation of the off-target site of the target sequence to assign a corresponding negative score; (10) setting the number of mismatched bases between the target sequence and the off-target site to be too small, for example: less than two gRNAs, to assign a corresponding negative score to the target sequence, and the smaller the number, the larger the absolute value of the negative score; (11) setting the target sequence of the target point with genomic sequence rearrangement (such as translocation, duplication, inversion) to assign a corresponding negative score; (12) setting the target sequence of the target point that can be methylated (for example, containing a CpG island) to assign a corresponding negative score, to reduce the decrease of targeting efficiency caused by methylation modification; (13) setting a corresponding negative score according to the interaction relationship between the target sequence and the sequence, for example, assigning a corresponding negative score to the target sequence with complementarity; (14) assigning a corresponding score to the target sequence according to the binding force between the target sequence and the target DNA (for example, the binding force between AT or GC); (15) assigning a corresponding score to the pair of target sequences according to the distance between the pair of target sequences, for example, assigning a corresponding negative score to the pair of target sequences with too small distance to affect the function of CRISPR-Cas system, such as Cas protein.
[0029] To achieve the above objectives, this application also provides a system for designing target sequences, the system comprising:
[0030] The judgment and reading module is used to determine that when there is a GG (or AG) base in the target sequence, read the base and the sequence preceding the base. The length of the sequence is mbp, where m≤50.
[0031] An editing module is used to edit the target sequence model using the sequence as the target sequence model to construct a target sequence, wherein the editing includes cutting, inputting sequence information, or reading sequence files;
[0032] The read and output module is used to analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0033] Preferably, in the above technical solution, the judgment and reading module includes:
[0034] The first reading unit is used to read the target sequence;
[0035] The first judgment unit is used to determine whether the target sequence contains a GG (or AG) base;
[0036] The first labeling unit is used to label the position of the GG (or AG) base when the target sequence contains a GG (or AG) base;
[0037] The second reading unit is used to read the base and the sequence preceding the base along the target sequence, with the base as the base point. The length of the sequence is mbp, where m≤50.
[0038] Preferably, in the above technical solution, the editing module includes:
[0039] The first editing unit is used to input the sequence information of the restriction endonuclease or read the sequence file of the restriction endonuclease to obtain the first preset target sequence, using the sequence as the target sequence model.
[0040] The second editing unit is used to take the first preset target sequence as the target sequence model, input the sequence information of the ligation vector or expression cassette or read the sequence file of the ligation vector or the sequence file of the expression cassette to obtain the second preset target sequence, wherein the expression cassette is composed of a promoter, a gene or sequence to be expressed, and a terminator.
[0041] The third editing unit is used to input sequence information for multi-fragment ligation or read sequence files for multi-fragment ligation using the second preset target sequence as the target sequence model, and obtain the third preset target sequence. The sequence for multi-fragment ligation includes, but is not limited to, the recognition or cleavage sequence of type II endonuclease, or the sequence that can be recognized and cleaved by exogenous or endogenous systems.
[0042] The first acquisition unit is used to acquire the constructed target sequence after editing is completed.
[0043] Preferably, in the above technical solution, the reading and output module includes:
[0044] The third reading unit is used to read the target sequence scoring criteria;
[0045] The fourth reading unit is used to read the target sequence in the created target sequence;
[0046] The first output unit is used to analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0047] To achieve the above objectives, the present invention also provides an electronic device, the electronic device comprising:
[0048] Memory, storing at least one instruction; and
[0049] The processor executes instructions stored in the memory to implement the target sequence design method as described in any one of claims 1 to 5.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] First, the target sequences designed using this method have the following advantages: 1. They can be designed using any reference genome or DNA sequence, or built-in genome or DNA sequences, resulting in diverse target sequence templates that allow for the design of specific target sequences for any given sequence; 2. By searching for GG (or AG) bases, target sequence templates are obtained. The designed target sequences, using the genome or DNA sequence of the target organism as a reference, avoid differences between the reference genome or DNA sequence and the target organism or target DNA sequence, thus solving the problem of inaccurate target sequence design or difficulty in ensuring specificity, leading to targeting failure or off-target effects; 3. In the target sequence design step, by adding restriction endonuclease recognition or cleavage sequences on both sides of the target sequence, it is easy to ligate the pre-set target sequence to a linker vector, expression cassette, or other sequences later, enabling rapid acquisition of the target sequence. 4. By adding linker vectors, expression cassettes, or multifunctional gRNAs to both sides of the target sequence, multiple gRNAs and / or donor DNAs can be rapidly obtained. Restriction endonuclease recognition or cleavage sequences can be added to both sides of the sequence containing multiple gRNAs and / or donor DNAs to facilitate linking with the vector or expression cassette, achieving modular expression of multiple gRNAs and / or donor DNAs. 5. The established scoring criteria are suitable for detecting the cleavage of most gRNA-guided Cas proteins. The target sequence scoring criteria can be used to quickly obtain the corresponding gRNA score, providing a reference for gRNA selection during gene editing.
[0052] Secondly, this system is created based on the method for designing target sequences, and has all the advantages of the method. Target sequences can be created conveniently according to the functional requirements of the target sequences. For example, corresponding linking fragments can be added to target sequences that need to be linked to vectors or fragments to facilitate linking with vectors or fragments, avoiding errors caused by manual addition that could lead to gene editing experiment failure. Attached Figure Description
[0053] Figure 1 This is a flowchart of an embodiment of the target sequence design method according to the present invention;
[0054] Figure 2 yes Figure 1 A flowchart of a specific implementation of step S100;
[0055] Figure 3 yes Figure 1 A flowchart of a specific implementation of step S200;
[0056] Figure 4 This system adds restriction enzyme sequences to the designed target sequence and generates a complementary sequence to the target sequence;
[0057] Figure 5 This is a diagram illustrating three different ways of expressing multiple gRNAs in one embodiment of the present invention;
[0058] Figure 6 yes Figure 1 A flowchart of a specific implementation of step S300 in the process;
[0059] Figure 7 This is a schematic diagram of a design target sequence system according to an embodiment of this application;
[0060] Figure 8 yes Figure 7 The diagram shows a specific implementation of the judgment and reading module.
[0061] Figure 9 yes Figure 7 The diagram shows a structural schematic of one specific embodiment of the editing module.
[0062] Figure 10 yes Figure 7 The diagram shows a specific embodiment of the read and output module.
[0063] Figure 11 This is a schematic diagram of the structure of an electronic device that implements the target sequence design method of the present invention.
[0064] Figure 5 Symbols and their names: 11 - Target sequence; 12 - SNR52 promoter; 13 - sgRNA backbone; 14 - Terminator; 15 - tRNA Gly . Detailed Implementation
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0066] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0067] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0068] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the target sequence design method of the present invention. The order of steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0069] The target sequence design method is applied to one or more electronic devices, which are devices that perform numerical calculations and / or information processing and / or model building according to pre-set or stored instructions. Their hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0070] The electronic device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0071] The electronic device may also include network devices and / or user devices. The network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0072] S100, when it is determined that there is a GG (or AG) base in the target sequence, read the base and the sequence before the base. The length of the sequence is mbp, where m≤50;
[0073] In the embodiments of the present invention, by locating the GG (or AG) bases on the target sequence, and then using the GG (or AG) bases as the starting point, a sequence of length mbp containing the GG (or AG) bases is read forward along the target sequence as a model for designing the target sequence. The target sequence designed by the model has improved guided cutting accuracy. Using specific sequence fragments on the target sequence as templates for target sequence design reduces the off-target rate during the target sequence guidance process.
[0074] S200, using the sequence as a target sequence model, editing the target sequence model to construct a target sequence, wherein the editing includes cutting, inputting sequence information, or reading a sequence file;
[0075] In the embodiments of the present invention, the sequence in S100 above is used as the target sequence model. According to the specific function of the target sequence, the target sequence is constructed in a variety of ways, such as cutting some bases in the target sequence model, adding specific sequence information, or reading existing sequence files and inserting them into the target sequence model. The target sequence designed in this way has more functions.
[0076] S300: Analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0077] Figure 2 This is a flowchart illustrating a specific implementation of step S100. In this embodiment, step S100 specifically includes the following steps:
[0078] S101, Read the target sequence;
[0079] S102, determine whether the target sequence contains GG (or AG) bases;
[0080] S103, when the target sequence contains a GG (or AG) base, mark the position of the GG (or AG) base;
[0081] S104, using the base as the starting point, read the base and the sequence preceding the base along the target sequence, the length of the sequence being mbp, where m≤50.
[0082] Figure 3 The flowchart illustrates a specific implementation of step S200. In this embodiment, step S200 specifically includes the following steps:
[0083] S201, using the sequence as the target sequence model, input the sequence information of the restriction endonuclease or read the sequence file of the restriction endonuclease to obtain the first preset target sequence;
[0084] S202, using the first preset target sequence as the target sequence model, input the sequence information of the restriction endonuclease or the sequence information of the expression cassette or read the sequence file of the restriction endonuclease or the sequence file of the expression cassette to obtain the second preset target sequence, wherein the expression cassette consists of a promoter, a gene or sequence to be expressed, and a terminator.
[0085] S203, using the second preset target sequence as the target sequence model, input sequence information for multi-fragment ligation or read sequence files for multi-fragment ligation to obtain the third preset target sequence, wherein the sequence for multi-fragment ligation includes, but is not limited to, recognition or cleavage sequences of type II endonucleases, or sequences that can be recognized and cleaved by exogenous or endogenous systems;
[0086] S204, after editing is complete, obtain the constructed target sequence.
[0087] like Figure 4 As shown, F is the designed target sequence, R is the complementary strand of the target sequence, where N represents the bases A / T / G / C, X represents the corresponding complementary base of N, and GATC and AAAC are the cleavage ends of the restriction endonuclease (BsmBI enzyme).
[0088] like Figure 5 As shown, in this embodiment of the invention, a sequence for multi-segment ligation is added to the target sequence to enable the ligation of multiple target sequences or to the sgRNA backbone, so as to generate multiple different gRNAs simultaneously through different expression modes. The number of gRNAs is not limited to 3, but can be greater than or equal to 9. The promoter that can be used is SNR52 or other sequences with similar properties, and the ligation fragment used is tRNAGly or other sequences with similar properties.
[0089] Figure 6 This is a flowchart illustrating a specific implementation of step S300. In this embodiment, step S300 specifically includes the following steps:
[0090] S301, Read the target sequence scoring criteria;
[0091] S302, Read the target sequence from the created target sequence;
[0092] S303, Analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0093] In this embodiment of the invention, the designed gRNA is assigned corresponding scores or negative scores based on its sequence characteristics, different base contents, off-target effects, etc., according to the target sequence scoring criteria. The target sequence and its score value are output. The target sequence with the high score is selected as the target sequence to guide the Cas9 protein to cut the target DNA in gene editing. After the target sequence with the high score is linked to the donor DNA, it is transferred into the organism. In the organism, the target sequence guides the Cas9 protein to the cutting site to cut the target DNA. The cut target DNA is repaired by homologous recombination using the donor DNA as a repair template to obtain the target DNA with directed mutation.
[0094] As a response to the aboveFigure 1 The implementation of the method shown in this application provides a structural schematic diagram of an embodiment for designing a target sequence system. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0095] like Figure 7 As shown, the system 400 for designing target sequences described in this embodiment includes:
[0096] The judgment and reading module 410 is used to judge when there is a GG (or AG) base in the target sequence, and read the base and the sequence before the base, wherein the length of the sequence is mbp, and m≤50;
[0097] The editing module 420 is used to edit the target sequence model using the sequence as the target sequence model to construct the target sequence, wherein the editing includes cutting, inputting sequence information, or reading sequence files;
[0098] The reading and output module 430 is used to analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0099] In the embodiments of the present invention, please refer to Figure 8 The diagram below shows a specific embodiment of the judgment and reading module 410, which includes:
[0100] The first reading unit 411 is used to read the target sequence;
[0101] The first judgment unit 412 is used to determine whether the target sequence contains a GG (or AG) base;
[0102] The first labeling unit 413 is used to label the position of the GG (or AG) base when the target sequence contains a GG (or AG) base;
[0103] The second reading unit 414 is used to read the base and the sequence preceding the base along the target sequence, with the base as the base point. The length of the sequence is mbp, where m≤50.
[0104] In the embodiments of the present invention, please refer to Figure 9 The diagram shows a specific embodiment of the editing module 420, which includes:
[0105] The first editing unit 421 is used to input the sequence information of the restriction endonuclease or read the sequence file of the restriction endonuclease to obtain the first preset target sequence, using the sequence as the target sequence model.
[0106] The second editing unit 422 is used to take the first preset target sequence as the target sequence model, input the sequence information of the ligation vector or expression cassette or read the sequence file of the ligation vector or the sequence file of the expression cassette to obtain the second preset target sequence, wherein the expression cassette is composed of a promoter, a gene or sequence to be expressed, and a terminator.
[0107] The third editing unit 423 uses the second preset target sequence as the target sequence model, inputs sequence information for multi-fragment ligation or reads sequence files for multi-fragment ligation, and obtains the third preset target sequence. The sequence for multi-fragment ligation includes, but is not limited to, recognition or cleavage sequences of type II endonucleases, or sequences that can be recognized and cleaved by exogenous or endogenous systems.
[0108] The first acquisition unit 424 is used to acquire the constructed target sequence after editing is completed.
[0109] In the embodiments of the present invention, please refer to Figure 10 This is a schematic diagram of a specific embodiment of the read and output module 430, which includes:
[0110] The third reading unit 431 is used to read the target sequence scoring criteria;
[0111] The fourth reading unit 432 is used to read the target sequence in the created target sequence;
[0112] The first output unit 433 is used to analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0113] Figure 11 This is a schematic diagram of the structure of an electronic device that implements the data determination method of the present invention. The electronic device 2 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0114] The electronic device 2 can also be, but is not limited to, any electronic product that can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device, such as personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, interactive network television (IPTV), smart wearable devices, robots, etc.
[0115] The electronic device 2 can also be a desktop computer, laptop, handheld computer, cloud server, or other computing device.
[0116] The network in which the electronic device 2 is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0117] In one embodiment of the present invention, the electronic device 2 includes, but is not limited to, a memory 21, a processor 22, and a computer program, such as a data determination program, stored in the memory 21 and executable on the processor 22.
[0118] Those skilled in the art will understand that the schematic diagram is merely an example of the electronic device 2 and does not constitute a limitation on the electronic device 2. It may include more or fewer components than shown in the diagram, or combine certain components, or different components. For example, the electronic device 2 may also include input / output devices, network access devices, buses, etc.
[0119] The processor 22 executes the operating system of the electronic device 2 and various installed applications. The processor 22 executes the applications to implement the steps in the above-described data determination method embodiments, for example... Figure 1 The steps S100, S200, and S300 are shown.
[0120] Alternatively, when the processor 22 executes the computer program, it implements the functions of each module / unit in the above-described device embodiments. For example, when it determines that there is a GG (or AG) base in the target sequence, it reads the base and the sequence preceding the base, the length of which is mbp, where m≤50; using the sequence as a target sequence model, it edits the target sequence model to construct a target sequence, wherein the editing includes cutting, inputting sequence information, or reading sequence files; it analyzes and judges the sequence characteristics of the target sequence according to the target sequence scoring criteria, assigns corresponding scores, and outputs the target sequence and its score value.
[0121] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 21 and executed by the processor 22 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device 2. For example, the computer program can be divided into a judgment and reading module 410, an editing module 420, and a reading and output module 430.
[0122] The memory 21 can be used to store the computer programs and / or modules. The processor 22 implements various functions of the electronic device 2 by running or executing the computer programs and / or modules stored in the memory 21 and calling the data stored in the memory 21. The memory 21 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0123] The memory 21 can be the external memory and / or internal memory of the electronic device 2. Further, the memory 21 can be a circuit with storage function in an integrated circuit that does not have a physical form, such as RAM (Random-Access Memory), FIFO (First-In-First-Out), etc. Alternatively, the memory 21 can also be a memory with a physical form, such as a memory module, a TF card (Trans-flash Card), etc.
[0124] If the modules / units integrated in the electronic device 2 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0125] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0126] Combination Figure 1 The memory 21 in the electronic device 2 stores multiple instructions to implement a data determination method. The processor 22 can execute the multiple instructions to achieve the following: when a target sequence has a GG (or AG) base, read the base and the sequence preceding the base, the length of which is mbp, where m≤50; using the sequence as a target sequence model, edit the target sequence model to construct a target sequence, wherein the editing includes cutting, inputting sequence information, or reading a sequence file; analyze and judge the sequence characteristics of the target sequence according to the target sequence scoring criteria, assign corresponding scores, and output the target sequence and its score value.
[0127] Specifically, the processor 22's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0128] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0129] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0130] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0131] The foregoing description of specific exemplary embodiments of the invention is for illustrative and explanatory purposes. These descriptions are not intended to limit the invention to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of the invention and its practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of the invention, as well as various different choices and variations. The scope of the invention is intended to be defined by the claims and their equivalents.
Claims
1. A method of designing a target sequence, characterized by, The method comprises the following steps: When there is a GG or AG base on the target sequence, reading the base and the sequence before the base, the length of the sequence being m bp, wherein m≤50, the target sequence including any input reference genome or DNA sequence, or using the built-in genome or DNA sequence; Using the sequence as a target sequence model, editing the target sequence model, the steps of constructing the target sequence being specifically: Using the sequence as a target sequence model, inputting the sequence information of a restriction enzyme or reading the sequence file of the restriction enzyme, obtaining a first preset target sequence; Using the first preset target sequence as a target sequence model, inputting the sequence information of a connecting vector or an expression cassette or reading the sequence file of the connecting vector or the sequence file of the expression cassette, obtaining a second preset target sequence, wherein the expression cassette is composed of a promoter, a gene or sequence to be expressed, and a terminator; Using the second preset target sequence as a target sequence model, inputting the sequence information for multi-fragment connection or reading the sequence file for multi-fragment connection, obtaining a third preset target sequence, wherein the sequence for multi-fragment connection includes a recognition or cutting sequence of a type II endonuclease, or a sequence that can be recognized and cut by an exogenous or endogenous system; After editing is completed, obtaining the constructed target sequence; According to a target sequence scoring standard, analyzing and judging the sequence characteristics of the target sequence, giving a corresponding score value, and outputting the target sequence and the score value, wherein the target sequence scoring standard includes: (1) according to the characteristics of the target gene position where the target sequence is located, such as a protein coding region, a non-coding element or a repeat region, giving different scores to the target sequence; (2) setting a target sequence that can be methylated for a target point, giving a corresponding negative score; (3) according to the random and uniform distribution of bases in the target sequence, giving a corresponding score to the target sequence.
2. The method of designing a target sequence according to claim 1, wherein, The specific steps of the step of judging when there is a GG or AG base on the target sequence, reading the base and the sequence before the base, the length of the sequence being m bp, include: Reading the target sequence; Judging whether the target sequence contains a GG or AG base; When the target sequence contains a GG or AG base, marking the position of the GG or AG base; Taking the base as a base point, reading the base and the sequence before the base along the target sequence, the length of the sequence being m bp, wherein m≤50.
3. The method of designing a target sequence according to claim 1, wherein, The specific steps of the step of according to a target sequence scoring standard, analyzing and judging the sequence characteristics of the target sequence, giving a corresponding score value, and outputting the target sequence and the score value include: Reading the target sequence scoring standard; Reading the target sequence in the created target sequence; According to the target sequence scoring standard, analyzing and judging the sequence characteristics of the target sequence, giving a corresponding score value, and outputting the target sequence and the score value.
4. A system for designing a target sequence, characterized in that, The system comprises: A judging and reading module for judging when there is a GG or AG base on the target sequence, reading the base and the sequence before the base, the length of the sequence being m bp, wherein m≤50. An editing module is configured to edit a target sequence model with the sequence as the target sequence model, to construct a target sequence, wherein the editing includes cleavage, input sequence information or reading sequence files; A reading and output module is configured to analyze and judge sequence features of the target sequence according to the target sequence scoring standard, to output the target sequence and the score value after assigning a corresponding score value.
5. The system for designing a target sequence according to claim 4, wherein, The judging and reading module includes: A first reading unit is configured to read a target sequence; A first judging unit is configured to judge whether the target sequence contains GG or AG bases; A first marking unit is configured to mark the position of the GG (or AG) base when the target sequence contains the GG or AG base; A second reading unit is configured to read the base and the sequence before the base along the target sequence with the base as the base point, and the length of the sequence is m bp, wherein m≤50.
6. The system for designing a target sequence according to claim 4, wherein, The editing module includes: A first editing unit is configured to input sequence information of a restriction enzyme or read sequence files of the restriction enzyme with the sequence as the target sequence model, to obtain a first preset target sequence; A second editing unit is configured to input sequence information of a connection vector or an expression cassette or read sequence files of the connection vector or sequence files of the expression cassette with the first preset target sequence as the target sequence model, to obtain a second preset target sequence, wherein the expression cassette is composed of a promoter, a gene or sequence to be expressed and a terminator; A third editing unit is configured to input sequence information for multi-fragment connection or read sequence files for multi-fragment connection with the second preset target sequence as the target sequence model, to obtain a third preset target sequence, wherein the sequence for multi-fragment connection includes a recognition or cleavage sequence of a type II endonuclease or a sequence that can be recognized and cleaved by an exogenous or endogenous system; A first obtaining unit is configured to obtain the constructed target sequence after editing is completed.
7. The system for designing a target sequence according to claim 4, wherein, The reading and output module includes: A third reading unit is configured to read a target sequence scoring standard; A fourth reading unit is configured to read a target sequence in the created target sequence; A first output unit is configured to analyze and judge sequence features of the target sequence according to the target sequence scoring standard, to output the target sequence and the score value after assigning a corresponding score value.
8. An electronic device, comprising: The electronic device includes: A memory is configured to store at least one instruction; and A processor is configured to execute the instruction stored in the memory to implement the method for designing a target sequence according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for designing point mutation model based on CRISPR-Cas9 technology
CN111508558A