Design and assembly method of an ultra-long flexible gRNA array and a multi-target editing system composed thereof

By adding VOX sequences to the gRNA array and selecting non-repetitive elements, the problem of difficulty and low efficiency of assembly of gRNA arrays in the prior art is solved, and efficient multi-target editing in Saccharomyces cerevisiae is achieved, and the stability and assembly simplicity of the array are improved.

CN116254265BActive Publication Date: 2025-06-24TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211513806.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-06-24
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

Existing multi-target editing technologies such as MAGE and CRISPR-based technologies have limitations in efficiency and gRNA array assembly, especially when synthesis of stable ultra-long gRNA arrays in yeast, there are problems with high repetition and difficult assembly.

Method used

By adding VOX sequences to the gRNA array and selecting non-repetitive elements to reduce the overall repetition level, an ultra-long flexible gRNA array was designed and assembled, and array design and assembly of 168 gRNAs and even more gRNAs was achieved.

Benefits of technology

The stability and simplicity of assembly of gRNA arrays are improved, and simultaneous editing of up to 104 different sites in Saccharomyces cerevisiae is achieved, and the editing sites within single bacteria are iteratively improved to 113.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0004230909900000061
    Figure GDA0004230909900000061
  • Figure GDA0004230909900000071
    Figure GDA0004230909900000071
  • Figure GDA0004230909900000141
    Figure GDA0004230909900000141
Patent Text Reader

Abstract

The present invention relates to the field of biotechnology, and in particular to a design and assembly method of an ultra-long flexible gRNA array and a multi-target editing system composed thereof. By designing the structure of the gRNA array and selecting non-repetitive elements therein, the present invention reduces the overall repetition degree, thereby reducing the synthesis difficulty of the gRNA array and improving the stability of the array. The present invention can achieve the design and assembly of arrays of 168 gRNAs or even more gRNAs. At the same time, the addition of vox sequences in the array makes the array flexible and variable, enabling random deletion and replication of rearrangement units in the ultra-long flexible gRNA array, and finally realizing regional multi-target editing of the genome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and particularly to a method for designing and assembling an ultra-long flexible gRNA array and a multi-target editing system composed thereof. Background Art

[0002] With the development of genome writing technology, multiplex genome engineering has been widely used in fields such as multiplex genome editing and rapid modification of metabolic pathways. As a powerful molecular biology operation means, multi-target editing can greatly improve the scope and efficiency of gene editing and transcriptional regulation. Multiplex automated genome engineering (MAGE) and CRISPR-Cas9-based multi-target editing technologies are often used for multi-target editing. MAGE is an automated and high-throughput method that relies on the annealing and binding of oligonucleotide chains at target sites to precisely introduce mutations. However, its efficiency is limited by the heterologous expression efficiency of single-stranded DNA binding proteins and the increase in the global genomic mutation rate caused by the inhibition of DNA mismatch repair. Different from the relatively high degree of host specificity of the MAGE system, the CRISPR-Cas9 system can perform efficient genomic editing in a variety of organisms. The CRISPR-Cas9-based multi-target editing technology has the characteristics of strong targeting and no host dependence, so this technology has been widely used in fields such as genome modification and metabolic pathway regulation. In addition, due to the simple structure of crRNA, the CRISPR-Cas12a system has also been used in multi-target editing. However, the assembly of a large number of repetitive structures in the gRNA array poses challenges for its application in the multi-target field.

[0003] The number of edits of the CRISPR-Cas9-based multi-target editing technology depends on the number of gRNAs in the system. The application of an ultra-long gRNA array can effectively increase the number and stability of multiple different target edits achieved by CRISPR-Cas9 in a single strain. Currently, there are two ways to express multiple gRNAs, including (i) constructing multiple plasmids containing single gRNA expression cassettes; (ii) polycistronic expression of gRNAs. The design of the gRNA array based on the polycistronic expression of gRNAs can effectively improve the editing scope and efficiency of the CRISPR-Cas9-based multi-target editing technology. However, the number of type III promoters in yeast is limited, and at the same time, the gRNA scaffold sequence cannot be changed, resulting in a very high degree of repetition of the generated gRNA array. Therefore, how to synthesize a stable ultra-long gRNA array for multi-target editing at different sites in a single bacterium is a key factor.

[0004] Currently, researchers can use a single gRNA targeting multi-copy sites to achieve tens of thousands of genome edits. Editing of multiple specific sites can achieve the replacement of stop codons in 33 essential genes and the deletion of 25 Porcine endogenous retroviruses (PERVs). By assembling and expressing different gRNA expression cassettes, the assembly of 22 gRNAs can be achieved in vivo and the expression of 13 genes in the large intestine can be inhibited.

[0005] However, the efficiency of existing multi-target editing technologies such as MAGE is limited by the heterologous expression efficiency of single-stranded DNA-binding proteins and the increased global genomic mutation rate caused by the inhibition of DNA mismatch repair. Multi-target editing based on the CRISPR system is limited by the length of the gRNA array. Due to the high degree of element repetition in the gRNA array, challenges are posed to the assembly of the array. At the same time, regional editing of the genome diversifies the genome editing regions, promoting the research of multi-target editing in aspects such as gene interaction and metabolic regulation. Summary of the Invention

[0006] In view of this, the present invention provides an ultra-long flexible gRNA array, its design and assembly methods, and a multi-target editing system composed thereof. By adding vox sequences and selecting non-repetitive elements in the gRNA array, the present invention reduces the overall repetition degree, thereby reducing the synthesis difficulty of the gRNA array and improving the stability of the array. The design and assembly of arrays with 168 gRNAs or even more gRNAs can be achieved, realizing

[0007] To achieve the above invention objectives, the present invention provides the following technical solutions:

[0008] A gRNA transcription unit, comprising a promoter, a tRNA, a gRNA of a target sequence, a VOX sequence, and a terminator;

[0009] The VOX sequence is as shown in SEQ ID NO:1.

[0010] In the present invention, the number of gRNAs of the target sequence is k, 3 ≤ k ≤ 6; the gRNAs of the target sequence include: a guide sequence complementary to the target sequence, a scaffold sequence, and a termination sequence; in some embodiments, the number of gRNAs of the target gene is 3, specifically including gRNA1, gRNA2, and gRNA3, and the scaffold sequences included in gRNA1, gRNA2, and gRNA3 are different and are arbitrarily selected from 3 synthetic non-repetitive gRNA scaffold libraries. Further, the scaffold sequence is selected from the sequences shown in SEQ ID NO:23 to 25, and the termination sequence is as shown in SEQ ID NO:26.

[0011] The number of tRNAs in each gRNA transcription unit is 4 to 7, and the tRNAs can be arbitrarily selected from a sequence library composed of 21 yeast tRNAs. Specifically, the sequences of the tRNAs are shown in SEQ ID NO: 2 to 22.

[0012] In the gRNA transcription unit, the promoter can be arbitrarily selected from a synthetic yeast type II promoter library [1] and the terminator can be arbitrarily selected from a sequence library composed of 12 synthetic terminators (the sequences are shown in SEQ ID NO: 27 to 38).

[0013] In some specific embodiments, when the number of gRNAs in each gRNA transcription unit is 3, the number of tRNAs is 4, specifically including 4 different tRNAs: tRNA1, tRNA2, tRNA3, and tRNA4, which are respectively selected from different tRNAs shown in SEQ ID NO: 2 to 22; the gRNA transcription unit from the 5' end to the 3' end is in turn: promoter - tRNA1 - gRNA1 - VOX sequence - tRNA2 - gRNA2 - VOX sequence - tRNA3 - gRNA3 - VOX sequence - tRNA4 - terminator, and the structural schematic diagram is as Figure 14 . In the gRNA transcription unit, gRNA1, gRNA2, and gRNA3 only represent three different gRNAs, and do not represent the order of the gRNAs in the transcription unit, and their order can be arbitrary.

[0014] The present invention also provides a gRNA array composed of the gRNA transcription units of the present invention.

[0015] In some embodiments, the gRNA array of the present invention includes x intermediate fragments, the x intermediate fragments altogether contain m primary fragments, each primary fragment contains 3 to 5 gRNA transcription units, and each transcription unit contains k gRNAs, and the m primary fragments are divided into X intermediate fragments evenly or unevenly; wherein, x ≥ 1, m ≥ 1, k ≥ 3, and x, m, and k are all integers.

[0016] The present invention provides a method for designing and assembling the gRNA array described above, including:

[0017] Step (1): Design and obtain n gRNAs according to the target gene, and divide them into n / k gRNA transcription units as described in any one of claims 1 to 5, each transcription unit contains k gRNAs, k ≥ 3, n ≥ 3, and n and k are all integers;

[0018] Step (2): Every 3 to 5 gRNA transcription units form a first-level fragment, and the n / k gRNA transcription units are evenly divided into m first-level fragments;

[0019] Step (3): The m first-level fragments are grouped evenly or unevenly, and each group of first-level fragments forms an intermediate fragment; The following design is carried out on the first-level fragments to obtain first-level assembly fragments:

[0020] Homologous arms and restriction enzyme cleavage sites are added to both ends of each first-level fragment. The same homologous arms are added to the 3' end of the previous first-level fragment and the 5' end of the next first-level fragment among adjacent two first-level fragments, so that the first-level fragments are integrated by enzyme digestion and Gibson assembly; The homologous arms added to the 5' end of the first first-level fragment and the 3' end of the Xth first-level fragment are vector homologous arms, and the remaining homologous arms are random homologous sequences;

[0021] Step (4): Synthesize m first-level assembly fragments respectively, and obtain the full-length sequence of the gRNA array through in vitro enzyme digestion and Gibson assembly.

[0022] Among them, in step (1), the n / k gRNA transcription units can be initiated by more than one promoter, or can be initiated by different promoters respectively. Preferably, different promoters are used for initiation, that is, the n / k gRNA transcription units are initiated by n / k different promoters respectively;

[0023] In step (2), the n / k gRNA transcription units are evenly divided into m first-level fragments; Every 3 to 5 gRNA transcription units form a first-level fragment, specifically it can be 3, 4 or 5, and the specific number can be determined according to the total number and full length of the gRNAs.

[0024] In step (3), the m first-level fragments are evenly or unevenly divided into x intermediate fragments; When the m first-level fragments cannot be equally divided into several intermediate fragments, they can be divided unevenly. For example, in the specific embodiment of the present invention, when the number of gRNAs is 168, the gRNA array contains 14 first-level fragments, each first-level fragment contains 4 gRNA transcription units, each transcription unit contains 3 gRNAs, that is, each first-level fragment contains 12 gRNAs; The 14 first-level fragments are divided into four combinations, that is, the first intermediate fragment includes: the 1st to 4th first-level fragments, the second intermediate fragment includes: the 5th to 8th first-level fragments; The third intermediate fragment includes: the 9th to 12th first-level fragments, and the fourth intermediate fragment includes: the 13th to 14th first-level fragments.

[0025] After grouping the gRNAs according to the above method of the present invention, the following design also needs to be carried out on the first-level fragments for the synthesis and assembly of the first-level fragments. Taking the grouping schemes with the number of gRNAs being 30 and 168 as examples for illustration:

[0026] (1) The gRNA array contains 30 gRNAs with a full length of 7,821 bp. The structure is as shown in Figure 15 . The number of gRNAs, n = 30, and k = 3, that is, each transcription unit contains 3 gRNAs. The gRNA array contains 10 gRNA transcription units. Every 5 gRNA transcription units form a first-level fragment, and this gRNA array contains two first-level fragments in total. The following designs are made for the two first-level fragments. See Figure 19 :

[0027] Add two homologous arms, A and B, to both ends of the first first-level fragment respectively, and add two homologous arms, B and C, to both ends of the other first-level fragment respectively. Among them, the A and C homologous arms are the homologous arms at both ends of the vector (such as the pCCI vector). At the same time, add appropriate restriction enzyme sites (such as EcoRI, XhoI, NotI) at the ends of the homologous arms. The two first-level assembly fragments obtained through the above design are synthesized by a gene synthesis company respectively.

[0028] Use restriction endonucleases to digest the two synthesized first-level assembly fragments (after synthesis by the gene company and constructed into plasmids, that is, the first-level assembly fragments are actually in the form of plasmids), and through Gibson assembly, use the homologous arm B to integrate the two first-level fragments, and use the homologous arms A and C at both ends to ligate the integrated fragment to the pCCI-LEU plasmid in vitro. The specific assembly process is shown in Figure 16 .

[0029] (2) When the number of gRNAs is 168, the full length of the array is 50,530 bp, containing 14 first-level fragments. Each first-level fragment contains 4 gRNA transcription units, and each transcription unit contains 3 gRNAs, that is, each first-level fragment contains 12 gRNAs; the 14 first-level fragments are divided into four combinations, P1, P2, P3, and P4. That is, the first-level fragments in P1 include: the 1st to 4th first-level fragments, the first-level fragments in P2 include: the 5th to 8th first-level fragments; the first-level fragments in P3 include: the 9th to 12th first-level fragments, and the first-level fragments in P4 include: the 13th to 14th first-level fragments.

[0030] Add homologous arms to both ends of the first-level fragments. Among them, add the homologous arms A and L of the vector (such as the pCCI vector) to the 5' end of the first first-level fragment and the 3' end of the 14th first-level fragment respectively, and add homologous arms with random sequences, such as homologous arms numbered B to N, at the remaining positions. At the same time, add appropriate restriction enzyme sites (such as EcoRI, XhoI, NotI) at the ends of all homologous arms. See Figure 18 , and the 14 first-level assembly fragments obtained through the above design are respectively labeled as:

[0031] The P1 intermediate assembly fragments include: 1-4 first-level assembly fragments; the P2 intermediate assembly fragments include: 5-8 first-level assembly fragments; the P3 intermediate assembly fragments include: 9-12 first-level assembly fragments; the P4 intermediate assembly fragments include: 13-14 first-level assembly fragments. The 14 first-level assembly fragments were respectively synthesized by a gene synthesis company.

[0032] During assembly, first, assemble using the intermediate assembly fragments as units. Digest the 1-4 first-level assembly fragments once by enzyme digestion and perform one Gibson assembly to obtain the intermediate assembly fragment P1, and ligate it into the pUC57 plasmid. Obtain the pUC57 recombinant plasmids containing P1, P2, P3, and P4 respectively according to this method; then perform one more enzyme digestion and one Gibson assembly on the four pUC57 recombinant plasmids to finally ligate the entire gRNA full-length sequence into the pCCI-LEU plasmid to complete the splicing of the entire gRNA array in vitro.

[0033] In a specific embodiment, the present invention verified 13 ligation sites (a total of 13 ligation sites in 14 first-level fragments), indicating that all fragments were correctly assembled (see Figure 19 ). After that, the plasmid was completely extracted from Escherichia coli, and the plasmid containing specific restriction enzyme sites was cut into multiple fragments of different lengths to determine whether the length of the plasmid was correct. By cutting with KpnI and XbaI, the plasmid was cut into fragments of different lengths. Agarose gel electrophoresis analysis showed that the length of the extracted plasmid and the sequence of the corresponding restriction enzyme sites were the same as the design ( Figure 20 ). At the same time, the present invention used a double enzyme digestion experiment with XhoI and BamHI to cut the gRNA array completely from the plasmid, and the specific recognition and cleavage of BamHI could linearize the complete plasmid. PFGE verification analysis showed that the gRNA array was the same length as the design and was completely ligated to the pCCI vector to form the plasmid pCCI-II-168gRNA containing 168 gRNAs ( Figure 21 ).

[0034] The intermediate fragments of P1-P4 were respectively obtained through one Gibson assembly and integration, and at the same time, the integrated intermediate fragments were respectively ligated into the pUC57 plasmid. Use restriction endonucleases to digest two first-level assembly fragments, perform Gibson assembly, integrate the two first-level fragments using the homologous arm B, and use the homologous arms A and C at both ends to ligate the integrated fragment into the pCCI-LEU plasmid in vitro. The specific assembly process is shown in Figure 18 .

[0035] The present invention also provides a recombinant vector comprising the gRNA transcription unit and the gRNA array described in the present invention. The present invention does not make special requirements on the backbone vector of the recombinant vector, and common types in the art can be used, including but not limited to PUC series vectors and pCCI series vectors. In some embodiments, the backbone vector of the recombinant vector is a PUC57 vector or a pCCI-LEU vector.

[0036] The present invention also provides a multi-target editing system, comprising a base editor, and the gRNA transcription unit, the gRNA array or the recombinant vector described in the present invention. In some embodiments, the multi-target editing system comprises a base editor and the gRNA array described in the present invention. The present invention does not have special limitations on the types of base editors, and common types in the art can be used. The gRNA array developed in the present invention can be adapted to a variety of Cas9 variants and applied to different scenarios, such as CRISPRa, CAISPRi, or for CRISPR / Cas9-based epigenetic modifications, etc. In some specific embodiments, the base editor is nCDA1Δ198-BE3, a base editor described in the literature "Engineering of high-precision base editors for site-specific single nucleotide replacement".

[0037] The present invention also provides the use of the gRNA transcription unit, the gRNA array, the recombinant vector, and the multi-target editing system in multi-target gene editing. In some embodiments of the present invention, the multi-target gene editing is carried out in yeast, specifically Saccharomyces cerevisiae.

[0038] In the present invention, there are no special requirements for the source of the vector homologous arms, and common vector types in the art can be used. For the specific sequence of the vector homologous arms, it depends on the selected vector type, and generally, sequences of 20 - 60 bp (preferably 60 bp) at both ends of the vector insertion site are selected. The sequences of the remaining homologous arms other than the vector homologous arms are random sequences of 20 bp that are not homologous to the genome.

[0039] The present invention has the following beneficial effects:

[0040] 1. The present invention utilizes the tRNA-gRNA structure, non-repetitive elements, and vox sequences to design and assemble an ultra-long flexible array with a length of 50,530 bp containing 168 specific gRNAs.

[0041] 2. By adding vox sequences to the gRNA array, the flexibility of the array is achieved.

[0042] 3. Using an ultra-long flexible gRNA array and a base editor, simultaneous editing of up to 104 different sites was achieved in Saccharomyces cerevisiae, and the number of edited sites in a single bacterium was increased to 113 through two rounds of iteration.

[0043] 4. By rearranging the ultra-long flexible gRNA array, multi-target editing of different regions of the genome was achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram showing the base editor editing glutamine and arginine codons into stop codons;

[0045] Figure 2 Schematic diagram showing the open reading frame structure of the base editor;

[0046] Figure 3 Schematic diagram showing the base editing of ADE2 in Saccharomyces cerevisiae by the base editor;

[0047] Figure 4 Schematic diagram showing the editing efficiency of the base editor under different expression intensities and integration forms;

[0048] Figure 5 Schematic diagram showing the optimization results of the base editor expression conditions;

[0049] Figure 6 Schematic diagram showing the structure of constructing a non-repetitive gRNA plasmid using Gibson assembly;

[0050] Figure 7 Schematic diagram showing the efficiency of non-repetitive gRNA base editing in yeast;

[0051] Figure 8 Schematic diagram showing the two-dimensional structure of gRNA;

[0052] Figure 9 Schematic diagram showing the effect of gRNA structure changes on base editing efficiency;

[0053] Figure 10 Schematic diagram showing the gRNA array structure for verifying the cleavage effect of 21 tRNAs;

[0054] Figure 11 Schematic diagram showing the effect of three gRNA arrays driving base editing;

[0055] Figure 12 Schematic diagram showing the effect of different expression modes of the base editor on the number of multi-target editing in a single bacterium;

[0056] Figure 13 Schematic diagram showing the recognition and cleavage of tRNA by RNase P and RNase Z;

[0057] Figure 14 Schematic diagram showing the gRNA transcription unit in the gRNA array;

[0058] Figure 15 Showing the design of a gRNA array containing 30 gRNAs;

[0059] Figure 16 Showing the assembly of a gRNA array containing 30 gRNAs, where the green line represents the sgRNA array, A, B, and C represent different homologous arms, A and C represent the homologous arms at both ends of the vector, and B is a homologous arm with a random sequence;

[0060] Figure 17 Showing the verification of a gRNA array containing 30 gRNAs;

[0061] Figure 18 Showing the assembly process of a gRNA array containing 168 gRNAs, where the green line represents the first - level fragment, the same letters represent homologous arms with the same sequence, different letters represent homologous arms with different sequences, A and L are the homologous arms at both ends of the vector, and the remaining homologous arms are random homologous sequences;

[0062] Figure 19 Showing the PCR verification of a gRNA array plasmid containing 168 gRNAs;

[0063] Figure 20 Showing the restriction enzyme digestion verification of a gRNA array plasmid containing 168 gRNAs with KpnI and XbaI;

[0064] Figure 21 Showing the restriction enzyme digestion verification of a gRNA array plasmid containing 168 gRNAs with BamHI and XhoI;

[0065] Figure 22 Showing the Sanger sequencing verification of 12 sites in the plasmid;

[0066] Figure 23 Showing the sequencing depth of all sites in the plasmid;

[0067] Figure 24 Showing the sequencing length of DNA fragments in ONT sequencing;

[0068] Figure 25 Showing the verification of the base sequence at position 39121 in the plasmid;

[0069] Figure 26 Showing the phenotypic changes of the strain after multi - target editing based on a gRNA array containing 30 gRNAs;

[0070] Figure 27 Showing the multi - target editing effect of a gRNA array containing 30 gRNAs;

[0071] Figure 28Show the base editing status of all target sites in yHX0366;

[0072] Figure 29 Show the colony phenotype verification of multi-target editing mediated by 168 gRNAs;

[0073] Figure 30 Characterize the editing status of 168 sites in 31 strains;

[0074] Figure 31 Characterize the editing efficiency of different sites in 31 strains;

[0075] Figure 32 Show the whole-genome sequencing analysis of the second-round multi-target editing strains;

[0076] Figure 33 Show the PCRtag verification of gRNA array rearrangement;

[0077] Figure 34 Show the whole-genome sequencing verification of gRNA array rearrangement;

[0078] Figure 35 Show the whole-genome sequencing verification of the induced editing strains. Detailed implementation manners

[0079] The present invention provides an ultra-long flexible gRNA array and its design and assembly methods. Those skilled in the art can draw on the content of this article and appropriately improve the process parameters to achieve. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art, and they are all regarded as included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously make changes or appropriate alterations and combinations to the methods and applications in this article without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.

[0080] The test materials used in the present invention are all ordinary commercially available products and can be purchased in the market.

[0081] The sequences involved in this article are shown in Table 1:

[0082] Table 1

[0083]

[0084]

[0085] The following further elaborates the present invention in conjunction with embodiments: Embodiment 1

[0086] 1.1 Selection of gRNA array elements

[0087] 1.1.1 Construction of the gRNA array element screening system

[0088] Evaluation of the editing effect of a multi-target editing system based on an ultra-long flexible gRNA array requires effective visualization means for characterization. Previously, for the CRISPR system, researchers developed techniques such as chromosome cutting, gene silencing, base editing, and gene drive. For the evaluation of multi-target editing and the evaluation of gRNA array element screening, there should be convenient and efficient detection means for the editing sites after the editing event occurs. The base editor uses a deaminase fused with the Cas9 protein to achieve base editing within a specific editing window. The change of bases in the DNA sequence can be detected by sequencing and does not change with the time series. In addition, under the action of cytidine deaminase, cytosine (C) will become thymine (T), so that the codons encoding glutamine (CAA; CAG), arginine (CGA) can be transformed into stop codons (TAA; TAG; TGA) under the action of the base editor, causing nonsense mutations and prematurely terminating the synthesis of the peptide chain, achieving the purpose of gene inactivation ( Figure 1 ). Therefore, in this study, the base editor was used as a tool for multi-target editing, and the effectiveness of the gRNA array linker element and the number of edits in a single strain were characterized by counting the gene inactivation and the number of base changes.

[0089] In the CBE base editor, it generally contains the Cas gene and the cytidine deaminase and base excision repair inhibitor fused with it. For the content of this study, the required base editor should have the characteristics of high editing efficiency and wide editing range, so as to be able to design gRNAs for more targets. CDA1 is an ortholog of AID from lamprey. Compared with the previously reported cytidine deaminase APOBEC1, this homolog has better C-to-T base editing activity in a specific DNA sequence. In addition, compared with the forms of Cas9 and nCas9, dCas9 is more suitable for the multi-target editing system because it does not cause double-strand breaks and the strains with multiple edits cannot grow. In the follow-up of this study, the base editor nCDA1Δ198-BE3 was used for base editing.

[0090] Construction process of the CRISPR base editing system: The nCDA1Δ198-BE3 optimized by Saccharomyces cerevisiae codons was synthesized by General Biosystems and restriction enzyme sites of KpnI and NotI were reserved at both ends of the gene. After digestion, it was ligated with the pRS413 vector obtained by PCR amplification to obtain the expression plasmid pRS413-P TEF1 -nCDA1Δ198-BE3-T CYC1。The deletion of ADE2 in Saccharomyces cerevisiae leads to the accumulation of the intermediate phosphoribosylaminoimidazole in the cell. After oxidation, this intermediate forms a red pigment, causing the colony color to change from white to red. The CRISPR targeting sequence (5'-TCAACTTAAGGCGAAGTTGT TGG -3', SEQ ID NO:39) targets the codon for glutamine (Gln) at position 233 of ADE2 in Saccharomyces cerevisiae. The C-to-T base editing can form a TAA stop codon, prematurely terminating the translation process of ADE2 and causing a change in the yeast colony color. The screening of the gRNA array components below is all based on counting the efficiency of colony turning red to characterize the effect of base editing.

[0091] Yeast carrying the plasmid with the CRISPR base editing system undergoes base editing in vivo. After 72 hours of culture, red colonies are formed on the corresponding screening plates, indicating the success of base editing ( Figure 3 ). Next, we tested the editing efficiency of the base editor under different expression intensities and integration forms, including expression on low-copy and high-copy plasmids and inducible expression integrated into the genome. The experimental results showed that the change in the expression form of the base editor did not result in an obvious change in the base editing efficiency ( Figure 4 ). Considering the plasmid instability and the different requirements for expression intensity due to different numbers of target sites, all subsequent base editors in this study were integrated between YOR072W and YOR073W on chrXV and induced for expression using galactose in the absence of glucose by the GAL1 promoter.

[0092] 1.1.2 Optimization of the gRNA sequence

[0093] The sequences of gRNAs in the gRNA array are relatively conserved. As a type II CRISPR system, CRISPR-Cas9 has its trans-activating tracrRNA hybridize with crRNA, thereby guiding the Cas9 protein to specifically cleave specific sites. In addition, the crRNA containing the target sequence fuses with tracrRNA, and the formed complete gRNA sequence can achieve specific cleavage of the genome. In previous studies, non-repetitive gRNA libraries were developed through methods such as biochemical modeling, biochemical characterization, and machine learning of the gRNA sequence, reducing the repetition degree of the gRNA array and making the array synthesis more difficult. However, no such attempts have been made in yeast.

[0094] We amplified the pRS42H vector using PCR and assembled the newly synthesized gRNA scaffold (handle) sequence into the vector using Gibson assembly, including a 20bp target sequence targeting ADE2 (5'-TCAACTTAAGGCGAAGTTGTTGG -3', SEQ ID NO:39) Figure 6 )。By inducing changes in yeast phenotypes after base editing, the editing effect and the effectiveness of the gRNA scaffold were determined. The newly constructed gRNA plasmid can be confirmed for its sequence by Sanger sequencing.

[0095] We constructed 28 non-repetitive gRNA scaffolds reported previously into vectors and induced the editing of ADE2 in yeast cells integrated with base editors. By counting the proportion of the number of strains with phenotypic changes in the total number of all strains, the effects of non-repetitive gRNAs were determined. For each non-repetitive gRNA, three parallel experiments were conducted. The results showed that most non-repetitive gRNAs did not function in yeast. However, for non-repetitive gRNAs numbered 1, 5, 8, 22, 34, and 37, they could bind to the base editor in yeast to play a role in targeting genes and base editing, but the efficiency was generally low (<40%). Figure 7 )。

[0096] Therefore, we compared the differences between non-repetitive gRNAs and the gRNAs used in the laboratory. Different from non-repetitive gRNAs, the gRNAs used in the laboratory have an additional 15-bp DNA sequence that is considered a terminator in the text. Figure 8 )。Therefore, we speculated that the addition of this sequence could correspondingly improve the overall editing efficiency. We selected three gRNAs with higher editing efficiencies numbered 1, 22, and 37 (the sequences are shown as SEQ ID NO:23 - 25 respectively) for subsequent research. The experimental results showed that after adding the 15-bp terminator sequence, the editing effect on the target was improved, and the efficiency of single editing was increased to more than 75%. Figure 9 )。This result indicates that the 76-bp scaffold sequence is particularly important for the binding and specific cleavage of gRNA to Cas9 in yeast. Therefore, in subsequent research, we will use the above three non-repetitive gRNAs as expression scaffolds to reduce the repetition degree of the entire gRNA array.

[0097] 1.1.4 Characterization of the Cleavage of gRNA Arrays by 21 tRNAs

[0098] Based on the above research results, tRNA was proven to be an effective gRNA flanking cleavage element, and the combination of tRNA-gRNA-tRNA can effectively release mature gRNA for the CRISPR system. Using the endogenous tRNA in Saccharomyces cerevisiae GlyCleaving the gRNA flanks enables editing of eight different sites. However, due to the large number of repeats of tRNA and the gRNA scaffold, the synthesis of larger-scale gRNA arrays encounters difficulties. There are 21 different types of tRNA in Saccharomyces cerevisiae for the transport of different amino acids, and ten of them have been shown to be able to effectively cleave the gRNA flanks.

[258] . To utilize more non-repetitive tRNA elements and reduce the repetition degree of the entire array, we need to characterize the cleavage ability of various tRNAs in yeast. Similar to the array structure designed for selecting cleavage elements before, we divided the 21 tRNAs into three groups and distributed them on both sides of the gRNA. A total of four target sequences (5'-TCAACTTAAGGCGAAGTTGT TGG -3', SEQ ID NO:39; 5'-CAAAGGCTGAACTACATTAC AGG -3', SEQ ID NO:40; 5'-TGTCGCTCAAAAGTTGGACT TGG -3', SEQ ID NO:41; 5'-ATCAAATCTTTTCCCGGTTG TGG -3', SEQ ID NO:42) targeting Gln at position 233, Gln at position 374, Gln at position 393, and Ile codon at position 244 of ADE2 were selected for the characterization of tRNA cleavage ability ( Figure 10 ).

[0099] Among them, due to the limitation of the number of tRNAs, the third gRNA array (tRNA(16-17)) only targets the first three sites and reuses one tRNA once Gly ( Figure 10 ).

[0100] All the arrays were introduced into yHX0362 and induced in galactose medium for 24 hours. After culturing on the plate for 72 hours, the proportion of colony color changes was counted to characterize the efficiency of base editing. The results showed that for the base editing induced by the gRNA arrays with three different tRNAs as the linking elements, the efficiency could reach more than 50% ( Figure 11 a). At the same time, we randomly selected 16 strains from each of the three groups of experiments for PCR amplification of ADE2 and used Sanger sequencing to determine the editing of all sites. As Figure 11As shown in b, all arrays were edited at their target sites, and the effects were not less than 50%. This indicates the effectiveness of the gRNA array structure and that for all tRNAs, they can be used as elements for gRNA flanking cleavage, expanding the number of non-repetitive elements in the gRNA array. At the same time, we compared the effects of driving base editor expression by two methods (constitutive expression in low-copy plasmids; inducible expression by chromosomal integration) on multi-target editing. It was found that compared with constitutive expression in low-copy plasmids, more sites were edited in a single strain when the base editor was induced by galactose ( Figure 12 ).

[0101] Therefore, by characterizing the editing effects of gRNA arrays composed of three different tRNAs, it was shown that all 21 tRNAs can be cleaved in the gRNA array transcript and can be used in subsequent array designs.

[0102] tRNA precursors are recognized and cleaved by RNase P and RNase Z in vivo, thereby removing the 5' and 3' terminal sequences. RNase Z makes an accurate cleavage at the 3' end of the tRNA at the junction with the gRNA without disrupting the gRNA structure. At the 5' end of the tRNA, the cleavage by RNase P occurs within the 5' leader sequence. At the same time, previous studies have also demonstrated that the presence of the leader sequence can effectively enhance the processing ability of RNase P. Therefore, the respective leader sequences will be added to the 5' ends of tRNAs during the gRNA array design process ( Figure 13 ). And all arrays are composed of transcription units as shown in Figure 14 . In this unit, the two ends of the gRNA are connected by different types of tRNAs, and a vox sequence is added at the 3' end for array rearrangement. Therefore, during the subsequent gRNA design process, each transcription unit will contain three gRNAs to ensure that all gRNAs can be effectively transcribed. Thus, we determined the basic structure of the gRNA array and the composition of various types of elements, thereby reducing the repetition degree of the entire array.

[0103] 1.2 Design and assembly of gRNA arrays containing 30 gRNAs

[0104] First, we designed a proof-of-concept experiment to verify the feasibility of the array design. We selected 30 target sites from five genes in the ADE2 and prodigiosin metabolic pathways for the design of the gRNA array. These target sites have at least one cytosine within the range of -20 to -15 bases from the PAM site to characterize the occurrence of editing events. Since SpCas9 was used, the PAM sequence of all target sites is NGG. Among them, gRNA target sites that can cause nonsense mutations were designed on each gene, which can prematurely terminate the translation of gene transcripts in the ADE2 or prodigiosin metabolic pathway, thereby causing the colonies to exhibit corresponding phenotypes. The design of the gRNA array was carried out according to the above principles. A total of 10 transcription units, 10 synthetic promoters, 21 tRNA elements, 3 gRNA scaffolds, and 10 synthetic terminators participated in the design of the entire gRNA array ( Figure 15 ).

[0105] The designed array has a full length of 7821 bp. To reduce the difficulty of array synthesis, we divided the entire array into two parts, with five transcription units as a group, and handed them over to a gene synthesis company for synthesis. At the same time, we added appropriate restriction enzyme sites (EcoRI, XhoI, NotI) to both ends of the synthesized fragments for the assembly of the fragments after synthesis. The two parts of the fragments were digested with restriction enzymes and assembled by Gibson assembly, and ligated to the pCCI-LEU plasmid in vitro using the 20-bp homologous arms at both ends of the fragments ( Figure 16 ). Although the repeatability of the gRNA array has been reduced, due to the limited number of its elements, there is still a certain degree of repetition. And the length of the subsequently synthesized gRNA array is relatively long, so the gRNA array will ultimately be assembled onto the pCCI vector. The ligated plasmid was introduced into E. coli for verification and amplification. By designing PCR primers at the junction, it was verified whether the plasmid was ligated correctly. The experimental results showed that the junctions between the fragments and the junctions between the fragments and the plasmid were all ligated correctly, indicating the correct assembly of the gRNA array ( Figure 17 ). Therefore, based on the array design principles in Section 5.2, we designed and successfully assembled a gRNA array targeting genes in the ADE2 and prodigiosin metabolic pathways and containing 30 gRNAs.

[0106] 1.3 Design and Assembly of an Ultra-Long Flexible gRNA Array Containing 168 gRNAs

[0107] Next, we attempted to verify the effectiveness of the design and assembly method for gRNA arrays containing hundreds of gRNAs. The yeast genome contains a large number of non-essential genes, and their individual deletion will not cause the death of yeast strains. Due to the influence of gene interactions, the combination of specific essential genes shows synthetic lethality in yeast and cannot be deleted. Previous studies have shown that through the PCR-mediated chromosomal deletion (PCD) technique, the 744,843 - 613,184 region of chrII containing 38 non-essential genes was completely deleted without affecting the growth state of the strain. Therefore, we selected this region as the targeting region for 168 gRNAs. Through programming, we searched for target sequences with cytosine in the interval from -20 to -15 from the PAM site (NGG sequence) within this region. And all target sequences were dispersed as evenly as possible throughout the entire chromosome. At the same time, for ADE2, we designed three gRNAs that can prematurely terminate gene translation to characterize the occurrence of base editing events. According to the previously set gRNA array design principle, we selected a combination of 56 non-repetitive synthetic promoters, 3 gRNA scaffolds, 21 tRNAs, and 12 synthetic terminators for the design of the gRNA array. The overall length of the designed gRNA array reached 50,530 bp, so we divided the array into 14 first-level fragments, each fragment containing 4 transcription units and a total of 12 gRNAs. The 14 first-level fragments were separated into four combinations, and each combination could be integrated through one Gibson assembly. During this process, the fragments were assembled onto the pUC57 plasmid to facilitate subsequent plasmid extraction and digestion. All first-level fragments were about 3600 bp in length, and due to the low degree of repetition, they could be directly synthesized by the company. After that, the four plasmids containing the intermediate fragments obtained were cut through the pre-retained restriction sites to obtain four linearized intermediate fragments. Similarly, the intermediate fragments could be assembled onto the pCCI vector through Gibson assembly using the 60-bp homologous arms at both ends, completing the splicing of the entire gRNA array in vitro ( Figure 18 ).

[0108] We reserved PCR sites at both the 5' and 3' ends of the first-level fragments to characterize the correct ligation of the fragments. Verification of 13 fragment ligation sites showed that all fragments were correctly assembled ( Figure 19) Subsequently, we extracted the plasmid intact from Escherichia coli and cut the plasmid containing specific restriction sites into multiple fragments of different lengths to determine whether the length of the plasmid was correct. After digestion with KpnI and XbaI, the plasmid was cut into fragments of different lengths. Agarose gel electrophoresis analysis showed that the length of the extracted plasmid and the sequence of the corresponding restriction sites were the same as the design. Figure 20 ) Meanwhile, we used a double digestion experiment with XhoI and BamHI to cut the gRNA array intact from the plasmid, and the specific recognition and cleavage of BamHI could linearize the intact plasmid. PFGE verification analysis showed that the gRNA array was the same length as the design and was completely ligated to the pCCI vector, forming the plasmid pCCI-II-168gRNA containing 168 gRNAs. Figure 21 )

[0109] Next, to further verify whether the sequence of the constructed plasmid was the same as the design, we selected 12 sites on the plasmid for Sanger sequencing verification. By comparing with the control sequence, it was shown that all 12 sites were on the plasmid and the bases had not changed. Figure 22 ) Meanwhile, we verified the extracted plasmid by third-generation sequencing (ONT sequencing). The nanopore sequencing method can achieve de novo sequencing of all DNA molecules, so it is suitable for sequencing DNA fragments with complex structures. The sequencing results showed that the sequencing depth covered all sequences of the plasmid. Figure 23 ) And the ONT sequencing quality remained at a high level. Figure 24 ) By comparing the assembled sequence with the control, we found that the adenine (A) at position 39,121 in the plasmid was replaced by guanine (G). Combining Sanger sequencing to test this site, the results showed that the base at this site had not changed. Figure 25 )

[0110] So far, we have designed and correctly assembled a complete ultra-long flexible gRNA array containing 168 gRNAs. Different from the previous array synthesis methods, the array with reduced repetition can be synthesized in segments by the company, thus reducing the assembly difficulty of the array, providing a simple gRNA array design method for multi-target editing, and further expanding the manipulation scope of multi-target editing.

[0111] 2.1 Characterization of multi-target editing based on a gRNA array containing 30 gRNAs

[0112] We utilized the designed gRNA arrays to edit different genomic loci and determined the efficiency of multi-target editing and the effects of editing at different targets. The gRNA arrays targeting five genes in the ADE2 and prodigiosin metabolic pathways, which were synthesized and assembled previously, were introduced into yHX0365. Strains with correct introduction were picked and cultured under galactose induction for 24 hours, and then spread onto the corresponding plates. gRNAs capable of performing base editing to form nonsense mutations were designed for the genes in the ADE2 and prodigiosin metabolic pathways in the gRNA arrays, causing corresponding phenotypic changes in the strains ( Figure 26 ). During the process of picking strains, we determined the sequencing strains based on the phenotypic changes of the strains. Among them, colonies with red color indicated that nonsense mutations had occurred in both the ADE2 gene and the genes in the prodigiosin metabolic pathway in their genomes, proving the occurrence of multi-target editing events. We picked 8 strains with red color and verified all their targets by PCR amplification and Sanger sequencing. The number of editing times induced singly in a single strain was counted, as well as the editing efficiency for different targets. The experimental results showed that at least 19 / 30 loci were edited in each strain. And in all strains, 29 / 30 loci had at least one editing event ( Figure 27 ). Meanwhile, we noticed that the average number of edits in a single bacterium accounted for more than 75% of the total number of target sites, suggesting that more gRNA expression could achieve multi-target editing in the strain. This experiment also showed that after single induction editing, strain yHX0366 could achieve editing of up to 27 / 30 different sites in a single bacterium, laying a foundation for further increasing the number of different sites edited in a single bacterium in the future ( Figure 28 ).

[0113] 2.2 Characterization of multi-target editing based on ultra-long flexible gRNA arrays

[0114] 2.2.1 Characterization of the effect of genomic multi-target editing in a single bacterium

[0115] To evaluate the editing of different sites by the expression of hundreds of different gRNAs in a single bacterium, we introduced the correctly assembled plasmid containing the gRNA array into yHX0362 to obtain strain yHX0463, and induced the expression of the base editor to edit 168 targets. The gRNA array contained gRNAs that could cause nonsense mutations in ADE2, so the change in the color of the strain would be used to characterize the base editing event ( Figure 29)。By counting the number of color-changing strains, we found that the efficiency of ADE2 nonsense mutation in this experiment (28.0%) was significantly lower than the base editing efficiency of the same target site guided by a single gRNA (>80%). We speculate that this is because the expression level of gRNA in single bacteria is relatively high, while the number of induced base editors cannot mediate sufficient base editing of all gRNAs at the corresponding sites. In other studies, it has also been shown that increasing the expression level of dCas9 can improve the effect of gene transcriptional repression mediated by multiple gRNAs. At the same time, through iterative induced base editing of the strains to extend the action time of the base editor, the number of site edits in single bacteria can also be increased, which is described in detail in Subsection 2.2.2.

[0116] We randomly selected 31 single colonies with red-colored strains for whole-genome sequencing ( Figure 30 ). At the same time, the starting strain yHX0362 and three strains (yHX0392, yHX0550, yHX0552) obtained by inducing yHX0463 in galactose medium for 24 hours were used as controls and also subjected to whole-genome sequencing.

[0117] Next, the data of whole-genome sequencing of the test strains were analyzed. Single nucleotide variations (SNVs) in the control yHX0362 were removed from all experimental strains as negative mutations. The change of C to T bases from -20 to -15 within all 20bp target sequences was recorded as the base editing at this site, so as to count the number of site edits in a single strain after single editing and the editing efficiency of 168 sites in 31 strains. The experimental results showed that multiple edits occurred in the genomes of all experimental strains, and the median number of edits was 77, and the average number of edits was 72. Among the yHX456 strains, after single induced editing, base editing occurred at 104 different sites on the genome ( Figure 30 ). And in all experimental strains, the average editing efficiency of all sites was 42.72% ( Figure 31 ), and at least one edit occurred at 145 sites, indicating that most of the designed gRNAs can bind and edit at the corresponding targets.

[0118] 2.2.2 Iterative Editing of the Multi-Target Editing System

[0119] The multi-target editing technology often fails to achieve editing of all target sites in a single experiment. In the MAGE technology, by increasing the number of MAGE cycle operations, the average number of mutated bases in a single cell can be increased from 3.1 to 5.6 [2]。During the experiment of codon replacement in cells, the researchers increased the editing of 6 / 47 sites in a single bacterium by adding one more round of transfection and editing of the cells.

[15] 。In this study, through the induced editing of the strain yHX0463 containing the plasmid with an ultra-long flexible gRNA array in the early stage, single-time editing of 102 and 104 target sites was achieved in the induced strains yHX0455 and yHX456.

[0120] To characterize the change in the number of edits in a single bacterium with the increase in the number of induced edits, we re-inoculated yHX0455 and yHX456 into a galactose-induced medium and induced for 24 hours. Since the color of the strain had changed in the previous round of editing, in this round of editing, we randomly selected the obtained colonies and used whole-genome sequencing to detect the base changes in the strain genome. Through the analysis of the base sequences of 168 sites, we found that through the second round of induced editing, edits of new target sites appeared in all randomly selected single bacteria ( Figure 30 ). At the same time, most of the newly added edited sites occurred in the sites that had been edited before. Among them, the sites corresponding to the 89-gRNA and 138-gRNA that had not been edited before each had one editing event in these 9 strains. And, in the strains obtained through the second round of editing based on yHX0456, 113 target sites in the genomes of yHX0401 and yHX0402 were edited, further increasing the number of edits at different sites in a single bacterium ( Figure 32 ). Therefore, by increasing the number of induced edits, the number of edits in a single bacterium can be further increased.

[0121] 2.2.3 Rearrangement of ultra-long flexible gRNA array and multi-target random editing

[0122] Since 34bp loxPsym sites are designed downstream of the non-essential genes of the synthetic chromosomes of Saccharomyces cerevisiae, we can use SCRaMbLE to rearrange the previously synthesized synV and synX in the laboratory, causing gene duplication, deletion, translocation, inversion, etc. [3,4,5] 。Through SCRaMbLE, the diversity of strain genotypes can be quickly constructed, thus obtaining strains with different phenotypes. In this study, vox sequences were added to the designed and constructed ultra-long flexible gRNA array, which can undergo site-specific recombination under the action of the Vika protein, quickly forming the diversity of the gRNA array. The rich multi-target editing combinations can provide basic tools for exploring gene interactions in the genome, strain transformation, etc. [6,7,8] 。Therefore, in this section, we will use the expression of the Vika protein to construct a gRNA array library and induce and edit it to characterize the diversity of edited sites in a single bacterium.

[0123] First, construct the constitutively expressed plasmid pRS415-CYC1-Vika-EBD-tCYC1. In this plasmid, Vika is fused and expressed with the estrogen-binding domain (EBD) under the drive of the weak CYC1 promoter. The effect of Vika-EBD on vox is affected by estradiol. This fusion protein will not be folded and remains in the cytoplasm without estradiol, so it cannot contact the nuclear DNA.

[0124] Introduce the constructed plasmid pCCI-II-168gRNA and pRS415-CYC1-Vika-EBD-tCYC1 into yHX0362, and obtain the correctly introduced strain yHX0549 through screening. After overnight culturing of yHX0549, transfer 500 μL of the bacterial liquid into 5 mL of medium, and add 1 μM of estradiol to induce the rearrangement of the gRNA array. After 3 hours, dilute the bacterial liquid and spread it on the plate medium. After culturing for three days, select strains to analyze the rearrangement of the gRNA array in them.

[0125] First, use the 13 PCR tags designed on the gRNA array to characterize the rearrangement of the array by PCR. Six strains (yHX0465, yHX0466, yHX0467, yHX0468, yHX0469, yHX0470) were selected to characterize the structural changes of the gRNA arrays they carried. The results showed that deletions occurred in different regions of the gRNA array, and the lengths of the deletions were diverse, forming a library of a certain scale ( Figure 33 ). To further characterize the structural variations in different gRNA arrays, we performed whole-genome sequencing on the above six strains. We counted the sequencing depth of the gRNA array, and the results showed that deletions and duplications occurred in the sequences after rearrangement of the gRNA array. Among them, the gRNA array in yHX0468 was duplicated in the region of 5248-22,305 and deleted in the region of 15,853-18,528, forming a complex structural change of duplication superimposed on deletion. The gRNA arrays in yHX469 and yHX470 underwent at least three deletion events in different regions, forming gRNA arrays with three different combinations of target sites ( Figure 34) Subsequently, we induced editing in six strains. Rearrangement of the gRNA array in the strains might lead to the deletion of the three ADE2-gRNAs used to characterize multi-target editing. Therefore, we randomly selected the corresponding strains with color changes and without color changes, and used whole-genome sequencing to analyze the editing sites in the selected strains. The results showed that the different gRNA arrays generated after rearrangement caused changes in multi-target editing preferences, thus forming strains with different editing patterns. Figure 35 ) Therefore, the rearrangement of the gRNA array can drive the base editor to edit in different regions, thereby forming a large library of strains with different genotypes.

[0126] Example 2 Comparison of the editing effects of different editing systems

[0127] The editing effects of the editing systems reported in the prior art and the editing system of the present invention are shown in Table 2:

[0128] Table 2

[0129]

[0130]

[0131] [1] de Boer C G, Vaishnav E D, Sadeh R, et al. Deciphering eukaryotic gene-regulatory logic with 100 million random promoters[J]. Nature biotechnology, 2020, 38(1): 56-65;

[0132] [2] Wang H H, Isaacs F J, Carr P A, et al. Programming cells by multiplex genome engineering and accelerated evolution[J]. Nature, 2009, 460(7257): 894-898;

[0133] [3] Wu Y, Zhu R Y, Mitchell L A, et al. In vitro DNA SCRaMbLE[J]. Nature Communications, 2018, 9(1): 1935;

[0134] [4]Jia B,Wu Y,Li B Z,et al.Precise control of SCRaMbLE in synthetichaploid and diploid yeast[J].Nature Communications,2018,9(1):1933;

[0135] [5]Shen M J,Wu Y,Yang K,et al.Heterozygous diploid and interspeciesSCRaMbLEing[J].Nature Communications,2018,9(1):1934;

[0136] [6]Roy K R,Smith J D,Vonesch S C,et al.Multiplexed precision genomeediting with trackablegenomic barcodes in yeast[J].Nature Biotechnology,2018,36(6):512-520.

[0137] [7]Kuzmin E,VanderSluis B,Wang W,et al.Systematic analysis of complexgenetic interactions[J].Science,2018,360(6386):eaao1729;

[0138] [8]Boone C,Bussey H,Andrews B J.Exploring genetic interactions andnetworks with yeast[J].Nature Reviews Genetics,2007,8(6):437-449;

[0139] [9]Liu B,Jing Z,Zhang X,et al.Large-scale multiplexed mosaic CRISPRperturbation in the whole organism[J].Cell,2022,185(16):3008-3024.e16;

[0140]

[10] Reis A C, Halper S M, Vezeau G E, et al. Simultaneous repression of multiple bacterial genes using nonrepetitive extra-long sgRNA arrays[J]. Nature Biotechnology, 2019, 37(11): 1294-1301;

[0141]

[11] Campa C C, Weisbach N R, Santinha A J, et al. Multiplexed genome engineering by Cas12a and CRISPR arrays encoded on single transcripts[J]. Nature Methods, 2019, 16(9): 887-893;

[0142]

[12] Chen Y, Hysolli E, Chen A, et al. Multiplex base editing to convert TAG into TAA codons in the human genome[J]. Nature communications, 2022, 13(1): 1-13.

[0143] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A gRNA transcription unit, characterized in that, The gRNA transcription unit consists of a promoter, tRNA, gRNA of the target sequence, VOX sequence, and terminator; the gRNA transcription unit from the 5'-end to the 3'-end is in sequence: promoter - tRNA1 - gRNA1 - VOX sequence - tRNA2 - gRNA2 - VOX sequence - tRNA3 - gRNA3 - VOX sequence - tRNA4 - terminator; The VOX sequence is as shown in SEQ ID NO:1; The number of gRNAs of the target sequence is k, where 3 ≤ k ≤ 6; the number of tRNAs is 4 - 7; The gRNAs of the target sequence include gRNA1, gRNA2, and gRNA3, and the scaffold sequences included in gRNA1, gRNA2, and gRNA3 are different; the tRNAs include tRNA1, tRNA2, tRNA3, and tRNA4 with different sequences; The sequences of the tRNAs are selected from the sequences shown in SEQ ID NO:2 - 22; The gRNAs of the target sequence include: a guide sequence complementary to the target sequence, a scaffold sequence, and a termination sequence; The scaffold sequences are respectively selected from the sequences shown in SEQ ID NO:23 - 25, and the termination sequence is as shown in SEQ ID NO:

26.

2. A gRNA array, characterized in that, Composed of the gRNA transcription unit described in claim 1.

3. The gRNA array according to claim 2, wherein It includes x intermediate fragments, the x intermediate fragments altogether contain m first-level fragments, each first-level fragment contains 3 - 5 gRNA transcription units, each transcription unit contains k gRNAs, and the m first-level fragments are divided into X intermediate fragments evenly or unevenly; wherein, x ≥ 1, m ≥ 1, k ≥ 3, and x, m, and k are all integers.

4. The method for designing and assembling the gRNA array according to claim 2 or 3, characterized in that, It includes: Step (1): Design and obtain n gRNAs according to the target sequence, and divide them into n / k gRNA transcription units as described in claim 1, each transcription unit contains k gRNAs, k ≥ 3, n ≥ 3, and n and k are all integers; Step (2): Every 3 - 5 gRNA transcription units form a first-level fragment, and divide the n / k gRNA transcription units evenly into m first-level fragments; Step (3): Group the m first-level fragments evenly or unevenly, and each group of first-level fragments forms an intermediate fragment; perform the following design on the first-level fragments to obtain a first-level assembly fragment: Add homologous arms and restriction enzyme cleavage sites at both ends of each first-level fragment, add the same homologous arms at the 3'-end of the previous first-level fragment and the 5'-end of the next first-level fragment among adjacent two first-level fragments, so that the first-level fragments are integrated through enzyme digestion and Gibson assembly; the homologous arms added at the 5'-end of the first first-level fragment and the 3'-end of the Xth first-level fragment are vector homologous arms, and the remaining homologous arms are random homologous sequences; Step (4): Synthesize m first-level assembly fragments respectively, and through in vitro enzyme digestion and Gibson assembly, obtain the full-length sequence of the gRNA array.

5. The design and assembly method according to claim 4, characterized in that, n = 30, k = 3, the gRNA array contains 10 gRNA transcription units, every 5 gRNA transcription units form a first-level fragment, and altogether contain two first-level fragments; Alternatively, n = 168, k = 3, the gRNA array comprises 56 gRNA transcription units, and every 4 gRNA transcription units form a first-level fragment, with a total of 14 first-level fragments.

6. A recombinant vector comprising the gRNA transcription unit according to claim 1 and the gRNA array according to claim 2 or 3.

7. A multi-target editing system, characterized in that, Comprising: A base editor; And the gRNA transcription unit according to claim 1, the gRNA array according to claim 2 or 3, the gRNA array obtained by the design and assembly method according to claim 4 or 5, or the recombinant vector according to claim 6.

8. The multi-target editing system according to claim 7, wherein, The base editor is nCDA1Δ198-BE3.

9. Use of the gRNA transcription unit according to claim 1, the gRNA array according to claim 2 or 3, the gRNA array obtained by the design and assembly method according to claim 4 or 5, the recombinant vector according to claim 6, and the multi-target editing system according to claim 7 or 8 in the preparation of multi-target gene editing products.

Citation Information

Patent Citations

  • Yeast for large-scale gene rearrangement and construction method thereof

    CN113046255A

  • Methods and compositions for editing nucleotide sequences

    CN114127285A