Method for preparing mutant library
By arranging and combining classic mutation sites and repairing gaps, long double-stranded nucleic acid libraries were synthesized, solving the problems of inaccurate mutation locations and complex operations in DNA shuffling technology, and improving the screening efficiency and functional detection accuracy of mutant proteins.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CODER THERAPEUTICS CO LTD
- Filing Date
- 2025-11-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing DNA shuffling techniques struggle to precisely control mutation locations and types during mutant library preparation, resulting in numerous invalid mutations, time-consuming and costly screening processes, and complex operations, thus limiting their application efficiency and versatility.
By arranging and combining classic mutation sites from the same protein, a long double-stranded nucleic acid library is synthesized. By repairing the gaps to form a complete long double-stranded nucleic acid, the dependence on random mutations is reduced and the efficiency of functional detection of mutant proteins is improved.
It reduces the difficulty of screening, increases the yield of functional proteins, reduces the generation of invalid mutations, simplifies the operation process, and reduces costs.
Smart Images

Figure PCTCN2025133886-FTAPPB-I100001 
Figure PCTCN2025133886-FTAPPB-I100002 
Figure PCTCN2025133886-FTAPPB-I100003
Abstract
Description
A method for preparing mutant libraries Technical Field
[0001] This application relates to the field of biomedicine, specifically to a method for preparing a mutant library. Background Technology
[0002] Mutation libraries, including mutant nucleic acid libraries and mutant protein libraries, play a crucial role in bioscience research and applications. They not only provide powerful tools for exploring biodiversity and conducting functional genomics research, but also play a central role in multiple fields such as drug discovery, protein engineering, disease mechanism research, and biotechnology applications. Through mutation libraries, scientists can identify key gene functional regions, optimize protein performance, discover new drug targets, and develop more effective vaccines and antibodies. Furthermore, mutation libraries serve as the basis for high-throughput screening experiments, accelerating the identification process of mutants with specific biological characteristics. In summary, mutation libraries are a key factor driving life science research and biotechnology innovation, providing valuable information and materials for a deeper understanding of the complexity of life processes and the development of new therapeutic strategies.
[0003] DNA shuffling is a commonly used method for preparing mutant libraries. It's a molecular biology approach that mimics natural evolution by introducing random mutations, fragmenting DNA, and then recombining these fragments to create new protein variants. The technique first constructs a DNA library containing numerous mutations, then cuts the DNA into small fragments and reassembles them using recombination techniques. Finally, a screening process selects mutants with desired characteristics. While DNA shuffling offers significant advantages in improving protein function and stability and accelerating the development of new proteins, it also has drawbacks. These include the difficulty in predicting and controlling the precise location and type of mutations, potentially leading to a large number of invalid mutations; the screening process can be time-consuming and costly; the PCR assembly method used in the assembly process requires the design and synthesis of different primers; and the complexity of the DNA shuffling operation, requiring specialized knowledge and equipment. These factors limit the efficiency and versatility of DNA shuffling in certain situations.
[0004] Due to the uncertainties inherent in DNA shuffling techniques, established classical protein mutations can be used instead of random mutations to prepare mutant libraries. However, introducing too many classical mutations results in a larger mutant library. While current DNA shuffling techniques have made some progress in preparing multiple DNA sequences, they still suffer from problems such as excessive time consumption, high costs, and the generation of random mutations. Therefore, finding new methods for preparing mutant libraries is of great significance for biological research. Summary of the Invention
[0005] The purpose of this application is to obtain a large number of potentially functionally altered mutant proteins by arranging and combining existing classical mutation sites from the same protein. Furthermore, after obtaining the long double-stranded nucleic acid library, the obtained mutant proteins can be functionally tested to identify those with functional changes. These functional changes can be the emergence of new functions, or the enhancement, weakening, or alteration of existing functions. This preparation method does not rely on random mutation as a necessary condition for the emergence of novel proteins, reducing the unpredictability of protein function, thereby lowering the screening difficulty and increasing the final yield of functional proteins.
[0006] This application provides a method for synthesizing a long double-stranded nucleic acid library from short single-stranded nucleic acids, characterized in that an intermediate is present during the synthesis process. The long double-stranded nucleic acids in the long double-stranded nucleic acid library are all homologous to the same original long double-stranded nucleic acid, preferably having about 80% homology, more preferably about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% homology.
[0007] In some embodiments, the long double-stranded nucleic acid library comprises two or more long double-stranded nucleic acids.
[0008] In some embodiments, the intermediate is a defective long double-stranded nucleic acid, specifically a defect referring to a gap in the nucleic acid. The gap refers to the absence of a phosphodiester bond between two adjacent nucleotides in the long double-stranded nucleic acid.
[0009] In some embodiments, the method further includes obtaining the complete long double-stranded nucleic acid by repairing the crack.
[0010] In some embodiments, the short single-stranded nucleic acid contains between 10 and 300 nucleotides. Specifically, the short single-stranded nucleic acid contains between 10 and 60 nucleotides or between 60 and 300 nucleotides.
[0011] In some embodiments, the length of the long double-stranded nucleic acid is more than twice that of the short single-stranded nucleic acid, preferably more than five times, and more preferably more than ten times.
[0012] In some embodiments, the intermediate has at least two slits.
[0013] In a preferred embodiment, the short single-stranded nucleic acid should have 100% homology with a portion of the long double-stranded nucleic acid.
[0014] In some embodiments, the short single-stranded nucleic acid is designed and synthesized by the following methods:
[0015] (1) Design a nucleotide sequence containing the long double-stranded nucleic acid;
[0016] (2) Each single strand of the double-stranded nucleotide sequence is divided into several short single-stranded nucleic acids, and each long double-stranded nucleic acid is divided into n short single-stranded nucleotides; and
[0017] (3) Synthesize the short single-stranded nucleic acid.
[0018] In step (2), the split points in the two chains are staggered from each other, preferably by at least 10 nucleotides.
[0019] In a specific implementation, the short single-stranded nucleic acids obtained after dividing the long double-stranded nucleic acid are named oligo-1, oligo-2, oligo-3, oligo-4, oligo-5, ..., oligo-(2n-1), oligo-2n, where n is a positive integer. Oligo-1, oligo-3, oligo-5, ..., oligo-(2n-1) are located on single strand -1, and oligo-2, oligo-4, oligo-6, ..., oligo-2n are located on single strand -2. Single strand -1 and single strand -2 are two single strands of the long double-stranded nucleic acid, respectively.
[0020] In a specific implementation, for the short single-stranded nucleic acid, oligo-1 corresponds to bases a1 to b1 of single-stranded oligo-1, oligo-2 corresponds to bases a2 to b2 of single-stranded oligo-2, oligo-3 corresponds to bases a3 to b3 of single-stranded oligo-1, ..., oligo-2n corresponds to base a1 to b3 of single-stranded oligo-2. 2n to b 2n The number of bases.
[0021] In a specific implementation, the nucleotides on single strand-1 and single strand-2 are arranged in opposite directions.
[0022] For example, the nucleotide sequence on single strand-1 is from the 5' end to the 3' end of single strand-1, more specifically, the 5' start position on single strand-1 is the first base; and the nucleotide sequence on single strand-2 is from the 3' end to the 5' end of single strand-2, more specifically, the 3' start position on single strand-2 is the first base.
[0023] For example, the nucleotide sequence on single strand-1 is from the 3' end to the 5' end of single strand-1, more specifically, the 3' start position on single strand-1 is the first base; and the nucleotide sequence on single strand-2 is from the 5' end to the 3' end of single strand-2, more specifically, the 5' start position on single strand-2 is the first base.
[0024] In a specific implementation, the single chain-1 has a total of m1 bases, the single chain-2 has a total of m2 bases, and for a1, a2, ..., a... 2n b1, b2, ..., b 2n a1 = a2 = 1; b (2n-1) =m1;b 2n =m2; a3=b1+1, a4=b2+1,…,a 2n =b (2n-2) +1.
[0025] In a specific implementation, a3 and a4, a5 and a6, ..., a 2n With a (2n-1) b1 and b2, b3 and b4, ..., b (2n-3) With b (2n-2) Or a2 and a3, a4 and a5, ..., a (2n-2) With a (2n-1) b2 and b3, b4 and b5, ..., b (2n-2) With b (2n-1) The difference in value represents the number of nucleotides that are misaligned between adjacent dividing points of single strand-1 and single strand-2. In a preferred embodiment, the number of misaligned nucleotides is 10 or more.
[0026] In a specific implementation, a1 and a4, a3 and a6, ..., a 2n With a (2n-3) b1 and b4, b3 and b6, ..., b 2n With b (2n-3) b 2n With b (2n-3) Or a3 and a4, a5 and a6, ..., a (2n-3) With a (2n-2) b1 and b2, b3 and b4, b5 and b6, ..., b (2n-3) With b (2n-2) The difference in value represents the number of nucleotides that are misaligned between adjacent dividing points of single strand-1 and single strand-2. In a preferred embodiment, the number of misaligned nucleotides is 10 or more.
[0027] In a specific embodiment, the aforementioned staggered sequences are overlapping sequences of the two short single-stranded nucleic acids. The arrangement of the overlapping sequences and the complementarity of the overlapping sequences between the two short single-stranded nucleic acids are crucial to the formation of the intermediate. The length of the overlapping sequences also affects the formation of the intermediate. In a preferred embodiment, the length of the overlapping sequences is at least 10, more preferably at least 12, even more preferably at least 14, and even more preferably at least 16.
[0028] In some embodiments, the long double-stranded nucleic acid is the original long double-stranded nucleic acid without any mutations.
[0029] In some embodiments, the long double-stranded nucleic acid is the original long double-stranded nucleic acid containing a mutation.
[0030] In some embodiments, the long double-stranded nucleic acid may contain both the original long double-stranded nucleic acid without mutation and the original long double-stranded nucleic acid with mutation. The two original long double-stranded nucleic acids may be in the same expression frame or not in the same expression frame.
[0031] In a preferred embodiment, the mutation is a classic site mutation.
[0032] In one embodiment, each mutation site contains 1-3 mutation directions, meaning the mutation directions can be any conventional nucleotide other than the original nucleotide at the mutation site. The conventional nucleotides refer to nucleotides on DNA containing adenine, guanine, cytosine, or thymine, and nucleotides on RNA containing adenine, guanine, cytosine, or uracil. For example, the original nucleotide is a nucleotide on DNA containing adenine, and the mutation directions at the mutation site can be selected from one, two, or three nucleotides containing guanine, cytosine, or uracil.
[0033] Furthermore, each mutation site contains 2-4 mutation types, where each mutation type refers to the mutation direction plus the original nucleotide. Let the number of mutation types be p.
[0034] In a specific implementation, the long double-stranded nucleic acid contains X mutation sites, where X is an integer greater than or equal to 2. More specifically, the X mutation sites are distributed on two or more short single-stranded nucleic acids that do not pair complementaryly. Preferably, the mutation sites are distributed on two or more short single-stranded nucleic acids of single-strand-1. Single-strand-1 can be the coding strand of the long double-stranded nucleic acid or the template strand of the long double-stranded nucleic acid.
[0035] In this application, the types of long double-stranded nucleic acids contained in the long double-stranded nucleic acid library are related to the total number of mutation site types contained in the long double-stranded nucleic acids and the mutation types contained in the mutation sites. Specifically, the types of long double-stranded nucleic acids contained in the long double-stranded nucleic acid library are related to the values of X and p.
[0036] In one specific embodiment, let the number of mutation types contained in the mutation sites on the long double-stranded nucleic acid be p1, p2, p3, ..., p X , where p1-p X It is 2, 3, or 4 and p1-p X They may be completely identical or not completely identical. The long double-stranded nucleic acid library contains long double-stranded nucleic acids of type p1×p2×p3×……×p X In a specific implementation, the number of mutation sites contained in each of the short single-stranded nucleic acids is x. More specifically, the number of mutation sites contained in each of the short single-stranded nucleic acids on single-strand-1 is x1, x2, x3, x4, ..., x n The single-stranded-1 can be the coding strand of the long double-stranded nucleic acid or the template strand of the long double-stranded nucleic acid.
[0037] In this application, the type of long double-stranded nucleic acids contained in the long double-stranded nucleic acid library is related to the number of mutation sites contained in each short single-stranded nucleic acid and the mutation type contained in the mutation sites. Specifically, the type of long double-stranded nucleic acids contained in the long double-stranded nucleic acid library is related to the values of x and p.
[0038] In one specific implementation, let the number of mutation types contained in the mutation site on the same short single-stranded nucleic acid be p1, p2, p3, ..., px. n , where p1-px n It is 2, 3, or 4 and p1-px n They may be completely identical or not completely identical. The same short-chain single-stranded nucleic acid can form P molecules with different sequences. n The aforementioned short-chain single-stranded nucleic acid, P n =p1×p2×p3×……×px n The long double-stranded nucleic acid library contains long double-stranded nucleic acids of type P1×P2×P3×……×P n .
[0039] In a more specific embodiment, the mutation sites on the same short-stranded single-stranded nucleic acid contain the exact same number of mutation types. The long-stranded double-stranded nucleic acid library contains long-stranded double-stranded nucleic acids of type p. x1 ×px2 ×p x3 ×p x4 ×……×p xn (Equation 1), where p is 2, 3, or 4. The p values in different terms of Equation 1 can be exactly the same or different.
[0040] In a more specific embodiment, the mutation type of the mutation site is represented by degenerate bases, where the number of specific bases referred to by the degenerate bases is equal to p.
[0041] In one embodiment, each of the short single-stranded nucleic acids on single-stranded 2 contains mutation sites that are completely complementary to those on single-stranded 1.
[0042] In one embodiment, the short single-stranded nucleic acid on the single-stranded 2 does not contain a mutation site.
[0043] In one embodiment, the number of mutation sites contained in each of the single strands-2 is greater than 0 and less than the number of mutation sites contained in each of the single strands-1.
[0044] In some embodiments, the long double-stranded nucleic acid library contains approximately 10 types of long double-stranded nucleic acids. 20 The following are preferred, approximately 10 15 Below, more preferably about 10 13 Below, approximately 10 is even more preferred. 11 Below, approximately 10 is even more preferred. 10 the following.
[0045] Furthermore, in the protein library encoded by the long double-stranded nucleic acid library, amino acid mutations correspond to the codon positions where the mutation sites on the long double-stranded nucleic acids are located. In a preferred embodiment, the amino acid mutation sites on the protein are classical mutation sites.
[0046] In a specific implementation, the number of mutation sites contained in each protein is y, where the value of y is the number of mutated codons contained in the long double-stranded nucleic acid encoding the protein.
[0047] In one embodiment, each mutation site on the protein contains (q-1) mutation directions, meaning the mutation direction can be any of the natural amino acids except the original amino acid at the mutation site. The natural amino acids refer to the 20 amino acids that exist in nature and constitute the basic building blocks of proteins, typically one or more of the following: glycine, alanine, valine, leucine, isoleucine, phenylalanine, tryptophan, tyrosine, aspartate, histidine, asparagine, glutamate, lysine, glutamine, methionine, arginine, serine, threonine, cysteine, and proline. For example, the original amino acid is glycine, and the mutation direction contained at the mutation site can be selected from any natural amino acid other than glycine.
[0048] In another embodiment, each mutation site contains q mutation types, where each mutation type refers to the mutation direction plus the original amino acid.
[0049] In one specific embodiment, let the number of mutation types contained in the mutation site on the same protein be q1, q2, q3, ..., q y , where q1-q y Let q1-q be any integer between 2 and 20 (inclusive) and q1-q y They may be completely identical or not completely identical. The protein library contains proteins of type q1×q2×q3×……×q y .
[0050] In a preferred embodiment, the value of q is selected from 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20.
[0051] In some embodiments, the original long double-stranded nucleic acid comprises one or more of the following: a protein coding sequence, a regulatory sequence, and a coding region of a non-coding RNA.
[0052] In some embodiments, the protein is a wild-type protein and / or a mutant protein.
[0053] In some implementations, the control sequence is one or more of the following: promoter, enhancer, silencer, insulator, and response element.
[0054] In some embodiments, the original long double-stranded nucleic acid comprises a nucleotide sequence encoding one or more of the following proteins: catalytic proteins, reporter proteins, therapeutic proteins, and preventative proteins.
[0055] In some embodiments, the original long double-stranded nucleic acid contains a nucleotide sequence encoding a reporter protein.
[0056] In some embodiments, the reporter protein is selected from one or more of the following: fluorescent protein (FP), luciferase, β-galactosidase (β-Galactosidase or LacZ), chloramphenicol acetyltransferase (CAT), secretory alkaline phosphatase (SEAP), mCherry, mNeonGreen, HaloTag, SNAPTag, and variants thereof.
[0057] In some embodiments, the original long double-stranded nucleic acid comprises a nucleotide sequence encoding a fluorescent protein, the fluorescent protein comprising the amino acid sequence shown in SEQ NO ID:16, and the original long double-stranded nucleic acid comprising the nucleotide sequence shown in SEQ ID NO:1.
[0058] In some implementations, the classic site mutation is derived from an existing mutant.
[0059] For example, the classic site mutation originates from Emerald. The site is one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, N149, M153, and I167. The mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: S65T, S72A, N149K, M153T, and I167T.
[0060] For example, the classic site mutation originates from Topaz. The site is one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, K79, and T203. The mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: S65G, S72A, K79R, and T203Y.
[0061] For example, the classic site mutation originates from ECFP. The site is one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S65, Y66, N146, M153, and V163. The mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: F64L, S65T, Y66W, N146I, M153T, and V163A.
[0062] For example, the classic site mutation originates from EBFP. The site is one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, Y66, and Y145. The mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: F64L, Y66H, and Y145F.
[0063] In some implementations, the classical site mutation occurs in certain regions of the protein that are associated with a specific function.
[0064] For example, the classic site mutation occurs in a region of the protein that improves chromophore formation. The site is the F64 site of the amino acid sequence shown in SEQ ID NO:16. The mutation is selected from the F64L mutation occurring in the amino acid sequence shown in SEQ ID NO:16.
[0065] For example, the classic site mutation occurs in a region related to protein folding. The site is one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S72, and V163. The mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: F64L, S72A, and V163A.
[0066] For example, the classic site mutation occurs in a region related to protein solubility. The site is one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: M153 and V163. The mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: M153T and V163A.
[0067] In some embodiments, the mutation occurs at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S65, Y66, S72, K79, Y145, N146, N149, M153, V163, I167, R168, and T203.
[0068] In some embodiments, the mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: F64L, S65T, S65G, S65C, S65A, Y66H, Y66Q, Y66R, Y66C, Y66W, S72A, K79R, Y145F, N146I, N149K, M153T, V163A, I167T, R168H, T203Y, T203K, and T203S.
[0069] In some embodiments, the original long double-stranded nucleic acid contains a nucleotide sequence encoding a catalytic protein.
[0070] The catalytic protein can be any catalyst that currently or in the future exists, either in its pure protein form or in combination with other substances. Examples include amylase, protease, lipase, ATP synthase, hexokinase, DNA polymerase, RNA polymerase, protein kinase, lactate dehydrogenase, PET hydrolase, and their variants.
[0071] In some embodiments, the original long double-stranded nucleic acid comprises a nucleotide sequence encoding polyethylene terephthalate hydrolase (PET hydrolase), the PET hydrolase comprising the amino acid sequence shown in SEQ ID NO:17, and the original long double-stranded nucleic acid comprising the nucleotide sequence shown in SEQ ID NO:18.
[0072] In some implementations, the classic site mutation is derived from an existing mutant.
[0073] In some embodiments, the mutation site is selected from one or more of the following sites in the amino acid sequence shown in SEQ ID NO:17: E9, A30, P53, S110, T160, G178, S184, T191, and A214.
[0074] In some embodiments, the mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:17: E9A, E9Q, E9P, E9D, A30K, A30N, A30T, A30M, A30I, A30Q, A30H, A30P, A30L, A30E, A30D, A30V, A30Y, A30S, A30F, P53A, P53G, P53V, P53R, P53L, S110N, S110K, S110T, S110R, S110H, S110Q, S110P, S110D, S1 10E, S110A, S110G, T160K, T160R, T160Q, T160P, T160I, G178R, G178S, G178T, G178A, G178C, G178W, S184K, S184N, S184T, S184P, S184Q, T191K, T191N, T191Q, T191H, T191P, T191E, T191D, T191A, A214N, A214T, A214S, A214D, and A214G.
[0075] In some embodiments, the method of repairing the crack on the intermediate is to form a phosphate diester bond at the crack.
[0076] In one embodiment, when the intermediate is preferably a linear nucleic acid, the repair method includes the following steps:
[0077] i. Attach the intermediate to the carrier;
[0078] ii. Transfer the vector into the host body.
[0079] The cracks in the intermediate are repaired within the host body by the host's built-in repair system.
[0080] In some embodiments, the vector is an expression vector. For example, the expression vector is a plasmid.
[0081] In another embodiment, where the intermediate is preferably a circular nucleic acid, more preferably a plasmid, the repair method includes transferring the plasmid into the host.
[0082] In some embodiments, the host is an organism capable of accommodating a foreign genome, preferably one that can also provide conditions for the amplification of the foreign genome.
[0083] In a preferred embodiment, the host is a model organism. For example, the model organism is *Escherichia coli*. Another example is yeast.
[0084] In some embodiments, the number of slits is 2-50.
[0085] In some embodiments, the number of slits is 51 or more.
[0086] In some embodiments, the method further includes repairing cracks on the intermediate before the vector or plasmid is transferred into the host. In a preferred embodiment, the repair includes the use of a ligase, preferably the use of T4 ligase.
[0087] For example, the repair occurs before the intermediate is attached to the carrier.
[0088] For example, the repair occurs after the intermediate is attached to the carrier.
[0089] On the other hand, this application also relates to the use of the aforementioned synthesis method in the synthesis of long double-stranded nucleic acid libraries.
[0090] In some embodiments, the long double-stranded nucleic acid consists of two or more strands.
[0091] In a preferred embodiment, the long double-stranded nucleic acid library contains approximately 10 types of long double-stranded nucleic acids. 20 The following are preferred, approximately 10 15 Below, more preferably about 10 13 Below, approximately 10 is even more preferred. 11 Below, approximately 10 is even more preferred. 10 the following.
[0092] In some embodiments, the long double-stranded nucleic acid is a linear nucleic acid.
[0093] In some embodiments, the long double-stranded nucleic acid is a circular nucleic acid, preferably a plasmid.
[0094] In some embodiments, the long double-stranded nucleic acid is a nucleic acid with a complex double-stranded structure, preferably a triangular Y-shaped double-stranded structure or a cross-shaped double-stranded structure.
[0095] On the other hand, this application also relates to a long double-stranded nucleic acid library synthesized by the aforementioned method.
[0096] In some embodiments, the long double-stranded nucleic acids in the long double-stranded nucleic acid library may be mutants of the original long double-stranded nucleic acid. The long double-stranded nucleic acids are homologous to the original long double-stranded nucleic acid, preferably having about 80% homology, more preferably about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% homology.
[0097] In a preferred embodiment, the long double-stranded nucleic acid library contains approximately 10 types of long double-stranded nucleic acids. 20 The following are preferred, approximately 10 15 Below, more preferably about 10 13 Below, approximately 10 is even more preferred. 11 Below, approximately 10 is even more preferred. 10 the following.
[0098] On the other hand, this application also relates to protein libraries encoded by the aforementioned long double-stranded nucleic acid libraries.
[0099] In some embodiments, the proteins in the protein library may be original proteins and / or mutants thereof, wherein the original proteins are encoded by the original long double-stranded nucleic acid. The proteins are homologous to the amino acid sequence of the original protein, preferably having about 80% homology, more preferably about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% homology.
[0100] In a preferred embodiment, the original protein is a fluorescent protein.
[0101] In another preferred embodiment, the original protein is a PET hydrolase.
[0102] On the other hand, this application also relates to the screening of the protein library, specifically the screening of functional proteins from the protein library.
[0103] In a specific implementation, the original protein is a fluorescent protein.
[0104] In some embodiments, the screening method is selected from fluorescence microscopy, flow cytometry, and microplate readers. The fluorescence microscope is selected from ordinary fluorescence microscopy, stereofluorescence microscopy, and confocal fluorescence microscopy.
[0105] In a specific implementation, the original protein is a PET hydrolase.
[0106] In some embodiments, the screening method is carried out via an enzyme-substrate reaction, and the functional activity of the PET hydrolase is reflected by characterizing the hydrolysis efficiency of the substrate by the PET hydrolase.
[0107] On the other hand, this application also protects mutant fluorescent proteins obtained by screening according to the aforementioned screening method.
[0108] In one embodiment, the fluorescent protein comprises an amino acid substitution at the T203 site of the amino acid sequence shown in SEQ ID NO:16. The substituted amino acid is tyrosine.
[0109] In a further embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, V163, I167, and T203. The substituted amino acids are: S65C / S65G, S72A, V163A, I167T, and T203Y.
[0110] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S65, S72, K79, M153, V163, I167, and T203. The substituted amino acids are: F64L, S65C, S72A, K79R, M153T, V163A, I167T, and T203Y.
[0111] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, K79, Y145, N149, M153, V163, I167, and T203. The substituted amino acids are: S65G, S72A, K79R, Y145F, N149K, M153T, V163A, I167T, and T203Y.
[0112] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, N149, M153, V163, I167, and T203. The substituted amino acids are: S65G, S72A, N149K, M153T, V163A, I167T, and T203Y.
[0113] In one embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65 and N149. The substituted amino acids are: S65C / S65T / S65A and N149K.
[0114] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S65, K79, Y145, N149, I167, and R168. The substituted amino acids are: F64L, S65C, K79R, Y145F, N149K, I167T, and R168H.
[0115] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, K79, Y145, N149, V163, and T203. The substituted amino acids are: S65T, S72A, K79R, Y145F, N149K, V163A, and T203S.
[0116] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, N146, N149, M153, V163, and I167. The substituted amino acids are: S65A, N146I, N149K, M153T, V163A, and I167T.
[0117] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, S72, Y145, N146, N149, M153, and V163. The substituted amino acids are: S65A, S72A, Y145F, N146I, N149K, M153T, and V163A.
[0118] In one embodiment, the fluorescent protein comprises an amino acid substitution at the Y66 site of the amino acid sequence shown in SEQ ID NO:16. The substituted amino acid is histidine.
[0119] In a further embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: Y66, K79, and V163. The substituted amino acids are: Y66H, K79R, and V163A.
[0120] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, Y66, S72, K79, Y145, N146, V163, and I167. The substituted amino acids are: F64L, Y66H, S72A, K79R, Y145F, N146I, V163A, and I167T.
[0121] In a specific embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S65, Y66, K79, M153, and V163. The substituted amino acids are: S65A, Y66H, K79R, M153T, and V163A.
[0122] In one embodiment, the fluorescent protein comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S72, K79, N146, V163, and I167. The substituted amino acids are: S72A, K79R, N146I, V163A, and I167T.
[0123] On the other hand, this application also includes mutant PET hydrolases obtained by screening according to the aforementioned screening method.
[0124] In some embodiments, the mutant PET hydrolase comprises amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:17: E9, A30, P53, S110, T160, G178, S184, T191, and A214.
[0125] For example, the substituted amino acids are: E9D, P53G, S110R, T160R, G178A, S184K, T191K and A214T.
[0126] For example, the substituted amino acids are: A30T, P53A, S110A, S184P and A214G.
[0127] For example, the substituted amino acids are: E9A, A30D, P53G, S110A, G178C, S184P, T191A, and A214G.
[0128] For example, the substituted amino acids are: A30T, P53A, T160I, G178S, S184T, T191P and A214S.
[0129] This application fragments protein-coding sequences containing and / or without mutations into short single-stranded fragments with overlapping sequences, wherein the mutation is preferably a classic site mutation selected in the prior art. Homologous fragments contain both short single-stranded nucleic acids containing and without mutations. Under various permutations and combinations, the number of mutated nucleic acids generated increases exponentially. At this point, many high-throughput methods suffer from drawbacks such as random mutation generation, low efficiency, and excessive time consumption. The synthetic method described in this application has the following advantages: (1) It saves time and produces a large number of nucleic acids with different mutation sites in the same reaction; (2) It causes fewer random mutations. After the long double-stranded nucleic acid with cracks is formed, the cracks are repaired. Unlike when polymerase and ligase are involved, the mismatches that occur during the pairing process are quickly fixed on the long double-stranded nucleic acid. Nucleic acids with cracks still have the opportunity to remove the wrong pairing and pair correctly; (3) It has a lower cost. The reaction uses fewer primers, polymerases, ligases, etc., which can reduce the production cost of mutant libraries to a certain extent; (4) It has little or no bias when synthesizing libraries. The proportion of each synthesized long double-stranded nucleic acid in the final product is not much different.
[0130] Other aspects and advantages of this application will readily be apparent to those skilled in the art from the detailed description below. Only exemplary embodiments of this application are shown and described in the following detailed description. As will be appreciated by those skilled in the art, the content of this application enables them to make modifications to the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application pertains. Accordingly, the descriptions in the accompanying drawings and specification of this application are merely exemplary and not restrictive. Attached Figure Description
[0131] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention can be better understood by referring to the exemplary embodiments and drawings described in detail below. A brief description of the drawings is as follows:
[0132] Figure 1 shows a flowchart of the synthesis of long double-stranded nucleic acid libraries from short single-stranded nucleic acids.
[0133] Figure 2 shows the difference in the effectiveness of short single-stranded nucleic acids, whether or not they have been purified by PAGE, when assembling into intermediates.
[0134] Figure 3 illustrates two methods for assembling the intermediate and the carrier according to this application. Figure 3A shows the Gibson assembly method; Figure 3B shows the chain substitution assembly method.
[0135] Figure 4 shows a more detailed schematic diagram of the Gibson assembly, in which the sequence of the intermediate is derived from the coding sequence of green fluorescent protein (GFP).
[0136] Figure 5 illustrates the different behaviors of different types of short single-stranded nucleic acids in the synthesis of long double-stranded nucleic acids. Figure 5A shows the flowchart of the synthesis and verification of long double-stranded nucleic acids from short single-stranded nucleic acids; Figure 5B shows the synthesis process of four different types of short single-stranded nucleic acids; Figure 5C shows the agarose gel electrophoresis results of the intermediates synthesized from the four different types of short single-stranded nucleic acids; Figure 5D shows the length and corresponding synthesis yield of the long double-stranded nucleic acids synthesized from the four different types of short single-stranded nucleic acids.
[0137] Figure 6 shows the results of fluorescence detection of colonies on the plate in Example 5.
[0138] Figure 7 shows the agarose gel electrophoresis results of the intermediates assembled from short single-stranded nucleic acids with different overlapping sequence lengths in Example 6.
[0139] Figure 8 shows a schematic diagram of the assembly of short single-stranded nucleic acids containing special fragments into intermediates. Figure 8A shows the circularity analysis (left) and predicted self-pairing structures (right) of nucleic acids containing self-pairing hairpin structures in Example 6 performed on mfold.org; Figure 8B shows a schematic diagram of the partitioning of tandem repeat sequences; Figure 8C shows the distribution of the number of repeats and the corresponding number of clones in the repeat sequence sequencing results, where 6 repeats is the correct number of repeats.
[0140] Figure 9 shows an agarose gel electrophoresis image of the product obtained by assembling short single-stranded nucleic acids containing a specific sequence using the PCA method in Example 7.
[0141] Figure 10 shows the performance of short single-stranded nucleic acids in Example 8 when synthesizing two long double-stranded nucleic acids of different lengths. Figure 10A shows a schematic diagram of the formation of intermediates from short single-stranded nucleic acids separated from long double-stranded nucleic acids of different lengths (gfp and bacterial rhodopsin); Figure 10B shows the agarose gel electrophoresis results of the intermediates formed after annealing and assembling the short single-stranded nucleic acids of gfp and bacterial rhodopsin in the same reaction; Figure 10C shows a statistical graph of the sequencing results of the assembled and amplified synthesized intermediates and sampled for sequencing.
[0142] Figure 11 shows the performance of short single-stranded nucleic acids in Example 8 when synthesizing two long double-stranded nucleic acids of the same length and high homology. Figure 11A shows a schematic diagram of the formation of intermediates from short single-stranded nucleic acids divided on homologous long double-stranded nucleic acids (nr2f1 and nr2f2); Figure 11B shows the number of identical bases on corresponding homologous fragments in the two fragments nr2f1 and nr2f2, with the homologous fragments being 60 nt in length. Figure 11C shows the agarose gel electrophoresis results of the intermediates formed after annealing and assembling the short single-stranded nucleic acids of nr2f1 and nr2f2 in the same reaction; Figure 11D shows the sequencing results of the synthesized intermediates after assembly, amplification, and sequencing.
[0143] Figure 12 shows the results of assembling homologous genes using the PCA method in Example 8. Figure 12A shows the agarose gel electrophoresis images of the products obtained from synthesizing homologous genes nr2f1 and nr2f2 in the same reaction; Figure 12B shows the sequence alignment results of the final products with the standard sequences of nr2f1 and nr2f2, respectively.
[0144] Figure 13A shows a simplified flowchart for preparing a mutant protein nucleic acid library; Figure 13B shows the fluorescence observation of bacterial colonies under a stereofluorescence microscope; Figure 13C shows the results of fluorescence detection of bacterial cells under a confocal microscope; Figure 13D shows the results of fluorescence detection in the FITC channel of a flow cytometer; Figures 13E and 13F show the results of fluorescence measurement of bacterial cells using a microplate reader, with three independent colonies in each group, and the error bars representing the standard deviation.
[0145] Figure 14 shows the sequence and phenotype of the mutant fluorescent protein. Figure 14A shows the degenerate bases designed at 12 selected fluorescent protein mutation sites; Figure 14B shows the results of scanning an E. coli library containing approximately 800 clones of the mutant fluorescent protein using stereofluorescence microscopy; Figure 14C shows the mutations occurring on different fluorescent proteins, with mutations present on the control fluorescent proteins Emerald, Topaz, and EBFP marked in green, orange, and blue, respectively; Figures 14D and 14E show the fluorescence detection results of bacterial cells under confocal microscopy; Figure 14F shows the fluorescence detection results under the FITC channel of a flow cytometer, with NC as the negative control, representing bacterial cells infected with pUC19. Median fluorescence intensity was used to compare the brightness of different samples, and the median fluorescence intensity was normalized to 1.
[0146] Figure 15 shows the chain displacement assembly method for assembling the intermediate and the carrier in this application.
[0147] Figure 16 shows an exemplary schematic diagram of the 16S rRNA synthesized in Example 12.
[0148] Figure 17 shows the detection pattern of the long double-stranded nucleic acid containing the specific sequence synthesized in Example 12. Figure 17A shows the electrophoresis pattern of the synthesized intermediate containing 100% AT. Figure 17B shows the electrophoresis pattern of the long double-stranded nucleic acid fragment containing the specific sequence synthesized using the PCA method; the left image shows the electrophoresis pattern of synthesized 100% AT and 100% GC, the middle image shows the electrophoresis pattern of synthesized DNA encoding rRNA, and the right image shows the electrophoresis pattern of encoding repetitive sequences.
[0149] Figure 18 shows the length distribution of human genes and the assembly diagram and electrophoresis diagram of intermediates containing special sequences. Figure 18A shows the proportion of human genes with lengths below 1200 bp and below 3000 bp, respectively. Figure 18B shows the assembly diagram and corresponding electrophoresis diagram of intermediates containing self-paired sequences, high-GC sequences, and repetitive sequences.
[0150] Figure 19 shows the electrophoresis diagram of intermediates synthesized from short single-stranded nucleic acids of different lengths in Example 13.
[0151] Figure 20 shows the effects of the length of the synthesized sequence, the number of cleavages in the intermediate, and the length distribution of the human gene in Example 14. Figure 20A shows the electrophoresis diagram of the intermediate when synthesizing long double-stranded nucleic acids of 300-3000 bp using the MOSAIC method. Figure 20B shows the electrophoresis diagram of the intermediate when synthesizing long double-stranded nucleic acids of 9900 bp using the MOSAIC method. Figure 20C shows the ability of intermediates containing different numbers of cleavages to form long double-stranded nucleic acids through host repair function after treatment with or without ligase. Figure 20D shows the length distribution of the human gene.
[0152] Figure 21 shows the detection results of the 9.9kb plasmid synthesized in Example 14. Figure 21A shows the distribution of the seven pairs of detection primers used in Example 14 on the plasmid, the electrophoresis diagrams of nucleic acid fragments amplified by different primers, and the assembly efficiency of the clones formed by the 9.9kb plasmid. Figure 21B shows the number and distribution of mutation types contained in three randomly selected successfully synthesized 9.9kb plasmids after sequencing.
[0153] Figure 22 shows the form and type of codons (top) and amino acids (bottom) at each selected mutation site in the fluorescent protein mutant library.
[0154] Figure 23 shows the NGS sequencing results of the prepared fluorescent protein mutant library.
[0155] Figure 24 shows the gel electrophoresis images of the purified mutants Y1, Y2, G2, and G4. The top image shows the results after purification without concentration normalization. The bottom image shows the results after concentration normalization.
[0156] Figure 25 shows a comparison of the fluorescence intensity of mutants G2 and G4 after concentration normalization with the fluorescent protein Emerald.
[0157] Figure 26 shows the preparation process of the PET hydrolase mutant library and the detection results of the new mutants. Figure 26A shows the amino acid mutation sites and mutation types on the PET hydrolase, as well as the total number of possible mutations. Figure 26B shows the NGS results of the prepared mutant library. Figure 26C shows the fluorescence spectra of wild-type PET hydrolase and the negative control after induction of FDBz expression. Figure 26D shows the mutations included in the newly obtained mutants. Figure 26E shows the gel electrophoresis results of the purified PET hydrolase mutants after concentration normalization. Figure 26F shows the hydrolytic ability (i.e., enzyme activity) of the purified mutants against FDBz after concentration normalization.
[0158] Figure 27 shows the preparation and detection process of the PET hydrolase mutant library and the detection results of the new mutants. Figure 27A shows the protein structure and mutation sites of the PET hydrolase (left) and the preparation and screening process of the PET hydrolase (right). Figure 27B shows the process of PET hydrolase hydrolyzing FDBz. Figure 27C shows the overall fluorescence intensity enhancement caused by the reaction of the bacterial mixture and FDBz after the transformation of the mutant library. Figure 27D shows the absorption rate of different mutants for different wavelengths of light. Figure 27E shows the hydrolytic ability of different mutants for FDBz.
[0159] Figure 28 shows the protein structure and mutation sites of the fluorescent protein. Detailed Implementation
[0160] The following specific embodiments illustrate the implementation of the invention. Those skilled in the art can easily understand other advantages and effects of the invention from the content disclosed in this specification.
[0161] Terminology Definition
[0162] Unless otherwise defined in this application, the scientific and technical terms used herein shall have the meanings commonly understood by one of ordinary skill in the art. Furthermore, unless the context requires otherwise, singular terms shall include plural terms, and plural terms shall include singular terms. Generally, the cell and tissue culture, molecular biology, immunology, microbiology, genetics, and proteins and nucleic acids described herein are concepts recognized by those skilled in the art.
[0163] In this application, the term "nucleic acid" can be understood both macroscopically and microscopically. Macroscopically, nucleic acid can be understood as nucleic acid molecules and mixtures containing nucleic acid molecules. The mixture may contain solvents (e.g., enzyme-free water), other necessary components (e.g., additives to stabilize the nucleic acid structure, adjuvants to aid subsequent reactions), and other possible impurities. Microscopically, nucleic acid is primarily understood as the molecular structure of nucleic acids. Nucleic acids are mainly composed of nucleotide chains, nucleotide rings, or other possible structures formed by the dehydration condensation of nucleotides to form phosphodiester bonds, where the phosphodiester bonds exist between the phosphate group and the pentose group of the nucleotide. In this application, the term "pentose phosphate chain" refers to a portion of the nucleic acid containing the phosphodiester bonds. The pentose phosphate chain contains the phosphate group, pentose group, and necessary linking structures (such as covalent bonds) of the nucleic acid, but does not contain the bases of the nucleic acid. Since the bases cannot form covalent bonds with each other, but only form hydrogen bonds with their opposite bases when the nucleic acid is double-stranded, the region where the base is located can be referred to as the base side.
[0164] In this application, the term "nucleotide sequence" refers to the sequence of nucleotides that make up a nucleic acid molecule (including DNA and RNA molecules). This sequence is typically represented by the base sequence on one side of the nucleic acid. The starting point of the nucleotide sequence can vary depending on the context. For multiple nucleic acids that are 100% homologous to the same double-stranded nucleic acid, their orientation can be determined from the context.
[0165] In this application, the term "length" generally refers to the length of nucleic acids, including short-chain nucleic acids, long-chain nucleic acids, and single-chain and double-chain nucleic acids. The length can be expressed as the number of bases in a single-chain nucleic acid; or as the total number of bases or base pairs between the two nucleotides at both ends of a double-chain nucleic acid. The two nucleotides at both ends can be on the same single strand or on different single strands. The double-chain nucleic acid may or may not contain dangling sequences. A dangling sequence refers to a segment on the double strand that lacks complementary pairing; it is called a dangling sequence because it usually extends beyond the double strand and is located at the end of the nucleic acid. In this application, the unit of length can be bp, nt, or mer / mers, where bp represents the number of base pairs in the complementary region of the double-chain nucleic acid; nt and mer / mers represent the number of bases in the single-chain nucleic acid. A base pair refers to a pairing formed by a one-to-one correspondence between bases in regions of high homology between the two single strands of a double-chain nucleic acid. Within the base pairs, the pairing between bases can be completely complementary, or a base on one chain can pair with a position on another chain where there is no base, or both of the above can coexist.
[0166] In this application, the terms "short-chain nucleic acid" and "long-chain nucleic acid" are generally distinguished by the length of the nucleic acid.
[0167] In this application, the term "short-chain nucleic acid" generally refers to a nucleic acid molecule that can be synthesized using only existing and potentially future de novo synthesis methods in the art. The de novo synthesis method refers to a nucleic acid synthesis method starting from a single nucleotide, currently mainly including phosphoramide chemical synthesis, photochemical synthesis, electrochemical synthesis, inkjet printing, and other de novo synthesis methods. The maximum length of the short-chain nucleic acid should be the limit of nucleic acids with high fidelity that can be synthesized by existing and potentially future de novo synthesis methods in the art. In some embodiments, the maximum length may be 300 bp. The minimum length of the short-chain nucleic acid should be the minimum length at which it can stably exist; when the short-chain nucleic acid is shorter than the minimum length, it will be subject to significant degradation or other potential risks. In some embodiments, the minimum length may be 10 bp.
[0168] In this application, the term "long nucleic acid" generally refers to a nucleic acid molecule that cannot be synthesized solely by de novo synthesis or whose synthesis by de novo synthesis is costly. The long nucleic acid is typically assembled from the short nucleic acid molecules. The minimum length of the long nucleic acid is generally greater than, but may be less than, the maximum length of nucleic acid molecules that can be synthesized by the de novo synthesis method. In some embodiments, the minimum length may be 300 bp, 100 bp, or 200 bp. In some embodiments, the minimum length should be greater than 60 bp.
[0169] In this application, the terms "double-stranded nucleic acid" and "single-stranded nucleic acid" are generally distinguished based on the structure of the nucleic acid.
[0170] In this application, the term "single-stranded nucleic acid" should be interpreted broadly, generally referring to a nucleic acid molecule from start to finish without any other nucleic acid strands. It is worth noting that single-stranded nucleic acids can also form reverse complementary structures, such as stem-loop structures, in regions of high homology, but this should not be used to classify them as double-stranded nucleic acids. The single-stranded nucleic acid can be a linear nucleic acid, a circular nucleic acid, a single-stranded nucleic acid containing a stem-loop structure, or any other possible structure.
[0171] In this application, the term "double-stranded nucleic acid" generally refers to a nucleic acid composed of two or more single-stranded nucleic acids. The double-stranded nucleic acid typically includes regions of high homology capable of forming reverse complementary regions. The term "double-stranded nucleic acid" should be interpreted broadly; for example, it is not limited to linear double-stranded nucleic acids composed of two reverse complementary strands, but may also include circular double-stranded nucleic acids composed of two reverse complementary strands, and may also include more complex nucleic acid structures containing double-stranded nucleic acids. Examples of complex nucleic acid structures may be, for example, a tri-Y double-stranded structure formed by the pairwise complementarity of three DNA single strands, or a cross-shaped double-stranded structure formed by the pairwise complementarity of four DNA single strands. Within the reverse complementary regions of the double-stranded nucleic acid, all bases may or may not have corresponding complementary bases; that is, the reverse complementarity rate may be 100% or less than 100%. The reverse complementarity ratio can be the ratio of the number of complementary base pairs in the entire region to the total number of theoretically existing bases on a single strand within that region (including non-existent bases that pair with bases on the other strand). The theoretically existing bases include the actual number of bases on the single strand and the number of bases skipped due to insufficient reverse complementarity (i.e., the aforementioned "non-existent bases"). The reverse complementarity ratio can be calculated manually or with computer assistance. An example of computer-aided calculation is a BLAST system. The double-stranded nucleic acid may contain only the region forming the reverse complementarity, or it may contain structures other than the reverse complementarity region, such as dangling sequences. A dangling sequence refers to a nucleotide chain extending from the reverse complementarity region.
[0172] In this application, the term "short-chain single-stranded nucleic acid" refers to a nucleic acid or nucleic acid molecule that meets both the characteristics of a short-chain nucleic acid and a single-stranded nucleic acid. The terms "short-chain single-stranded nucleic acid," "oligonucleotide," and "oligonucleic acid" are used interchangeably. The term "long-chain double-stranded nucleic acid" refers to a nucleic acid or nucleic acid molecule that meets both the characteristics of a long-chain nucleic acid and a double-stranded nucleic acid. It should be noted that the number of long-chain double-stranded nucleic acids in this application refers to the number of long-chain double-stranded nucleic acids with different sequences; that is, long-chain double-stranded nucleic acids with the same sequence should not be considered as different long-chain double-stranded nucleic acids. It is worth noting that if the sequence difference is due to mutations during synthesis, whether it should be considered as different long-chain double-stranded nucleic acids should be determined based on the specific context.
[0173] In this application, the term "classical site mutation" or "classical mutation site" refers to currently known and potentially future known mutation sites that affect protein function, including nucleotide mutation sites and amino acid mutation sites. The "classical mutation site" or "classical site mutation" can refer to a mutation site and its mutation type found in a functional mutant protein known in the art. Whether "classical site mutation" or "classical mutation site" specifically refers to an amino acid mutation or a nucleotide mutation in this application can be determined from the context.
[0174] In this application, the term "misaligned" generally refers to the presence of a dangling sequence between two short single-stranded nucleic acids with inversely complementary regions, in addition to the inversely complementary regions. This dangling sequence may or may not be inversely complementary to the other short single-stranded nucleic acid. Furthermore, on each pair of inversely complementary short single-stranded nucleic acids, one or both short single-stranded nucleic acids may have the dangling sequence. In the intermediate, after the misaligned sequences form an inversely complementary double strand, they can be referred to as overlapping sequences.
[0175] In this application, the term "long-chain double-stranded nucleic acid library" refers to a composition containing at least two long-chain double-stranded nucleic acids. In this application, the long-chain double-stranded nucleic acids in the library should be nucleic acids with high homology to the same original long-chain double-stranded nucleic acid, for example, having 80% homology, 85% homology, 90% homology, 95% homology, 96% homology, 97% homology, 98% homology, or 99% homology. The term "long-chain double-stranded nucleic acid" is interpreted broadly, and can refer to a mixture containing long-chain double-stranded nucleic acids. This mixture may contain a solvent (e.g., enzyme-free water), other necessary components (e.g., additives to stabilize nucleic acid structures, adjuvants to aid subsequent reactions), and other possible impurities. The long double-stranded nucleic acids in the long double-stranded nucleic acid library can be linear nucleic acids or circular nucleic acids; they can be classic linear structures; they can be the tri-Y double-stranded structure or the cross-shaped double-stranded structure; they can contain the pendant sequence or not; they can have the ability to encode proteins or not.
[0176] In this application, the term "mutation" refers to an event that occurs in the amino acid sequence corresponding to a protein or the nucleotide sequence encoding the protein, resulting in an alteration of the amino acid sequence or nucleotide sequence. The alteration can be a deletion, substitution, and / or insertion in the amino acid sequence or nucleotide sequence. Therefore, the alteration can change the function of the protein or it can have no effect on the function of the protein.
[0177] In this application, the term "site" refers to the position of a specific amino acid or nucleotide in an amino acid sequence or nucleotide sequence. The site can be represented by the numerical value of the position of the amino acid or nucleotide on the original protein or the original long double-stranded nucleic acid. Specifically, the first amino acid at the N-terminus of the amino acid sequence is the amino acid at site 1, the amino acid at the C-terminus is the amino acid at site 2, and so on; similarly, the three nucleotides corresponding to the amino acid at site 1 are nucleotides 1, 2, and 3, the three nucleotides corresponding to the amino acid at site 2 are nucleotides 4, 5, and 6, and so on. It is important to note that regardless of any mutations in the amino acid or nucleotide sequence corresponding to the protein, resulting in changes to the actual site of a specific amino acid or nucleotide, the site always remains the same as its position on the original protein or the original long double-stranded nucleic acid.
[0178] In this application, the term "protein library" refers to the protein encoded by nucleic acid molecules in the long double-stranded nucleic acid library. The protein may be a functional protein or a non-functional protein. The protein may be a post-translationally processed protein, a protein that has not undergone post-translational processing, or a protein undergoing post-translational processing. Post-translational processing refers to the modification process a protein undergoes after translation from messenger RNA to protein. Post-translational processing may involve adding modifying groups to specific amino acids or altering the properties of the protein through proteolytic cleavage.
[0179] In this application, the term "crack" generally refers to a structure present on the pentose phosphate chain of a nucleic acid that causes a discontinuity in the nucleic acid chain. The crack may be formed due to the absence of a phosphodiester bond.
[0180] In this application, the term "intermediate" generally refers to a long double-stranded nucleic acid with a cleavage that occurs during the synthesis of the long double-stranded nucleic acid from the short single-stranded nucleic acid. Specifically, compared to the long double-stranded nucleic acid, the intermediate has a cleavage on the pentose phosphate chain of the nucleic acid, preferably at the interface of the short single-stranded nucleic acid. The short single-stranded nucleic acid can be formed into the intermediate through an annealing process well known in the art, or other possible methods. In the context of this application, the term "intermediate" should be interpreted broadly, meaning that any long double-stranded nucleic acid containing a cleavage can be an intermediate. Specific examples include intermediates formed by annealing the short single-stranded nucleic acid, intermediates formed by annealing the short single-stranded nucleic acid and then treating them with a ligase to ligate part of the cleavage, intermediates not connected to a vector, intermediates connected to a vector, intermediates containing a pendant sequence, and intermediates not containing a pendant sequence.
[0181] In this application, the terms "linear nucleic acid" and "circular nucleic acid" are generally distinguished by the presence or absence of linked ends. "Linear nucleic acid" typically refers to nucleic acid without linked ends. "Circular nucleic acid" refers to nucleic acid containing linked ends, which can be composed of a start and an end point on the nucleic acid, or any two points on the nucleic acid. The term "linked" should be interpreted in a narrower sense, meaning linked by covalent bonds on the pentose phosphate chain of the nucleic acid.
[0182] In this application, the term "vector" or "expression vector" generally refers to any molecule used to transfer nucleic acid information to a host (e.g., plasmids or viruses, virus-like particles, polycations, peptide vectors, liposomes, and / or hybridization vectors). The term "vector" includes nucleic acid molecules capable of transporting another nucleic acid linked to them. One type of vector is a "plasmid," which refers to a circular double-stranded DNA molecule into which a target DNA fragment can be inserted. Another type of vector is a viral vector, in which a target DNA fragment can be inserted into a viral genome. Some vectors are capable of autonomous replication in the host cells to which they are introduced (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors). Generally, expression vectors useful in recombinant nucleic acid technologies are typically in the form of plasmids. The terms "plasmid" and "vector" are used interchangeably herein because plasmids are the most commonly used form of vector. However, the disclosure of this application may include other forms of expression vectors, such as viral vectors (e.g., replication-defective retroviruses, adenoviruses, and adeno-associated viruses), which have equivalent functions.
[0183] In this application, the term "and / or" should be understood as any one of the multiple elements used for connection or any combination of elements.
[0184] In this application, the term "comprising" or "including" generally means including the explicitly specified features, but does not exclude other elements.
[0185] In this application, the term “selected from” generally refers to the selection of objects and all combinations thereof. For example, “selected from A, B and C” means all combinations of A, B and C, such as A, B, C, A+B, A+C, B+C, or A+B+C.
[0186] In this application, the term "about" generally refers to a variation within a range of 0.5% to 10% above or below a specified value, such as a variation within a range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% above or below a specified value.
[0187] Invention Details
[0188] In this application, since the short single-stranded nucleic acid synthesis intermediates are assembled through complementary pairing between short single-stranded nucleic acids, this method is named Molecular Self-Assembly Induced Cloning (MOSAIC).
[0189] intermediate
[0190] The most important feature of the synthesis method involved in this application is the formation of the intermediate.
[0191] The intermediate refers to a defective intermediate state formed during the synthesis of a long double-stranded nucleic acid from the short single-stranded nucleic acid, and is therefore called an intermediate. Except for the lack of phosphodiester bonds (creating a gap) between adjacent nucleotides, the intermediate can be considered to have the same characteristics as the long double-stranded nucleic acid. It should be noted that the lack of phosphodiester bonds between adjacent nucleotides does not mean that every pair of adjacent nucleotides is missing, but only a portion of the adjacent nucleotides, and a very small portion at that. Preferably, the number of adjacent nucleotides lacking phosphodiester bonds is 10% or less of the total nucleotides.
[0192] The intermediate is typically formed by annealing the short single-stranded nucleic acids that make up the long double-stranded nucleic acid in the same mixing system. In biology, especially in nucleic acid synthesis, annealing generally refers to a process used to pair two complementary sequences, including heating and cooling.
[0193] During annealing, the purpose of heating is to break all hydrogen bonds between bases, causing all short-chain single-stranded nucleic acids in the system to unfold; this process is called denaturation. Heating temperatures can reach approximately 80°C and above, approximately 85°C and above, approximately 90°C and above, approximately 91°C and above, approximately 92°C and above, approximately 93°C and above, approximately 94°C and above, approximately 95°C and above, approximately 96°C and above, approximately 97°C and above, approximately 98°C and above, approximately 99°C and above, and approximately 100°C and above.
[0194] After denaturation breaks the hydrogen bonds between bases, the system needs to be cooled. The purpose of this process is to allow different short-chain single-stranded nucleic acids to form inter-base hydrogen bonds, thus enabling the dispersed short-chain single-stranded nucleic acids to form a cohesive whole through these bonds. During cooling, due to complementary overlapping sequences between the short-chain single-stranded nucleic acids, the overlapping sequences between the denatured and unfolded short-chain single-stranded nucleic acids will form inter-base hydrogen bonds with their corresponding overlapping sequences. Simultaneously, in the inventors' design, most short-chain single-stranded nucleic acids have overlapping sequences with at least two other short-chain single-stranded nucleic acids. That is, if only two short-chain single-stranded nucleic acids are complementary, a dangling sequence will be formed. This dangling sequence usually overlaps with another short-chain single-stranded nucleic acid, thus forming an inter-base hydrogen bond with that other short-chain single-stranded nucleic acid. This process gradually increases the number of complementary base pairs. This process repeats until the resulting product has no dangling sequences or the length of the dangling sequences is insufficient to support a stable whole with another short-chain single-stranded nucleic acid. This process may or may not involve the breaking of other chemical bonds, but it typically does not involve the breaking of phosphodiester bonds. The resulting short-chain single-stranded nucleic acid, formed by hydrogen bonds, is the intermediate described in this application. This process can be slow cooling or rapid cooling. More specifically, it can be a cooling rate of approximately 0.01°C per second, approximately 0.1°C per second, approximately 1°C per second, approximately 2°C per second, approximately 3°C per second, approximately 4°C per second, or approximately 5°C per second. The final cooling temperature can be a relatively high temperature, such as approximately 40-60°C (e.g., approximately 40°C, approximately 41°C, approximately 42°C, approximately 43°C, approximately 44°C, approximately 45°C, approximately 46°C, approximately 47°C, approximately 48°C, approximately 49°C, approximately 50°C, approximately 51°C, approximately 52°C, approximately 53°C, approximately 54°C, approximately 55°C, approximately 56°C, approximately 57°C, approximately 58°C, approximately 59°C, or approximately 60°C); or it can be a relatively low temperature. For example, approximately 15-40°C (e.g., approximately 15°C, approximately 16°C, approximately 17°C, approximately 18°C, approximately 19°C, approximately 20°C, approximately 21°C, approximately 22°C, approximately 23°C, approximately 24°C, approximately 25°C, approximately 26°C, approximately 27°C, approximately 28°C, approximately 29°C, approximately 30°C, approximately 31°C, approximately 32°C, approximately 33°C, approximately 34°C, approximately 35°C, approximately 36°C, approximately 37°C, approximately 38°C, approximately 39°C, approximately 40°C).
[0195] In some embodiments, the annealing process may further include a pre-cooling process between heating and cooling, wherein the final temperature of the pre-cooling may be lower than the heating temperature and higher than the final cooling temperature.
[0196] During the formation of the intermediate, since it only involves the formation of hydrogen bonds, and hydrogen bonds are non-covalent bonds, they are less stable than covalent bonds and are easily broken by external forces or dynamically changed. Furthermore, the fewer hydrogen bonds a component has, the less stable it is. Therefore, the intermediate connected by hydrogen bonds can also be an unstable structure, especially those components where the homology between overlapping sequences is less than 100%. In overlapping sequences with less than 100% homology, not every pair of corresponding bases has a hydrogen bond. Therefore, hydrogen bonds are relatively fewer compared to overlapping sequences with 100% homology, and the structure is relatively unstable. These overlapping sequences are more prone to hydrogen bond breakage, leading to unpairing of the overlapping sequences. In this case, the two nucleic acids formed by the breakage have dangling sequences that can be complementary to other nucleic acids again. These dangling sequences can potentially complement nucleic acids containing 100% homology to form intermediates. Alternatively, they may continue to complement nucleic acids containing less than 100% homology to form intermediates. In this case, the intermediate is more likely to repeat the previous hydrogen bond breakage process until all overlapping sequences in the intermediate have 100% homology or the break in the intermediate is repaired. Through this process of hydrogen bond breakage and reconnection, the intermediate can spontaneously form a high-fidelity nucleotide sequence.
[0197] The intermediate can be considered as a long double-stranded nucleic acid lacking a partial phosphodiester bond, which is typically located at the interface between the short single-stranded nucleic acid and another adjacent short single-stranded nucleic acid on the same single strand. Therefore, the intermediate should be considered to have approximately the same length and molecular weight as the long double-stranded nucleic acid. Thus, methods characterizing nucleic acid length or molecular weight can be used to detect the formation of the intermediate. The method for characterizing nucleic acid length can be selected from one or more of the following: gel electrophoresis, real-time quantitative PCR, ultracentrifugation, capillary electrophoresis, nanoparticle tracking analysis, nucleic acid sequencing, spectroscopic methods, mass spectrometry, atomic force microscopy, optical microscopy, flow cytometry, nuclear magnetic resonance, and electron microscopy, preferably gel electrophoresis. The gel electrophoresis can be agarose gel electrophoresis or polyacrylamide gel electrophoresis.
[0198] In some embodiments, the intermediate for the same long double-stranded nucleic acid can be one, two, or more. In a preferred embodiment, the presence of two or more intermediates occurs when the long double-stranded nucleic acid is a circular nucleic acid or a long linear nucleic acid.
[0199] The intermediate is usually present in the form of a mixture after its formation. The mixture may contain unreacted short single-stranded nucleic acids, or it may contain any intermediate state from which the intermediate is formed. Therefore, the mixture can be separated and purified to obtain a purer intermediate, or it can be directly proceeded to the downstream steps without separation and purification.
[0200] The separation method can be based on the following properties of nucleic acids: molecular weight, charge, density, specific binding to a specific ligand, and isoelectric point, preferably molecular weight. The molecular weight-based separation method can be selected from one or more of the following: gel electrophoresis and mass spectrometry, preferably gel electrophoresis. The gel electrophoresis can be agarose gel electrophoresis or polyacrylamide gel electrophoresis, preferably agarose gel electrophoresis.
[0201] The purification method is typically performed after the separation and generally includes: (1) separating bands containing the same molecular weight as the intermediate using an instrument (e.g., an instrument emitting ultraviolet light); (2) removing the electrophoretic medium from the system; and (3) removing other impurities. In (1), the separation method can be physical separation, such as cutting with a knife, or chemical separation, or biological separation; in (2), the method for removing the electrophoretic medium can be heating to dissolve, enzymatic digestion, or transferring the target band to another material before purification; in (3), the removal method can be using a filter column or adsorption column, or adding a washing solution. The purification method can be performed using a commercial kit or by performing the purification in-house.
[0202] The evaluation criteria for the intermediates are also related to the synthetic methods involved in this application. The evaluation criteria for the intermediates can be yield, which is the ratio of the mass of the intermediate obtained after isolation and / or purification to the total mass of the short single-stranded nucleic acid from which the synthesis began. The evaluation criteria for the intermediates can also be accuracy, which refers to the percentage of the final synthesized product that completely conforms to the expected sequence. This completely conforming product can be a product that is completely identical to the long double-stranded nucleic acid sequence.
[0203] Short single-stranded nucleic acids
[0204] Short single-stranded nucleic acids are the smallest units that make up intermediates or long double-stranded nucleic acids in this application. The short single-stranded nucleic acid that makes up the long double-stranded nucleic acid should be completely identical in sequence to a certain segment of the long double-stranded nucleic acid to be synthesized.
[0205] The short single-stranded nucleic acid can be designed by dividing the sequence of the long double-stranded nucleic acid to be synthesized into an appropriate number of parts, and the division method should meet the following conditions.
[0206] (1) Except for the short single-stranded nucleic acids at one end of a single strand and the other end of the opposing single strand, which may have overlapping sequences with only one other short single-stranded nucleic acid, the remaining short single-stranded nucleic acids should have overlapping sequences with at least two other short single-stranded nucleic acids besides themselves. The number of overlapping sequences or the number of hydrogen bonds formed should maintain a relatively stable complementary pairing state. Preferably, the overlapping sequence is greater than or equal to 10 base pairs or the number of hydrogen bonds formed is greater than or equal to 20. More preferably, the overlapping sequence is greater than or equal to 16 base pairs or the number of hydrogen bonds formed is greater than or equal to 32.
[0207] (2) All short single-stranded nucleic acids may be of the same length or of different lengths. The length of all short single-stranded nucleic acids should be greater than or equal to the overlapping sequence of one short single-stranded nucleic acid or greater than or equal to the sum of the overlapping sequences of one short single-stranded nucleic acid and the other short single-stranded nucleic acids, preferably equal to the overlapping sequence of one short single-stranded nucleic acid or the sum of the overlapping sequences of one short single-stranded nucleic acid and the other short single-stranded nucleic acids.
[0208] (3) The length of the short single-stranded nucleic acid can be between 10 and 300 nucleotides, preferably between 20 and 300 nucleotides, more preferably between 30 and 300 nucleotides, even more preferably between 40 and 300 nucleotides, and even more preferably between 54 and 300 nucleotides, and even more preferably between 60 and 300 nucleotides. The short single-stranded nucleic acid at one end of one single strand and the other end of the opposing single strand should be greater than or equal to 10 nucleotides, and the remaining short single-stranded nucleic acids should be greater than or equal to 20 nucleotides.
[0209] The long double-stranded nucleic acid composed of the short single-stranded nucleic acid can be a nucleic acid with a classic linear double-stranded structure or a non-classical structure, such as a triangular Y-shaped or cross-shaped structure. In some embodiments, the long double-stranded nucleic acid is a long double-stranded nucleic acid with a classic linear structure, and the short single-stranded nucleic acid may have overlapping sequences with one or two other short single-stranded nucleic acids. In some embodiments, the long double-stranded nucleic acid is a triangular Y-shaped or cross-shaped structure, and the short single-stranded nucleic acid may have overlapping sequences with one, two, three, or even four or more other short single-stranded nucleic acids.
[0210] The short single-stranded nucleic acid may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The short single-stranded nucleic acid may or may not contain modifications, and examples of modifications may be selected from one or more of the following: methylation, phosphorylation, acetylation, adenylation, uridineation, glycosylation, sulfur modification, pseudouracilization, and ribourethridineization.
[0211] The short single-stranded nucleotide may contain classical nucleotides that make up an organism. For example, the classical nucleotide may be selected from one or more of the following bases: adenine, guanine, cytosine, thymine, and uracil. The short single-stranded nucleic acid may contain non-classical nucleotides. For example, the non-classical nucleotide may be selected from one or more of the following bases: methylcytosine (e.g., 5-methylcytosine), methyladenine (e.g., N6-methyladenine), pseudouracil, ribouracil, xanthine, and hypoxanthine.
[0212] The short single-stranded nucleic acid can be a conventional short single-stranded nucleic acid or a short single-stranded nucleic acid with a special structure. For example, the short single-stranded nucleic acid can contain sequences that are complementary to other sequences on the same short single-stranded nucleic acid, called self-pairing sequences, and the short single-stranded nucleic acid containing self-pairing sequences can form a stem-loop structure. The content of guanine and cytosine on the short single-stranded nucleic acid can be arbitrary, for example, 0-about 10%, about 10%-about 20%, about 20%-about 30%, about 30%-about 40%, about 40%-about 50%, about 50%-about 60%, about 60%-about 70%, about 70%-about 80%, about 80%-about 90%, about 90%-about 100%.
[0213] The short single-stranded nucleic acid can be synthesized de novo using a single nucleotide as a starting material after designing the target sequence. The synthesis method can be selected from one or more of the following: solid-phase phosphoramidite method, photochemical deprotection method, electrochemical deprotection method, inkjet printing method, integrated circuit control method, high-throughput parallel synthesis based on sorting principle, enzymatic synthesis method, mixed enzyme-mediated enzymatic reaction, TdT-dNTP cross-linked enzymatic reaction, and dibasic monomer synthesis method. The short single-stranded nucleic acid can also be obtained by finding an existing nucleic acid fragment containing the target sequence and obtaining the fragment through a truncation method, which can be selected from one or more of the following: restriction endonuclease digestion method, PCR method, gene gun method, and magnetic bead capture method. The short single-stranded nucleic acid can also be synthesized using a single nucleic acid and nucleotide that are 100% homologous to the target sequence but shorter than the target sequence as starting material, and the synthesis method can refer to the above-described de novo synthesis methods. The short single-stranded nucleic acid can also be synthesized using two or more nucleic acids that are 100% homologous to the target sequence but shorter than the target sequence as starting material.
[0214] After synthesis, the short single-stranded nucleic acid can be purified before downstream steps, or it can be directly proceeded to downstream steps without purification. The method for purifying the short single-stranded nucleic acid can be selected from one or more of the following: high-performance affinity adsorption (HAP), polyacrylamide gel electrophoresis (PAGE), enhanced polyacrylamide gel electrophoresis (ULTRAPAGE), high-performance liquid chromatography (HPLC), high-performance liquid chromatography combined with capillary electrophoresis (HPLC-CE), reversed-phase purification filter cartridge (RPC), specific adsorption of whole sequence (ePAGE), oligonucleotide purification column (OPC), reversed-phase chromatography (RP-HPLC), ion exchange chromatography (IE-HPLC), with polyacrylamide gel electrophoresis (PAGE) being the preferred method.
[0215] Long double-stranded nucleic acid
[0216] In this application, the intermediate is repaired to form the long double-stranded nucleic acid, which can be accomplished by forming phosphodiester bonds between the nucleotides on both sides of the cleavage on the intermediate. The phosphodiester bonds can be formed by enzymatic reactions (e.g., polymerases, ligases, etc.), or not by enzymatic reactions, or by in vivo nucleic acid repair, or by other possible methods, preferably by in vivo nucleic acid repair and / or enzymatic reactions.
[0217] The in vivo nucleic acid repair refers to introducing the intermediate into a host with phosphodiester bond synthesis function, and using the host's built-in repair system to repair the cracks on the intermediate. The host may contain cells in the division phase. The host may be a multicellular organism or an in vitro cell line derived from a multicellular organism, or a single-celled organism, preferably a single-celled organism. The single-celled organism may be a single-celled animal, a single-celled plant, or a single-celled microorganism, preferably a single-celled microorganism. The single-celled microorganism may be selected from one or more of the following: *Escherichia coli*, *Radiata-resistant Cocci*, *Bacillus subtilis*, thermophilic archaea, *Saccharomyces cerevisiae*, and *Bacillus buddingus*, preferably *Escherichia coli* and *Bacillus buddingus*.
[0218] In the nucleic acid repair method, the intermediate can be directly transferred into the host, or the intermediate can be first linked to a vector before being transferred into the host. The vector can be a biological or non-biological organism that helps the intermediate repair phosphodiester bonds within the host or helps the intermediate enter the host; preferably, it is an expression vector. The expression vector can be selected from one or more of the following: plasmids, bacteriophages, artificial chromosomes, transposons, introns, and viruses; plasmids are preferred.
[0219] The ligation method can be through ligation of homologous fragments on the vector and the intermediate, or through ligation of complementary pendant sequences on the vector and the intermediate. The ligation method using homologous fragments can be performed using homologous recombinases or without homologous recombinases. The ligation method using pendant sequences can be performed using ligases or without ligases. The method of introducing the homologous fragment into the intermediate can be by adding the target homologous fragment to the long double-stranded nucleic acid used as a design template during the design of the short single-stranded nucleic acid. The homologous fragment can be added to both ends of the long double-stranded nucleic acid or to one end. The homologous fragment can also be added after the intermediate is formed, for example, through a nucleotide-based synthetic method. The method for introducing the pendant sequence into the intermediate can be as follows: when designing the short single-stranded nucleic acid, the target pendant sequence is added to the long double-stranded nucleic acid, which serves as the design template; or when designing the short single-stranded nucleic acid, the target pendant sequence and its complementary pairing sequence are added to the long double-stranded nucleic acid, which serves as the design template, and the complementary pairing sequence of the pendant sequence is subjected to enzymatic digestion or digestion (e.g., enzymatic hydrolysis, acidic hydrolysis, alkaline hydrolysis, pyrolysis, chemical hydrolysis, microbial fermentation, photolysis, oxidizing agents, reducing agents, electrochemical methods, etc.) after the intermediate is formed.
[0220] In the host, after the intermediate is repaired to form the long double-stranded nucleic acid, the long double-stranded nucleic acid can be amplified within the host or not. The amplification of the long double-stranded nucleic acid can be performed within the host or after isolating the long double-stranded nucleic acid from the host. The isolation method can be selected from one or more of the following: chemical extraction, enzymatic digestion, ultrasonic disruption, freeze-thaw reaction, sodium dodecyl sulfate (SDS) method, hexadecyltrimethylammonium chloride (CTAB) method, commercial kits, magnetic bead method, gel electrophoresis purification, and density gradient centrifugation. The amplification method can be selected from one or more of the following: polymerase chain reaction (PCR), cycle-mediated isothermal amplification (LAMP), single-primer isothermal amplification (SPIA), sequence-dependent amplification (NASBA), multiple substitution amplification (MDA), CRISPR-Cas9 system, and in vitro transcription (IVT).
[0221] The enzymatic reaction refers to the formation of phosphodiester bonds through the ligation activity of a polymerase or the ligase itself. The polymerase may be selected from one or more of the following: DNA polymerase, RNA polymerase, reverse transcriptase, terminal deoxynucleotidyl transferase, polyadenylate polymerase, uracil-DNA glycosidase, and Qβ-RNA-dependent RNA polymerase. The ligase may be selected from one or more of the following: T4 DNA ligase, E. coli DNA ligase, T7 DNA ligase, RNase H ligase, ligase I, ligase II, ligase III, and ligase IV, preferably T4 DNA ligase.
[0222] In the aforementioned nucleic acid repair method, in vivo nucleic acid repair and enzymatic reactions can be combined. Preferably, the enzyme used in the enzymatic reaction can be a ligase. In specific embodiments, the enzymatic reaction can occur before the intermediate or the vector containing the intermediate is transferred into the host, or after the intermediate or the vector containing the intermediate is transferred into the host; it can occur before the intermediate is ligated to the vector, or it can occur after the intermediate is ligated to the vector. In a preferred embodiment, the combined reaction can occur when the number of nicks is greater than or equal to 41, 51, 61, 71, 81, 91, or 101.
[0223] The long double-stranded nucleic acid can be a long double-stranded nucleic acid without a pendant sequence, or it can be a long double-stranded nucleic acid containing one or more pendant sequences. Therefore, the intermediate may or may not contain pendant sequences. The long double-stranded nucleic acid can be a linear nucleic acid or a circular nucleic acid. Therefore, the intermediate can also be a linear nucleic acid or a circular nucleic acid.
[0224] The long double-stranded nucleic acid can be a conventional long double-stranded nucleic acid, or it can be a long double-stranded nucleic acid containing multiple fragments with the same sequence or the short single-stranded nucleic acid.
[0225] The long double-stranded nucleic acid can be of any length, preferably greater than or equal to 500 bases or base pairs, 800 bases or base pairs, 1000 bases or base pairs, 1200 bases or base pairs, 1500 bases or base pairs, 1800 bases or base pairs, 2000 bases or base pairs, 2500 bases or base pairs, 3000 bases or base pairs, 3300 bases or base pairs, 4000 bases or base pairs, 6000 bases or base pairs, 9000 bases or base pairs, or 9900 bases or base pairs.
[0226] The long double-stranded nucleic acid may contain two or more sequences. In some embodiments, the long double-stranded nucleic acid contains two or more sequences, which may be on the same long double-stranded nucleic acid or on different long double-stranded nucleic acids; the sequences may be contained in different expression frames or in the same expression frame to form a fusion gene or fusion protein. The sequences may be protein-coding sequences (including DNA and RNA), non-coding RNA (e.g., transfer RNA, ribosomal RNA, etc.) or sequences expressing said non-coding RNA, or regulatory sequences (e.g., promoters, enhancers, silencers, insulators, and response elements, etc.). In some embodiments, the sequences are protein-coding sequences, and the long double-stranded nucleic acid may also contain nonsense sequences, which are sequences that do not encode proteins. Examples of nonsense sequences may include sequences from one or more of the following: introns, intergenic regions, non-coding RNA or sequences expressing said non-coding RNA, regulatory sequences, repetitive sequences, transposons, and pseudogenes.
[0227] In this application, the long double-stranded nucleic acid can be evaluated by fidelity. The fidelity can refer to the proportion of the actual sequence of the synthesized long double-stranded nucleic acid that has a higher homology to a standard sequence than a specific value when the sequence of the synthesized long double-stranded nucleic acid is sampled and tested. This specific value can be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5%. The actual sequence of the long double-stranded nucleic acid can be obtained by sequencing. The standard sequence can be a sequence of the long double-stranded nucleic acid recognized by those skilled in the art, preferably a standard sequence from an authoritative website (e.g., NCBI) gene database.
[0228] In this application, the long double-stranded nucleic acid can be evaluated by an error-free rate. The error-free rate can refer to the proportion of times the actual sequence of the synthesized long double-stranded nucleic acid shows 100% homology with a standard sequence when the sequence of the synthesized long double-stranded nucleic acid is sampled and tested. The actual sequence of the long double-stranded nucleic acid can be obtained by sequencing. As noted by those skilled in the art, "error-free" and "all correct" are used interchangeably.
[0229] In this application, when the synthesized long double-stranded nucleic acid is a linear nucleic acid, the intermediate usually needs to be assembled with a vector and transferred into the host to complete the synthesis or even amplification of the long double-stranded nucleic acid.
[0230] In some embodiments, the assembly can be performed using the Gibson assembly method. In some embodiments, the assembly can be performed using the chain substitution method.
[0231] In some embodiments, the assembly can be performed by direct, nonspecific ligation using a ligase.
[0232] The method for preparing the long double-stranded nucleic acid library protected in this application is applicable to any protein. After determining the original long double-stranded nucleic acid, the amino acid sequence encoding the protein, and the selected mutation sites and mutation types, the long double-stranded nucleic acid is divided into short single-stranded nucleic acids. The mutation sites are distributed to the short single-stranded nucleic acids, and the corresponding short single-stranded nucleic acids are synthesized. The long double-stranded nucleic acid library is then synthesized according to the synthesis method protected in this application, thereby obtaining the long double-stranded nucleic acid library. For the synthesis of the long double-stranded nucleic acid library described in this application, at a minimum, only the sequence of the original long double-stranded nucleic acid or the amino acid sequence of its encoded protein, and the required mutations are needed.
[0233] mutant protein
[0234] In a specific embodiment, the original long double-stranded nucleic acid can be a protein coding sequence. The protein can be selected from one or more of the following: catalytic proteins, reporter proteins, therapeutic proteins, and preventative proteins. Since this application also involves the screening process for functional proteins, from the perspective of ease of screening, selecting catalytic proteins and reporter proteins for the synthesis testing of the long double-stranded nucleic acid library is a better choice. However, it is worth noting that the establishment of nucleic acid libraries or protein libraries related to therapeutic and preventative proteins is also important for the fields of biosynthesis and other fields (such as drug screening).
[0235] It is worth noting that the protein encoded by the original long double-stranded nucleic acid is not necessarily a protein that originally exists in nature (i.e., a wild-type protein), but may also be an artificially modified protein. The artificial modification can be the modification of the protein, and examples of such modifications include, but are not limited to, phosphorylation, glycosylation, acetylation, methylation, ubiquitination, SUMOylation, biotinylation, lipidation, nitration, disulfide bond formation, protein mutation (such as the insertion, deletion, and / or substitution of amino acids), chemical cross-linking, and tagging.
[0236] Examples of catalytic proteins may include, but are not limited to, amylases, proteases, lipases, ATP synthases, hexokinases, DNA polymerases, RNA polymerases, protein kinases, lactate dehydrogenases, PET hydrolases, and their variants. Examples of reporter proteins may include, but are not limited to, fluorescent proteins (FP), luciferase, β-galactosidase (β-galactosidase or LacZ), chloramphenicol acetyltransferase (CAT), secretory alkaline phosphatase (SEAP), mCherry, mNeonGreen, HaloTag, SNAPTag, and their variants. Examples of therapeutic proteins may include, but are not limited to, antibodies, recombinant enzymes, hormones, growth factors, coagulation factors, cytokines, and therapeutic vaccines. Examples of prophylactic proteins may include, but are not limited to, prophylactic vaccines.
[0237] In a specific embodiment, the original protein encoded by the original long double-stranded nucleic acid is a fluorescent protein (FP). It is worth noting that fluorescent protein can refer to a class of proteins, that is, proteins that emit visible fluorescence when excited by light of a specific wavelength can be called fluorescent proteins; the visible fluorescence can be visible to the naked eye or visible with the aid of instruments. The fluorescent protein may possess the following characteristics: self-luminescence, spectral properties, stability, integrability, and diversity. Fluorescent proteins have wide applications in cell biology, molecular biology, neurobiology, developmental biology, and bioengineering, such as for live-cell imaging, protein-protein interaction analysis, and real-time monitoring of gene expression. Therefore, constructing libraries of fluorescent proteins and screening for functional fluorescent proteins is of great significance to the development of these fields.
[0238] In a specific embodiment, the original protein encoded by the original long-chain double-stranded nucleic acid is polyethylene terephthalate hydrolase (PET hydrolase). The PET hydrolase can hydrolyze specific chemical bonds. In this application, the PET hydrolase is used to hydrolyze fluorescein dibenzoate (FDBz). After being hydrolyzed by the PET hydrolase, FDBz releases fluorescein that emits green fluorescence, as shown in Figure 27B.
[0239] Mutation type
[0240] In a specific implementation, for the long double-stranded nucleic acid library encoding the original long double-stranded nucleic acid, the fluorescent protein should be understood as a protein containing the amino acid sequence shown in SEQ ID NO:16. SEQ ID NO:16 may be the backbone sequence of the fluorescent protein, and any modifications thereto may alter its function, potentially affecting fluorescence or other functions (such as protein folding or protein solubility); it may enhance or weaken protein function.
[0241] The modifications occurring on the fluorescent protein backbone sequence may include, but are not limited to, phosphorylation, glycosylation, acetylation, methylation, ubiquitination, SUMOylation, biotinylation, lipidation, nitration, disulfide bond formation, protein mutation (such as amino acid insertion, deletion, and / or substitution), chemical cross-linking, and tagging, with protein mutation being preferred. The protein mutation may be amino acid insertion, deletion, or substitution, with substitution being preferred. The substitution can occur at classical sites on the fluorescent protein backbone sequence, which can be derived from currently available functional mutants, such as S65, S72, N149, M153, and I167 on Emerald, S65, S72, K79, and T203 on Topaz, F64, S65, Y66, N146, M153, and V163 on ECFP, and F64, Y66, and Y145 on EBFP; or it can be derived from regions associated with specific functions, such as F64 located in chromophore generation regions, F64, S72, and V163 located in protein folding-related regions, and M153 and V163 located in protein solubility-related regions. In a specific implementation, the classical site mutation can occur at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S65, Y66, S72, K79, Y145, N146, N149, M153, V163, I167, and T203. Additionally, a random mutation at the R168 site was also observed during the screening process. More specifically, the classical site mutation can be selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: F64L, S65T, S65G, S65C, S65A, Y66H, Y66Q, Y66R, Y66C, Y66W, S72A, K79R, Y145F, N146I, N149K, M153T, V163A, I167T, R168H, T203Y, T203K, and T203S.
[0242] In a specific implementation, for the long double-stranded nucleic acid library encoding the original long double-stranded nucleic acid PET hydrolase, the PET hydrolase should be understood as a protein containing the amino acid sequence shown in SEQ ID NO:17. SEQ ID NO:17 can be the main sequence of the PET hydrolase, and any modifications thereto may alter its function, potentially affecting catalytic function or other functions (such as protein folding or protein solubility); it may enhance or weaken protein function.
[0243] The modifications occurring on the PET hydrolase backbone sequence may include, but are not limited to, phosphorylation, glycosylation, acetylation, methylation, ubiquitination, SUMOylation, biotinylation, lipidation, nitration, disulfide bond formation, protein mutation (e.g., amino acid insertion, deletion, and / or substitution), chemical cross-linking, and tagging, preferably protein mutation. The protein mutation may be an amino acid insertion, deletion, or substitution, preferably substitution. The substitution may occur at classical sites on the PET hydrolase backbone sequence, which may originate from currently obtained functional mutants; or from regions related to a specific function. In a specific embodiment, the classical site mutation may occur at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:17: E9, A30, P53, S110, T160, G178, S184, T191, and A214; more specifically, the classical site mutation may be selected from those occurring in SEQ ID NO:17. One or more of the following mutations in the amino acid sequence shown in NO:17: E9A, E9Q, E9P, E9D, A30K, A30N, A30T, A30M, A30I, A30Q, A30H, A30P, A30L, A30E, A30D, A30V, A30Y, A30S, A30F, P53A, P53G, P53V, P53R, P53L, S110N, S110K, S110T, S110R, S110H, S110Q, S110P, S110D, S1 10E, S110A, S110G, T160K, T160R, T160Q, T160P, T160I, G178R, G178S, G178T, G178A, G178C, G178W, S184K, S184N, S184T, S184P, S184Q, T191K, T191N, T191Q, T191H, T191P, T191E, T191D, T191A, A214N, A214T, A214S, A214D, and A214G.
[0244] Compared to the methods commonly used in the prior art for constructing libraries by inducing random mutations, the mutation sites selected in this application are all existing mutation sites that have been proven to have an impact on function, resulting in a higher hit rate in screening.
[0245] When designing the short single-stranded nucleic acids, the classical site mutation may or may not appear in any single short single-stranded nucleic acid, or it may appear once, twice, three times, four times, or even the entire short single-stranded nucleic acid may contain the classical site mutation. It is important to note that the classical site mutation usually refers to an amino acid mutation; therefore, its expression in the short single-stranded nucleic acid should be determined through codon conversion, which can refer to the classical correspondence of one amino acid to three nucleotides. Since the same amino acid may correspond to multiple codons, the expression of the classical site mutation in the short single-stranded nucleic acid should be more diverse than its expression in the amino acid sequence.
[0246] Filtering methods
[0247] After synthesizing the long double-stranded nucleic acid library, this application also relates to a screening scheme for the protein library encoded by the long double-stranded nucleic acid library, the main purpose of which is to screen for proteins with functional changes. These functional changes can be enhancements or reductions; they can be reductions or additions of a particular function.
[0248] For the protein libraries obtained above, the functions of different types of proteins usually vary.
[0249] In one embodiment, the protein is a fluorescent protein, whose function is typically to receive excitation light and emit fluorescence. Therefore, for a fluorescent protein, the functional change can be a change in the spectral range or a change in light intensity. The change in spectral range can be a change in the spectral range of the excitation light that can be received, specifically a redshift or a violetshift; the change in spectral range can also be a change in the spectral range of the emitted light, specifically a redshift or a violetshift. The change in light intensity can be a change in the minimum excitation light intensity required to excite the fluorescent protein to emit detectable emission light, specifically an increase in the minimum light intensity (i.e., a decrease in the photosensitivity of the fluorescent protein) or a decrease in the minimum light intensity (i.e., an increase in the photosensitivity of the fluorescent protein); the change in light intensity can also be a change in the intensity of the emitted light by the fluorescent protein under the same excitation light intensity, specifically an increase in the emission light intensity or a decrease in the emission light intensity. Detection of functional changes in fluorescent proteins is typically performed using one or more of the following methods: fluorescence microscopy, single-photon counting, flow cytometry, fluorescence resonance energy transfer (FRET), and bimolecular fluorescence complementation (BiFC).
[0250] In another embodiment, the protein is a PET hydrolase, whose function is typically to hydrolyze ester bonds. Therefore, for PET hydrolase, the functional change can be a change in hydrolase activity, either increased or decreased. Hydrolase activity can be characterized by the amount of hydrolysis products produced or hydrolysis substrate consumed per unit time. Increased hydrolysis product production or decreased hydrolysis substrate consumption per unit time indicates increased hydrolase activity per unit volume; conversely, decreased hydrolysis product production or increased hydrolysis substrate consumption per unit time indicates decreased hydrolase activity per unit volume. In a specific embodiment, the hydrolysis substrate can be fluorescein dibenzoate (PDB2), and the corresponding hydrolysis product is fluorescein. Therefore, the amount of hydrolysis products produced per unit time can be characterized by detecting the luminescence intensity of the fluorescein product, thereby characterizing the change in PET hydrolase activity. Besides screening proteins with altered functions, the protein libraries obtained by the synthesis method described in this application can also be used for screening in other aspects, such as screening based on physical properties, chemical properties, or biological properties. The physical properties can be selected from one or more of the following: molecular weight, isoelectric point, solubility, stability, absorbance coefficient, light scattering, thermal stability, mechanical stability, hydrodynamic radius, secondary structure, tertiary structure, and quaternary structure. The chemical properties can be selected from one or more of the following: reactivity of amino and carboxyl groups, side chain reactivity, phosphorylation, glycosylation, oxidation, reduction, alkylation, acid hydrolysis, enzymatic hydrolysis, incorporation of non-natural amino acids, coordination, hydrophobic interaction, ion exchange, and affinity chromatography. The biological properties can be selected from one or more of the following: enzyme activity, structural function, transport function, signal transduction, immune response, gene regulation, cell recognition, growth factor activity, cell cycle regulation, apoptosis regulation, cell differentiation regulation, metabolic pathway participation, molecular chaperone function, neurotransmission, photosensitivity, and hormone activity.
[0251] Without being limited by any theory, the embodiments described below are merely for illustrating the various technical solutions of the present invention and are not intended to limit the scope of the present invention.
[0252] Example
[0253] The water used to dissolve the nucleic acids in this application was DEPC water purchased from Invitrogen, the primers were purchased from Sangon Biotech, and all enzymes (unless otherwise noted) were purchased from New England Laboratories. The flowchart for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids is shown in Figure 1. This synthetic method is named MOSAIC.
[0254] As shown in Examples 1-4, the simplified flow of the MOSAIC method described in this application is as follows:
[0255] (1) Synthesize and purify short-chain single-stranded nucleic acids;
[0256] (2) The purified short-chain single-stranded nucleic acid was annealed and assembled into an intermediate, which was then separated and purified.
[0257] (3) After the separated intermediate is connected to the carrier, it is transformed into the host and the host's built-in repair system is used to repair the cracks on the intermediate.
[0258] Example 1: Synthesis and Purification of Short Single-Stranded Nucleic Acids
[0259] Except for the short single-stranded nucleic acids used in library synthesis, which were purchased from Integrated DNA Technologies, all short single-stranded nucleic acids used in this application were artificially synthesized nucleic acid molecules. Specifically, the artificial synthesis method was column synthesis based on the phosphoramide method. All short single-stranded nucleic acids required purification after synthesis.
[0260] To purify short single-stranded nucleic acids using PAGE purification, samples to be used in the same mixture system are purified in batches in one reaction. The specific steps are as follows.
[0261] Prepare 50 ml of 8% denaturing polyacrylamide gel [8% acrylamide-bisacrylamide (19:1), 8M urea, 1xTBE (89mM Tris-borate, 89mM boric acid, 2mM EDTA, pH 8), 0.4% APS, 0.04% TEMED]. Prepare short-chain single-stranded nucleic acid loading samples by adding an equal volume of 2xTBE, 8M urea loading buffer to a synthetically produced short-chain single-stranded nucleic acid mixture. Vortex the samples and incubate at 95°C for 5 minutes to denature the DNA, then maintain at 50°C until ready for loading. After pre-running the gel at 300V (i.e., preheating) for 15 minutes, perform gel electrophoresis at 300V for 1 hour in a 50°C water bath. Cut the target bands from the gel using a UV lamp. Place the strip into a 1.5 ml centrifuge tube, crush the gel, and add 800 μl of 1x TE (10 mM Tris-Cl, 1 mM EDTA, pH 8). Freeze at -20°C until solidified, then shake overnight at 1,900 rpm in a shaker at 4°C. Centrifuge at 12,000 rpm for 2 min and collect the supernatant. Resuspend the precipitate in 700 μl of 1x TE, centrifuge at 12,000 rpm for 2 min, and collect the supernatant. Repeat this step once. Collect all the supernatant into a 10 ml centrifuge tube, add 3 volumes of n-butanol, vortex, centrifuge at 7,800 rpm for 2 min, collect the aqueous phase, and repeat this step until the aqueous phase is less than 500 μl. Transfer the aqueous phase to a new 1.5 ml tube, add 1 / 10 volume of 3M NaOAc and 2 volumes of 100% ethanol. Incubate at -20°C for 30 min, then centrifuge at 12,000 rpm for 20 min at 4°C. Discard the supernatant and air dry for 10 min to obtain the purified DNA sample. Resuspend the DNA in 400 μl of DEPC water and transfer to a 10K filter column. Centrifuge at 12,000 rpm for 15 min, then add 400 μl of DEPC water and centrifuge twice at 12,000 rpm for 15 min. Determine the concentration of the remaining solution using a nanodrop microvolume spectrophotometer.
[0262] As shown in Figure 2, the short single-stranded nucleic acids purified by PAGE showed better assembly results than the unpurified ones. The assembly method is shown in Example 2. The intermediates of the purified short single-stranded nucleic acid assembly formed clearly visible bands on the gel, while the unpurified ones did not have any bands.
[0263] Example 2: Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids
[0264] In this application, short-chain single-chain nucleic acids are mainly synthesized as intermediates through an annealing process.
[0265] In a 1x TE buffer containing 10 mM MgCl2, the short single-stranded nucleic acids required for the synthesis of the long double-stranded nucleic acid were mixed and adjusted to a total concentration of approximately 2 μM (calculated using mass concentration and the average molecular weight of the short single-stranded nucleic acids) or 20 μL of solution containing 2–10 μg of short single-stranded nucleic acids were used. The short single-stranded nucleic acid mixture was then annealed in a gradient thermal cycler to form the intermediate. The annealing procedure was as follows: 95 °C for 5 min, 70 °C for 30 min, and then cooled to 16 °C.
[0266] After further optimization by the inventors, the following annealing procedure can be applied to the synthesis of all long double-stranded nucleic acids: react at 95°C for 5 min, react at 80°C for 20 min, and then decrease the temperature by 1°C every 15 min until it is cooled to 16°C.
[0267] Example 3: Separation and purification of intermediates from annealing products
[0268] A 1% agarose gel was prepared using 0.5×TBE buffer supplemented with 10 mM MgCl2 as the solvent and pre-stained with SYBR Safe (Thermo Scientific). Annealed samples were loaded and electrophoresed at 90 V in an ice-water bath. The gel was imaged using Typhoon. Target gel bands were then cut, carefully crushed in a Freeze'N Squeeze column (Bio-Rad), and directly centrifuged at 1,000 rpm for 2 minutes at room temperature. Concentrations were measured using a nanodrop.
[0269] Example 4: Repairing cracks on the intermediate body
[0270] Vector preparation: The vector was pUC19 plasmid, purchased from / from Tiangen. The plasmid vector was generated using PCR with 2x Taq Master Mix (Sangon Biotech) according to the manufacturer's instructions. The PCR product was digested with the restriction endonuclease DpnI, and after electrophoresis (100V, 30 min), it was extracted from an agarose gel and recovered using the GeneJET PCR Purification Kit (Thermos Scientific) according to the manufacturer's instructions. If the vector is used for the strand substitution method shown in Figure 3B, it needs to be digested with a nick enzyme (37℃ for 1 h; the enzyme type is determined by the recognition sequence).
[0271] Insert preparation: To reduce the truncation rate of sequencing results, the obtained product can be pre-ligated. The reaction system includes 17 μl of intermediate, 2 μl of 10x T4 ligase buffer, and 1 μl of T4 polynucleotide kinase (NEB). After reacting at 37°C for 2 h, 1 μl of T4 ligase is added, and the reaction is carried out at 25°C for 2 h, followed by a final reaction at 52°C for 1 h.
[0272] Recombination: Gibson assembly was performed using the HiFi DNAAssembly Kit according to the instructions. For strand displacement recombination, the vector and the treated intermediate were mixed in a 0.2 ml tube and reacted at 50°C for 20 min. The ratio of vector to intermediate was 1:1.
[0273] Transformation: Gently mix 10 μl of reaction solution with 100 μl of competent Escherichia coli DH5α (TIANGEN). Incubate the mixture on ice for 20 minutes, heat shock at 42°C for 1 minute, place on ice for 3 minutes, then add 800 μl of LB bacterial liquid medium and incubate at 37°C for 40 minutes. Treat the incubated solution as follows: (1) Spread the solution on LB agar plates containing ampicillin or bleomycin, and send them for sequencing after the clones grow to detect gene synthesis; (2) Add the solution directly to 5 mL of LB liquid medium and incubate for 8 hours, and then perform plasmid extraction.
[0274] The schematic diagram of the Gibson assembly method is shown in Figure 3A and Figure 4. Gibson assembly mainly involves connecting homologous fragments between the vector and the intermediate.
[0275] The chain substitution recombination assembly method is shown in Figure 3B and Figure 15, and is briefly described as follows: (1) Double-strand cleavage is performed on the two nickase cleavage sites at the ends of the vector; (2) Single-stranded nucleic acid longer than the overlapping sequence is used to extend the cleavage to form the overhanging sequences at both ends of the vector; (3) The overhanging sequences of the intermediate are replaced with the shorter strand on the vector to assemble a circular double-stranded nucleic acid; (4) Subsequent tear repair work is performed.
[0276] Transformation method of yeast: The product of vector and fragment ligation and competent yeast were co-incubated at 30℃ for one hour. After incubation, the yeast was centrifuged at 2000rpm for 2min, resuspended in 100μL of pure water, spread on selective plates, and cultured at 30℃ for two days. The surviving clones were observed, and the surviving clones were selected, expanded, plasmids were extracted and transformed into E. coli, and then E. coli were sent for sequencing to detect whether gene synthesis occurred.
[0277] Example 5: Synthesis of a single conventional long-chain double-stranded nucleic acid
[0278] As shown in Figure 5A, the MOSAIC method described in this application mainly includes two stages: (1) Designing the target gene sequence into short single-stranded nucleic acids with overlapping sequences and synthesizing these short single-stranded nucleic acids, and then assembling them into intermediates through annealing hybridization. In one case, the intermediates, except for containing a cleft and possibly containing homologous ends that are assembled with the vector, should have other characteristics (such as the base sequence) consistent with the target gene. (2) After assembling the intermediates with the vector, transforming them into a host (such as Escherichia coli or yeast) to repair the cleft and amplify the target gene, and then randomly selecting clones for DNA sequencing to verify the synthesis yield of the target gene.
[0279] This application first tests the feasibility of MOSAIC by synthesizing the green fluorescent protein (GFP) encoding gene sequence. As shown in Figures 4 and 5B, the inventors first divided an 810 bp DNA sequence containing the GFP coding sequence and two homologous ends, generating 26 60 nt short single-stranded nucleic acids with overlapping sequences and two 30 nt short single-stranded nucleic acids (26 x 60-mers and 2 x 30-mers), where the overlap between each group of nucleic acids is 30. Simultaneously, the GFP coding sequence was filled with nonsense sequences to achieve the designed 810 bp fragment. Next, 28 short single-stranded nucleic acids of GFP were synthesized and purified in batches using the method of Example 1. The purified short single-stranded nucleic acids were annealed to assemble them into intermediates and characterized using the method of Example 3. The characterization results are shown in Figure 5C.
[0280] After ligation into the vector and transformation into *E. coli*, 16 single transformants were randomly selected for sequencing. Five of the 16 sequenced clones were confirmed to have gene sequences completely identical to the target gene. Furthermore, as shown in Figure 6, under the induction of isopropyl β-d-thiogactopyranoside, the inventors observed green fluorescent colonies on the plate, further confirming the successful assembly and expression of the gfp-encoding gene.
[0281] Example 6: Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids with Overlapping Sequences of Different Lengths
[0282] To investigate whether the length of overlapping sequences affects the assembly of intermediates, the inventors designed six groups of short single-stranded nucleic acids with different overlapping lengths, including those containing multiple different lengths. These were then annealed as described in Example 2 to assemble into intermediates, and characterized using agarose gel electrophoresis at 120V for 1 hour with 3% magnesium chloride. The characterization results and the overlapping sequence lengths of the short single-stranded nucleic acids composing the different intermediates are shown in Figure 7. The right side of Figure 7 shows the long double-stranded nucleic acid divided into multiple short single-stranded nucleic acids, with the length of each overlapping region listed below its location in each design. For example, in Group 1, the first overlapping region is 64 nt long, the second is 16 nt long, and so on. Correspondingly, the first short single-stranded nucleic acid separated from the upper single strand is 80 nt long, with overlapping sequences of 64 nt and 16 nt with the short single-stranded nucleic acids on the left and right sides. Since the purpose of this experiment was to investigate whether the length of overlapping sequences affects assembly, it was not strictly stipulated that the assembled intermediates should not contain overhanging sequences. As shown in Figure 7, the inventors set the minimum overlapping sequence to be 16nt and the maximum to be 40nt.
[0283] The gel electrophoresis results in Figure 7 show that all short single-stranded nucleic acids were successfully assembled into intermediates, so the reduction of overlapping sequences does not affect the assembly of intermediates.
[0284] Example 7: Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids with Special Sequences
[0285] To further evaluate the capabilities of MOSAIC, the inventors performed assembly tests on some short single-stranded nucleic acids containing specific sequences. As shown in Figures 5B, 8A, and 8B, the inventors synthesized three structures: (1) a 630 bp fragment (20 × 60-mers and 2 × 30-mers) with 74.3% self-paired hairpin structure (Figures 5B and 8A); (2) a 1650 bp fragment (54 × 60-mers and 2 × 45-mers) with a GC content of approximately 77% (Figure 5B); and (3) a 330 bp human microsatellite region (10 × 60-mers and 2 × 30-mers) composed of six tandem repeat sequences of the same 45-bp sequence (Figures 5B and 8B).
[0286] As shown in Figures 5C and 8C, using MOSAIC, we successfully assembled all three special fragments. After more accurate calculation of the sequencing results, the error-free rates were 62.5% (5 / 8), 66.7% (4 / 6), and 17.6% (3 / 17), respectively.
[0287] Meanwhile, the inventors used the PCA (polymerase cycle assembly) method to assemble the above three short single-stranded nucleic acids containing special sequences and characterized the products, the results of which are shown in Figure 9. As shown in Figure 9, the gfp sequence and self-paired sequence in Example 5 could be successfully assembled into the final product; high GC did not assemble or generated fragmented products; and repetitive sequences generated a series of products of indefinite length. A comparison of Figure 9 with Figures 5C and 8C shows that the long double-stranded nucleic acid synthesis method protected in this application has advantages in synthesizing special sequences.
[0288] Example 8: Testing the mismatch probability when synthesizing multiple sequences
[0289] To further explore whether MOSAIC can be adapted to a wider range of applications, the inventors hope to apply the MOSAIC system to a multi-gene synthesis platform, that is, to synthesize a large number of gene sequences from different sources in a single reaction. Before performing multi-gene synthesis, the inventors want to explore whether there are any interferences between different gene synthesizers that would lead to a decrease in yield or error-free rate.
[0290] The purpose of this embodiment is to test whether cross-pairing will occur between short single-stranded nucleic acids originating from different target long double-stranded nucleic acids when synthesizing multiple long double-stranded nucleic acids in the same reaction, resulting in chimeric sequences in the product due to interference. This embodiment tests the cross-pairing from two perspectives.
[0291] Testing long double-stranded nucleic acids of different lengths
[0292] The inventors investigated the cross-pairing of two genes (gfp and bacterial rhodopsin) of different lengths synthesized in parallel using MOSAIC. Figure 10A shows the short single-stranded nucleic acids after dividing the 810-bp gfp fragment (26×60-mers and 2×30-mers) and the 1,160-bp bacterial rhodopsin fragment (28×80-mers and 2×40-mers). These short single-stranded nucleic acids were mixed together in the same reaction system for annealing assembly as described in Example 2, and the assembled products were electrophoresed to observe whether the target intermediate was generated. As shown in Figure 10B, two bands of the same length as the target sequence appeared on the agarose gel, proving that there was no obvious cross-pairing between the two sequences. The products were then purified, and the purified assembly products were mixed for downstream transformation and sequencing. As shown in Figure 10C, the sequencing results showed one or more point mutations in some sequences, but no large-fragment sequence changes, further confirming the successful parallel synthesis of the two genes with a low probability of cross-pairing.
[0293] Testing the coding sequence of homologous genes
[0294] Considering that short single-stranded nucleic acid sequences derived from highly homologous fragments are more likely to have high similarity and are more likely to cross-pair, the inventors also evaluated the applicability of the MOSAIC method for assembling highly homologous sequence fragments.
[0295] As shown in Figure 11A, the two nr2f1 and nr2f2 gene fragments, each 930 bp in length and with a sequence similarity of up to 85%, were divided into 32 short single-stranded nucleic acids (30 × 60-mers and 2 × 30-mers), and 16-nt dangling sequences were added to each end.
[0296] Precise sequence alignment was performed on corresponding homologous fragments in the two fragments to determine the degree of homology between the corresponding short single-stranded nucleic acids in the two fragments. Figure 11B shows the number of identical bases on corresponding homologous fragments in the nr2f1 and nr2f2 fragments. As shown in Figure 11B, the nr2f1 and nr2f2 fragments have high homology, and the homologous sites are relatively dispersed, with almost all short single-stranded nucleic acids showing more than 70% homology.
[0297] The synthesized and purified short single-stranded nucleic acids were annealed and assembled in the same reaction, and the assembled products were electrophoresed to observe whether the target intermediate was generated. As shown in Figure 11C, a band of the same length as the target sequence appeared on the agarose gel. The product was then purified, and the purified assembly product was ligated to a vector for downstream transformation and sequencing. As shown in Figure 11D, some sequences in the sequencing results had one or more point mutations, but no cross-pairing sequence formation was observed, thus confirming that the two genes were synthesized in parallel and that the probability of cross-pairing was low.
[0298] Taking homologous gene coding sequences as an example, this study investigates the advantages of the MOSAIC method compared to traditional synthesis methods.
[0299] Since each short single-stranded nucleic acid in a homologous sequence has high homology, the probability of mismatch is much higher than that of other sequences. Therefore, the inventors decided to use homologous genes as an example to illustrate the advantages of the synthesis method protected in this application in reducing sequence mismatch compared to existing methods.
[0300] The inventors used the PCA method to assemble and characterize the short single-stranded nucleic acids derived from the homologous genes nr2f1 and nr2f2 in the same reaction, and then sequenced the products. The results are shown in Figure 12. As shown in Figure 12A, the PCA method successfully synthesized sequences of the same length as the final product. As shown in Figure 12B, the inventors compared each sequenced sequence with the standard sequences of nr2f1 and nr2f2, and displayed the comparison results of a portion of the sequences horizontally in Figure 12B, where mismatches are shown as white vertical lines. Among the obtained sequences, there were no sequences that completely matched the standard sequences of nr2f1 and nr2f2, i.e., the error rate was 0. As can be seen from Figure 12B, mismatches that appeared in the nr2f1 comparison results did not appear in the nr2f2 sequence, and vice versa. Therefore, it can be determined that the homologous genes synthesized by the PCA method had a large number of cross-pairings.
[0301] Therefore, it is evident that using MOSAIC for gene sequence synthesis has the advantage of a low mismatch rate.
[0302] Example 9: Preparation of a fluorescent protein mutant library
[0303] Leveraging the MOSAIC method's ability to synthesize multiple target nucleic acids in parallel and its convenience in independently designing each short single-stranded nucleic acid, it is used to prepare mutant libraries. Short single-stranded nucleic acids containing mutation sites self-assemble and combine to form long double-stranded nucleic acid products containing different mutation sites, thus forming a mutant library. In this embodiment, fluorescent proteins are used as an example to demonstrate the effectiveness of the methods for preparing long double-stranded nucleic acid libraries and protein libraries in this application, and their ability to screen for new functional mutant proteins.
[0304] As shown in Figures 13A, 14A, 22, and 28, the inventors selected 12 classic mutation sites on wild-type green fluorescent protein (GFP) and designed corresponding mutation sequences. They then incorporated short single-stranded nucleic acids containing mutations into wild-type GFP without mutations, and assembled the resulting fluorescent protein nucleic acid library. The specific steps are as follows.
[0305] The fluorescent protein gene coding sequence is divided into 60nt short single-stranded nucleic acids, of which 30nt is an overlapping sequence. At the same time, there is a 30nt short single-stranded nucleic acid at each end of the two non-complementary single strands of the long double-stranded nucleic acid. The short single-stranded nucleic acid may or may not contain mutations at the classic mutation sites. In this application, the GFP protein is used as the original long double-stranded nucleic acid, and its nucleic acid sequence is shown in SEQ ID NO:1. The inventors selected the codon sites corresponding to 12 classical amino acid mutation sites. These amino acid sites are: F64, S65, Y66, S72, K79, Y145, N146, N149, M153, V163, I167, and T203. The codon sites and corresponding mutation forms are shown in Figures 14A, 22, 28, and Table 1. Table 1 shows the short single-stranded nucleic acids designed on the GFP protein coding strand. The underlined triplet codons in Table 1 are the codons corresponding to the classical amino acid mutation sites. N, S, Y, etc., are degenerate bases, and their corresponding bases are shown in Table 2. It should be noted that the specific number of bases corresponding to the degenerate bases is the same as the mutation type of the mutation site on the short single-stranded nucleic acid.
[0306] In one embodiment, the short single-stranded nucleic acid on the corresponding template strand can be the original sequence without mutations, divided from the 5' end to the 3' end according to the rule that the first short single-stranded nucleic acid is 30nt and each subsequent short single-stranded nucleic acid is 60nt. The short single-stranded nucleic acid from the template strand and the corresponding short single-stranded nucleic acid containing degenerate bases from the coding strand are not 100% complementary. After the formed intermediate enters the host to repair the crack, it can form a long double-stranded nucleic acid with or without mutation sites after semi-conservative replication of double-stranded nucleic acids.
[0307] In another embodiment, the short single-stranded nucleic acid on the corresponding template strand is 100% complementary to the short single-stranded nucleic acid on the coding strand, that is, the short single-stranded nucleic acid on the template strand also contains degenerate bases, and the degenerate bases are complementary to those on the coding strand, so the intermediate formed is 100% complementary.
[0308] It should be noted that Figure 14A and Table 1 show cases where the corresponding sites contain the original sequence. For example, in Table 1, F4 contains TTN corresponding to site F64. TTN is actually a set of four codons: TTA, TTT, TTG, and TTC, corresponding to the amino acids phenylalanine (F) and leucine (L), including both mutated and non-mutated forms at the corresponding amino acid sites. Furthermore, since each codon in Table 1 includes both mutated and non-mutated amino acids, the actual number of nucleic acids corresponding to each short single-stranded nucleic acid containing a mutated site should increase exponentially with the number of mutated sites, and the number of encoded polypeptides should also increase exponentially. For example, assuming each mutated codon corresponds to two amino acids (including those from the original sequence), then F4 containing four mutated codons should be able to encode 2^4 polypeptides, or 16 polypeptides, corresponding to a greater number of nucleic acid types.
[0309] Table 1. Short single-stranded nucleic acid sequences used to prepare fluorescent protein mutant libraries.
[0310] Table 2 Degenerate Base Correspondence Table
[0311] Short-chain single-stranded nucleic acids were mixed and annealed according to the steps described in Example 2 to assemble the full-length intermediate containing the fluorescent protein mutant. After gel purification, the intermediate was further cloned into an expression vector using the ClonExpress Ultra One-Step Cloning Kit (Vazyme, C115-01). Then, either of the following was performed: (1) the cloning mixture was transformed into DH5α and spread onto selective medium, or (2) it was directly added to 5 mL of LB liquid medium and cultured for 8 hours before plasmid extraction. The clones were scanned using a stereofluorescence microscope (Carl ZEISS SteREO Discovery.V12) to screen for clones that showed fluorescence. Plasmids were extracted and sequenced from 10 of the fluorescent single colonies. The extracted plasmids were re-transformed into DH5α to obtain more and more purified single colonies for detecting bacterial brightness.
[0312] As shown in Figures 13B and 14B, some mutant fluorescent proteins with functional changes compared to wild-type fluorescent proteins were screened using stereofluorescence microscopy.
[0313] As shown in Figure 14C, the sequence of the newly screened mutant protein differs from that of the wild type and existing fluorescent protein mutants, confirming that the protein library involved in this application has screened new fluorescent protein mutants. In addition to the 12 classic amino acid site mutations, a mismatch-formed mutation site, R168, also appeared during the screening. The classic site mutations contained in the newly screened mutant proteins and their corresponding sources are shown in Figure 14C, with mutation sites from the same source displayed in the same color. The mutation site sources shown in Figure 14C demonstrate the importance of Emerald's S65, S72, N149, M153, and I167; Topaz's S65, S72, K79, and T203; ECFP's F64, S65, Y66, N146, M153, and V163; and EBFP's F64, Y66, and Y145 in screening fluorescent protein mutants.
[0314] The plasmid obtained from LB liquid medium in Example 4 was subjected to NGS sequencing, and the results are shown in Figure 23. As shown in Figure 23, of the 1.4M reads measured, approximately 564,000 reads were in the fluorescent protein mutant library predicted in this application, and most protein variants appeared less than 5 times in the sequencing, indicating a relatively uniform distribution of the sequencing data.
[0315] Example 10 Detection of bacterial brightness
[0316] Microplate reader and flow cytometry
[0317] Single colonies obtained in Example 9 were cultured overnight at 37°C in LB liquid medium supplemented with 100 μg / ml carbenicillin. All cultures were then transferred to 4°C overnight to accumulate fluorescent proteins. For measurements using a microplate reader, 1 ml of bacterial cell suspension was harvested from each culture and washed with 1 ml of PBS. Cells suspended in 100 μl of PBS were added to black 96-well clear-bottomed microplates (Corning), with three replicates for each sample. Cell concentration was measured using OD600, and fluorescence was measured using a λex / λem = 488 / 520 nm detection setting on a TECAN Infinite M200 pro multimode microplate reader. Cells infected with pUC19 were used as a negative control, and background fluorescence values were subtracted. Brightness data were plotted in GraphPad Prism.
[0318] As shown in Figures 13E and 13F, the fluorescence intensity of the mutant proteins Y1, Y2, G3, and G4 screened in this application is stronger than that of existing fluorescent protein mutants.
[0319] When performing measurements using flow cytometry, 1 ml of cell pellet was collected and washed with 1 ml of PBS. Fluorescence levels were quantified using a Beckman Coulter CytoFlex S flow cytometer. Graphs were generated using FlowJo software.
[0320] As shown in Figures 13D and 14F, the fluorescence intensity of different mutant proteins emitted light at different wavelengths was recorded and plotted as curves.
[0321] Confocal Detection
[0322] Single colonies were cultured overnight at 37°C in LB broth supplemented with 100 μg / mL carbenicillin. All cultures were then transferred to 4°C overnight to accumulate fluorescent protein. 1 mL of bacterial cell suspension was pelleted and resuspended in 1 mL PBS. 5 μl of the resuspended cells were placed in an observation dish and fixed on a 10 mm x 10 mm agarose mat (1% agarose prepared with PBS). Cells were imaged using a Nikon AX microscope equipped with a x100 / 1.4 NADIC objective. Cells infected with pUC19 were used as a negative control.
[0323] As shown in Figures 13C and 14E, the fluorescent protein mutants screened in this application were demonstrated to have the ability to emit fluorescence under a confocal microscope.
[0324] Example 11: Effect of the number of cracks in the intermediate on nucleic acid synthesis
[0325] In Example 5, when synthesizing GFP, the inventors noticed that many fragments were truncated (62.5%, 10 / 16) when using the Gibson assembly method, while no fragment truncation occurred when using the chain substitution method. Therefore, it is hypothesized that the number of cracks on the intermediate may be positively correlated with the truncation rate.
[0326] To verify this hypothesis, the inventors implemented two strategies to reduce the number of cleavages on the intermediate (in the initial design, the intermediate contained 26 cleavages): (1) designing longer short single-stranded nucleic acids, using 180-nt short single-stranded nucleic acids with 90-nt overlap (9×180mers and 2×90mers, 12 cleavages); (2) pre-ligating the intermediate by contacting it with a ligase (such as T4 ligase) after its formation and before it is assembled into the vector, and ligating some of the cleavages first.
[0327] An example of the aforementioned pre-connection method and the method without pre-connection is shown below:
[0328] The short single-stranded nucleic acid was divided into two subpools of 1650bp.
[0329] The unligated group is the aforementioned (1): 3 μg of short single-stranded nucleic acid mixtures in each subpool were placed in 20 μl of 1×TE buffer containing 10 mM MgCl2 and annealed in a thermal cycler (95℃ for 10 min, 70℃ for 30 min, 65℃ for 30 min, 60℃ for 30 min, 10℃ ∞).
[0330] The ligation group, as described in (2) above, consisted of a mixture of short-chain single-stranded nucleic acids totaling approximately 3.3 μg, annealed in a thermal cycler using T4 polynucleotide kinase (NEB, M0201) (37°C for 10 h, 95°C for 10 min, 70°C for 2 h, 16°C at ∞). Then, 18 μl of the assembled fragment was treated with NEB T4 DNA ligase (NEB, M0202) in a total volume of 20 μl overnight at 4°C. The assembly product or the ligation product of the two subpools was mixed at a 1:1 molar ratio and directly transformed into the BY4742 strain according to the protocol in Example 4. The transformation mixture was inoculated onto SC-his plates and cultured at 30°C for two days, followed by PCR analysis.
[0331] As shown in Table 3, both methods effectively reduced the truncation rate to 12% (2 / 17) or 0% (0 / 18), respectively. Since altering the length of short single-stranded nucleic acids is relatively more complex, pre-ligation of intermediates is a better method for reducing the truncation rate. Therefore, this method was widely applied in this application.
[0332] Table 3. Effects of different treatment strategies on gene truncation rate
[0333] Example 12 Synthesis of long double-stranded nucleic acids containing specific sequences
[0334] Regarding the synthesis effect of short single-stranded nucleic acids containing special sequences in Example 7, the inventors supplemented the synthesis experiments with 16S rRNA (a very complex nucleic acid containing many self-pairing structures, an exemplary structure of which is shown in Figure 16), AT sequences, and short single-stranded nucleic acids of different lengths designed for repetitive sequences.
[0335] As shown in Figures 17 and 18B, the inventors synthesized the following four structures: (1) a 1590bp fragment containing the 16S rRNA encoding gene (Figure 18B, top); (2) a 1050bp fragment with 100% GC content (Figure 17B, left and Figure 18B, middle); (3) a 1050bp fragment with 100% AT content (Figure 17A and Figure 17B, left); and (4) a 330bp human microsatellite region containing six tandem repeats of the same 45bp sequence (the reaction contains 10 short single-stranded nucleic acids of 57-66-mers, 1 of 41-mers, and 1 of 44-mers, Figure 18B, bottom).
[0336] As shown in Figures 17 and 18B, using MOSAIC, the inventors successfully assembled all four special fragments. After more accurate calculation of the sequencing results, the error-free synthesis rates were 50% (5 / 10, determined by nanopore sequencing), 30% (3 / 10, determined by nanopore sequencing), 28% (7 / 25), and 16% (4 / 25), respectively.
[0337] Meanwhile, the inventors used the PCA (polymerase cycle assembly) method to assemble the above four short single-stranded nucleic acids containing specific sequences and characterized the products, the results of which are shown in Figure 17B. As shown in Figure 17B, the PCA method cannot effectively synthesize the desired nucleic acid fragments. Therefore, compared with the traditional PCA method, the MOSAIC method has advantages in synthesizing the above-mentioned specific sequences.
[0338] Example 13 Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids of Different Lengths
[0339] To clarify the differences between short single-stranded nucleic acids of different lengths in the synthesis of long double-stranded nucleic acids and to explore the most suitable length of short single-stranded nucleic acids for the synthesis of long double-stranded nucleic acids, the inventors designed short single-stranded nucleic acids of different lengths for the full-length GFP sequence synthesized in Example 5. That is, the full-length GFP sequence was synthesized in different reactions, and the length of the short single-stranded nucleic acids used in different reactions was different. However, the length of the short single-stranded nucleic acids used in the same reaction was the same except for the head and tail.
[0340] The inventors designed short single-stranded nucleic acids of lengths of 20nt, 40nt, 60nt, 80nt, 100nt, and 120nt (with overlapping sequences of lengths of 10nt, 20nt, 30nt, 40nt, 50nt, and 60nt) for the same full-length GFP sequence, and prepared long double-stranded nucleic acids according to the methods in Examples 1-4. The electrophoresis results of the intermediates synthesized from the short single-stranded nucleic acids of different lengths are shown in Figure 19.
[0341] As shown in Figure 19, except for 20 nt, all short single-stranded nucleic acids of various lengths can be used to synthesize intermediates of the correct length. Given that 60 nt of short single-stranded nucleic acids offers the best cost-effectiveness in commercial synthesis, the inventors have determined it to be the preferred length for the short single-stranded nucleic acid building blocks in the MOSAIC synthesis method.
[0342] Example 14 Maximum gene length synthesized using the MOSAIC method
[0343] In this embodiment, the inventors explored the maximum length of sequences that can be synthesized by the nucleic acid synthesis method protected in this application.
[0344] The inventors used short single-stranded nucleic acids with a length of 60 nt to synthesize long double-stranded nucleic acids ranging from 300 bp to 9900 bp. As shown in Figure 20A, the yield of the corresponding intermediate gradually decreased as the length of the long double-stranded nucleic acid increased. As shown in Figure 20B, the intermediate could still be successfully formed when the long double-stranded nucleic acid was 9900 bp.
[0345] Regarding crack repair, the inventors also adjusted the repair strategy to improve the success rate of long-chain double-stranded nucleic acid synthesis.
[0346] In linear nucleic acids, without the use of ligases according to Examples 1-4, the inventors can stably synthesize long double-stranded nucleic acids of no more than 1200 bp with fewer than 40 nicks. When a ligase (preferably T4 ligase) is used for the intermediate, the inventors are able to successfully synthesize long double-stranded nucleic acids of 3000 bp.
[0347] After research and analysis, the inventors discovered that the main obstacle to the synthesis of long double-stranded nucleic acids (LLCs) lies in the number of cleavages exceeding the repair capacity of the host. Therefore, the inventors investigated the effects of the number of cleavages on the intermediate and the use of ligases on the synthesis of LCCs, the results of which are shown in Figure 20C. As shown in Figure 20C, an increase in the number of cleavages reduces the success rate of LCC synthesis. When the number of cleavages reaches 70 or 100, LCC synthesis becomes impossible. However, the use of ligases can improve the success rate of LCC synthesis when the number of cleavages increases.
[0348] As shown in Figures 18A and 20D, more than 88% of human gene coding sequences are less than 3000 bp, which is within the synthesis range of MOSAIC.
[0349] In summary, the MOSAIC method can synthesize long double-stranded nucleic acids with a maximum length of 3000 bp, indicating that most human gene coding sequences can be synthesized using the MOSAIC method.
[0350] In the circular nucleic acid, the inventors successfully synthesized a 9.9 kb plasmid. Specifically, the inventors divided the required short single-stranded nucleic acid into three portions, forming three 3300 bp reaction systems, and assembled them into three intermediates. These intermediates were then purified separately, mixed in a 1:1:1 molar ratio, and transformed into the BY4742 strain according to the method in Example 4. The transformation mixture was inoculated onto SC-his plates and incubated at 30°C for two days, followed by PCR analysis.
[0351] For the synthesized 9.9kb plasmid, the inventors designed seven pairs of primers to synthesize different fragments. The positions of the seven primer pairs are shown in Figure 21A. Six clones were selected for PCR analysis, and the results are shown in Figure 21A. As shown in Figure 21A, in the six selected clones, all seven primer pairs synthesized fragments of the corresponding lengths.
[0352] Finally, the inventors randomly selected three successfully synthesized 9.9kb plasmids and used Sanger sequencing to obtain their full sequences. As shown in Figure 27B, all three plasmids contain mutations that do not affect their normal function, including base deletions, substitutions, and insertions. The number of different mutation types in each plasmid varies.
[0353] Example 15: Exemplary PCA Synthesis Procedure
[0354] Mix 1 μL of 1.2 μM short-chain single-stranded nucleic acid mixture, 10 μL of 2×PrimeSTAR polymerase mixture (Takara Bio), and 9 μL of DEPC water in a PCR reaction tube and perform the reaction using the following procedure: (1) React at 98℃ for 2 min; (2) React at 98℃ for 10 s, react at 65-72℃ (determine the specific reaction temperature according to the sequence) for 20 s, and react at 72℃ for 1 min; (3) Repeat step (2) 54 times; (4) After reacting at 72℃ for 10 min, store the reaction tube at 4℃.
[0355] Example 16 Expression, purification and characterization of fluorescent protein mutants
[0356] To determine whether the stronger fluorescence intensity of Y1, Y2, G3, and G4 obtained in Examples 9 and 10 was due to their inherent fluorescence properties rather than high expression levels, the inventors screened, purified, and normalized the protein concentration of the four mutants, and measured their fluorescence intensity at the same concentration.
[0357] The expression and purification of the mutant were accomplished through the following procedures.
[0358] (1) The mutant was fused with the N-terminal His6 tag, the TEV cleavage site and the GGGGSGGGGS linker peptide (the sequence of which is shown in SEQ ID NO:19), transformed into Escherichia coli BL21(DE3) strain (purchased from TransGen Biotech, catalog number CD601), and then cultured in LB liquid medium for expression.
[0359] (2) After the bacteria grew to OD600 = 0.8, 1 mM isopropyl β-D-1-thiogalactoside (IPTG, purchased from Solarbio Life Sciences, catalog number I8070) was added to the culture system, and then expression was induced for 16 hours under the conditions of 16℃ and 160 rpm shaking culture.
[0360] (3) Collect the cultured bacteria by centrifugation at 4000 rpm for 10 minutes, then resuspend them in lysis buffer A (20 mM HEPES, 100 mM NaCl, 2.5 mM MgCl2, 10% glycerol, 1 mM DTT, pH 7.4), and then sonicate (40 W, 3 seconds on, 6 seconds off).
[0361] (4) After the bacterial lysate treated by sonication was centrifuged at 16,000g for 30 minutes at 4°C, the suspension was incubated with a Ni-NTA adsorption column (nickel column, purchased from LABLEAD, catalog number N30210) equilibrated with lysis buffer A at 4°C for 1 hour.
[0362] (5) Wash the adsorption column with lysis buffer A containing 25 mM imidazole, and then elute the adsorbed protein with lysis buffer A containing 200 mM imidazole.
[0363] (6) The eluted proteins were digested with 50 U / mL His6-tagged fused TEV protease (purchased from Beyotime, catalog number P2307) at 4°C for 8 hours, and then used... Ultra (10kDa MWCO) exchanges the digested protein into an imidazole-free lysis buffer.
[0364] (7) Incubate the mixture obtained in the previous step with a Ni-NTA adsorption column at 4°C for 3 hours.
[0365] (8) The flowthrough containing the purified protein was exchanged by ultrafiltration into storage buffer A (10mM Tris-HCl, 100mM NaCl, 1mM EDTA, 1mM DTT and 50% glycerol, pH=8).
[0366] After purification, the concentration of the protein mutant was adjusted to 0.1 mg / mL, and it was detected using a Thermo Varioskan LUX microplate reader under 488 nm absorption and 510 nm emission conditions.
[0367] The results of fluorescence intensity detection after normalizing the expression level and concentration of the mutants are shown in Figures 24 and 25. As shown in Figures 24 and 25, G2 and G4 still showed enhanced fluorescence intensity after concentration normalization, indicating that their stronger fluorescence intensity is due to their inherent fluorescence properties rather than due to high expression levels.
[0368] Example 17 Preparation of PET hydrolase mutant library
[0369] The inventors then applied MOSAIC-based library preparation technology to PET hydrolase, an enzyme that hydrolyzes polyethylene terephthalate (PET), a common plastic and a major source of plastic pollution. The inventors first analyzed the overall research landscape of PET hydrolases, with a recent study identifying a novel, potent PET hydrolase, Mipa-P. Several potentially beneficial mutation sites in this protein were also revealed. The inventors believed that MOSAIC could help assess whether these sites have synergistic effects. Based on the Mipa-P backbone and mutations at nine sites, the inventors designed a mutant library of over 40 million variants (as shown in Figure 26A) by dividing the PET hydrolase coding sequence into short-chain single-stranded nucleic acids, following the method in Example 9. Wild-type Mipa-P (WT), Mipa-PM4 (P1, the most significantly enhanced Mipa-P variant), and Mipa-PM19 (P2, a mutant modified from P1 to obtain extreme thermostability) were synthesized as positive controls to evaluate the performance of the novel mutants.
[0370] The amino acid sequence of the PET hydrolase is shown in SEQ ID NO:17, and the mutation sites and types contained therein are shown in Figure 26A. Specifically, the mutations are: E9A, E9Q, E9P, E9D, A30K, A30N, A30T, A30M, A30I, A30Q, A30H, A30P, A30L, A30E, A30D, A30V, A30Y, A30S, A30F, P53A, P53G, P53V, P53R, P53L, S110N, S110K, S110T, S110R, S110H, S110Q, S110P, S110D, S110E, S11 The genera and subtypes listed are: 0A, S110G, T160K, T160R, T160Q, T160P, T160I, G178R, G178S, G178T, G178A, G178C, G178W, S184K, S184N, S184T, S184P, S184Q, T191K, T191N, T191Q, T191H, T191P, T191E, T191D, T191A, A214N, A214T, A214S, A214D, and A214G. Including cases where no mutation occurs at the site, theoretically, the library can encode 26,127,360 mutant proteins.
[0371] After successfully constructing a library of PET hydrolase mutants, NGS results identified approximately 700,000 specified variants, which were relatively evenly distributed across the regions. Considering the sequencing depth, it is reasonable to infer that the entire library construction was unbiased. The NGS results are shown in Figure 26B. As shown in Figure 26B, out of approximately 7 million sequencing reads, about 1 million reads contained amino acid sequences belonging to the predicted library, and almost all mutants appeared less than 3 times, indicating a uniform sequence distribution in the mutant library.
[0372] Example 18 Screening and detection of PET hydrolases
[0373] To achieve high-throughput screening of PET hydrolase mutants, the inventors developed a fluorescence detection method using fluorescein dibenzoate (FDBz) as a fluorescent reporter group (as shown in Figures 27A and 27B). Specifically, the degradation of FDBz by PET hydrolase releases fluorescein, which emits green fluorescence (as shown in Figure 27B), thereby enabling high-throughput evaluation and screening via fluorescence-activated cell sorting (FACS).
[0374] The inventors first demonstrated that PET hydrolase can hydrolyze FDBz and produce fluorescence. The specific steps were as follows: First, *E. coli* containing wild-type PET hydrolase was induced to express the protein, and then incubated with FDBz for one hour. After centrifugation of the incubation solution, the bacteria were resuspended in fresh LB medium, and the results are shown in Figure 26C. As shown in Figure 26C, obvious green fluorescence was observed in the wild-type strain, while it was not observed in the negative control, proving that FDBz can be taken up by cells and subsequently degraded by PET hydrolase, emitting green fluorescence.
[0375] The inventors incubated the culture containing the mutant library with FDBz and then sorted it by flow cytometry (FACS), collecting the cell population with the brightest fluorescence at 1% (as shown in the right image of Figure 27A). After this sorting step, most of the non-functional variants were eliminated. Subsequently, the inventors incubated the 200 single clones obtained in the previous step with FDBz again, thereby successfully isolating four new variants (V1-V4), whose mutation types are shown in Figure 26D.
[0376] V1 contains mutations of E9D, P53G, S110R, T160R, G178A, S184K, T191K, and A214T. V2 contains mutations of A30T, P53A, S110A, S184P, and A214G. V3 contains mutations of E9A, A30D, P53G, S110A, G178C, S184P, T191A, and A214G. V4 contains mutations of A30T, P53A, T160I, G178S, S184T, T191P, and A214S.
[0377] The specific steps of flow cytometry (FACS) sorting are as follows.
[0378] (1) The mutant was cultured in LB liquid medium in Escherichia coli BL21(DE3) strain.
[0379] (2) When the bacteria grew to OD600 = 0.6, 0.5 mM IPTG and 250 μM fluorescein dibenzoate (FDBz, Aladdin, catalog number F597560) were added, and expression was induced for 18 hours at 18℃ and 160 rpm.
[0380] (3) The bacteria were washed three times with cold PBS and diluted to OD600 = 0.2.
[0381] (4) After the bacterial solution was filtered through a 5μm filter membrane, the bacterial single-cell suspension was sorted using a fluorescence cell sorter, and the data was analyzed using Flowjo software.
[0382] Example 19 Expression, purification and characterization of PET hydrolase mutant
[0383] The inventors expressed and purified seven proteins (WT, P1, P2, and V1 to V4). After verifying the concentration and purity, the performance of the ten proteins was measured in vitro, and the results are shown in Figures 26F, 27D, and 27E. As shown in Figures 26F, 27D, and 27E, V1-V4 all exhibited higher enzyme activity compared to WT, with V1 showing the most significant increase, exhibiting 31% higher activity than WT. Meanwhile, despite some differences, the enzyme activities of V2 to V4 were comparable to those of P1 and P2.
[0384] The expression and purification can be referred to in Example 16, and the specific steps are as follows.
[0385] (1) The mutant was fused with the C-terminal His6 tag and the fusion protein was expressed.
[0386] (2) Collect bacteria and sterilize them at 4°C using BeyoLytic TM Bacterial active protein extraction reagent (purchased from Beyotime, catalog number P0013Q) was used for lysis for 1 hour.
[0387] (3) After centrifuging the lysis buffer at 4°C and 16,000g for 30 minutes, take the supernatant and incubate the Ni-NTA adsorption column equilibrated with lysis buffer B (40mM Tris-HCl, 150mM NaCl, pH 8.0) at 4°C for 1 hour.
[0388] (4) Wash and elute the adsorption column with lysis buffer B containing 30 mM and 300 mM imidazole, respectively.
[0389] (5) Through Ultra (10kDa molecular weight cutoff) replaced the eluted protein with storage buffer B (40mM Tris-HCl, 150mM NaCl, 50% glycerol, pH 8.0).
[0390] Both purified and unpurified PET hydrolases can be characterized using any of the following methods.
[0391] (1) 1 mg / L PET hydrolase mutant was mixed with 250 μM FDBz in lysis buffer B and incubated at 37 °C for 3 hours. The reaction was detected by a Thermo Varioskan LUX microplate reader with an excitation wavelength of 495 nm and an emission wavelength of 520 nm.
[0392] (2) Flow cytometry can also be used to screen for PET hydrolase mutants. Protein mutants were expressed using *E. coli* strain BL21(DE3) in LB liquid medium. When the bacteria reached OD600 = 0.6, 0.5 mM IPTG and 250 μM fluorescein benzoate (FDBz, Aladdin, catalog number F597560) were added, and expression was induced for 18 hours at 18°C and 160 rpm. The bacteria were washed three times with cold PBS, and the bacterial suspension was diluted to OD600 = 0.2. After filtration through a 5 μm filter, single-cell suspensions were sorted using BD FACS Aria SORP. Data were analyzed using Flowjo software.
[0393] The foregoing detailed description is provided by way of explanation and example and is not intended to limit the scope of the appended claims. Various variations of the embodiments listed herein will be apparent to those skilled in the art and are reserved within the scope of the appended claims and their equivalents.
Claims
1. A method for synthesizing a long double-stranded nucleic acid library from short single-stranded nucleic acids, the method comprising the following steps: a) The short single-stranded nucleic acids are complementary to form an intermediate, which is a long double-stranded nucleic acid with a cleavage, wherein the cleavage refers to the absence of a phosphodiester bond between two adjacent nucleic acids in the long double-stranded nucleic acid; and b) Repair the cracks on the intermediate to obtain the long double-stranded nucleic acid; The long double-stranded nucleic acids in the long double-stranded nucleic acid library are all homologous to the same original long double-stranded nucleic acid, preferably with 80% homology, and more preferably with 90% homology.
2. The method according to claim 1, wherein the long double-stranded nucleic acid library comprises two or more long double-stranded nucleic acids.
3. The method according to claim 1 or 2, wherein the short single-stranded nucleic acid has a length of 10-300 nucleotides.
4. The method according to any one of claims 1-3, wherein the length of the long double-stranded nucleic acid is more than twice that of the short single-stranded nucleic acid, preferably more than five times, and more preferably more than ten times.
5. The method according to any one of claims 1-4, wherein the short single-stranded nucleic acid is designed and synthesized by the following method: (1) Design a nucleotide sequence containing the long double-stranded nucleic acid; (2) Each single strand of the double-stranded nucleotide sequence is divided into several short single-stranded nucleic acids, and each long double-stranded nucleic acid is divided into n short single-stranded nucleotides; and (3) Synthesize the short single-stranded nucleic acid.
6. The method according to any one of claims 1-5, wherein the long double-stranded nucleic acid is the original long double-stranded nucleic acid without mutation and / or containing mutation.
7. The method according to claim 6, wherein the mutation is a classic site mutation.
8. The method according to any one of claims 1-7, wherein: (1) The long double-stranded nucleic acid contains X mutation sites, where X is an integer greater than or equal to 2; (2) The mutation types contained in the X mutation sites are p1, p2, p3, ..., p X , where p1 to p X Each is independently selected from 2, 3, or 4; (3) the long double-stranded nucleic acid library contains long double-stranded nucleic acids of type p1×p2×p3×……×p X ; Preferably, the X mutation sites are distributed on two or more short single-stranded nucleic acids that do not undergo complementary pairing.
9. The method according to any one of claims 1-8, wherein: (1) The number of mutation sites contained in each of the short single-stranded nucleic acids on one single strand of the long double-stranded nucleic acid are x1, x2, x3, x4, ..., x n (1) Each of them is an integer greater than or equal to 0; (2) The number of mutation types contained in the mutation sites on the same short-chain single-stranded nucleic acid are p1, p2, p3, ..., ... p1, p2, p3, ..., p1, p2, p3, p1, p2, p3, ... n , where p1-px n Each is independently selected from 2, 3, or 4; (3) the same short single-stranded nucleic acid can form P with different sequences. n The aforementioned short-chain single-stranded nucleic acid, P n =p1×p2×p3×……×px n (4) The long double-stranded nucleic acid library contains long double-stranded nucleic acids of type P1×P2×P3×……×P n .
10. The method according to any one of claims 1-9, wherein the long double-stranded nucleic acid library contains long double-stranded nucleic acids of type 10. 10 the following.
11. The method according to any one of claims 6-10, wherein the original long double-stranded nucleic acid comprises one or more of the following: a protein coding sequence, a regulatory sequence, and a coding region of a non-coding RNA.
12. The method according to claim 11, wherein the protein is a wild-type protein or a mutant protein.
13. The method according to claim 11, wherein the regulation sequence is one or more of the following: a promoter, an enhancer, a silencer, an insulator, and a response element.
14. The method according to any one of claims 6-12, wherein the original long double-stranded nucleic acid comprises a nucleotide sequence encoding one or more of the following proteins: catalytic proteins, reporter proteins, therapeutic proteins, and preventative proteins.
15. The method of claim 14, wherein the original long double-stranded nucleic acid comprises a nucleotide sequence encoding a reporter protein.
16. The method according to claim 14 or 15, wherein the reporter protein is selected from one or more of the following: fluorescent protein (FP), luciferase, β-galactosidase (β-galactosidase or LacZ), chloramphenicol acetyltransferase (CAT), secretory alkaline phosphatase (SEAP), mCherry, mNeonGreen, HaloTag, SNAPTag and variants thereof.
17. The method according to any one of claims 6-16, wherein the original long double-stranded nucleic acid comprises a nucleotide sequence encoding a fluorescent protein, the fluorescent protein comprises the amino acid sequence shown in SEQ NO ID:16, and the original long double-stranded nucleic acid comprises the nucleotide sequence shown in SEQ ID NO:
1.
18. The method of claim 17, wherein the mutation occurs at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: F64, S65, Y66, S72, K79, Y145, N146, N149, M153, V163, I167, R168, and T203.
19. The method according to claim 17 or 18, wherein the mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:16: F64L, S65T, S65G, S65C, S65A, Y66H, Y66Q, Y66R, Y66C, Y66W, S72A, K79R, Y145F, N146I, N149K, M153T, V163A, I167T, R168H, T203Y, T203K, and T203S.
20. The method of claim 14, wherein the original long double-stranded nucleic acid comprises a nucleotide sequence encoding a catalytic protein.
21. The method of claim 20, wherein the catalytic protein is selected from one or more of the following: amylase, protease, lipase, ATP synthase, hexokinase, DNA polymerase, RNA polymerase, protein kinase, lactate dehydrogenase, PET hydrolase, and variants thereof.
22. The method according to claim 20 or 21, wherein the original long double-stranded nucleic acid comprises a nucleotide sequence encoding polyethylene terephthalate hydrolase (PET hydrolase), the PET hydrolase comprising the amino acid sequence shown in SEQ ID NO:17, and the original long double-stranded nucleic acid comprising the nucleotide sequence shown in SEQ ID NO:
18.
23. The method of claim 22, wherein the mutation occurs at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:17: E9, A30, P53, S110, T160, G178, S184, T191, and A214.
24. The method according to claim 22 or 23, wherein the mutation is selected from one or more of the following mutations occurring in the amino acid sequence shown in SEQ ID NO:17: E9A, E9Q, E9P, E9D, A30K, A30N, A30T, A30M, A30I, A30Q, A30H, A30P, A30L, A30E, A30D, A30V, A30Y, A30S, A30F, P53A, P53G, P53V, P53R, P53L, S110N, S110K, S110T, S110R, S110H, S110Q, S110P, S110D, S1 10E, S110A, S110G, T160K, T160R, T160Q, T160P, T160I, G178R, G178S, G178T, G178A, G178C, G178W, S184K, S184N, S184T, S184P, S184Q, T191K, T191N, T191Q, T191H, T191P, T191E, T191D, T191A, A214N, A214T, A214S, A214D, and A214G.
25. The method according to any one of claims 5-24, wherein in step (2), the split points in the two strands are staggered by at least 10 nucleotides.
26. The method according to any one of claims 1-25, wherein in step b), the repair method is to form a phosphate diester bond at the crack.
27. The method according to claim 26, wherein the intermediate is a linear nucleic acid, and the repair method comprises: i. Attach the intermediate to the carrier; ii. Transfer the vector into the host body.
28. The method according to any one of claims 1-27, wherein the number of the cracks is 2-50.
29. The method according to any one of claims 1-27, wherein the number of cracks is 51 or more.
30. The method according to any one of claims 27-29, further comprising treating the intermediate with a ligase prior to step i.
31. Use of the method according to any one of claims 1-30 in the synthesis of long double-stranded nucleic acid libraries.
32. A long double-stranded nucleic acid library prepared by the method according to any one of claims 1-30.
33. The protein library encoded by the nucleic acid library of claim 32.
34. The protein library according to claim 33, wherein the protein is a fluorescent protein and / or a mutant thereof.
35. The protein library according to claim 33, wherein the protein is a PET hydrolase and / or a mutant thereof.
36. A mutant fluorescent protein, wherein the fluorescent protein comprises an amino acid substitution at the T203 site of the amino acid sequence shown in SEQ ID NO:16, wherein the substituted amino acid is tyrosine.
37. A mutant fluorescent protein, wherein the fluorescent protein comprises an amino acid substitution at the Y66 site of the amino acid sequence shown in SEQ ID NO:16, wherein the substituted amino acid is histidine.
38. A mutant fluorescent protein comprising amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:16: S72, K79, N146, V163 and I167, wherein the substituted amino acids are: S72A, K79R, N146I, V163A and I167T.
39. A mutant PET hydrolase comprising amino acid substitutions at one or more of the following sites in the amino acid sequence shown in SEQ ID NO:17: E9, A30, P53, S110, T160, G178, S184, T191 and A214.
40. The PET hydrolase according to claim 39, wherein the substituted amino acids are: E9D, P53G, S110R, T160R, G178A, S184K, T191K and A214T.
41. The PET hydrolase according to claim 39, wherein the substituted amino acids are: A30T, P53A, S110A, S184P and A214G.
42. The PET hydrolase according to claim 39, wherein the substituted amino acids are: E9A, A30D, P53G, S110A, G178C, S184P, T191A and A214G.
43. The PET hydrolase according to claim 39, wherein the substituted amino acids are: A30T, P53A, T160I, G178S, S184T, T191P and A214S.