DNA Site-Directed Mutagenesis Method and Protein Mutant Phenotype Screening Method Based Thereon
Through the improved overlap extension PCR method and DEEplus framework, low-cost, high-precision, and high-throughput DNA site-directed mutant construction and protein phenotype screening are achieved, solving the automation problems of mutant construction and screening in the prior art, and improving experimental efficiency and accuracy.
Patent Information
- Application Number
- CN202510387613.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing DNA site-directed mutation methods cannot achieve low-cost, high-precision, high-throughput and automated mutant construction and phenotypic screening. The traditional overlap extension PCR method depends on the complex steps of intermediate product purification, making it difficult to quickly, high-throughput and automation.
Using an improved overlap extension PCR method, DNA site-directed mutations were achieved through three rounds of PCR cycles, and protein mutant phenotype screening was used using automated primer design scripts and deep learning models, which were simplified into a DEEplus framework, reducing DNA purification and enzyme digestion steps, and improving experimental efficiency and accuracy.
The success rate of DNA site-directed mutations is 99.9% and the accuracy rate is 100%. It can build thousands to tens of thousands of mutants in a short period of time, reducing costs, ensuring the automation of experiments and high throughput, and is suitable for accurate research on protein phenotypes.
Smart Images

Figure CN119876126B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for site-directed DNA mutagenesis and a method for screening protein mutant phenotypes based thereon, belonging to the technical field of gene site-directed mutagenesis. Background Art
[0002] With the rapid development of whole-genome sequencing technology and the emergence of artificial intelligence-driven protein design technology, etc., the research paradigm of cytogenetics has gradually become more refined from the overall gene and protein levels to the nucleotide and amino acid levels. The cytological phenotype study of single nucleotide variations (and the resulting single amino acid missense mutations) has received increasing attention, which relies on genetic operations such as site-directed mutagenesis or random mutagenesis on the coding sequence of proteins (i.e., DNA) to obtain mutants. For example, common in vitro DNA (such as plasmids) site-directed mutagenesis techniques include the overlap extension polymerase chain reaction (PCR) method, cassette mutagenesis method, inverse PCR circularization method, etc.; in vitro DNA random mutagenesis techniques include PCR based on degenerate primers or sequential error-prone PCR, etc. Furthermore, by transiently transfecting DNA into cells, the phenotypes of protein mutants are explored. These mutation construction techniques, as the cornerstone of cell biology and molecular biology, play a crucial role in the development of life sciences.
[0003] However, none of the existing methods can achieve low-cost, high-throughput construction of saturated mutants on the basis of meeting precise DNA site-directed mutagenesis. For example, both the cassette mutagenesis method and the inverse PCR cyclization method require the constructed DNA fragment to be transformed into competent bacteria, and the positive clones are screened by sequencing monoclonal colonies (for example, recently Lacoste et al. constructed 3,448 mutants by this method and transfected them into cell lines to observe phenotypes) [1]. Although such methods ensure the accuracy of mutagenesis, they are costly, time-consuming, and have cumbersome procedures, making it difficult to achieve high-throughput mutant construction in a short time. Although some methods have been improved on this basis, such as Ma et al. recently invented the SMuRF technology to achieve high-throughput and high-precision mutant construction [2], its core steps still rely on high-cost and time-consuming techniques such as purification of intermediate products and enzymatic digestion, so the goal of rapid, accurate, and low-cost mutant construction cannot be achieved. Random mutagenesis (such as error-prone PCR) technology can achieve the construction of a large number of mutant DNAs in a short time, but the mutagenesis is not directional and controllable, which causes great difficulties for the precise study of gene and protein functions, and also makes such technologies unable to meet the requirements of saturated and high-precision mutant construction. In addition, recently, a study used the method of whole DNA sequence synthesis to achieve high-throughput construction of mutants [3], but this method is extremely costly and does not have the universality of application, and cannot meet the urgent needs of the current scientific research field.
[0004] The traditional splicing by overlap extension PCR (SOEing-PCR) is based on the method of introducing mutations in the overlapping region of the reverse primer pair. The site-directed mutant DNA fragment is constructed through three rounds of PCR, so it does not require the steps of transformation and monoclonal colony sequencing, and has a lower cost and shorter time consumption. It is a relatively commonly used precise mutant construction technology at present. However, this method depends on the purification of intermediate product fragments, and the steps are still relatively complex, making it difficult to achieve the goal of rapid, high-throughput, and automated mutant construction; moreover, the amplification efficiency of this method is affected by the DNA sequence of the overlapping region, and the error rate and failure rate are still relatively high, making it difficult to provide effective guarantee for high-precision and saturated mutant construction.
[0005] Therefore, there is a need to develop new methods to meet the goals of low-cost, high-precision, high-throughput, and automated mutant construction and protein phenotype screening and analysis.
[0006] References:
[0007] [1] Lacoste J et al. Pervasive mislocalization of pathogenic coding variants underlying human disorders. Cell. 2024 Nov 14.
[0008] [2] Ma K et al. Saturation mutagenesis-reinforced functional assays for disease-related genes. Cell. 2024 Nov 14.
[0009] [3] Beltran A et al. Site-saturation mutagenesis of 500 human protein domains. Nature. 2025 Jan. Summary of the Invention
[0010] Problems to be Solved by the Invention
[0011] The traditional overlap extension PCR method (SOEing-PCR) relies on the purification of intermediate product fragments, and the steps are still relatively complex, making it difficult to achieve the goal of rapid, high-throughput, and automated mutation construction.
[0012] To achieve the goals of low-cost, high-precision, high-throughput, and automated DNA (or protein) mutant construction and phenotype screening, the present invention improves the key technologies and methods of SOEing-PCR on the basis of overlap extension PCR, and constructs a new DNA site-directed mutagenesis method.
[0013] Furthermore, a protein mutant phenotype screening method is constructed based on this DNA site-directed mutagenesis method. In the present invention, this protein mutant phenotype screening method is also referred to as the DEEplus (Directed efficient engineering of protein by linear unbiased SOEing) framework, abbreviated as DEEplus.
[0014] Solutions for Solving the Problems
[0015] [1]. A DNA site-directed mutagenesis method, the method comprising:
[0016] First reaction: Using the first upstream universal primer pF1 as the upstream primer and the downstream mutant primer pMut-R as the downstream primer, and using the polynucleotide containing the DNA to be mutated as the template, perform polymerase chain reaction to obtain the first amplicon;
[0017] Second reaction: Using the upstream mutant primer pMut-F as the upstream primer and using the first downstream universal primer pR1 as the downstream primer, and using the polynucleotide containing the DNA to be mutated as the template, perform polymerase chain reaction to obtain the second amplicon;
[0018] Third reaction: Using the diluted first amplicon and the diluted second amplicon together as the template, using the second upstream universal primer pF2 as the upstream primer and the second downstream universal primer pR2 as the downstream primer, perform polymerase chain reaction to obtain the site-directed mutated DNA;
[0019] Wherein, the downstream of the first amplicon and the upstream of the second amplicon contain the site to be mutated, and there is an overlapping part;
[0020] The first upstream universal primer pF1 includes an upstream complementary binding part and a first non-complementary binding part located upstream of the upstream complementary binding part; the first downstream universal primer pR1 includes a downstream complementary binding part and a second non-complementary binding part located upstream of the downstream complementary binding part; neither the first non-complementary binding part nor the second non-complementary binding part contains a sequence complementary to the template;
[0021] The second upstream universal primer pF2 includes a polynucleotide having the same nucleotides as at least a part of the first non-complementary binding part in a continuous manner; the second downstream universal primer pR2 includes a polynucleotide having the same nucleotides as at least a part of the second non-complementary binding part in a continuous manner.
[0022] [2]. According to the DNA site-directed mutation method described in [1], wherein, the downstream mutant primer pMut-R and the upstream mutant primer pMut-F contain the mutant site nucleotides; and except for the mutant site nucleotides, the downstream mutant primer pMut-R and the upstream mutant primer pMut-F are respectively complementary to the regions of the antisense strand and the sense strand of the DNA to be mutated containing the site to be mutated; and / or, there is an overlapping region between the region of the DNA to be mutated complementary to the downstream mutant primer pMut-R and the region of the DNA to be mutated complementary to the upstream mutant primer pMut-F.
[0023] [3]. The DNA site-directed mutagenesis method according to [1] or [2], wherein the upstream complementary binding portion is complementary to the 5'-end of the sense strand of the DNA to be mutated; and the downstream complementary binding portion is complementary to the 5'-end of the antisense strand of the DNA to be mutated.
[0024] [4]. The DNA site-directed mutagenesis method according to any one of [1] - [3], wherein, in the third reaction, the first amplicon and the second amplicon are each diluted by at least 15 times; and / or, in each 20 μL of the reaction system of the third reaction, the dosages of the diluted first amplicon and the diluted second amplicon are each within the range of 0.5 - 3 μL.
[0025] [5]. The DNA site-directed mutagenesis method according to any one of [1] - [4], wherein, in the reaction system of the third reaction, the concentrations of the second upstream universal primer pF2 and the second downstream universal primer pR2 are each at least 0.2 pM.
[0026] [6]. The DNA site-directed mutagenesis method according to any one of [1] - [5], wherein, in the reaction systems of the first reaction and the second reaction, the dosages of the template are each within the range of 0.5 - 5 ng; and / or, in the reaction systems of the first reaction and the second reaction, the concentrations of the downstream mutant primer pMut-R and the upstream mutant primer pMut-F are each within the range of 0.01 - 0.5 pM; and / or, in the reaction systems of the first reaction and the second reaction, the concentrations of the first upstream universal primer pF1 and the first downstream universal primer pR1 are each within the range of 0.01 - 0.3 pM; and / or, in the first reaction and the second reaction, the number of cycles of the polymerase chain reaction is each within the range of 20 - 30 times.
[0027] [7]. A high-throughput screening method for protein mutant phenotypes, which uses the DNA site-directed mutagenesis method according to any one of [1] - [6] to perform site-directed mutagenesis on the polynucleotide encoding the protein.
[0028] [8]. The method according to [7], wherein the method comprises:
[0029] The step of designing primers, using an automated primer design script to obtain the primers used in the DNA site-directed mutagenesis method;
[0030] The step of obtaining the DNA encoding the protein mutant, using the primers obtained in the step of designing primers, and through the DNA site-directed mutagenesis method, obtaining the DNA encoding the protein mutant containing the mutation site; and
[0031] Steps for phenotypic detection and screening of protein mutants, and perform phenotypic detection and screening on the DNA encoding the protein mutant obtained in the step of obtaining the DNA encoding the protein mutant.
[0032] [9]. According to the method described in [7] or [8], wherein, in the step of designing primers, an automated primer design script is used to design the primer sequences used in the DNA site-directed mutagenesis method.
[0033]
[10] . According to the method described in any one of [7]-[9], wherein, in the step of phenotypic detection and screening of the protein mutant, the DNA encoding the protein mutant is transfected into cells and imaged with a microscope to obtain an image of the phenotype of the protein mutant, and based on a deep learning model of a convolutional neural network, the image of the phenotype of the protein mutant is analyzed to screen and classify the phenotype of the protein mutant.
[0034] Effects of the Invention
[0035] The DNA site-directed mutagenesis method provided by the present invention optimizes and improves the experimental process of the PCR part. Through a simple three-round PCR cycle (see details in Figure 1 ), the construction of DNA site-directed mutagenesis can be achieved, with a success rate of about 99.9% and an accuracy rate of 100%.
[0036] Based on the DNA site-directed mutagenesis method provided by the present invention, a protein mutant phenotype screening method (DEEplus framework) is constructed. It can be combined with an automated primer design script to achieve saturated scanning of the entire coding sequence of the target protein and batch primer design, and can complete the primer construction work of thousands or even tens of thousands of mutants in a short time. Then, through the DNA site-directed mutagenesis method of the present invention, a large number of site-directed mutants are constructed in batches. The above DNA mutants can be directly transfected into cell lines in batches, and through high-content automated imaging and an automated classification model based on deep learning, precise and high-throughput analysis and screening of the phenotypes of mutant proteins are completed. Therefore, the DNA site-directed mutagenesis method provided by the present invention or the protein mutant phenotype screening method based on it has a low use cost, makes up for the problem of high sequencing costs of mutation construction methods such as whole-sequence synthesis or monoclonal sequencing, and retains accuracy.
[0037] The DNA site-directed mutagenesis method provided by the present invention or the protein mutant phenotype screening method based on it also does not depend on steps such as DNA purification or restriction enzyme digestion. Therefore, it can ensure the automation and high throughput of the entire experimental process, thereby greatly improving the experimental rate.
[0038] In addition, compared with random mutations (such as error-prone PCR), the mutant sequences constructed by the DNA site-directed mutagenesis method of the present invention or the protein mutant phenotype screening method based thereon are precise and directional, and the mutation types are uniform. Therefore, it is suitable for the precise study of protein phenotypes, ensuring the reliability and repeatability of experimental results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 : Flow chart of the DNA site-directed mutagenesis method provided by the present invention (i.e., the improved high-throughput SOEing-PCR). Using the coding region containing the target protein (Protein of interest, POI) as the DNA to be mutated, such as a circular plasmid containing the coding gene tuba1a of human α-tubulin TUBA1A (i.e., a polynucleotide containing the DNA to be mutated) as a template, the first half sequence (i.e., the first amplicon) and the second half sequence (i.e., the second amplicon) containing the target mutation are obtained through two rounds of PCR reactions (PCR cycle 1 (cycle1) and PCR cycle 2 (cycle2), corresponding to the first reaction and the second reaction respectively). These two sequences contain partial overlaps near the mutation site. Furthermore, through the third round of PCR reaction (PCR cycle 3 (cycle3), corresponding to the third reaction), using the sequences obtained from the first two rounds of PCR as templates, the final full-length target fragment is amplified. This linear fragment contains the target mutation and can be directly used for subsequent transfection and other operations.
[0040] Figure 2 : High-throughput screening framework diagram of DEEplus protein mutant phenotypes. For the coding region of a given protein (for example, mostly the corresponding cDNA sequence of humans), primer pair sequences for all single nucleotide variations (SNVs) are designed using an automated primer design script. These primers can be directly used for high-throughput improved SOEing-PCR experiments (abbreviated as high-throughput SOEing-PCR, which is also the DNA site-directed mutagenesis method provided by the present invention; the cycle is about 1 day). Furthermore, the mutation information is linearly corresponded to the reaction wells one by one, and automated sample addition and PCR reactions are carried out based on a machine to achieve high-throughput improved SOEing-PCR to rapidly construct the target product (about 1 week). Then, each target fragment is directly transfected into a cell line, and through high-content microscopic imaging, phenotype pictures of all mutation sites are obtained (about 1 - 3 weeks). Finally, a convolutional neural network training classification model is used to screen, classify, and identify the phenotypes of all mutations (about 1 week).
[0041] Figure 3:Schematic diagram of the principle for designing forward (reverse) primers for mutation sites. First, 25 bases upstream and 25 bases downstream of the mutation site are selected as the starting forward (reverse) primers. This pair of primers contains the target mutation site. Then, the length of the overlapping region with the template (fragment 1) is shortened to make its Tm = 55 - 60 °C and the 3' end is G or C. Finally, the 5' end of the primer is deleted so that the overlapping region of pMut-F and pMut-R is 24 bases.
[0042] Figure 4 :Partial Sanger sequencing results of the constructed tuba1a mutants. After the mutants of tuba1a were constructed using the high-throughput improved SOEing-PCR, the target mutation sites (48 are listed in the figure) were identified using first-generation Sanger sequencing. The positions of the target bases, the original bases, and the mutated bases are listed above each sequencing peak diagram. The amino acid mutations caused by the base mutations are in parentheses. The sequencing results are all correct and there are no miscellaneous peaks.
[0043] Figure 5 :Partial display of the SNV map of tuba1a. The cytological phenotypes of all SNV mutants (a total of 2683 missense mutations) of tuba1a were obtained using DEEplus screening framework and labeled according to the scores (0 - 1). The horizontal axis is the amino acid position of tuba1a, and the vertical axis is the types of amino acids mutated into other amino acids. Circles represent missense mutations, and diamonds represent the wild-type amino acid types at this site. This figure shows the phenotypic distribution or scoring of missense mutations at amino acid positions 2 - 82 of tuba1a. Detailed implementation manners
[0044] The following will detail various exemplary embodiments, features, and aspects of the present invention. The special word "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.
[0045] In addition, to better illustrate the present invention, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present invention can also be implemented without some specific details. In other instances, methods, means, equipment, and steps well-known to those skilled in the art are not described in detail to highlight the gist of the present invention.
[0046] Unless otherwise stated, the units used in this specification are all international standard units, and the numerical values and numerical ranges appearing in the present invention should be understood to include the systematic errors inevitable in industrial production.
[0047] In this specification, the meaning expressed by using "can" includes both the meaning of performing a certain process and the meaning of not performing a certain process.
[0048] In this specification, the "some specific / preferred embodiments", "some other specific / preferred embodiments", "embodiments", etc. mentioned refer to the specific elements related to the embodiment (for example, features, structures, properties, and / or characteristics) included in at least one of the embodiments described herein, and may or may not exist in other embodiments. Additionally, it should be understood that the elements may be combined in various embodiments in any suitable manner.
[0049] In this specification, the numerical range represented by "numerical value A to numerical value B" refers to a range that includes the endpoint numerical values A and B.
[0050] In this specification, the terms "polynucleotide" and "nucleic acid" used interchangeably refer to a polymeric form of nucleotides (ribonucleotides or deoxyribonucleotides) of any length. Thus, this term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, unnatural, or derivative nucleobases.
[0051] In the technical field, "G", "C", "A", "T", and "U" generally represent the bases guanine, cytosine, adenine, thymine, and uracil, respectively. However, it is also generally known in the art that each of "G", "C", "A", "T", and "U" also generally represents a nucleotide containing guanine, cytosine, adenine, thymine, and uracil as bases, which is a common way in representing deoxyribonucleic acid sequences and / or ribonucleic acid sequences. Therefore, in the context of the present invention, the meanings represented by "G", "C", "A", "T", and "U" include the above various possible situations. However, it should be understood that the term "ribonucleotide" or "nucleotide" may also refer to a modified nucleotide or an alternative replacement moiety. Those skilled in the art will recognize that guanine, cytosine, adenine, and uracil can be replaced by other moieties without substantially changing the base-pairing properties of an oligonucleotide (including a nucleotide having such a replacement moiety).
[0052] As used in this specification, "hybridizable" or "complementary" or "substantially complementary" means that a nucleic acid (e.g., RNA, DNA) contains a nucleotide sequence that enables the nucleic acid to non-covalently bind (i.e., form Watson-Crick base pairs and / or G / U base pairs), "anneal", or "hybridize" with another nucleic acid in a sequence-specific, antiparallel manner (i.e., the nucleic acid specifically binds to the complementary nucleic acid) under appropriate in vitro and / or in vivo temperature and solution ionic strength conditions. Standard Watson-Crick base pairing includes: adenine (A) pairs with thymidine (T), adenine (A) pairs with uracil (U), and guanine (G) pairs with cytosine (C). In addition, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization between a DNA molecule and an RNA molecule: guanine (G) can also pair with uracil (U). For example, in the case of base pairing between a tRNA anticodon and a codon in mRNA, G / U base pairing is at least partially responsible for the degeneracy of the genetic code.
[0053] Hybridization and washing conditions are well known and are exemplified in Sambrook, J., Fritsch, E.F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly in Chapter 11 and Table 11.1 of that reference; and in Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). The conditions of temperature and ionic strength determine the "stringency" of the hybridization.
[0054] In the present invention, "moderate stringency conditions", "medium-high stringency conditions", "high stringency conditions" or "very high stringency conditions" describe the conditions for nucleic acid hybridization and washing. Guidance for performing hybridization reactions can be found in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6, which is incorporated herein by reference. Both aqueous and non-aqueous methods are described in this reference and either can be used. For example, specific hybridization conditions are as follows: (1) Low stringency hybridization conditions are in 6× sodium chloride / sodium citrate (SSC) at about 45°C, then at least 50°C, washed twice in 0.2× SSC, 0.1% SDS (for low stringency conditions, the wash temperature can be increased to 55°C); (2) Moderate stringency hybridization conditions are in 6× SSC at about 45°C, then at 60°C, washed once or more in 0.2× SSC, 0.1% SDS; (3) High stringency hybridization conditions are in 6× SSC at about 45°C, then at 65°C, washed once or more and preferably in 0.2× SSC, 0.1% SDS; (4) Very high stringency hybridization conditions are 0.5M sodium phosphate, 7% SDS, at 65°C, then at 65°C, washed once or more in 0.2× SSC, 1% SDS.
[0055] Hybridization requires two nucleic acids to contain complementary sequences, but mismatches between bases are possible. The conditions suitable for hybridization between two nucleic acids depend on the length and degree of complementarity of the nucleic acids, which are well-known variables in the art.
[0056] According to the present invention, the terms "upstream" and "downstream" are relative terms that define the linear positions of at least two elements in a nucleic acid molecule (whether single-stranded or double-stranded) oriented in the 5' to 3' direction.
[0057] "Amplicon" refers to a segment of DNA or RNA that is the source and / or product of a natural or artificial amplification or replication event. In the context of this report, "amplification" refers to the generation of one or more copies of a genetic fragment, DNA to be mutated, or template, specifically the generation of an amplicon. As the product of an amplification reaction, an amplicon can be used interchangeably with common laboratory terms such as PCR products.
[0058] DNA Site-Directed Mutagenesis Method
[0059] In some aspects of the present invention, there is provided a DNA site-directed mutagenesis method, the method comprising:
[0060] First reaction: In the first reaction, the first upstream universal primer pF1 is used as the upstream primer, and the downstream mutant primer pMut-R is used as the downstream primer. Using the polynucleotide containing the DNA to be mutated as the template, polymerase chain reaction is carried out to obtain the first amplicon.
[0061] Second reaction: In the second reaction, the upstream mutant primer pMut-F is used as the upstream primer, and the first downstream universal primer pR1 is used as the downstream primer. Using the polynucleotide containing the DNA to be mutated as the template, polymerase chain reaction is carried out to obtain the second amplicon.
[0062] Third reaction: In the third reaction, the diluted first amplicon and the diluted second amplicon are used together as the template, the second upstream universal primer pF2 is used as the upstream primer, and the second downstream universal primer pR2 is used as the downstream primer. Polymerase chain reaction is carried out to obtain the site-directed mutated DNA.
[0063] Wherein, the downstream of the first amplicon and the upstream of the second amplicon contain the site to be mutated, and there is an overlapping part.
[0064] The first upstream universal primer pF1 includes an upstream complementary binding part and a first non-complementary binding part located upstream of the upstream complementary binding part; the first downstream universal primer pR1 includes a downstream complementary binding part and a second non-complementary binding part located upstream of the downstream complementary binding part; neither the first non-complementary binding part nor the second non-complementary binding part contains a sequence complementary to the template.
[0065] The second upstream universal primer pF2 includes a polynucleotide having the same nucleotides as at least a part of the first non-complementary binding part in sequence; the second downstream universal primer pR2 includes a polynucleotide having the same nucleotides as at least a part of the second non-complementary binding part in sequence.
[0066] The DNA site-directed mutagenesis method provided by the present invention does not rely on DNA purification, monoclonal sequencing, full-sequence synthesis or restriction enzyme digestion, reduces the physical cost, and has a broader and more universal application prospect. Moreover, the DNA site-directed mutagenesis method provided by the present invention can achieve rapid, automated and high-throughput completion of the reaction.
[0067] Downstream Mutation Primer pMut-R and Upstream Mutation Primer pMut-F
[0068] In some embodiments, the downstream mutation primer pMut-R and the upstream mutation primer pMut-F contain a mutated site nucleotide (i.e., in the DNA to be mutated, the nucleotide after mutation at the site to be mutated, also referred to as the mutated base). Through this mutated site nucleotide, after three rounds of PCR, site-directed mutation of the DNA to be mutated is achieved, the mutated site nucleotide is obtained, and thus the corresponding site-directed mutated DNA is obtained.
[0069] In some embodiments, the downstream mutation primer pMut-R and the upstream mutation primer pMut-F are DNA-specific primers (i.e., complementary to the DNA to be mutated). Except for the mutated site nucleotide they contain, they complementarily bind to the region near the site to be mutated in the DNA to be mutated, that is, they complementarily bind to the region of the DNA to be mutated containing the site to be mutated.
[0070] In some specific embodiments, the upstream mutation primer pMut-F complementarily binds to the sense strand of the DNA to be mutated. In some specific embodiments, the downstream mutation primer pMut-R complementarily binds to the antisense strand of the DNA to be mutated.
[0071] As Figure 3 shown, the downstream mutation primer pMut-R and the upstream mutation primer pMut-F can be obtained by the following exemplary method: First, select approximately 25 bases upstream and approximately 25 bases downstream of the site to be mutated as the starting forward (reverse) primer. This pair of primers contains the target mutated site. Then, shorten the length of the overlapping region with the template (fragment 1) so that its Tm = 55 - 60 °C and the 3'-end is G or C. Finally, delete the 5'-end of the primer so that the overlapping region of pMut-F and pMut-R is approximately 24 bases.
[0072] In this specification, the Tm value (Melting Temperature) of a primer refers to the temperature required for the double-stranded structure formed by the primer and its complementary DNA sequence to unwind (separate into single strands) under specific conditions.
[0073] In some embodiments, the regions of the DNA to be mutated that are complementary to the downstream mutation primer pMut-R and the upstream mutation primer pMut-F have an overlapping region, for example, overlapping at least 20 nucleotides, such as overlapping 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, 25 nucleotides.
[0074] In some embodiments, the region of the DNA to be mutated that is complementary to the downstream mutation primer pMut-R and the region of the DNA to be mutated that is complementary to the upstream mutation primer pMut-F completely overlap, that is, the region of the DNA to be mutated that is complementary to the downstream mutation primer pMut-R and the region of the DNA to be mutated that is complementary to the upstream mutation primer pMut-F are the same region in the DNA to be mutated.
[0075] In some embodiments, the lengths of both the downstream mutation primer pMut-R and the upstream mutation primer pMut-F are between 25 and 45 nucleotides.
[0076] In some embodiments, the Tm values of both the downstream mutation primer pMut-R and the upstream mutation primer pMut-F are between 55 and 60 °C.
[0077] In some embodiments, the 3'-ends of both the downstream mutation primer pMut-R and the upstream mutation primer pMut-F are G or C.
[0078] First Upstream Universal Primer pF1 and First Downstream Universal Primer pR1
[0079] In some embodiments, the first upstream universal primer pF1 includes an upstream complementary binding portion and a first non-complementary binding portion located upstream of the upstream complementary binding portion, wherein the upstream complementary binding portion is complementary to the 5'-end of the DNA to be mutated (sense strand).
[0080] In some embodiments, the length of the upstream complementary binding portion is between 15 and 25 nucleotides.
[0081] In some embodiments, the length of the first non-complementary binding portion is between 25 and 35 nucleotides, for example, about 30 nucleotides.
[0082] In some embodiments, the first downstream universal primer pR1 includes a downstream complementary binding portion and a second non-complementary binding portion located upstream of the downstream complementary binding portion, wherein the downstream complementary binding portion is complementary to the 5'-end of the DNA to be mutated (antisense strand).
[0083] In some embodiments, the length of the downstream complementary binding portion is between 15 and 25 nucleotides.
[0084] In some embodiments, the length of the second non-complementary binding portion is between 25 and 35 nucleotides, for example, about 30 nucleotides.
[0085] In some embodiments, neither the first non-complementary binding portion nor the second non-complementary binding portion contains a sequence complementary to the template, so that the second upstream universal primer pF2 and the second downstream universal primer pR2, which contain some consecutive nucleotides thereof, also cannot bind to the template.
[0086] Second Upstream Universal Primer pF2 and Second Downstream Universal Primer pR2
[0087] In some embodiments, the second upstream universal primer pF2 comprises a polynucleotide having the same nucleotides as at least a portion of the first non-complementary binding portion in sequence. Thus, during the annealing process in the polymerase chain reaction (the third reaction), the second upstream universal primer pF2 can complementarily bind to the portion corresponding to the first non-complementary binding portion in the first amplicon and cannot complementarily bind to the template.
[0088] In some embodiments, the length of the second upstream universal primer pF2 is 20 - 25 nucleotides.
[0089] In some embodiments, the second downstream universal primer pR2 comprises a polynucleotide having the same nucleotides as at least a portion of the second non-complementary binding portion in sequence. Thus, during the annealing process in the polymerase chain reaction (the third reaction), the second downstream universal primer pR2 can complementarily bind to the portion corresponding to the second non-complementary binding portion in the second amplicon and cannot complementarily bind to the template.
[0090] In some embodiments, the length of the second downstream universal primer pR2 is 20 - 25 nucleotides.
[0091] It can be seen that the second upstream universal primer pF2 and the second downstream universal primer pR2 are respectively partially identical in sequence to the first upstream universal primer pF1 and the first downstream universal primer pR1, but cannot complementarily bind to the sequence of the starting template. Therefore, they can specifically bind to the template used in the third round of PCR and will not bind to the starting template, ensuring that only the full-length fragment carrying the target mutation can be amplified, while the starting template will not be amplified at all, achieving the specificity of amplification.
[0092] First Reaction and Second Reaction
[0093] The first reaction can also be referred to as the first round of amplification, the first round of PCR, the first round of PCR reaction, or PCR cycle 1 (cycle 1). Similarly, the second reaction can also be referred to as the second round of amplification, the second round of PCR, the second round of PCR reaction, or PCR cycle 2 (cycle 2).
[0094] In some embodiments, in the reaction systems of the first reaction and the second reaction, the amounts of the template used are 0.5 - 5 ng respectively, preferably 1 - 3 ng, such as 1 ng, 2 ng, or 3 ng.
[0095] In some embodiments, in the reaction systems of the first reaction and the second reaction, the concentrations of the downstream mutant primer pMut-R and the upstream mutant primer pMut-F are 0.01 - 0.5 pM respectively, preferably 0.1 - 0.4 pM, such as 0.2 pM, 0.25 pM, 0.3 pM.
[0096] In some embodiments, in the reaction systems of the first reaction and the second reaction, the concentrations of the first upstream universal primer pF1 and the first downstream universal primer pR1 are 0.01 - 0.3 pM respectively, preferably 0.05 - 0.2 pM, such as 0.05 pM, 0.1 pM, 0.15 pM, 0.2 pM.
[0097] In some embodiments, in the first reaction and the second reaction, the number of cycles of the polymerase chain reaction (denaturation, annealing, extension, and these three steps are one cycle) is 20 - 30 times respectively, preferably 25 times.
[0098] Those skilled in the art can determine the annealing temperature and extension time according to the Tm values of the specific primers (such as the downstream mutant primer pMut-R and the upstream mutant primer pMut-F) in the first reaction and the second reaction, and the amplicon length.
[0099] In some embodiments, in the first reaction and the second reaction, excessive increase in the template amount, increase in the final primer concentration, or increase in the number of amplification cycles may all interfere with subsequent experiments, resulting in an increase in the experimental failure rate or an increase in non-specific amplification.
[0100] It can be understood that in the reaction systems of the first reaction and the second reaction, in addition to the above-mentioned template and primers, it also contains DNA polymerase, dNTP, and PCR buffer. Those skilled in the art can select appropriate DNA polymerase, dNTP, and PCR buffer as long as they can complete the present invention.
[0101] In some embodiments, in some exemplary embodiments, a high-fidelity enzyme premix is used, the final amplification system is 10 μL, the template amount is 2 ng / reaction, the final concentration of the pF1 / pR1 primers is 0.1 pM, the final concentration of the pMut-R / pMut-F primers is 0.25 pM, the annealing temperature and extension time are determined by the Tm values of pMut-F / R and the amplicon length, and the number of amplification cycles is 25.
[0102] In some embodiments, the lengths of the first amplicon and the second amplicon are generally both less than 10 kb.
[0103] Third Reaction
[0104] The third reaction can also be referred to as the third round of amplification, the third round of PCR, or PCR cycle 3.
[0105] In some embodiments, in the reaction system of the third reaction, the diluted first amplicon and the diluted second amplicon are used as templates. For example, the template can be a mixture obtained by uniformly mixing the diluted first amplicon and the diluted second amplicon. Alternatively, the diluted first amplicon and the diluted second amplicon can be separately added to the reaction system. In some embodiments, the first amplicon and the second amplicon are each diluted by at least 15-fold, preferably at least 20-fold, more preferably at least 100-fold. In some embodiments, the first amplicon and the second amplicon are each diluted by less than 150-fold, preferably less than 120-fold.
[0106] In some embodiments, in each 20 μL of the reaction system of the third reaction, the amounts of the diluted first amplicon and the diluted second amplicon are each in the range of 0.5 - 3 μL, preferably 0.5 - 2.5 μL, such as 0.5 μL, 1 μL, 2 μL, 2.5 μL.
[0107] In some embodiments, in the reaction system of the third reaction, the concentrations of the second upstream universal primer pF2 and the second downstream universal primer pR2 are each at least 0.2 pM, preferably at least 0.3 pM, more preferably at least 0.4 pM.
[0108] Insufficient dilution multiples of the first two rounds of PCR products (the first amplicon and the second amplicon), or too low final concentrations of the second upstream universal primer pF2 and the second downstream universal primer pR2 may both lead to amplification failure or an increase in non-specific fragments.
[0109] In some embodiments, in the third reaction, the number of cycles of the polymerase chain reaction (denaturation, annealing, extension, and these three steps of reaction constitute one cycle) is at least 25 times, preferably at least 30 times. In some embodiments, the number of cycles of the polymerase chain reaction is less than 33 times.
[0110] Those skilled in the art can determine the annealing temperature and extension time based on the Tm values of the specific primers (such as the second upstream universal primer pF2 and the second downstream universal primer pR2) in the third reaction and the length of the amplicon.
[0111] Polynucleotide Containing DNA to be Mutated
[0112] In some embodiments, the DNA to be mutated can be any DNA.
[0113] In some specific embodiments, the DNA to be mutated contains the DNA encoding the protein of interest (POI) to be mutated.
[0114] In some specific embodiments, the DNA to be mutated comprises the DNA encoding the target protein to be mutated, and at least one of an enhancer, a promoter, a fluorescent tag, and a terminator operably linked thereto. That is, the DNA to be mutated comprises an expression cassette encoding the target protein.
[0115] In some embodiments, the length of the DNA to be mutated is generally less than 10 kb.
[0116] In some specific embodiments, the polynucleotide comprising the DNA to be mutated is an expression vector, such as a plasmid, which comprises the DNA to be mutated, such as the DNA encoding the target protein or the expression cassette encoding the target protein.
[0117] Exemplary Embodiments
[0118] See Figure 1, the present invention optimizes and improves the experimental procedure of overlap extension PCR, making it independent of DNA purification steps, and efficiently and specifically amplifying the target fragment (with a success rate of approximately 99.9%). In this exemplary embodiment, a circular plasmid containing the target protein coding region (here taking human α-tubulin tuba1a as an example) is used as the starting template, and reverse primer pairs capable of complementary pairing are designed near the target mutation site. This primer pair contains the mutation site (referred to as pMut-R and pMut-F respectively). pMut-R and the universal primer pF1 upstream of the promoter or enhancer serve as the forward and reverse primers for the first round of PCR (cycle1), and pMut-F and the universal primer pR1 downstream of the terminator serve as the forward and reverse primers for the second round of PCR (cycle2), respectively amplifying the starting template. A high-fidelity enzyme premix is used, the final amplification system is 10 μL, the template amount is 2 ng / reaction, the final concentration of pF1 / pR1 primers is 0.1 pM, the final concentration of pMut-R / pMut-F primers is 0.25 pM, the annealing temperature and extension time are determined by the Tm value of pMut-F / R and the length of the amplified fragment, and the number of amplification cycles is 25. Excessive increase in template usage, increase in the final primer concentration, or increase in the number of amplification cycles may all interfere with subsequent experiments, resulting in an increased failure rate of the experiment or an increase in non-specific amplification. Furthermore, using the products of the first two rounds of PCR as templates, they are diluted 20-fold (up to 100-fold) respectively with double-distilled water (ddH2O). After mixing, 1 μL is taken and added to a total of 20 μL of the third-round amplification system. The third-round amplification (cycle3) uses a high-fidelity enzyme premix, and the universal forward and reverse primers pF2 and pR2 are added to make their final concentrations 0.4 pM (or higher) respectively for specific amplification of the full-length target fragment. The annealing temperature is 53 - 56°C (the Tm value of pF2 / pR2 is 56°C), the extension time is determined by the length of the full-length fragment, and the number of amplification cycles is 30 or more. Insufficient dilution multiple of the products of the first two rounds of PCR or too low final concentration of pF2 / pR2 may both lead to amplification failure or an increase in non-specific fragments. In addition, pF2 / pR2 are partially identical in sequence to pF1 / pR1 respectively, but completely different from the sequence of the starting template. Therefore, they can specifically bind to the template used in the third-round PCR, but not to the starting template, ensuring that only the full-length fragment carrying the target mutation can be amplified, while the starting template will not be amplified at all, achieving the specificity of amplification. The inventors identified the mutation sequences obtained from the previous experiments by random Sanger sequencing and proved that the obtained full-length target sequences are completely correct (N > 100); and through subsequent cell transfection experiments and phenotypic analysis, it was proved that the ratio of successfully amplifying the full-length target sequence using this method is 99.9% (2680 / 2683).
[0119] High-throughput screening method for protein mutant phenotypes
[0120] In some aspects of the present invention, a high-throughput screening method for protein mutant phenotypes is provided, which uses the DNA site-directed mutagenesis method described above to perform site-directed mutagenesis on the polynucleotide encoding the protein.
[0121] To achieve high-throughput and automated phenotypic screening of protein mutants, the present invention combines automated primer design, the DNA site-directed mutagenesis method provided by the present invention, high-content imaging, and deep learning to design a high-throughput screening framework for protein mutant phenotypes (DEEplus) with a wide range of application scenarios. Compared with the currently commonly used mutant construction and phenotypic screening strategies, the time efficiency of this framework is increased by about 10 times ( Figure 2 )
[0122] In some specific embodiments, the high-throughput screening method for protein mutant phenotypes includes the following steps:
[0123] A step of designing primers, using an automated primer design script to obtain the primers used in the DNA site-directed mutagenesis method;
[0124] A step of obtaining the DNA encoding the protein mutant, using the primers obtained in the step of designing primers, and through the DNA site-directed mutagenesis method, obtaining the DNA encoding the protein mutant containing the mutation site; and
[0125] A step of detecting and screening the phenotypes of protein mutants, performing phenotypic detection and screening on the DNA encoding the protein mutant obtained in the step of obtaining the DNA encoding the protein mutant.
[0126] All links of the high-throughput screening method for protein mutant phenotypes provided by the present invention can be completed by instruments (such as a multi-channel automated workstation, a PCR instrument), achieving automation and high throughput, and significantly improving the experimental speed in the construction of a large number of mutants.
[0127] Steps for Designing Primers
[0128] For a given protein of interest (POI), using an automated primer design script, quickly obtain the reverse primer pairs corresponding to all mutation sites (downstream mutation primer pMut-R and upstream mutation primer pMut-F), as well as the first upstream universal primer pF1 and the first downstream universal primer pR1, the second upstream universal primer pF2 and the second downstream universal primer pR2, which can be directly used in the subsequent DNA site-directed mutagenesis method (performed by automated PCR reaction).
[0129] The automated primer design script can directly design all primer sequences applicable to the DNA site-directed mutagenesis method provided by the present invention.
[0130] In some embodiments, in the automated primer design script, the parameters of the primer design script can be fine-tuned, such as increasing the length of the primer overlap region or directly increasing the primer length, etc., so that although the designed primers are different, they can still be used in the DNA site-directed mutagenesis method provided by the present invention.
[0131] Steps for Obtaining DNA Encoding Protein Mutants
[0132] Applying the DNA site-directed mutagenesis method of the present invention, that is, the optimized SOEing-PCR, with the assistance of an automated instrument (such as the Biomek ® NXp automated workstation), high-throughput synthesis of mutant DNA products (site-directed mutagenesis DNA) can be carried out.
[0133] In some specific embodiments, the reaction products of the DNA site-directed mutagenesis method can be operated in a microtiter plate and can be directly docked with subsequent steps of protein mutant phenotype detection and screening, such as batch cell transfection and high-content imaging, greatly shortening the time cost.
[0134] Steps for Detecting and Screening Phenotypes of Protein Mutants
[0135] In the steps of protein mutant phenotype detection and screening, the DNA encoding the protein mutant obtained in the step of obtaining the DNA encoding the protein mutant is subjected to phenotype detection and screening.
[0136] In some embodiments, the DNA encoding the protein mutant obtained in the step of obtaining the DNA encoding the protein mutant can be directly used for batch cell transfection; subsequently, using a high-content microscope, automated imaging and original picture collection are carried out for all reaction wells (each well specifically corresponds to a mutant).
[0137] It can be understood that those skilled in the art can also use other confocal microscopes (including super-resolution microscopes, etc.), as long as they can equally achieve the phenotype imaging of mutants.
[0138] It can be understood that there are no particular limitations on the cell line, transfection reagent, transfection ratio, etc. used in transfection, as long as they can complete the transient transfection operation of the reaction products of the DNA site-directed mutagenesis method.
[0139] In some alternative embodiments, before transfection, the reaction products of the DNA site-directed mutagenesis method can also be subjected to DNA purification, for example, using methods known to those skilled in the art for DNA purification.
[0140] In some embodiments, a deep learning model based on a convolutional neural network (CNN) is used. First, it is trained using small-scale samples, and then the phenotypes of all protein mutants are screened and classified to achieve an automated phenotype identification process (see Figure 2 ).
[0141] It can be seen that in the high-throughput screening method for protein mutant phenotypes provided by the present invention, the DNA site-directed mutagenesis method provided by the present invention can be combined with automated primer design, high-content imaging, and an automated image recognition and classification model based on deep learning, enabling the design, construction, and high-throughput phenotype identification of protein mutants from scratch. It is applicable to the construction of human single nucleotide variant (SNV) maps, human protein missense mutation maps, the saturated screening of protein mutants, and the phenotype screening and identification in protein directed evolution.
[0142] In some exemplary embodiments, for a protein with a length of 500 amino acids (cDNA length of 1500 bases), using an automated primer design script, all primer sequences corresponding to approximately 3000 single nucleotide variants (SNVs) can be obtained within one minute (since codons are degenerate, only SNV sites that result in amino acid missense mutations are retained here). The forward and reverse (see details in Figure 1 ) primers corresponding to each mutation are synthesized in a well plate and spatially correspond one by one. Then, through automated pipetting by an instrument and the reaction in the DNA site-directed mutagenesis method provided by the present invention, the construction of all thousands of mutant DNAs can be completed within several days (at most one week). Furthermore, the SNV mutants are transfected into the corresponding cell lines in batches, and automated image acquisition is performed using a high-content microscope. The phenotype acquisition process of all mutants can be completed within several weeks (on average, the acquisition rate is about 200 - 500 mutants per day). Finally, the small-scale image samples obtained from the previous acquisition are used for model training and optimization, and are used for the automated analysis of all samples, including distinguishing different phenotypes of individual cells (such as Functional, Medium, or Pathogenic), and then scoring and evaluating the biological functions of different mutants. This process takes about several days. Therefore, the DEEplus protein mutant phenotype screening framework can achieve automated, batch, and saturated identification and analysis of mutants, and its physical cost and time cost are much lower than those of traditional deep mutational scanning, and it is more suitable for the construction of human single nucleotide variant (SNV) maps, large-scale mutant screening, or protein directed evolution processes based on mutations.
[0143] Exemplary Embodiments
[0144] For the protein of interest (POI), first obtain its human cDNA sequence (e.g., from the National Center for Biotechnology Information - NCBI website), and clone it into a mammalian cell transient transfection plasmid vector (such as the pCDNA3.0 vector), with an enhancer and a promoter upstream and a terminator downstream. If necessary, insert a fluorescent tag into the protein coding region (such as Figure 1 ). Input this sequence (enhancer - promoter - protein coding region - fluorescent tag - terminator) into an automated primer design script (Python 3 environment needs to be installed) to design the corresponding primer sequences for all single nucleotide variations (SNVs), and return them to the user in tabular form. This table lists all SNV sites, the corresponding amino acid variations, and the forward and reverse primer sequences suitable for SOEing - PCR. Subsequently, synthesize each forward primer in a well plate, with each well corresponding to a mutation, and synthesize the reverse primers in another well plate so that the mutation in each well corresponds one - to - one with the well plate of the forward primers. Furthermore, use an automated pipetting workstation (such as Biomek ® NXp workstation) to batch - configure the reaction system in the DNA site - directed mutagenesis method ( Figure 1 ), and use an ordinary PCR instrument to complete the reaction in the DNA site - directed mutagenesis method. Since the products obtained from the first two rounds of PCR have a one - to - one positional relationship in the forward and reverse primer well plates, the configuration of the third - round PCR system can be directly completed by the automated pipetting workstation, greatly improving the experimental efficiency. After the third - round PCR is completed, a linear DNA fragment containing the mutation site can be obtained, with a concentration of approximately 100 ng / μL (the actual concentration depends on the number of amplification cycles in the third round and the high - fidelity enzyme premix used, and there may be slight deviations between different mutant samples, but generally it does not have a significant impact on the cytological phenotype), and it can be directly used for transient transfection of cells without purification. For transfection, polyethyleneimine (PEI) can be used as a large - scale transfection reagent, and the volume ratio of mutant PCR product to PEI is 5:1, that is, for every 1 μL of mutant PCR product transfected, 0.2 μL of PEI solution (1 mg / mL) is used, and transient transfection operations are carried out according to the standard transfection procedure, so that each well of cells is transfected with a specific mutant protein. Furthermore, according to actual needs, adopt live - cell imaging or immunostaining methods, and use a high - content microscope to automatically image the cells in the culture plate (such as P96 - 1.5 - H - N, CellVis). The imaging conditions (such as the number of channels, magnification, number of pictures collected per well, etc.) can be optimized and adjusted according to the actual situation to obtain high - quality images at a high rate and high resolution for subsequent analysis and other processes. Finally, use a small number of samples (such as 100 mutants) to train and optimize a convolutional neural network model so that it can automatically identify the transfected positive cells in the original image and classify or score the protein phenotypes, thereby batch - screening and identifying the mutant phenotypes in a short period.
[0145] Example
[0146] The embodiments of the present invention will be described in detail below in conjunction with examples. However, those skilled in the art will understand that the following examples are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention. For those not specified in the examples, they are carried out under conventional conditions or conditions recommended by the manufacturer. For reagents or instruments not specified by the manufacturer, they are all conventional products that can be obtained commercially.
[0147] Tubulin is a basic component of the cytoskeletal microtubules and is crucial for the behavior and function of eukaryotes. Various mutations of tubulin can lead to a variety of human diseases. Establishing a human single nucleotide variant (SNV) map of tubulin is of great significance for understanding the function of microtubules and treating human genetic diseases. Therefore, in this example, taking human α-tubulin TUBA1A as an example, DEEplus was used to systematically and comprehensively explore the cytological phenotypes generated by all variants.
[0148] Step 1: Obtaining the starting template sequences in the first reaction and the second reaction
[0149] For human α-tubulin TUBA1A as the protein of interest (POI), first obtain its human cDNA sequence (such as obtained from the website of the National Center for Biotechnology Information - NCBI, gene number NM_006009.4), and clone it into a mammalian cell transient transfection plasmid vector (such as pCDNA3.0 vector, Addgene: #12452), with an enhancer and a promoter upstream and a terminator downstream), and insert a fluorescent tag into the protein coding region (such as Figure 1 )
[0150] Specifically, in this example, a circular plasmid containing the target protein coding region (that is, the DNA to be mutated, specifically human α-tubulin tuba1a in this example) (that is, a polynucleotide containing the DNA to be mutated) is used as the starting template (pCDNA3.0 vector, 6882 bp in length), and its enhancer-promoter-protein coding region (including the fluorescent tag)-terminator sequence is as follows (SEQ ID NO:1):
[0151]
[0152] Among them, the single-underlined part is the CMV enhancer / promoter sequence; the lowercase letter part is the cDNA sequence of tuba1a (including gfp11), and the base to be mutated (the fourth c) is bold and underlined in the subsequent examples; the double-underlined part is the bGH terminator sequence.
[0153] Step 2: Primer design
[0154] Using an automated primer design script, a total of 2,683 SNV sites that cause different missense mutations and their primer sequences ( Figure 2 ), including downstream mutation primers pMut-R and upstream mutation primers pMut-F for each site, as well as first upstream universal primer pF1, first downstream universal primer pR1, second upstream universal primer pF2, and second downstream universal primer pR2, were obtained.
[0155] Specifically, a partial sequence (enhancer-promoter-protein coding region (including fluorescent tag)-terminator) in the starting template sequence obtained in Step 1 above was input into the automated primer design script (Python 3 environment needs to be installed) to design the corresponding primer sequences for all single nucleotide variations (SNVs), and returned to the user in tabular form. This table lists all SNV sites, the corresponding amino acid variations, and the forward and reverse primer sequences applicable to the DNA site-directed mutagenesis method of the present invention. Subsequently, each forward primer was synthesized in a well plate, with each well corresponding to a mutation, and the reverse primer was synthesized in another well plate, so that the mutation in each well corresponded one-to-one with the well plate of the forward primer.
[0156] Among them, the primer design script (for the specific design principle, see Figure 3 ), according to the input DNA sequence, designed the forward and reverse primers required for all base mutations in the coding region in sequence. The starting length of each pair of primers was 50 nucleotides. First, the overlapping region with the template (i.e., fragment 1) was shortened so that its Tm = 55 - 60 °C and the 3'-terminal base was G or C; then, the 5'-ends of the forward and reverse primers (i.e., fragment 2) were shortened so that the overlapping bases of pMut-R and pMut-F were 24 nucleotides, and a pair of primers required for mutating this base could be obtained.
[0157] Step 3: Obtaining DNA encoding protein mutants
[0158] In this example, the inventors optimized and improved the experimental procedure of overlap extension PCR so that it does not depend on the DNA purification step and efficiently and specifically amplifies the target fragment (success rate is about 99.9%) ( Figure 1 ).
[0159] In this example, each primer obtained in Step 2 was synthesized into the corresponding well plate, and an automated pipetting workstation (such as Biomek ® NXp automated pipetting workstation) was used to configure the reaction systems for the first to third reactions of the DNA site-directed mutagenesis method. A conventional PCR instrument was used for the amplification reaction. The construction of all 2,683 mutants was completed within one week. The mutant PCR products were stored frozen at -20 °C and could be stored for several months without significantly reducing the transfection efficiency.
[0160] Specifically, the reaction systems of the first to third reactions for batch configuration of the DNA site-directed mutagenesis method using a Biomek ® NXp automated pipetting workstation are completed, and the reactions of the DNA site-directed mutagenesis method of the present invention are completed using a PCR instrument (such as Eppendorf, Mastercycler Nexus). Since the products obtained in the first two rounds of PCR have a one-to-one corresponding positional relationship in the forward and reverse primer plates, the configuration of the third-round SOEing-PCR system can be directly completed by an automated pipetting workstation, greatly improving the experimental efficiency. After the third-round SOEing-PCR is completed, a linear DNA fragment containing the mutation site can be obtained, with a concentration of about 100 ng / μL (the actual concentration depends on the number of amplification cycles in the third round and the high-fidelity enzyme premix used, and there may be slight deviations between different mutant samples, but generally it does not have a significant impact on the cytological phenotype), and it can be directly used for transient transfection of cells without purification.
[0161] Taking one site as an example, the site-directed mutagenesis method of the DNA of the present invention will be described as follows:
[0162] For example, if it is desired to mutate the fourth nucleotide c to a in the target protein coding region of the starting template sequence in step one, a pair of reverse primers that can complementarily pair are designed near the target mutation site. This pair of primers contains the mutation site (referred to as the downstream mutant primer pMut-R and the upstream mutant primer pMut-F respectively). pMut-R and the first upstream universal primer pF1 upstream of the promoter or enhancer are used as the forward and reverse primers for the first reaction (PCR cycle1), and pMut-F and the first downstream universal primer pR1 downstream of the terminator are used as the forward and reverse primers for the second reaction (PCR cycle2). The starting template is amplified respectively. Using a high-fidelity enzyme premix (Vazyme, #P525-01), the final amplification system is 10 μL, the template amount is 2 ng / reaction, the final concentration of the pF1 / pR1 primers is 0.1 pM, the final concentration of the pMut-R / pMut-F primers is 0.25 pM, the annealing temperature and extension time are determined by the Tm values of pMut-F / R and the length of the amplified fragment (in this example, the specific annealing temperature is: 58°C), and the number of amplification cycles is 25.
[0163] The primers and PCR conditions used in the first reaction and the second reaction are specifically as follows:
[0164] pF1 (SEQ ID NO:2):
[0165]
[0166] Among them, the bold part is the pF2 binding sequence, i.e., the first non-complementary binding part; the underlined part is the part binding to the template, i.e., the upstream complementary binding part.
[0167] pR1 (SEQ ID NO:3):
[0168]
[0169] Among them, the bold part is the pR2 binding sequence, i.e., the second non-complementary binding part; the underlined part is the part binding to the template, i.e., the downstream complementary binding part.
[0170] pMut-F (SEQ ID NO:4):
[0171] GCCACCATG A GTGAGTGCATCTCCATCCACGTTG
[0172] Among them, the underlined part is the mutated base.
[0173] pMut-R (SEQ ID NO:5):
[0174] GGAGATGCACTCAC T CATGGTGGCGGATCCGAGC
[0175] Among them, the underlined part is the mutated base.
[0176] The PCR reaction conditions are as follows:
[0177] Pre-denaturation: 95°C, 3 minutes (min)
[0178] Denaturation: 95°C, 15 seconds (sec)
[0179] Annealing: 58°C, 15 sec
[0180] Extension: 72°C, 30 sec (the first reaction); 72°C, 90 sec (the second reaction)
[0181] The number of amplification cycles is 25
[0182] Final extension: 72°C, 5 min
[0183] In the research of this example, it was found that excessive increase in the template amount, increase in the final primer concentration, or increase in the number of amplification cycles may all interfere with subsequent experiments, resulting in an increase in the experimental failure rate or an increase in non-specific amplification. Furthermore, the products of the first two rounds of PCR (the first reaction and the second reaction) were used as templates, diluted 20-fold respectively with double-distilled water (ddH2O), mixed well, and then 1 μL was taken and added to the third-round amplification system of the third reaction with a total volume of 20 μL.
[0184] For the third reaction (PCR cycle 3), a high-fidelity enzyme premix (Vazyme, #P525-01) is used, and the second upstream universal primer pF2 and the second downstream universal primer pR2 are added, with their final concentrations being 0.4 pM respectively, for specific amplification of the full-length target fragment. The annealing temperature is 53 - 56 °C (the Tm value of pF2 / pR2 is 56 °C), the extension time is determined by the length of the full-length fragment, and the number of amplification cycles is 30 or more.
[0185] The primers used in the third reaction of this example are as follows:
[0186] pF2 (SEQ ID NO:6): CCATGGAGCTCCAAATAATG
[0187] pR2 (SEQ ID NO:7): GATTTGACGTCATGAGAGGC
[0188] In the research of the present invention, it is found that insufficient amplification multiples of the first two rounds of PCR products before dilution, or too low final concentrations of pF2 / pR2 may both lead to amplification failure or an increase in non-specific fragments.
[0189] The PCR reaction conditions (for the third reaction) are as follows:
[0190] Pre-denaturation: 95 °C, 3 min
[0191] Denaturation: 95 °C, 15 sec
[0192] Annealing: 53 °C, 15 sec
[0193] Extension: 72 °C, 2 min
[0194] The number of amplification cycles is 33
[0195] Final extension: 72 °C, 5 min
[0196] In this example, pF2 / pR2 are partially identical in sequence to pF1 / pR1 respectively, but are completely different from the sequence of the starting template. Therefore, they can specifically bind to the template used in the third round of PCR, rather than binding to the starting template, ensuring that only the full-length fragment carrying the target mutation can be amplified, while the starting template will not be amplified at all, achieving the specificity of amplification.
[0197] As confirmed by subsequent experiments, in this example, the mutant sequences obtained from the previous experiments were identified by random Sanger sequencing, which proved that the target full-length sequences obtained by using this method were completely correct (N>100); and through subsequent cell transfection experiments and phenotypic analysis, it was proved that the ratio of successfully amplifying the target full-length sequences by using this method was 99.9% (2680 / 2683).
[0198] Step 4. Phenotype detection and screening of protein mutants
[0199] As described above, the products obtained by the DNA site-directed mutagenesis method provided by the present invention can be directly used for transient transfection of cells and subsequent phenotype detection and screening of protein mutants without purification.
[0200] Polyethyleneimine (PEI) can be used as a large-scale transfection reagent for transfection. For the different mutant PCR products obtained in Step 3, the volume ratio of PEI used is 5:1, that is, for every 1 μL of mutant PCR product transfected, 0.2 μL of PEI solution (1 mg / mL) is used. The transient transfection operation is carried out according to the standard transfection procedure, that is, the PCR product is diluted with 5 μL of opti-MEM medium, and the PEI solution is diluted with 5 μL of opti-MEM medium. Then the two components are mixed and allowed to stand for 15 min, and then added to the single-well cells in a 96-well cell culture plate (CellVis, #P96-1.5H-N), so that each well of cells is transfected with a specific mutant protein. Furthermore, according to actual needs, live cell imaging or immunostaining is adopted, and a high-content microscope is used to automatically image the cells in the culture plate. The imaging conditions (such as the number of channels, magnification, number of pictures collected per well, etc.) can be optimized and adjusted according to the actual situation to obtain high-quality images at a high speed and high resolution, so as to be used for subsequent analysis and other processes. After that, a small amount of samples (such as 100 mutants) are used for training and optimization of the convolutional neural network model, so that it can automatically identify the transfected positive cells in the original image and classify or score the protein phenotypes, so as to batch complete the screening and identification of mutant phenotypes in a short period.
[0201] In this example, by randomly selecting 100 mutant PCR products for first-generation Sanger sequencing, it was proved that the full length of the target sequences was completely correct, the target mutations were correctly introduced, and there were no heterozygous peaks ( Figure 4 )
[0202] Furthermore, an appropriate amount of HeLa cells (from ATCC) were seeded in a 96-well cell culture plate for imaging, with a density of 50 - 70% at the time of transfection. Each mutant was transfected into the cells in each well via PEI and cultured for another day. At this time, the cell density was approximately 80 - 90%, suitable for high-content imaging. Each cell culture plate was placed in a high-content microscope (AI dual-rotating disk high-content imaging microscope, MetaXpress) for automated imaging. The magnification was 60X (water lens), and the autofocus mode was used. Images were taken in the red / green dual channels (the green channel was for GFP-labeled tubulin, and the red channel was for microtubule dye). 36 images were taken for each well, and a total of 190,000 - 200,000 original images were obtained. This process took 2 - 3 weeks, with approximately 200 mutants transfected and imaged per day, and the success rate was 99.9% (2680 / 2683).
[0203] After that, based on a deep learning model of artificial intelligence (see Redmon J et al. You Only LookOnce: Unified, Real-Time Object Detection. 2016 IEEE Conference on ComputerVision and Pattern Recognition. 2016.), the inventors identified and classified the phenotypes of all 2683 mutants, and scored each mutant according to its similarity to the wild-type protein (i.e., whether it could assemble into microtubules). The lower the score, the smaller the impact of the mutation on the protein function, while the higher the score, the greater the impact of the mutation on the protein, and thus it might lead to human diseases ( Figure 5 ). These results were also completely consistent with the cytological phenotypes of tuba1a mutations reported in previous literature, indicating that the results obtained by this method were highly accurate and reliable, and thus had important reference significance for disease research. In contrast, the physical costs (main sources: high-fidelity enzymes for PCR, cell culture plates for imaging, and microscope usage costs) and time costs of this method were much lower than those of traditional molecular cloning and saturated mutant screening. Therefore, it was more suitable for large-scale, batch, and high-throughput protein mutant research. The comparison between the method provided by the present invention and the existing methods is specifically shown in Table 1 below.
[0204] Table 1
[0205]
[0206] "×" indicates that this method is not applicable to the statistical analysis of this parameter.
[0207] It should be noted that although the technical solutions of the present invention are introduced with specific examples, those skilled in the art can understand that the present invention should not be limited thereto.
[0208] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for site-directed mutagenesis of DNA, the method comprising: The first reaction, using a first upstream universal primer as the upstream primer and a downstream mutation primer as the downstream primer, and using a polynucleotide containing the DNA to be mutated as a template, performing a polymerase chain reaction to obtain a first amplicon; The second reaction, using an upstream mutation primer as the upstream primer and a first downstream universal primer as the downstream primer, and using a polynucleotide containing the DNA to be mutated as a template, performing a polymerase chain reaction to obtain a second amplicon; The third reaction, using the diluted first amplicon and the diluted second amplicon together as a template, using a second upstream universal primer as the upstream primer and a second downstream universal primer as the downstream primer, performing a polymerase chain reaction to obtain the site-directed mutated DNA; Wherein, the downstream of the first amplicon and the upstream of the second amplicon contain the site to be mutated and there is an overlapping part; The first upstream universal primer includes an upstream complementary binding part and a first non-complementary binding part located upstream of the upstream complementary binding part; the first downstream universal primer includes a downstream complementary binding part and a second non-complementary binding part located upstream of the downstream complementary binding part; neither the first non-complementary binding part nor the second non-complementary binding part contains a sequence complementary to the template; The second upstream universal primer includes a polynucleotide having the same nucleotides as at least a part of the first non-complementary binding part in sequence; the second downstream universal primer includes a polynucleotide having the same nucleotides as at least a part of the second non-complementary binding part in sequence; Wherein, the downstream mutation primer and the upstream mutation primer contain mutant site nucleotides; and except for the mutant site nucleotides, the downstream mutation primer and the upstream mutation primer are respectively complementary to the regions of the antisense strand and the sense strand of the DNA to be mutated containing the site to be mutated; The regions of the DNA to be mutated complementary to the downstream mutation primer and the regions of the DNA to be mutated complementary to the upstream mutation primer have overlapping regions with each other.
2. The DNA site-directed mutagenesis method according to claim 1, wherein, The upstream complementary binding part is complementary to the 5'-end of the sense strand of the DNA to be mutated; the downstream complementary binding part is complementary to the 5'-end of the antisense strand of the DNA to be mutated.
3. The DNA site-directed mutagenesis method according to claim 1, wherein, In the third reaction, the first amplicon and the second amplicon are respectively diluted by at least 15 times; and / or, In each 20 μL reaction system of the third reaction, the dosages of the diluted first amplicon and the diluted second amplicon are respectively 0.5 - 3 μL.
4. The DNA site-directed mutagenesis method according to claim 1, wherein, In the reaction system of the third reaction, the concentrations of the second upstream universal primer and the second downstream universal primer are respectively at least 0.2 pM.
5. The DNA site-directed mutagenesis method according to claim 1, wherein In the reaction systems of the first reaction and the second reaction, the dosages of the template are respectively 0.5 - 5 ng; and / or, in the reaction systems of the first reaction and the second reaction, the concentrations of the downstream mutation primer and the upstream mutation primer are respectively 0.01 - 0.5 pM; And / or, in the reaction system of the first reaction and the reaction system of the second reaction, the concentrations of the first upstream universal primer and the first downstream universal primer are 0.01 - 0.3 pM respectively; And / or, in the first reaction and the second reaction, the number of cycles of the polymerase chain reaction is 20 - 30 times respectively.
6. A high-throughput screening method for protein mutant phenotypes, which uses the DNA site-directed mutagenesis method according to any one of claims 1 - 5 to perform site-directed mutagenesis on the polynucleotide encoding the protein.
7. The method according to claim 6, wherein, The method includes: The step of designing primers, using an automated primer design script to obtain the primers used in the DNA site-directed mutagenesis method; The step of obtaining the DNA encoding the protein mutant, using the primers obtained in the step of designing primers, through the DNA site-directed mutagenesis method, to obtain the DNA encoding the protein mutant containing the mutation site; and The step of detecting and screening the protein mutant phenotype, performing phenotype detection and screening on the DNA encoding the protein mutant obtained in the step of obtaining the DNA encoding the protein mutant.
8. The method according to claim 7, wherein In the step of designing primers, use an automated primer design script to design the primer sequences used in the DNA site-directed mutagenesis method.
9. The method according to claim 7, wherein, In the step of detecting and screening the protein mutant phenotype, transfect the DNA encoding the protein mutant into cells and image with a microscope to obtain an image of the phenotype of the protein mutant, And, Based on a deep learning model of a convolutional neural network, analyze the image of the phenotype of the protein mutant, and screen and classify the phenotype of the protein mutant.
Citation Information
Patent Citations
Rapid bidirectional multilocus gene mutation method
CN102690807A