Modified strains for producing recombinant silk
By weakening or eliminating the activity of YPS1-1 and YPS1-2 proteases, the problem of recombinant protein degradation in Pythagorean yeast strains was solved and the yield of protein was improved.
Patent Information
- Application Number
- CN201780095487.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-10-03
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2037-10-03
AI Technical Summary
During culture of yeast Pilgrimia strains, recombinantly expressed proteins may degrade before collection, resulting in a reduced yield of full-length recombinant proteins.
By attenuating or eliminating the activity of YPS1-1 and YPS1-2 proteases, the Pilgrimia strains are modified to reduce the degradation of the recombinant proteins.
It effectively reduces the degradation of recombinant proteins and improves the yield of full-length recombinant proteins.
Smart Images

Figure CN111315763B_ABST
Abstract
Description
[0001] Sequence Listing
[0002] This application contains a sequence listing, which has been submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy was created on October 25, 2017, is named 37324PCT_CRF_sequencelisting.txt, and is 388,944 bytes in size. Technical Field
[0003] The present disclosure relates to strain optimization methods for producing proteins or metabolites from cells or increasing their yields. The present disclosure also relates to compositions obtained by those methods. In particular, the present disclosure relates to yeast cells selected or genetically engineered to reduce the degradation of recombinant proteins expressed by yeast cells, and methods of culturing yeast cells for producing useful compounds. Background of the Invention
[0005] The methylotrophic yeast Pichia pastoris is widely used for the production of recombinant proteins. Pichia pastoris grows to high cell densities, provides tightly controlled methanol-inducible transgene expression, and efficiently secretes heterologous proteins in defined media.
[0006] However, during the cultivation of Pichia pastoris strains, recombinantly expressed proteins may be degraded before they are collected, resulting in the formation of a protein mixture containing recombinantly expressed protein fragments and a reduced yield of full-length recombinant protein. Therefore, tools and engineered strains for mitigating protein degradation in Pichia pastoris are needed. Summary of the invention
[0007] In some embodiments, provided herein are Pichia pastoris microorganisms in which the activity of the YPS1-1 protease and the YPS1-2 protease has been attenuated or eliminated, wherein the microorganism expresses a recombinant polypeptide.
[0008] In some embodiments, the YPS1-1 protease comprises a polypeptide sequence at least 95% identical to SEQ ID NO:67. In some embodiments, the YPS1-1 protease comprises SEQ ID NO:67. In some embodiments, the YPS1-1 protease is encoded by the YPS1-1 gene. In some embodiments, the YPS1-1 gene comprises a polynucleotide sequence at least 95% identical to SEQ ID NO:1. In some embodiments, the YPS1-1 gene comprises at least 15, 20, 25, 30, 40 or 50 consecutive nucleotides of SEQ ID NO:1. In some embodiments, the YPS1-1 gene comprises SEQ ID NO:1. In some embodiments, the YPS1-1 gene is at the locus PAS_chr4_0584 of the microorganism.
[0009] In some embodiments, the YPS1-2 protease comprises a polypeptide sequence at least 95% identical to SEQ ID NO:68. In some embodiments, the YPS1-2 protease comprises SEQ ID NO:68. In some embodiments, the YPS1-2 protease is encoded by the YPS1-2 gene. In some embodiments, the YPS1-2 gene comprises a polynucleotide sequence at least 95% identical to SEQ ID NO:2. In some embodiments, the YPS1-2 gene comprises at least 15, 20, 25, 30, 40 or 50 consecutive nucleotides of SEQ ID NO:2. In some embodiments, the YPS1-2 gene comprises SEQ ID NO:2. In some embodiments, the YPS1-2 gene is at the locus PAS_chr3_1157 of the microorganism.
[0010] In some embodiments, the YPS1-1 gene or the YPS1-2 gene or both have been mutated or knocked out.
[0011] In some embodiments, the microorganism expresses a recombinant protein. In some embodiments, the recombinant protein comprises at least one blocking polypeptide sequence from a silk protein. In some embodiments, the recombinant protein comprises a filamentous polypeptide. In some embodiments, the filamentous polypeptide comprises one or more repeating sequences {GGY-[GPG-X 1 ] n1 -GPS-(A) n2} n3 (SEQ ID NO: 514), wherein X 1=SGGQQ (SEQ ID NO: 515) or GAGQQ (SEQ ID NO: 516) or GQGPY (SEQ ID NO: 517) or AGQQ (SEQ ID NO: 518) or SQ; n1 is 4 to 8; n2 is 6 to 20; n3 is 2 to 20. In some embodiments, the filamentous polypeptide comprises a polypeptide sequence encoded by SEQ ID NO:462.
[0012] In some embodiments, the activity of one or more additional proteases in the microorganism has been reduced or eliminated. In some embodiments, the one or more additional proteases comprise YPS1-5, MCK7 or YPS1-3.
[0013] In some embodiments, the YPS1-5 gene is at locus PAS_chr3_0688 of the microorganism.
[0014] In some embodiments, the MCK7 protease is encoded by a MCK7 gene comprising a polynucleotide sequence at least 95% identical to SEQ ID NO: 7. In some embodiments, the MCK7 gene comprises at least 15, 20, 25, 30, 40, or 50 consecutive nucleotides of SEQ ID NO: 7. In some embodiments, the MCK7 gene comprises SEQ ID NO: 7. In some embodiments, the MCK7 gene is at locus PAS_chr1-1_0379 of the microorganism.
[0015] In some embodiments, the YPS1-3 protease is encoded by a YPS1-3 gene comprising a polynucleotide sequence at least 95% identical to SEQ ID NO: 3. In some embodiments, the YPS1-3 gene comprises at least 15, 20, 25, 30, 40, or 50 consecutive nucleotides of SEQ ID NO: 3. In some embodiments, the YPS1-3 gene comprises SEQ ID NO: 3. In some embodiments, the YPS1-3 gene is at locus PAS_chr3_0299 of the microorganism.
[0016] In some embodiments, the one or more additional proteases comprise a polypeptide sequence at least 95% identical to a polypeptide sequence selected from the group consisting of SEQ ID NOs: 68-130. In some embodiments, the one or more additional proteases comprise a polypeptide sequence selected from the group consisting of SEQ ID NOs: 68-130. In some embodiments, the one or more additional proteases are encoded by a polynucleotide sequence at least 95% identical to a polynucleotide sequence selected from the group consisting of SEQ ID NOs: 3-66. In some embodiments, the one or more additional proteases are encoded by a polynucleotide sequence comprising at least 15, 20, 25, 30, 40, or 50 consecutive nucleotides of a polynucleotide sequence selected from the group consisting of SEQ ID NOs: 3-66.
[0017] In some embodiments, the microorganism comprises a 3X, 4X, or 5X protease knockout.
[0018] According to some embodiments of the present invention, the present invention also provides an engineered Pichia pastoris microorganism, which comprises YPS1-1 and YPS1-2 activities reduced due to mutation or deletion of a YPS1-1 gene comprising SEQ ID NO: 1 and a YPS1-2 gene comprising SEQ ID NO: 2, wherein the microorganism further comprises a recombinant expression protein comprising a polypeptide sequence encoded by SEQ ID NO: 462.
[0019] Also provided herein, in some embodiments, are cell cultures comprising a protease-reduced microorganism as described herein.
[0020] According to some embodiments, provided herein is a cell culture comprising a microorganism in which YPS1-1 and YPS1-2 activities have been attenuated or eliminated as described herein, wherein the microorganism recombinantly expresses a protein, wherein degradation of the recombinantly expressed protein is less compared to a cell culture comprising an otherwise identical Pichia pastoris microorganism in which YPS1-1 and YPS1-2 activities have not been attenuated or eliminated.
[0021] In some embodiments, provided herein is a method for producing a recombinant protein with reduced degradation, the method comprising: culturing a microorganism whose YPS1-1 and YPS1-2 activities have been attenuated or eliminated as described herein in a culture medium under conditions suitable for expression of the recombinantly expressed protein; and isolating the recombinant protein from the microorganism or the culture medium.
[0022] In some embodiments, the recombinant protein is secreted from the microorganism, and wherein isolating the recombinant protein comprises collecting a culture medium comprising the secreted recombinant protein. In some embodiments, the level of degradation of the recombinant protein is reduced compared to the recombinant protein produced by an otherwise identical microorganism wherein the YPS1-1 and the YPS1-2 protease activities are not attenuated or eliminated.
[0023] Also provided herein is a method for modifying Pichia pastoris to reduce the degradation of a recombinantly expressed protein, comprising knocking out or mutating genes encoding YPS1-1 and YPS1-2 proteins. In some embodiments, the method for modifying Pichia pastoris to reduce the degradation of a recombinantly expressed protein further comprises knocking out or mutating one or more additional genes encoding YPS1-3 proteins, YPS1-5 proteins, or MCK7 proteins. In some embodiments, the method for modifying Pichia pastoris to reduce the degradation of a recombinantly expressed protein further comprises knocking out one or more genes encoding a protein comprising a polypeptide selected from the group consisting of SEQ ID NO: 68-130.
[0024] In some embodiments, the recombinantly expressed protein comprises a poly A sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive alanine residues (SEQ ID NO: 519). In some embodiments, the recombinantly expressed protein comprises a filamentous polypeptide. In some embodiments, the filamentous polypeptide comprises one or more repeating sequences {GGY-[GPG-X 1 ] n1 -GPS-(A) n2} n3 (SEQ ID NO: 514), wherein X 1 =SGGQQ (SEQ ID NO: 515) or GAGQQ (SEQ ID NO: 516) or GQGPY (SEQ ID NO: 517) or AGQQ (SEQ ID NO: 518) or SQ; n1 is 4 to 8; n2 is 6 to 20; and n3 is 2 to 20. In some embodiments, the recombinantly expressed protein comprises the polypeptide sequence encoded by SEQ ID NO: 462. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The foregoing and other objects, features and advantages will become apparent from the following description of specific embodiments of the present invention as illustrated in the accompanying drawings in which like reference characters refer to the same parts in the different views. The accompanying drawings are not necessarily drawn to scale, emphasis instead being placed upon illustrating the principles of various embodiments of the present invention.
[0026] Figure 1is a plasmid map for KU 70 deletion with zeocin resistance marker.
[0027] Figure 2 is a plasmid map of a plasmid containing the nourseothricin marker used with homology arms for targeted protease gene deletion.
[0028] Figure 3A and Figure 3B is a cassette for protease knockout with homology arms targeting the desired protease gene flanked by a nourseoin resistance marker.
[0029] Figure 4 Representative Western blots of proteins isolated from individual KO strains are shown to illustrate protein degradation from these strains.
[0030] Figure 5 Representative Western blots of proteins isolated from double KO strains are shown to illustrate protein degradation from these strains.
[0031] Figure 6 Representative Western blots of proteins isolated from 2X, 3X, 4X, and 5X protease KO strains subcultured in BMGY or YPD to show protein degradation in these strains. DETAILED DESCRIPTION
[0032] The details of various embodiments of the invention are set forth in the following description. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims.
[0033] definition
[0034] Unless otherwise defined herein, scientific and technical terms related to the present invention should have the meanings commonly understood by those of ordinary skill in the art. Further, unless the context otherwise requires, singular terms should include plural, and plural terms should include the singular. The terms "a" and "an" include plural references, except where the context otherwise specifies. Typically, the nomenclature used in conjunction with the following and the following techniques are well-known and commonly used in the art: biochemistry, enzymology, molecule and cell biology, microbiology, genetics and protein and nucleic acid chemistry and hybridization as described herein.
[0035] The following terms, unless otherwise specified, shall be understood to have the following meanings:
[0036] The term "polynucleotide" or "nucleic acid molecule" refers to a polymeric form of nucleotides of at least 10 bases in length. The term includes DNA molecules (e.g., cDNA or genomic DNA or synthetic DNA) and RNA molecules (e.g., mRNA or synthetic RNA), as well as analogs of DNA or RNA containing non-natural nucleotide analogs, non-native internucleoside bonds, or both. Nucleic acids can be in any topological conformation. For example, nucleic acids can be single-stranded, double-stranded, triple-stranded, quadruple-stranded, partially double-stranded, branched, hairpin-shaped, circular, or in a padlocked conformation.
[0037] Unless otherwise specified, and as an example of all sequences described herein in the general format "SEQ ID NO:", "a nucleic acid comprising SEQ ID NO: 1" refers to a nucleic acid, at least a portion of which has the following sequence: (i) the sequence of SEQ ID NO: 1, or (ii) a sequence complementary to SEQ ID NO: 1. The choice between the two is determined by the context. For example, if the nucleic acid is used as a probe, the choice between the two depends on the requirement that the probe is complementary to the desired target.
[0038] An "isolated" RNA, DNA or mixed polymer is one that is substantially separated from other cellular components that naturally accompany the original polynucleotide in its natural host cell, such as ribosomes, polymerases and genomic sequences with which it is naturally associated.
[0039] An "isolated" organic molecule (e.g., silk protein) is an organic molecule that is substantially separated from the cellular components (membrane lipids, chromosomes, proteins) of the host cell from which it originates or from the culture medium in which the host cell is cultured. The term does not require that the biomolecule be separated from all other chemical substances, although some isolated biomolecules can be purified to near homogeneity.
[0040] The term "recombinant" refers to a biological molecule (e.g., a gene or protein) that: (1) has been removed from its naturally occurring environment, (2) is not associated with all or part of a polynucleotide with which the gene is found in nature, (3) is operably linked to a polynucleotide to which it is not linked in nature, or (4) does not occur in nature. The term "recombinant" can be used with respect to cloned DNA isolates, chemically synthesized polynucleotide analogs or polynucleotide analogs biologically synthesized by heterologous systems, as well as proteins and / or mRNA encoded by such nucleic acids.
[0041] In this article, if a heterologous sequence is placed adjacent to an endogenous nucleic acid sequence so that the expression of the endogenous nucleic acid sequence is altered, the endogenous nucleic acid sequence (or the encoded protein product of the sequence) in the genome of an organism is considered to be a "recombinant". In this context, a heterologous sequence is a sequence that is not naturally adjacent to an endogenous nucleic acid sequence, whether the heterologous sequence itself is endogenous (derived from the same host cell or its progeny) or exogenous (derived from a different host cell or its progeny). For example, for the original promoter of a gene in the genome of a host cell, the promoter sequence can be replaced (e.g., by homologous recombination) so that the gene has an altered expression pattern. The gene will now become a "recombinant" because it is separated from at least some of the sequences that are naturally flanked by it.
[0042] A nucleic acid is also considered "recombinant" if it contains any modifications that do not naturally occur in the corresponding nucleic acid in the genome. For example, an endogenous coding sequence is considered "recombinant" if it contains artificially introduced insertions, deletions, or point mutations, such as those introduced by human intervention. "Recombinant nucleic acid" also includes nucleic acids that are integrated into a host cell chromosome at a heterologous site and nucleic acid constructs that exist as episomes.
[0043] As used herein, the phrase "degenerate variant" of a reference nucleic acid sequence includes nucleic acid sequences that can be translated according to the standard genetic code to provide the same amino acid sequence as the amino acid sequence translated from the reference nucleic acid sequence. The term "degenerate oligonucleotide" or "degenerate primer" is used to refer to an oligonucleotide that is capable of hybridizing to a target nucleic acid sequence that is not necessarily identical in sequence, but is homologous to each other in one or more specific segments.
[0044] In the context of nucleic acid sequences, the term "percentage of sequence identity" or "identical" refers to the residues in two sequences that are identical when the maximum corresponding alignment is performed. The length of the sequence identity comparison may be at least about 9 nucleotides, usually at least about 20 nucleotides, more usually at least about 24 nucleotides, usually at least about 28 nucleotides, more usually at least about 32 nucleotides, and preferably at least about 36 or more nucleotides. There are many different algorithms known in the art that can be used to measure nucleotide sequence identity. For example, polynucleotide sequences can be compared using FASTA, Gap or Bestfit, which are programs in the Wisconsin Package version 10.0 of Genetics Computer Group (GCG), Madison, Wis. FASTA provides an alignment and a percentage of sequence identity in the region of the best overlap between the query sequence and the search sequence. Pearson, Methods Enzymol. 183: 63-98 (1990) (incorporated herein by reference in its entirety). For example, the percent sequence identity between nucleic acid sequences can be determined using FASTA with its default parameters (wordlength of 6 and scoring matrix of NOPAM factors) or using Gap with its default parameters as provided in GCG version 6.1 (incorporated herein by reference). Alternatively, sequences can be compared using the computer programs BLAST (Altschul et al., J. Mol. Biol. 215:403-410 (1990); Gish and States, Nature Genet. 3:266-272 (1993); Madden et al., Meth. Enzymol. 266:131-141 (1996); Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997); Zhang and Madden, Genome Res. 7:649-656 (1997)), in particular blastp or tblastn (Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997)).
[0045] The terms "substantial homology" or "substantial similarity" when referring to nucleic acids or fragments thereof mean that there is nucleotide sequence identity over at least about 75%, 80%, 85%, preferably at least about 90%, and more preferably at least about 95%, 96%, 97%, 98% or 99% of the nucleotide bases when optimally aligned with appropriate nucleotide insertions or deletions of another nucleic acid (or its complementary strand) as measured by any recognized sequence identity algorithm, such as FASTA, BLAST or Gap discussed above.
[0046] Alternatively, there is substantial homology or similarity when a nucleic acid or its fragment hybridizes with another nucleic acid, with a chain of another nucleic acid, or with its complementary chain under stringent hybridization conditions. "Stringent hybridization conditions" and "stringent wash conditions" depend on many different physical parameters in the context of nucleic acid hybridization experiments. As will be readily appreciated by those skilled in the art, nucleic acid hybridization will be affected by conditions such as salt concentration, temperature, solvent, base composition of hybridizing material, length of complementary region, and number of nucleotide base mismatches between hybridizing nucleic acids. Personnel with ordinary skill in the art know how to change these parameters to obtain specific hybridization stringency.
[0047] Generally speaking, "stringent hybridization" means hybridization under a specific set of conditions at a temperature lower than the thermal melting point (T) of a specific DNA hybrid. m ) is carried out at a temperature about 25°C lower than the T of a specific DNA hybridization. m The temperature is about 5°C lower. m It is the temperature at which 50% of the target sequence hybridizes with the probe that matches perfectly. See Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1989) p. 9.51, which is hereby incorporated by reference. For purposes herein, "stringent conditions" for solution phase hybridization are defined as aqueous hybridization (i.e., without formamide) at 65°C for 8-12 hours in 6xSSC (wherein 20xSSC contains 3.0M NaCl and 0.3M sodium citrate), 1% SDS, followed by two washes at 65°C for 20 minutes in 0.2xSSC, 0.1% SDS. It will be appreciated by those skilled in the art that hybridization at 65°C will proceed at different rates, depending on various factors, including the length and percent identity of the sequences being hybridized.
[0048] Nucleic acids (also referred to as polynucleotides) of the present invention may include sense and antisense strands of RNA, cDNA, genomic DNA, and synthetic forms of the aforementioned and mixed polymers. As will be readily appreciated by those skilled in the art, they may be chemically or biochemically modified or may contain non-natural or derived nucleotide bases. Such modifications include, for example, labeling, methylation, replacement of one or more naturally occurring nucleotides by analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are synthetic molecules that simulate the ability of polynucleotides to bind to a specified sequence through hydrogen bonds and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages replace phosphate linkages in the main chain of the molecule. Other modifications may include, for example, analogs in which the ribose ring contains a bridging moiety or other structure, such as those found in "locked" nucleic acids.
[0049] The term "mutated", when applied to a nucleic acid sequence, refers to a nucleotide in a nucleic acid sequence that may be inserted, deleted or changed compared to a reference nucleic acid sequence. A single change (point mutation) may be made at a locus, or multiple nucleotides may be inserted, deleted or changed at a single locus. In addition, one or more changes may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences may be mutated by any method known in the art, including but not limited to mutagenesis techniques, such as "error-prone PCR" (a process of performing PCR under conditions where the fidelity of replication of a DNA polymerase is low, thereby obtaining a high mutation rate over the entire length of the PCR product; see, for example, Leung et al., Technique, 1: 11-15 (1989) and Caldwell & Joyce, PCR Methods Applic. 2: 28-33 (1992)); and "oligonucleotide directed mutagenesis" (a process of enabling site-specific mutations to be produced in any interested cloned DNA segment; see, for example, Reidhaar-Olson and Sauer, Science 241: 53-57 (1988)).
[0050] As used herein, the term "weakening" generally refers to functional deletion, including mutation, partial or complete deletion, insertion or other changes implemented to a gene sequence or a sequence controlling transcription of a gene sequence, which reduces or inhibits the generation of a gene product or makes a gene product lose its function. In some examples, functional deletion is described as a knockout mutation. Weakening also includes placing a gene under the control of a promoter with lower activity by changing the nucleotide sequence, downwardly regulating, expressing an interfering RNA, a ribozyme or an antisense sequence of a gene of interest, or an amino acid sequence change achieved by any other technology known in the art. In an example, the sensitivity of a specific enzyme to feedback inhibition or inhibition caused by a composition that is not a product or a reactant (non-pathway specific feedback) is reduced so that the enzyme activity is not affected by the presence of a compound. In other examples, an enzyme that has been changed to have lower activity can be referred to as an attenuated enzyme.
[0051] As used herein, the term "deletion" refers to the removal of one or more nucleotides from a nucleic acid molecule or one or more amino acids from a protein, with the regions on both sides joined together.
[0052] As used herein, the term "knockout" means a gene whose expression or activity level has been reduced to zero. In some instances, knockout of a gene is achieved by deleting part or all of its coding sequence. In other instances, knockout of a gene is achieved by introducing one or more nucleotides into its open reading frame, thereby resulting in the translation of a nonsense or otherwise dysfunctional protein product.
[0053] As used herein, the term "vector" means such a nucleic acid molecule that can transport another nucleic acid that has been connected to it. One type of vector is a "plasmid", which generally refers to a circular double-stranded DNA loop to which an additional DNA segment can be connected, but also includes linear double-stranded molecules, such as those obtained by polymerase chain reaction (PCR) amplification or treating a circular plasmid with a restriction enzyme. Other vectors include cosmids, bacterial artificial chromosomes (BACs) and yeast artificial chromosomes (YACs). Another type of vector is a viral vector, in which an additional DNA segment can be connected to the viral genome (discussed in more detail below). Certain vectors can replicate autonomously in the host cell into which they are introduced (e.g., a vector with a replication origin that works in the host cell). Other vectors can be integrated into the genome of the host cell after being introduced into the host cell, thereby being replicated together with the host genome. In addition, some preferred vectors can guide the expression of genes operably connected to them. Such vectors are referred to herein as "recombinant expression vectors" (or simply "expression vectors").
[0054] An "operably linked" or "operably linked" expression control sequence refers to a linkage in which the expression control sequence is in close proximity to a gene of interest to control the gene of interest, as well as an expression control sequence that acts in trans or at a distance to control the gene of interest.
[0055] The term "expression control sequence" refers to a polynucleotide sequence necessary for affecting the expression of a coding sequence operably linked to them. Expression control sequences are sequences that control transcription, post-transcriptional events, and translation of nucleic acid sequences. Expression control sequences include appropriate transcription initiation, termination, promoter, and enhancer sequences; effective RNA processing signals, such as splicing and polyadenylation signals; sequences that stabilize cytoplasmic mRNA; sequences that improve translation efficiency (e.g., ribosome binding sites); sequences that improve protein stability; and, when necessary, sequences that improve protein secretion. The nature of such control sequences varies depending on the host organism; in prokaryotes, such control sequences typically include promoters, ribosome binding sites, and transcription termination sequences. The term "control sequence" is intended to include at least all components whose presence is essential for expression, and may also include additional components whose presence is advantageous, such as leader sequences and fusion partner sequences.
[0056] The term "regulatory element" refers to any element that affects the transcription or translation of a nucleic acid molecule. These include, by way of example, but not limited to, regulatory proteins (e.g., transcription factors), chaperone proteins, signaling proteins, RNAi molecules, antisense RNA molecules, microRNAs, and RNA aptamers. The regulatory element may be endogenous to the host organism. The regulatory element may also be exogenous to the host organism. The regulatory element may be a synthetically produced regulatory element.
[0057] As used herein, the term "promoter", "promoter element" or "promoter sequence" refers to a DNA sequence that, when linked to a nucleotide sequence of interest, is capable of controlling the transcription of a nucleotide sequence of interest into mRNA. The promoter is typically (but not necessarily) located 5' (i.e., upstream) of the nucleotide sequence of interest that is transcribed into mRNA controlled by the promoter, and provides a site for specific binding of RNA polymerase and other transcription factors for initiating transcription. The promoter may be endogenous to the host organism. The promoter may also be exogenous to the host organism. The promoter may be a synthetically produced regulatory element.
[0058] Promoters that can be used to express the recombinant genes described herein include constitutive and inducible / repressible promoters. When multiple recombinant genes are expressed in the engineered organisms of the present invention, different genes can be controlled by different promoters or the same promoter in different operons, or the expression of two or more genes can be controlled by a single promoter as part of an operon.
[0059] As used herein, the term "recombinant host cell" (or simply "host cell") is intended to refer to a cell into which a recombinant vector has been introduced. It should be understood that such terms are intended to refer not only to the specific subject cell, but also to the progeny of such cells. Because certain modifications may appear in subsequent generations due to mutations or environmental influences, such progeny may not actually be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein. Recombinant host cells can be isolated cells or cell lines grown in culture or can be cells that reside in living tissues or organisms.
[0060] As used herein, the term "peptide" refers to a short polypeptide, e.g., typically less than about 50 amino acids in length, more typically less than about 30 amino acids in length. The term as used herein includes analogs and mimetics that mimic structure and therefore mimic biological function.
[0061] The term "polypeptide" encompasses naturally occurring and non-naturally occurring proteins and fragments, mutants, derivatives and analogs thereof. A polypeptide may be monomeric or polymeric. Further, a polypeptide may comprise a plurality of different domains, each of which has one or more different activities.
[0062] The term "isolated protein" or "isolated polypeptide" is a protein or polypeptide that, by virtue of its source or derivation: (1) is not associated with naturally associated components that accompany it in its native state, (2) is present in a purity not found in nature, where purity can be judged by the presence of other cellular material (e.g., lacking other proteins from the same species), (3) is expressed by cells from a different species, or (4) does not occur in nature (e.g., it is a fragment of a polypeptide found in nature, or it includes amino acid analogs or derivatives or linkages other than standard peptide bonds not found in nature). Thus, a chemically synthesized polypeptide or a polypeptide synthesized in a cellular system different from the cell from which it naturally originates would be "separated" from its naturally associated components. A polypeptide or protein may also be rendered substantially free of naturally associated components using protein purification techniques well known in the art. As defined, "isolated" does not necessarily require that the protein, polypeptide, peptide or oligopeptide so described has been physically removed from its native environment.
[0063] The term "polypeptide fragment" refers to a polypeptide having a deletion, such as an amino-terminal and / or carboxyl-terminal deletion, compared to a full-length polypeptide. In a preferred embodiment, a polypeptide fragment is a continuous sequence, wherein the amino acid sequence of the fragment is identical to the corresponding positions in the naturally occurring sequence. The length of the fragment is generally at least 5, 6, 7, 8, 9 or 10 amino acids, preferably at least 12, 14, 16 or 18 amino acids, more preferably at least 20 amino acids, more preferably at least 25, 30, 35, 40 or 45 amino acids, even more preferably at least 50 or 60 amino acids, and even more preferably at least 70 amino acids.
[0064] If a nucleic acid sequence encoding a protein has a similar sequence to a nucleic acid sequence encoding the second protein, then the protein has "homology" or is "homologous" to a second protein. Alternatively, a protein has homology to a second protein if it has a "similar" amino acid sequence to the second protein. (Thus, the term "homologous proteins" is defined to mean that two proteins have similar amino acid sequences.) As used herein, homology between two regions of amino acid sequence (particularly with respect to predicted structural similarity) is interpreted to imply functional similarity.
[0065] When "homologous" is used for a protein or peptide, it should be recognized that non-identical residue positions often differ by conservative amino acid substitutions. A "conservative amino acid substitution" is a conservative amino acid substitution in which an amino acid residue is replaced by another amino acid residue with a side chain (R group) having similar chemical properties (e.g., charge or hydrophobicity). In general, conservative amino acid substitutions will not substantially change the functional properties of the protein. In the case where two or more amino acid sequences differ from each other by conservative substitutions, the percentage of sequence identity or degree of homology may be adjusted upward to correct for the conservative nature of the substitution. The manner in which such adjustments are made is well known to those skilled in the art. See, for example, Pearson, 1994, Methods Mol. Biol. 24: 307-31 and 25: 365-89 (incorporated herein by reference).
[0066] Twenty conventional amino acids and their abbreviations follow conventional usage. See Immunology-A Synthesis (Golub and Gren, Sinauer Associates, Sunderland, Mass., 2nd edition, 1991), which is incorporated herein by reference. Twenty conventional amino acids, non-natural amino acids (e.g., α-, α-disubstituted amino acids, N-alkyl amino acids) and other unconventional amino acid stereoisomers (e.g., D-amino acids) may also be suitable components of the polypeptides of the present invention. Examples of unconventional amino acids include: 4-hydroxyproline, γ-carboxyglutamate, ε-N,N,N-trimethyllysine, ε-N-acetyllysine, O-phosphoserine, N-acetylserine, N-formylmethionine, 3-methylhistidine, 5-hydroxylysine, N-methylarginine and other similar amino acids and imino acids (e.g., 4-hydroxyproline). In the polypeptide symbols used herein, according to standard usage and convention, the left-hand end corresponds to the amino terminal and the right-hand end corresponds to the carboxyl terminal.
[0067] The following six groups each contain amino acids that are conservative substitutions for each other: 1) alanine (S), threonine (T); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), alanine (A), valine (V); and 6) phenylalanine (F), tyrosine (Y), tryptophan (W).
[0068] The sequence homology of polypeptides, sometimes also referred to as the percentage of sequence identity, is usually measured using sequence analysis software. See, for example, the Sequence Analysis Software Package of the Genetics Computer Group (GCG), University of Wisconsin Biotechnology Center, 910 University Avenue, Madison, Wis. 53705. Protein analysis software uses homology metrics assigned to various substitutions, deletions, and other modifications (including conservative amino acid substitutions) to match similar sequences. For example, GCG contains programs such as "Gap" and "Bestfit", which can be used with default parameters to determine the sequence homology or sequence identity between closely related polypeptides (e.g., homologous polypeptides from different species of an organism) or between a wild-type protein and its mutant protein. See, for example, GCG version 6.1.
[0069] When comparing a particular polypeptide sequence to a database containing a large number of sequences from different organisms, a useful algorithm is the computer program BLAST (Altschul et al., J. Mol. Biol. 215:403-410 (1990); Gish and States, Nature Genet. 3:266-272 (1993); Madden et al., Meth. Enzymol. 266:131-141 (1996); Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997); Zhang and Madden, Genome Res. 7:649-656 (1997)), especially blastp or tblastn (Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997)).
[0070] The preferred parameters for BLASTp are: expectation: 10 (default); filter: seg (default); gap opening cost: 11 (default); gap extension cost: 1 (default); maximum alignments: 100 (default); word length: 11 (default); number of descriptions: 100 (default); penalty matrix: BLOWSUM62.
[0071] Preferred parameters for BLASTp are: expectation: 10 (default); filter: seg (default); gap opening cost: 11 (default); gap extension cost: 1 (default); maximum alignment: 100 (default); word length: 11 (default); number of descriptions: 100 (default); penalty matrix: BLOWSUM62. The length of polypeptide sequences compared for homology will generally be at least about 16 amino acid residues, typically at least about 20 residues, more typically at least about 24 residues, typically at least about 28 residues, and preferably more than about 35 residues. When searching a database containing sequences from a large number of different organisms, it is preferred to compare amino acid sequences. Database searches using amino acid sequences can be measured by algorithms known in the art other than blastp. For example, FASTA (a program in GCG version 6.1) can be used to compare polypeptide sequences. FASTA provides an alignment and percentage of sequence identity for the region of best overlap between the query sequence and the search sequence. Pearson, Methods Enzymol. 183:63-98 (1990) (incorporated herein by reference). For example, the percent sequence identity between amino acid sequences can be determined using FASTA as provided in GCG Version 6.1 (incorporated herein by reference) with its default parameters (word length of 2, PAM250 scoring matrix).
[0072] Throughout the specification and claims, the word "comprise" or variations such as "comprises" or "comprising", will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.
[0073] Although exemplary methods and materials are described below, methods and materials similar or equivalent to those described herein can also be used in the practice of the present invention and will be apparent to those skilled in the art. All publications and other references mentioned herein are incorporated herein by reference in their entirety. In the event of a conflict, the present specification including definitions shall prevail. The materials, methods, and examples are illustrative only and are not intended to be limiting.
[0074] Overview
[0075] Provided herein are recombinant strains and methods of producing recombinant strains to increase the yield of full-length desired products in target cells, for example, by reducing protease degradation.
[0076] In some embodiments, to attenuate protease activity in Pichia pastoris, the genes encoding these enzymes are inactivated or mutated to reduce or eliminate activity. This can be accomplished by mutating or inserting the gene itself or by modifying genetic regulatory elements. This can be accomplished by standard yeast genetics techniques. Examples of such techniques include gene replacement by double homologous recombination, in which homologous regions flanking the inactivated gene are cloned in a vector flanked by a selectable marker gene (e.g., an antibiotic resistance gene or a gene that complements an auxotrophic yeast strain).
[0077] Alternatively, PCR amplification can be performed to the homologous region, and it is connected to a selectable marker gene by overlapping PCR. Subsequently, by methods known in the art such as electroporation, such DNA fragments are converted into Pichia pastoris. Then by standard techniques (such as PCR or Southern blotting on genomic DNA), the transformant grown under selective conditions is subjected to gene disruption event analysis. In alternative experiments, gene inactivation can be achieved by single homologous recombination, in which case, for example, the 5' end of the ORF of the gene is cloned on a promoterless vector also containing a selectable marker gene. After such vectors are linearized by digestion with restriction enzymes that only cut the vectors in the homologous fragments of the target gene, such vectors are converted into Pichia pastoris. By PCR or Southern blotting on genomic DNA, integration at the target gene site has been confirmed. In this way, the replication of the gene fragment cloned on the vector is realized in the genome, generating two copies of the target gene locus: the first copy, in which the ORF is incomplete, thereby resulting in the expression of the shortened inactive protein (if expressed); and the second copy, which is not used to drive the promoter for transcription.
[0078] Alternatively, transposon mutagenesis is used to inactivate the target gene. Libraries of such mutants can be screened by PCR for insertion events in the target gene.
[0079] The functional phenotype (i.e., defect) of the engineered / knockout strain can be evaluated using techniques known in the art. For example, the defect of the protease activity of the engineered strain can be detected using any of a variety of methods known in the art, such as determination of the hydrolysis activity of a chromogenic protease substrate, band shift of a substrate protein of a selected protease, etc.
[0080] The weakening of protease activity described herein can be achieved by mechanisms other than knockout mutations. For example, the desired protease can be weakened via amino acid sequence changes in the following manner: changing the nucleic acid sequence, placing the gene under the control of a promoter with lower activity, downward regulation, expressing interfering RNA, ribozymes or antisense sequences targeting the gene of interest, or any other technology known in the art. In preferred strains, the protease activity of the protease encoded at PAS_chr4_0584 (YPS1-1) and PAS_chr3_1157 (YPS1-2) (for example, comprising a polypeptide of SEQ ID NO:66 and 67) is weakened by any of the above methods. In some aspects, the present invention relates to methylotrophic yeast strains, especially Pichia pastoris strains, wherein YPS1-1 and YPS1-2 genes (for example, as shown in SEQ ID NO:1 and SEQ ID NO:2) have been inactivated. In some embodiments, additional protease encoding genes can also be knocked out according to the method provided herein to further reduce the protease activity of the desired protein product expressed by the strain.
[0081] Production of recombinant strains
[0082] Provided herein are transformed strains to reduce the method for activity, for example, using a carrier to deliver a recombinant gene or knock out or otherwise weaken an endogenous gene as required. These carriers can take the form of a carrier backbone containing a replication origin and a selective marker (usually antibiotic resistance, but many other methods are possible), or allow a linear fragment to be incorporated into the target cell chromosome. The carrier should correspond to the selected organism and insertion method.
[0083] Once the elements of the vector are selected, the construction of the vector can be performed in many different ways. In one embodiment, a DNA synthesis service or a method for preparing each vector individually can be used.
[0084] Once the DNA for each vector is obtained (including additional elements required for insertion and manipulation), it must be assembled. There are many possible assembly methods, including (but not limited to) restriction enzyme cloning, blunt end ligation, and overlapping assembly [see, e.g., Gibson, DG et al., Enzymatic assembly of DNA molecules up to several hundred kilobases. Nature methods, 6(5), 343–345 (2009) and GeneArt Kit (http: / / tools.invitrogen.com / content / sfs / manuals / geneart_seamless_cloning_and_assembly_man.pdf)]. Overlapping assembly provides a way to ensure that all elements are assembled in the correct position and that no undesired sequences are introduced.
[0085] The vectors generated above can be inserted into target cells using standard molecular biology techniques, such as molecular cloning. In one embodiment, the target cells have been engineered or selected so that they already contain the genes required to produce the desired product, although this can also be done during or after further vector insertion.
[0086] Depending on the organism and the type of library element (plasmid or genomic insert), several known methods for inserting a vector containing the DNA to be incorporated into a cell can be used. These may include, for example, transformation of microorganisms capable of taking up and replicating DNA from the native environment, transformation by electroporation or chemical means, transduction with viruses or bacteriophages, mating of two or more cells, or conjugation from different cells.
[0087] Several methods for introducing recombinant DNA into bacterial cells are known in the art, including but not limited to transformation, transduction, and electroporation, see Sambrook et al., Molecular Cloning: A Laboratory Manual (1989), 2nd ed., Cold Spring Harbor Press, Plainview, NY. Non-limiting examples of commercial kits and bacterial host cells for transformation include NovaBlue Singles TM (EMD Chemicals Inc., NJ, USA), Max DH5α TM 、One BL21(DE3) Escherichia coli cells, One BL21 (DE3) pLys Escherichia coli cells (Invitrogen Corp., Carlsbad, Calif., USA), XL1-Blue competent cells (Stratagene, CA, USA). Non-limiting examples of commercial kits and bacterial host cells for electroporation include Zappers TM Electrocompetent cells (EMD Chemicals Inc., NJ, USA), XL1-Blue electroporation-competent cells (Stratagene, CA, USA), ElectroMAX TM Agrobacterium tumefaciens LBA4404 cells (Invitrogen Corp., Carlsbad, Calif., USA).
[0088] Several methods for introducing recombinant nucleic acids into eukaryotic cells are known in the art. Exemplary methods include transfection, electroporation, liposome-mediated nucleic acid delivery, microinjection into host cells, see Sambrook et al., Molecular Cloning: A Laboratory Manual (1989), Second Edition, Cold Spring Harbor Press, Plainview, NY. Non-limiting examples of commercial kits and reagents for transfecting recombinant nucleic acids into eukaryotic cells include Lipofectamine TM 2000、Optifect TM Reagents, calcium phosphate transfection kit (Invitrogen Corp., Carlsbad, Calif., USA), Transfection reagents, Transfection reagent (Stratagene, CA, USA). Alternatively, the recombinant nucleic acid can be introduced into insect cells (e.g., sf9, sf21, High Five TM )middle.
[0089] The transformed cells are separated so that each clone can be tested individually. In one embodiment, this is accomplished by spreading the culture on one or more flat plates containing a selection agent (or lacking a selection agent) that will ensure that only the transformed cells survive and reproduce. The specific agent can be an antibiotic (if the library contains an antibiotic resistance marker), a metabolite of the deletion (for nutritional deficiency supplementation) or other selection methods. The cells are grown into single colonies, each of which contains a single clone.
[0090] Bacterial colonies are screened for the production of desired proteins, metabolites or other products or for a reduction in protease activity. In one embodiment, screening identifies recombinant cells with the highest (or sufficiently high) product production titer or efficiency. This includes a reduction in the proportion of degradation products or an increase in the total amount of the desired polypeptide of the full length collected from the cell culture.
[0091] This assay can be performed by growing each clone (one per hole) in a multi-well culture plate. Once the cells have reached a suitable biomass density, they are induced with methanol. After a period of time (generally 24-72 hours of induction), the culture is harvested by rotating in a centrifuge to precipitate the cells and remove the supernatant. The supernatant from each culture can then be tested for protease activity and / or protein degradation.
[0092] Silk Sequence
[0093] In some embodiments, the modified strains with reduced protease activity described herein recombinantly express filamentous polypeptide sequences. In some embodiments, the filamentous polypeptide sequence is 1) a block copolymer polypeptide composition produced by mixing and matching the repeat domains derived from the silk polypeptide sequence, and / or 2) recombinant expression of the block copolymer polypeptides with sufficiently large size (about 40kDa) to form useful fibers by secretion by industrially scalable microorganisms. Large (about 40kDa to about 100kDa) block copolymer polypeptides (comprising sequences of almost all published amino acid sequences from spider silk polypeptides) engineered by silk repeat domain fragments can be expressed in modified microorganisms described herein. In some embodiments, the silk polypeptide sequence is matched and designed to produce highly expressed and secreted polypeptides that can form fibers. In some embodiments, the knockout of protease genes in the host modified strains or the reduction of protease activity reduce the degradation of filamentous polypeptides.
[0094] The composition for expression and secretion of block copolymer is provided herein in several embodiments, and the block copolymer is engineered from the combined mixture of silk polypeptide domains spanning silk polypeptide sequence space, wherein the block copolymer has minimum degradation. In some embodiments, the method for secreting block copolymer with minimum degradation in expandable microorganisms (e.g., yeast, fungi and gram-positive bacteria) is provided herein. In some embodiments, the block copolymer polypeptide comprises 0 or more N-terminal domains (NTD), 1 or more repeat domains (REP) and 0 or more C-terminal domains (CTD). In some aspects of the embodiment, the block copolymer polypeptide is>100 amino acids of a single polypeptide chain. In some embodiments, the block copolymer polypeptide comprises a domain that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a block copolymer polypeptide sequence disclosed in International Publication No. WO / 2015 / 042164, “Methods and Compositions for Synthesizing Improved Silk Fibers” (incorporated by reference in its entirety).
[0095] Several types of primitive spider silk have been identified. The mechanical properties of each naturally spun type are believed to be closely related to the molecular composition of that silk. See, e.g., Garb, JE et al., Untangling spider silk evolution with spidroin terminal domains, BMC Evol. Biol., 10:243 (2010); Bittencourt, D. et al., Protein families, natural history and biotechnological aspects of spidersilk, Genet. Mol. Res., 11:3 (2012); Rising, A. et al., Spider silk: recent advances in proteins. recombinant production, structure-function relationships and biomedical applications, Cell. Mol. Life Sci., 68:2, pg.169-184 (2011); and Humenik, M. et al., Spider silk: understanding the structure-function relationship of a natural fiber, Prog. Mol. Biol. Transl. Sci., 103, pg. 131-85 (2011). For example:
[0096] Acoustic gland (AcSp) filaments tend to have high toughness, which is the result of a combination of reasonably high strength and reasonably high ductility. AcSp filaments are characterized by large block ("monoblock repeat") sizes, which are often doped with motifs of polyserine and GPX. Tubular gland (TuSp or cylindrical) filaments tend to have large diameters, with moderate strength and high ductility. TuSp filaments are characterized by their polyserine and polythreonine content, as well as short bundles of polyalanine. Large ampullate (MaSp) filaments tend to have high strength and moderate ductility. MaSp filaments can be one of two subtypes: MaSp1 and MaSp2. MaSp1 filaments are generally less ductile than MaSp2 filaments and are characterized by polyalanine, GX and GGX motifs. MaSp2 filaments are characterized by polyalanine, GGX and GPX motifs. Small ampullate (MiSp) filaments tend to have moderate strength and moderate ductility. MiSp filaments are characterized by GGX, GA, and poly A motifs, and often contain a spacer element of about 100 amino acids. Flagellar (Flag) filaments tend to have high extensibility and moderate strength. Flag filaments are usually characterized by GPG, GGX, and a short spacer motif.
[0097] The properties of each silk type can vary from species to species, and spiders with different lifestyles (e.g., sedentary web spinners vs. vagabond hunters) or evolutionarily older spiders can produce silks with properties different from those described above (for descriptions of spider diversity and classification, see Hormiga, G. and Griswold, CE, Systematics, phylogeny, and evolution of orb-weaving spiders, Annu. Rev. Entomol. 59, pg. 487-512 (2014); and Blackedge, TA et al., Reconstructing web evolution and spider diversification in the molecular era, Proc. Natl. Acad. Sci. USA, 106: 13, pg. 5229-5234 (2009)). However, synthetic block copolymer polypeptides with sequence similarity and / or amino acid composition similarity to the repeating domains of the original silk proteins can be used to produce consistent silky fibers that reproduce the properties of the corresponding natural silk fibers on a commercial scale.
[0098] In some embodiments, a list of putative silk sequences can be compiled by searching GenBank for relevant terms, such as "spidroin", "fibroin", "MaSp", and those sequences can be pooled with additional sequences obtained by independent sequencing work. The sequences are then translated into amino acids, duplicate entries are filtered, and manually split into domains (NTD, REP, CTD). In some embodiments, the candidate amino acid sequences are reverse translated into DNA sequences optimized for expression in Komagataella yeast. The DNA sequences are each cloned into an expression vector and transformed into Komagataella yeast. In some embodiments, various silk domains that have been shown to be successfully expressed and secreted are then assembled in a combinatorial manner to construct silk molecules that can form fibers.
[0099] Silk polypeptides characteristically consist of a repeat domain (REP) flanked by non-repeat regions (e.g., C-terminal and N-terminal domains). In embodiments, the length of the C-terminal and N-terminal domains is between 75 and 350 amino acids. The repeat domains exhibit a hierarchical architecture, such as Figure 1As shown. The repeat domain contains a series of blocks (also called repeat units). The blocks are repeated throughout the silk repeat domain, sometimes perfectly repeated, sometimes imperfectly repeated (constituting a quasi-repeat domain). The length and composition of the blocks vary between different silk types and between different species. Table 1 lists examples of block sequences from selected species and silk types, and further examples are given in the following literature: Rising, A. et al., Spider silk proteins: recent advances in recombinant production, structure-function relationships and biomedical applications, Cell Mol. Life Sci., 68: 2, pg169-184 (2011), and Gatesy, J. et al., Extreme diversity, conservation, and convergence of spider silk fibroin sequences, Science, 291: 5513, pg. 2603-2605 (2001). In some cases, blocks can be arranged in a regular pattern to form a larger macro-repeat that appears multiple times (usually 2 to 8 times) in the repeat domain of the silk sequence. Repeated blocks within a repeat domain or macro-repeat, as well as macro-repeat repeated within a repeat domain, can be separated by spacer elements. In some embodiments, the block sequence comprises a glycine-rich region, followed by a poly A region. In some embodiments, short (about 1 to 10) amino acid motifs appear multiple times within the block. For the purposes of the present invention, blocks from different natural silk polypeptides can be selected without reference to a cyclic arrangement (i.e., blocks similar to other aspects identified between silk polypeptides may not be aligned due to a cyclic arrangement). Therefore, for example, for the purposes of the present invention, the "block" of SGAGG (SEQ ID NO: 494) is the same as GSGAG (SEQ ID NO: 495) and is the same as GGGSA (SEQ ID NO: 496); they are all cyclic arrangements of each other. The particular arrangement chosen for a given silk sequence may be dictated, among other things, by convenience (usually beginning with a G). Silk sequences obtained from the NCBI database can be divided into blocks and non-repetitive regions.
[0100] Table 1: Sample block sequences
[0101]
[0102]
[0103]
[0104]
[0105] The fiber-forming block copolymer polypeptide from block and / or macro-repetitive domain according to some embodiments of the present invention is described in International Publication No. WO / 2015 / 042164 (incorporated by reference). According to the domain (N-terminal domain, repetitive domain and C-terminal domain), the natural silk sequence obtained from protein database (such as GenBank) or by de novo sequencing is decomposed. The N-terminal domain and C-terminal domain sequence selected for synthesis and assembly into fiber include natural amino acid sequence information and other modifications described herein. The repetitive domain is decomposed into repetitive sequence, and the repetitive sequence contains representative blocks, and the block is generally 1 to 8 according to the type of silk, and the block captures key amino acid information, while the size of the DNA encoding amino acids is reduced to a fragment that is easy to synthesize. In some embodiments, the appropriately formed block copolymer polypeptide includes at least one repetitive domain containing at least 1 repetitive sequence, and is optionally flanked by N-terminal domain and / or C-terminal domain.
[0106] In some embodiments, the repetitive domain comprises at least one repetitive sequence. In some embodiments, the repetitive sequence is 150 to 300 amino acid residues. In some embodiments, the repetitive sequence comprises multiple blocks. In some embodiments, the repetitive sequence comprises multiple macro repeats. In some embodiments, blocks or macro repeats are divided into multiple repetitive sequences.
[0107] In some embodiments, repetitive sequence starts with glycine, and can not end with phenylalanine (F), tyrosine (Y), tryptophan (W), cysteine (C), histidine (H), asparagine (N), methionine (M) or aspartic acid (D), to meet DNA assembly requirements. In some embodiments, some of repetitive sequence can be changed compared with the original sequence. In some embodiments, repetitive sequence can be changed, for example, by adding serine (to avoid terminating at F, Y, W, C, H, N, M or D) to the C-terminus of polypeptide. In some embodiments, repetitive sequence can be modified by filling in the homologous sequence of another block in incomplete block. In some embodiments, repetitive sequence can be modified by rearranging the order of block or macro repeat body.
[0108] In some embodiments, non-repetitive N- and C-terminal domains can be selected for synthesis. In some embodiments, the N-terminal domain can be removed, for example, by a leader signal sequence identified by SignalP (Peterson, TN et al., SignalP 4.0: discriminating signal peptides from transmembrane regions, Nat. Methods, 8: 10, pg. 785-786 (2011).
[0109] In some embodiments, the N-terminal domain, repeat sequence, or C-terminal domain sequence can be from Agelenopsis aperta, Aliatypus gulosus, Aphonopelmaseemanni, Aptoschius sp. AS217, Aptoschius sp. AS220, Araneus diadematus, Araneus gemmoides, Araneus ventricosus, Argiope amoena, Argiope argentata, Argiope bruennichi, Argiope trifasciata, Atypoides riversi, Avicularia juruensis, Bothriocyrtum californicum, Deinopis spinosa, Diguetia canities), black fishing spider (Dolomedestenebrosus), Euagrus chisoseus, nursery web spider (Euprosthenops australis), masticated flag spider (Gasteracantha mammosa), Hypochilus thorelli, Kukulcania hibernalis, black widow spider (Latrodectus hesperus), Megahexura fulva, Metepeira grandiosa, golden orb weaver (Nephila antipodiana), club astilbe (Nephila clavata), astilbe spider (Nephila clavipes), Madagascar astilbe (Nephila madagascariensis), spotted astilbe (Nephila pilipes), Nephilengyscruentata, Parawixia bistriata, green lynx spider (Peucetiaviridans), primitive carnivorous spider (Plectreurys tristis, Poecilotheria regalis, Tetragnatha kauaiensis, or Uloborus diversus.
[0110] In some embodiments, the silk polypeptide nucleotide coding sequence can be operably connected to an α mating factor nucleotide coding sequence. In some embodiments, the silk polypeptide nucleotide coding sequence can be operably connected to another endogenous or heterologous secretion signal coding sequence. In some embodiments, the silk polypeptide nucleotide coding sequence can be operably connected to a 3X FLAG nucleotide coding sequence. In some embodiments, the silk polypeptide nucleotide coding sequence is operably connected to other affinity tags such as 6 to 8 His residues (SEQ ID NO:520).
[0111] Filamentous polypeptide
[0112] In certain embodiments, the Pichia yeast strain disclosed herein has been modified to express a filamentous polypeptide. WO2015 / 042164, in particular paragraphs 114 to 134 (incorporated herein by reference), provides a method for producing a preferred embodiment of a filamentous polypeptide. Disclosed herein is a synthetic protein copolymer based on a recombinant spider silk protein fragment sequence derived from, for example, MaSp2 from the species Argiope bruennichi. A filamentous polypeptide is described, comprising two to twenty repeating units, wherein the molecular weight of each repeating unit is greater than about 20 kDa. There are more than about 60 amino acid residues organized into many "quasi-repeating units" in each repeating unit of the copolymer. In some embodiments, the repeating unit of the polypeptide described in the present disclosure has at least 95% sequence identity with the MaSp2 dragline silk protein sequence.
[0113] In some embodiments, each "repeat unit" of the filamentous polypeptide comprises from two to twenty "quasi-repeat" units (ie, n 3 is 2 to 20). Quasi-repeats are not necessarily exact repeats. Each repeat may consist of quasi-repeats in series. Equation 1 shows the composition of repeating units according to the present disclosure and the composition of repeating units incorporated by reference from WO 2015 / 042164. Each filamentous polypeptide may have one or more repeating units as defined by equation 1.
[0114] {GGY-[GPG-X 1 ] n1 -GPS-(A) n2} n3 (SEQ ID NO:514) (Equation 1)
[0115] Variable component X 1 (referred to as a "motif") is any of the following amino acid sequences shown in Equation 2, and X 1 Varies randomly within each quasi-repeating unit.
[0116] X 1 =SGGQQ (SEQ ID NO: 515) or GAGQQ (SEQ ID NO: 516) or
[0117] GQGPY (SEQ ID NO: 517) or AGQQ (SEQ ID NO: 518) or SQ (Equation 2)
[0118] Referring again to equation 1, in equation 1, "GGY-[GPG-X 1 ] n1 -GPS" (SEQ ID NO: 521) is referred to as the "first region". The quasi-repeating unit is formed in part by repeating the first region in the quasi-repeating unit 4 to 8 times. That is, n 1 The value of represents the number of first region units repeated within a single quasi-repeating unit, n 1 The value of is any one of 4, 5, 6, 7 or 8. Use "(A) n2 The constituent element represented by "(SEQ ID NO: 522)" (i.e., poly A sequence) is referred to as the "second region" and is formed by repeating the amino acid sequence "A" n times in each quasi-repeating unit. 2 times (SEQ ID NO: 522). That is, n 2 The value of represents the number of second region units repeated within a single quasi-repeating unit, n 2 The value is any one of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. In some embodiments, the repeating unit of the polypeptide of the present disclosure has at least 95% sequence identity with the sequence comprising the quasi-repeating unit described in Equations 1 and 2. In some embodiments, the repeating unit of the polypeptide of the present disclosure has at least 80%, or at least 90%, or at least 95%, or at least 99% sequence identity with the sequence comprising the quasi-repeating unit described in Equations 1 and 2.
[0119] In another embodiment, 3 "long" quasi-repeating units are followed by 3 "short" quasi-repeating units. A short quasi-repeating unit is one in which n 1 =4 or 5 quasi-repeating units. The long quasi-repeating unit is defined as 1 =6, 7 or 8 quasi-repeating units. In some embodiments, all short quasi-repeating bodies have the same X at the same position within each quasi-repeating unit of the repeating unit. 1 In some embodiments, no more than 3 of the 6 quasi-repeating units have the same X 1 Motif.
[0120] In other embodiments, the repeating unit consists of quasi-repeating units that use the same X in the rows within the repeating unit. 1 In another embodiment, the repeating unit consists of quasi-repeating units, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 quasi-repeating units use the same X in a single quasi-repeating unit of the repeating unit. 1 No more than 2 times.
[0121] Thus, in some embodiments, provided herein are strains of yeast that recombinantly express a filamentous polypeptide with reduced degradation to increase the amount of the full-length polypeptide present in an isolated product from a cell culture. In some embodiments, the strain expressing the filamentous polypeptide is a Pichia pastoris strain comprising a PAS_chr4_0584 knockout and a PAS_chr3_1157 knockout.
[0122] Equivalents and scope
[0123] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein.The scope of the present invention is not intended to be limited to the above description, but rather is set forth in the appended claims.
[0124] In the claims, articles such as "a," "an," and "the" may mean one or more than one, unless otherwise indicated or otherwise clear from the context. Claims or descriptions that include "or" between one or more members of a group are deemed to comply if one, more than one, or all of the group members are present in, employed by, or otherwise related to a given product or process, unless otherwise indicated or otherwise clear from the context. The invention includes embodiments in which exactly one of the group members is present in, employed by, or otherwise related to a given product or process. The invention includes embodiments in which more than one or all of the group members are present in, employed by, or otherwise related to a given product or process.
[0125] It should also be noted that the term "comprising" is intended to be open ended, and allows but does not require the inclusion of additional elements or steps. When the term "comprising" is used herein, the term "consisting of" is also encompassed and disclosed thereby.
[0126] If a range is given, the endpoints are included. In addition, it should be understood that unless otherwise indicated or obvious from the context and the understanding of a person of ordinary skill in the art, the values expressed as ranges can adopt any specific value or sub-range within the range stated in different embodiments of the present invention, up to one tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise.
[0127] All cited sources, such as references, publications, databases, database entries and techniques cited herein, are incorporated by reference into this application, even if not explicitly stated in the citation. If the statements in the cited sources conflict with the statements in this application, the statements in this application shall prevail.
[0128] Section headings and table headings are not intended to be limiting.
[0129] Example
[0130] Below is the example of the specific embodiment of the present invention.These examples are provided only for illustrative explanation, and are not intended to limit the scope of the present invention in any way.For the numerals (for example, amount, temperature, etc.) used, try to ensure accuracy, but of course some experimental errors and deviations should be allowed.
[0131] Unless otherwise indicated, the implementation of the present invention will utilize conventional methods of protein chemistry, biochemistry, recombinant DNA technology and pharmacology within the technical scope of the art. Such techniques are fully explained in the literature. See, for example, TECreighton, Proteins: Structures and Molecular Properties (WH Freeman and Company, 1993); A Lehninger, Biochemistry (Worth Publishers, Inc., latest supplement); Sambrook et al., Molecular Cloning: A Laboratory Manual (2nd edition, 1989); Methods In Enzymology (S. Colowick and N. Kaplan, Academic Press, Inc.); Remington's Pharmaceutical Sciences, 18th edition (Easton, Pennsylvania: Mack Publishing Company, 1990); Carey and Sundberg, Advanced Organic Chemistry, 3rd edition (Plenum Press), Volumes A and B (1992).
[0132] Example 1: Production of recombinant yeast expressing 18B
[0133] First, we transformed a strain of Pichia pastoris to disable KU70 function to facilitate further editing and engineering. A HIS+ derivative of the Pichia pastoris strain GS115 (NRRL Y15851) was electroporated with a DNA cassette HIS consisting of homology arms flanking the bleomycin resistance marker and targeting the KU70 locus. The map of the cassette is shown in Figure 1 The sequences are given in Table 10. The transformants were plated on YPD agar plates supplemented with zeocin. This resulted in loss of function of KU70.
[0134] We then modified the strain to express a recombinant gene encoding a filamentous polypeptide. The HIS+ derivative of Komagataella phaffii strain GS115 (NRRL Y15851) was transformed with the recombinant vector (SEQ ID NO: 462) to express and secrete a filamentous polypeptide ("18B") (SEQ ID NO: 463). Transformation was accomplished by electroporation as described in PMID 15679083 (incorporated herein by reference).
[0135] Each vector includes an 18B expression cassette containing a polynucleotide sequence encoding a filamentous protein in a recombinant vector flanked by a promoter (pGCW14) and a terminator (tAOX1 pA signal). The recombinant vector also includes a major resistance marker for selecting bacterial and yeast transformants and a bacterial origin of replication. The first recombinant vector includes a targeting region that directs the 18B polynucleotide sequence to integrate directly at the 3′ end of the AOX2 locus in the Pichia pastoris genome. The resistance marker in the first vector confers resistance to G418 (also known as Geneticin). The second recombinant vector includes a targeting region that directs the 18B polynucleotide sequence to integrate directly at the 3′ end of the TEF1 locus in the Pichia pastoris genome. The resistance marker in the second vector confers resistance to hygromycin B.
[0136] Example 2: Generation of a library of single protease KO mutants
[0137] After successful transformation and secretion of 18B in a recombinant Pichia pastoris strain, 65 open reading frames (ORFs) encoding proteases were targeted for deletion (Table 2). Cells were transformed with a vector containing a DNA cassette with approximately 1150 bp homology arms flanking a nourseothricin resistance marker. Figure 2 The plasmid map containing the nourseothricin resistance marker is shown in and the sequence is given in Table 11.
[0138] The homology arms used for each target were amplified by the primers given in Table 7 and inserted into the nourseothricin resistance plasmid. The homology arms were inserted into the nourseothricin plasmid to generate a cassette containing the nourseothricin resistance marker flanked by 3' and 5' homology arms to the target protease, as shown in Table 7. Figure 3A and Figure 3B As shown. Figure 3A In FIG, the NourResistance Cassette is shown flanked by homology arms (HA1 and HA2). Figure 3B , detailed information of the nourseothricin marker is shown, including the promoter of the ILV5 gene from Saccharomyces cerevisiae (pILV5), the nourseothricin acetyltransferase gene from Streptomyces noursei (nat), and the poly A signal of the CYC1 gene from Saccharomyces cerevisiae.
[0139] The homology arms in each vector targeted one of the 65 desired protease loci as given in Table 2. Transformants were plated on YPD agar plates supplemented with nourseothricin and grown at 30°C for 48 hours.
[0140] Table 2 - Proteases targeted for deletion in Pichia pastoris strains
[0141]
[0142]
[0143] Example 3: Testing individual protease knockout clones for reduction in protein degradation
[0144] The resulting clones were inoculated into 400 μL of buffered glycerol complex medium (BMGY) in a 96-well plate and incubated at 30°C, 1,000 rpm for 48 hours. After 48 hours of incubation, 4 μL of each culture was used to inoculate 400 μL of BMGY in a 96-well plate and then incubated at 30°C for 48 hours. Guanidine thiocyanate was added to the cell culture to a final concentration of 2.5 M to extract the recombinant protein. After 5 minutes of incubation, the solution was centrifuged and the supernatant was sampled and analyzed by Western blotting.
[0145] Western blot data for representative clones for each protease knockout are shown in Figure 3. Individual protease deletions did not show a clear effect on the distribution of silk fragments detected by Western blot.
[0146] Example 4: Generation of a protease double knockout library
[0147] In addition to single KOs, pairwise combinations of different proteases were knocked out. These proteases were chosen in part because they are paralogs of each other that may have compensatory functions.
[0148] To generate double knockouts, nourseothricin resistance was eliminated from the single protease knockout strain generated in Example 2, and the second protease was deleted by transformation with a second nourseothricin resistance cassette as provided in Example 2. Transformants were plated on YPD agar plates supplemented with nourseothricin and incubated at 30° C. for 48 hours. The double protease knockouts tested are given in Table 3.
[0149] Table 3 - Protease double KO strains of Pichia pastoris expressing filamentous polypeptides
[0150]
[0151]
[0152] Example 5: Testing double protease knockout clones for reduction in protein degradation
[0153] The resulting clones were inoculated into 400 μL of buffered glycerol complex medium (BMGY) in a 96-well plate and incubated at 30°C, 1,000 rpm for 48 hours. After 48 hours of incubation, 4 μL of each culture was used to inoculate 400 μL of BMGY in a 96-well plate and then incubated at 30°C for 48 hours. Guanidine thiocyanate was added to the cell culture to a final concentration of 2.5 M to extract the recombinant protein. After 5 minutes of incubation, the solution was centrifuged and the supernatant was sampled and analyzed by Western blotting.
[0154] Figure 4 Representative results from different protease double knockout strains are shown. As shown, although there is protein degradation in all single knockout strains tested, the combination of PAS_chr4_0584+PAS_chr3_1157 protease knockouts (from bacterial strain 3 of Table 3) causes degradation products to be almost completely eliminated. Other protease combinations do not all cause the elimination of degradation products.
[0155] Example 6: Additional protease knockout strains
[0156] As shown in Examples 4 and 5, modified Pichia cells capable of producing a desired protein (e.g., 18B) were transformed to delete proteases at PAS_chr4_0584 and PAS_chr3_1157 to mitigate degradation of the desired protein. We further knocked out one or more additional proteases to increase the yield of the full-length product and minimize degradation.
[0157] For each additional knockout, an additional protease gene was deleted from a single protease KO (1X KO), double protease KO (2X KO), triple protease KO (3X KO), or quadruple protease KO (4X KO) by transformation with nourseothricin with homology arms targeting the desired gene as provided in Example 2. The protease genes knocked out in each strain are shown in Table 4:
[0158] Table 4: 2X to 5X KO strains
[0159]
[0160] The resulting cells were separated on selective media plates (by auxotrophic or antibiotic resistance markers), and individual clones were isolated for further testing. Individual clones were tested by liquid culture assays under conditions that produced product protein as follows: a colony of each isolated strain was inoculated into 400 μL of buffered glycerol complex medium (BMGY) in a 96-well plate and incubated at 30°C, 1000 rpm agitation for 48 hours. After incubation for 48 hours, 4 μL of each culture was used to inoculate 400 μL of BMGY or 400 μL of YPD (yeast extract powder peptone dextrose medium) in a 96-well plate, and then incubated at 30°C, 1,000 rpm for 48 hours.
[0161] Proteins expressed by the cells were isolated and analyzed for degradation as follows: Recombinant proteins were extracted by adding guanidine thiocyanate to the cell culture to a final concentration of 2.5 M. After 5 minutes of incubation, the solution was centrifuged and the supernatant was sampled and analyzed by Western blotting.
[0162] Figure 5 Shown are the results of Western blots of purified proteins from 2X KO, 3X KO, 4X KO and 5X KO strains inoculated in BMGY or YPD. As shown, deletion of additional protease genes from strains with PAS_chr4_0584+PAS_chr3_1157 protease knockouts (strain 3 from Table 3) resulted in further elimination of degradation products.
[0163] Other Implementations
[0164] It is to be understood that the words which have been used are words of description rather than limitation and that changes may be made within the purview of the appended claims without departing from the true scope and spirit of the invention in its broader aspects.
[0165] While the present invention has been described at some length and with some particularity from several of the described embodiments, it is not intended to be limited to any such details or embodiments or to any particular embodiment, but should be read with reference to the appended claims so that such claims will be interpreted as broadly as possible in light of the prior art so as to effectively encompass the intended scope of the invention.
[0166] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In the event of a conflict, the present specification, including definitions, will control. In addition, section headings, materials, methods, and examples are illustrative only and are not intended to be limiting.
[0167]
[0168]
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214]
[0215]
[0216]
[0217]
[0218]
[0219]
[0220]
[0221]
[0222]
[0223] Table 8: Forward and reverse primers used to amplify modified sequences
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237]
[0238]
[0239]
[0240]
[0241]
Claims
1. A Pichia pastoris microorganism, wherein the protease whose activity has been weakened or eliminated is a YPS1-1 protease and a YPS1-2 protease, wherein the YPS1-1 protease comprises the polypeptide sequence of SEQ ID NO:67, and the YPS1-2 protease comprises the polypeptide sequence of SEQ ID NO:68, wherein the microorganism expresses a recombinant protein, wherein the recombinant protein comprises the polypeptide sequence of SEQ ID NO:
463.
2. The microorganism of claim 1, wherein the YPS1-1 protease is encoded by the YPS1-1 gene.
3. The microorganism of claim 2, wherein the YPS1-1 gene comprises SEQ ID NO:
1.
4. The microorganism of claim 2, wherein the YPS1-1 gene is located at locus PAS_chr4_0584 of the microorganism.
5. The microorganism of claim 1, wherein the YPS1-2 protease is encoded by a YPS1-2 gene.
6. A microorganism as described in claim 5, wherein the YPS1-2 gene is located at the locus PAS_chr3_1157 of the microorganism.
7. The microorganism of claim 1, wherein the YPS1-1 gene or the YPS1-2 gene or both are mutated or knocked out.
8. The microorganism of claim 1, wherein the recombinant protein is encoded by SEQ ID NO:
462.
9. An engineered microorganism of Pichia pastoris comprising YPS1-1 and YPS1-2 activities reduced by mutation or deletion of a YPS1-1 gene comprising SEQ ID NO:1 and a YPS1-2 gene comprising SEQ ID NO:2, wherein the microorganism further comprises an expressed recombinant protein comprising a polypeptide sequence encoded by SEQ ID NO:
462.
10. A cell culture comprising the microorganism of any one of claims 1 to 9.
11. A cell culture comprising the microorganism of any one of claims 1 to 9, wherein degradation of the recombinant protein is lower than a cell culture comprising an otherwise identical Pichia pastoris microorganism in which YPS1-1 and YPS1-2 activities are not attenuated or eliminated.
12. A method for producing a recombinant protein with reduced degradation, comprising: Cultivating the microorganism according to any one of claims 1 to 9 in a culture medium under conditions suitable for expression of the recombinant protein; as well as The recombinant protein is isolated from the microorganism or the culture medium.
13. The method of claim 12, wherein the recombinant protein is secreted by the microorganism, and wherein isolating the recombinant protein comprises collecting culture medium containing the secreted recombinant protein.
14. The method of claim 12, wherein the level of degradation of the recombinant protein is reduced compared to the recombinant protein produced by an otherwise identical microorganism wherein the YPS1-1 and the YPS1-2 protease activities are not attenuated or eliminated.
15. A method for modifying Pichia pastoris to reduce the degradation of an expressed recombinant protein comprising the polypeptide sequence of SEQ ID NO: 463, comprising knocking out or mutating the genes encoding the YPS1-1 protein comprising the polypeptide sequence of SEQ ID NO: 67 and the YPS1-2 protein comprising the polypeptide sequence of SEQ ID NO: 68, so that their activities are weakened or eliminated.
Citation Information
Patent Citations
Methods and compositions for synthesizing improved silk fibers
WO2015042164A2