Directed evolution of a novel gluten-degrading protease for diagnostics and treatment of gluten sensitivity / intolerance
Novel proteases like N21, H3, and 2C11, developed through ancestral sequence reconstruction, efficiently degrade gluten peptides at stomach pH, addressing the challenge of gluten sensitivity and intolerance by reducing harmful peptide presence.
Patent Information
- Application Number
- PCT/US2025/021803
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
Existing proteases are ineffective in degrading protease-resistant gluten peptides, which cause immune responses in individuals with gluten sensitivity or intolerance, and there is a need for specific proteases that can effectively degrade these peptides to treat or diagnose gluten-related disorders.
Development of novel proteases, such as N21, H3, and 2C11, which are designed through ancestral sequence reconstruction and directed evolution to efficiently cleave gluten peptides at pH 2-4, utilizing fluorogenic peptide substrates for activity verification.
The novel proteases effectively degrade immunogenic gluten peptides, providing diagnostic tools and therapeutic options for gluten sensitivity and intolerance by reducing harmful peptide presence in the gastrointestinal tract.
Smart Images

Figure US2025021803_02102025_PF_FP_ABST
Abstract
Description
DIRECTED EVOLUTION OF A NOVEL GLUTEN-DEGRADING PROTEASE FORDIAGNOSTICS AND TREATMENT OF GLUTEN SENSITIVITY / INTOLERANCECROSS-REFERENCE TO RELATED APPLICATIONSThis application claims priority to U.S. Provisional Application 63 / 571,944 filed on March 29, 2024, which is incorporated herein by reference in its entirety.SEQUENCE LISTINGThe Instant Application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on March 26, 2025, is named “SEQ LIST— 107648059. xml” and is 54,200 bytes in size. The Sequence Listing does not go beyond the disclosure in the application as filed.BACKGROUND
[0001] Celiac disease (CD) is caused by an immune reaction to eating gluten, a mixture of proteins found in wheat, barley and rye, for example. In patients with celiac disease, ingesting gluten, and in particular gluten proteins, causes an immune response in the small intestine which can damage the small intestine over time.
[0002] Baking flour, for example, contains 10-13% protein. This protein includes 46 wt% gliadin (prolamin), 40 wt% glutenin, 9 wt% albumin and 5 wt% globulin. Gliadin is the major immunogenic protein. The gluten-free limit per FDA regulations is <20 ppm.
[0003] Once ingested, gluten undergoes a partial digestion by gastric-pancreatic and brush-border proteolytic enzymes which results in many peptides of different length (a few to more than 30 amino acids in length) which are resistant to further digestion due to the high content of proline residues which can render them protease-resistant. The use of exogenous proteolytic enzy mes for gluten detoxification has been suggested as a strategy for CD treatment.
[0004] What is needed are novel proteases and methods of specifically degrading gluten proteins which could provide diagnostics and treatments for gluten-insensitive and gluten-intolerant individuals.BRIEF SUMMARY
[0005] In one aspect, included herein is a protease of any of SEQ ID NOs: 6-14 or a proteolytically active variant thereof with greater than 85%, 90%, 95% or 98% sequence identity thereto, wherein the proteolytically active variant cleaves a peptide of SEQ ID NO: 17 at QP|QL to provide peptides of any of SEQ ID NOs. 18-24.
[0006] Also included are compositions comprising the protease such as pharmaceutical compositions, food supplements, and food products.
[0007] Further included are polynucleotides encoding the proteases, expression cassettes comprising the polynucleotides, and cells comprising the foregoing.
[0008] Also included is a method of treating a disorder associated with gluten intolerance, comprising administering to a subject in need thereof an effective amount of the proteases described herein.
[0009] Also included is a method of testing a subject suspected of having gluten intolerance or gluten sensitivity7, comprising administering to the subject a composition comprising a protease or variant thereof as described herein, waiting a period of time, and after the period of time, assessing one or more symptoms of gluten intolerance or gluten sensitivity in the subject, wherein a reduction in the one or more symptoms indicates the subject has gluten intolerance or gluten sensitivity.
[0010] In another aspect, a fluorogenic peptide substrate comprisesFluorophore! - SEQ ID NO: 4 - K. - Fluorophore 2, wherein Fluorophore 1 and Fluorophore 2 are active at pH 2-4.
[0011] In another aspect, a method of identifying an ancestral protease comprises identifying a modem protease sequence, wherein the modem protease is neprosin; querying the modem protease sequence in a sequence database to provide a list of similar sequences; selecting a subset of the list of similar sequences using a sequence similarity cut-off; aligning the subset of the list of similar sequences using a multiple sequence alignment tool to provide aligned sequences; prepanng a phylogenetic tree based on the aligned sequences; selecting a node from the phylogenetic tree and selecting the ancestral proteases from the node; and performing a fluorescence assay with the fluorogenic peptide substrate described herein to verify the activity' of the ancestral protease to degrade gluten peptides.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] FIG. 1 shows the fluorogenic substrate currently used in this study- (HiLyte™ Fluor 488)-SEQ ID NO: 4-(QXL®-520)-NH2 (presumed structure according to datasheet). The peptide sequence of this compound is derived from the natural sequence of wheat gliadin (UniProtKB Pl 8573, GDA9 WHEAT AA# 82-90).
[0013] FIG. 2 shows a schematic of cleavage of the fluorogenic peptide substrate.
[0014] FIG. 3 shows a schematic of ancestral sequence reconstruction (ASR).
[0001] FIG. 4 shows a method of preparing a bacterial neprosin-like tree, and an exemplary tree highlighting node 21.
[0016] FIG. 5 shows more detail of node 21, the last common ancestor betw een MaNEP, EaNEP and KaNEP.
[0017] FIG. 6 shows BL21(DE3) cells expressing N21 vs an inactive protein.
[0018] FIG. 7 is a schematic of AncNEP N21FL.
[0019] FIG. 8 is an Alphafold prediction of the structure of N21.
[0020] FIG. 9 shows that purified N21 is active at pH 2, not pH 4.
[0021] FIG. 10 shows a purified N21 band shift corresponds to activity delay at pH 2.
[0022] FIG. 11 shows a 96 well plate screening workflow for further evolving N21.
[0023] FIG. 12 shows a plate activity7heatmap highlighting wild type (WT) and a dead mutant.
[0024] FIG. 13 shows the N21 domain boundaries as assessed by LC / MS (data not shown).
[0025] FIG. 14 shows a comparison of the substrate cleavage activity' of N21, H3 and 2C11.
[0026] FIG. 15 shows the pH-dependence of H3 activity.
[0027] FIG. 16 show's a comparison of the pH profiles of activity for H3 and 2C11 .
[0028] FIG. 17 compares KM curves for N21, H3 and 2C11.
[0029] FIG. 18 shows H3 activation and stability'.
[0030] FIG. 19 shows the stability of H3 compared to TAK-062.
[0031] The above-described and other features will be appreciated and understood by those skilled in the art from the following detailed description, drawings, and appended claims.DETAILED DESCRIPTION
[0032] Described herein are novel specific proteases which degrade the immunogenic peptides from gluten, specifically enzymes which are active in the stomach at pH 2-4. While degradation with naturally-occurring proteases has been attempted, none so far have provided sufficient degradation. Gluten proteins resist degradation because they are generally- disordered and aggregate easily; they generally have low water solubility; they are often crosslinked with other gluten proteins; they have a high prohne / glutamine content which makes them pepsin-resistant; and they lack cleavage motifs for trypsin.
[0033] In order to design novel specific proteases that degrade the immunogenic peptides from gluten, gluten epitopes were identified. Epitopes relevant in celiac disease include: gliadin / glutenin digestion intermediates and Pro / Gln / Glu rich peptides that trigger T- cell proliferation. A ‘'33-mer” immunogenic peptide that resists digestion has the following sequence: LQLQPF-PQPQLPY-PQPQLPY-PQPQLPY-PQPQPF (SEQ ID NO: 1), which has repeating motifs of PQP / PQQP / PQLP / PYP / PFP (SEQ ID NO: 2). Another relevant sequence is PEQP / PELP (SEQ ID NO: 3).
[0034] As shown in FIG. 1, an approximately 2 kDa Anorogenic peptide substrate of SEQ ID NO: 4 (PQPQLPYPQK) based on a gluten epitope of SEQ ID NO: 5 (PQPQLPYPQ) that contains a Auorophore on the N-terminus and a quencher on the N-terminus was designed. The excitation wavelength for this peptide is at 485±20nm (5-FAM or HiLyte™ Fluor 488) and the emission is at 520-530nm. Emission is quenched by FRET (QXL®-520) prior to substrate cleavage. Emission increases after reaction (cleavage) and quencher release allowing measurement of an increase of Auorescence signal as an indicator of cleavage. FIG. 2 illustrates a schematic of cleavage of the Auorogenic peptide substrate.
[0035] Using this substrate, the inventors took the approach of using ancestral proteases (e g., extinct proteases that do not exist in nature today) which they calculated using ancestral sequence reconstruction as a starting point for directed evolution. It was hypothesized that those reconstructed ancestral proteins are more evolvable than modem proteins since they were back-calculated from all modem sequences that survived evolution, hence a selection for the most evolvable sequences. Consequently, the expectation was that ancestral proteins would provide a superior starting place for directed evolution. In addition, ancestral proteins have been found to have higher expression profiles, be more soluble and more stable than their modem counterparts. As described herein, a modem plant protease, neprosin, was used as a starting point to curate a bacterial sequence data set for ancestral sequence reconstruction (ASR). FIG. 3 is a schematic of ASR. In ASR, using relatedmodem sequences, an “ancestral” gene / protein is reconstructed using multiple sequence alignment (MSA) and the corresponding phylogenetic tree. FIG. 4 shows a method of preparing a neprosin-like tree, and an exemplary tree highlighting node 21. The approach is a probabilistic approach in which calculates the most probable sequence of a common ancestor between modem nodes. FIG. 5 shows more detail of node 21, the last common ancestor between MaNEP, EaNEP and KaNEP.
[0036] Using the approach described in the examples. Ancestral NEProsin Node 21 SMP(Single Most Probable sequence) Full Length (AncNEP N21FL), or simply N21, was identified. SEQ ID NO: 6 is N21, including the propeptide and sequence tag. SEQ ID NO: 7 is N21, including the propeptide. SEQ ID NO: 8 is the active N21 protease after autoactivation.AncNEP N21FL plus TEV-site and 6His tag (SEQ ID NO: 6) (M)ARAPKKLTPFSEFIESVKAAKHEEFKARPGAKVKDAEEFEEMRQHLLNLYEGVEVQHS FVDEDGQIFDCIPIEQQPSLRGSGAKVIATPPDLPPAAGASAKEAEESAKAVQPPLSPDRT DRFGNAMSCPDGTIPMRRVTLEELARFETLEDFFRKGPNGAGKRPP^REA. WWANWY HKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGGSGAGLQTVECGW QVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTS GGAQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQLATNATEIDYGGE TVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTPSQPSPSCYTIDVT NASASWNEYFFFGGPGGSNCENLYFQGSHHHHHH* (Propeptide, Fusion tag)AncNEP N21FL (SEQ ID NO: 7)(M)ARAPKKLTPFSEFIESVKAAKHEEFKARPGAKVKDAEEFEEMRQHLLNLYEGVEVQHS FVDEDGQIFDCIPIEQQPSLRGSGAKVIATPPDLPPAAGASAKEAEESAKAVQPPLSPDRT DRFGNAMSCPDGTIPMRRVTLEELARFETLEDFFRKGPNGAGKRPP^E S PY NAAY HKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGGSGAGLQTVECGW QVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTS GGAQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQLATNATEIDYGGE TVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTPSQPSPSCYTIDVT NASASWNEYFFFGGPGGSNCN21CD (SEQ ID NO: 8) REASAAPPAVAATHKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVG GSGAGLQTVECGWQVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSS WALGGALSPVSTSGGAQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQ LATNATEIDYGGETVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLT PSQPSPSCYTIDVTNASASWNEYFFFGGPGGSNC
[0037] As described in Example 2, N21 was then used as a starting point for further protease evolution. As shown in FIG. 11, a protocol was developed to further evolve the identified N21 sequence. Variants H3 (SEQ ID NOs. 9-11) and 2C11 (SEQ ID NOs. 12-14) were identified. H3 (SEQ ID NO: 9) has mutations F40^ C, V61^A, A171 — >E, A291^T. SEQ ID NO: 10 is H3, including the propeptide. SEQ ID NO: 11 is the active H3 protease.H3 FL plus TEV-site and 6His tag (SEQ ID NO: 9) (M)ARAPKKLTPFSEFIESVKAAKHEEFKARPGAKVKDAEECEEMRQHLLNLYEGVEVQHS FADEDGQIFDCIPIEQQPSLRGSGAKVIA TPPDLPPAA GASAKEAEESAKA VQPPLSPDRT DRFGNAMSCPDGTIPMRRVTLEELARFETLEDFFRKGPNGAGKRPP^EA EA YANAAYHKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGGSGAGLQTVECGW QVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTS GGTQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQLATNATEIDYGGE TVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTPSQPSPSCYTIDVT NASASWNEYFFFGGPGGSNCENLYFQGSHHHHHH*H3 FL (SEQ ID NO: 10) (M)ARAPKKLTPFSEFIESVKAAKHEEFKARPGAKVKDAEECEEMRQHLLNLYEGVEVQHS FADEDGQIFDCIPIEQQPSLRGSGAKVIATPPDLPPAAGASAKEAEESAKAVQPPLSPDRT DRFGNAMSCPDGTIPMRRVTLEELARFETLEDFFRKGPNGAGKRPP^EASEA YANAAY HKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGGSGAGLQTVECGW QVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTS GGTQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQLATNATEIDYGGE TVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTPSQPSPSCYTIDVT NASASWNEYFFFGGPGGSNCH3CD (SEQ ID NO: 11)REASEAPPAVAATHKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGG SGAGLQTVECGWQVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSW ALGGALSPVSTSGGTQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQL ATNATEIDYGGETVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTP SQPSPSCYTIDVTNASASWNEYFFFGGPGGSNC
[0038] 2C11 (SEQ ID NO: 12) has mutations F40^C, V61^A, AVA^E, A291^T; A 145^T, F153—>L. G290^A. S395^C. SEQ ID NO: 13 is 2C11, including the propeptide. SEQ ID NO: 14 is the active 2C11 protease.2C11 FL plus TEV-site and 6His tag (SEQ ID NO: 12) (M)ARAPKKLTPFSEFIESVKAAKHEEFKARPGAKVKDAEECEEMRQHLLNLYEGVEVQHS FADEDGQIFDCIPIEQQPSLRGSGAKVIATPPDLPPAAGASAKEAEESAKAVQPPLSPDRT DRFGNAMSCPDGTIPMRR VTLEEL TRFETLEDLFRKGPNGAGKRPP EAS LN YNN AATHKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGGSGAGLQTVECGW QVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTS GATQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQLATNATEIDYGGE TVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTPSQPSPCCYTIDVT NASASWNEYFFFGGPGGSNCENLYFQGSHHHHHH*2C11 FL (SEQ ID NO: 13)(M)ARAPKKLTPFSEFIESVKAAKHEEFKARPGAKVKDAEECEEMRQHLLNLYEGVEVQHS FADEDGQIFDCIPIEQQPSLRGSGAKVIATPPDLPPAAGASAKEAEESAKAVQPPLSPDRT DRFGNAMSCPDGTIPMRR VTLEEL TRFETLEDLFRKGPNGAGKRPP EN^NPYNN AATHKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGGSGAGLQTVECGW QVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTS GATQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQLATNATEIDYGGE TVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTPSQPSPCCYTIDVT NASASWNEYFFFGGPGGSNC2C11CD SEQ ID NO: 14REASEAPPAVAATHKYAHAYQNVDNLGGHSFLNVWNPAVGANQIFSLSQHWYVGG SGAGLQTVECGWQVYPGKYGNNKPVLFIYWTADNYNKTGCYNLDCSAFVQTNSSWALGGALSPVSTSGATQYEIELAYYLSGGNWWLYLNGTSASDAIGYYPATLFGGGQL ATNATEIDYGGETVGTTSWPPMGSGAFPSEGYRHAAYQRDIYYYPPSGGSQSASLTP SQPSPCCYTIDVTNASASWNEYFFFGGPGGSNC
[0039] Also included herein are polypeptides with greater than 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the N21, H3 and 2C11 proteases described herein (SEQ ID NOs. 6-14). As used herein, the terms ‘"identical” or percent sequence ""identity” in the context of two or more proteins, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues that are the same, when compared and aligned (introducing gaps, if necessary') for maximum correspondence. The percent identity can be measured using sequence comparison software or algorithms or by visual inspection. Various algorithms and software are known in the art that can be used to obtain alignments of amino acid sequences.
[0040] The percent sequence identity “X” of a first amino acid sequence to a second sequence amino acid is calculated as 100 times (Y / Z), where Y is the number of amino acid residues scored as identical matches in the alignment of the first and second sequences (as aligned by visual inspection or a particular sequence alignment program) and Z is the total number of residues in the second sequence. If the length of a first sequence is longer than the second sequence, the percent identity of the first sequence to the second sequence will be higher than the percent identity’ of the second sequence to the first sequence.
[0041] In an aspect, a sequence with a specified percentage of sequence identity includes conservative amino acid substitutions.
[0042] A “conservative amino acid substitution” is one in yvhich one amino acid residue is replaced with another amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, including basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e g., threonine, valine, isoleucine) and aromatic side chains (e.g., ty rosine, phenylalanine, tryptophan, histidine). For example, substitution of a phenylalanine for a tyrosine is a conservative substitution. In an aspect, the transmembrane domain comprises only conservative amino acid substitutions.
[0043] In an aspect, the polypeptides yvith greater than 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the N21, H3 and 2C11 proteases in their active form after cleavage of the pro-peptide are active to cleave the 33-mer peptide of SEQ ID NO: 17 at a pH of 2-4. Inan aspect, a solution of 1 pM N21, H3 or 2C11 protease or variant thereof and 1 mg / mL SEQ ID NO: 17 peptide, when incubated at a pH of between 2.5 at 37°C for 24 hours, provides cleavage of 80% or more of a peptide of SEQ ID NO: 17 at QP|QL to provide peptides of SEQ ID NOs. 18-24.
[0044] Also included are polynucleotides encoding N21, H3 and 2C11 or a fragment thereof, such as a proteolytically active fragment thereof, expression vectors comprising the polynucleotides and cells comprising the expression vectors or polynucleotides.
[0045] As used herein, “isolated nucleic acids” are those that have been removed from their normal surrounding nucleic acid sequences in the genome or in cDNA sequences. Such isolated nucleic acid sequences may comprise additional sequences useful for promoting expression and / or purification of the encoded protein, including but not limited to polyA sequences, modified Kozak sequences, and sequences encoding epitope tags, export signals, and secretory signals, nuclear localization signals, and plasma membrane localization signals. It will be apparent to those of skill in the art, based on the teachings herein, what nucleic acid sequences will encode the polypeptides of the invention.
[0046] In a further aspect, provided are nucleic acid expression vectors comprising the isolated nucleic acid operatively linked to a suitable control sequence. “Recombinant expression vector” includes vectors that operatively link a nucleic acid coding region or gene to any control sequences capable of effecting expression of the gene product. “Control sequences” operably linked to the nucleic acid sequences of the invention are nucleic acid sequences capable of effecting the expression of the nucleic acid molecules. The control sequences need not be contiguous with the nucleic acid sequences, so long as they function to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between a promoter sequence and the nucleic acid sequences and the promoter sequence can still be considered “operable linked” to the coding sequence. Other such control sequences include, but are not limited to, polyadenylation signals, termination signals, and ribosome binding sites. Such expression vectors can be of any type known in the art, including but not limited plasmid and viral-based expression vectors. The control sequence used to drive expression of the disclosed nucleic acid sequences in a mammalian system may be constitutive (driven by any of a variety of promoters, including but not limited to, CMV, SV40, RSV, actin, EF) or inducible (driven by any of a number of inducible promoters including, but not limited to, tetracycline, ecdysone, steroid-responsive). The construction of expression vectors for use in transfecting prokaryotic cells is also well known in the art, and thus can be accomplished via standard techniques.
[0047] Recombinant host cells comprising the nucleic acid expression vectors are also included. The host cells can be either prokaryotic or eukaryotic. The cells can be transiently or stably transfected or transduced. Such transfection and transduction of expression vectors into prokaryotic and eukaryotic cells can be accomplished via any technique known in the art, including but not limited to standard bacterial transformations, calcium phosphate coprecipitation, electroporation, or liposome mediated-, DEAE dextran mediated-, poly cationic mediated-, or viral mediated transfection.
[0048] A method of producing a polypeptide according to the disclosure is an additional part of the disclosure. The method comprises the steps of (a) culturing a host cell cells comprising the nucleic acid expression vectors under conditions conducive to the expression of the polypeptide, and (b) optionally, recovering the expressed polypeptide. The expressed polypeptide can be recovered from the cell free extract, cell pellet, or recovered from the culture medium. Methods to purify recombinantly expressed polypeptides are well known to the man skilled in the art.
[0049] The N21, H3 and 2C 11 proteases and variants thereof can be administered alone or formulated into a composition with an excipient, such as a pharmaceutical composition, a food supplement, or a food product.
[0050] The pharmaceutical composition may comprise in addition to the polypeptides, nucleic acids, etc. of the disclosure (a) a lyoprotectant; (b) a surfactant: (c) a bulking agent; (d) a tonicity adjusting agent; (e) a stabilizer; (f) a preservative and / or (g) a buffer.
[0051] In some embodiments, the buffer in the pharmaceutical composition is a Tris buffer, a histidine buffer, a phosphate buffer, a citrate buffer or an acetate buffer. The pharmaceutical composition may also include a lyoprotectant, e.g. sucrose, sorbitol or trehalose. In certain embodiments, the pharmaceutical composition includes a preservative e.g., benzalkonium chloride, benzethonium, chlorohexidine, phenol, m-cresol, benzyl alcohol, methylparaben, propylparaben, chlorobutanol, o-cresol, p-cresol, chlorocresol, phenylmercurie nitrate, thimerosal, benzoic acid, and various mixtures thereof. In other embodiments, the pharmaceutical composition includes a bulking agent, like glycine. In yet other embodiments, the pharmaceutical composition includes a surfactant e.g., polysorbate- 20, polysorbate-40, polysorbate-60, polysorbate-65, polysorbate-80, polysorbate-85, poloxamer-188, sorbitan monolaurate, sorbitan monopalmitate, sorbitan monostearate, sorbitan monooleate, sorbitan trilaurate, sorbitan tristcarate, sorbitan trioleaste, or a combination thereof. The pharmaceutical composition may also include a tonicity adjustingagent, e.g., a compound that renders the formulation substantially isotonic or isoosmotic with human blood. Exemplary tonicity adjusting agents include sucrose, sorbitol, glycine, methionine, mannitol, dextrose, inositol, sodium chloride, arginine and arginine hydrochloride. In other embodiments, the pharmaceutical composition additionally includes a stabilizer, e.g., a molecule which, when combined with a protein of interest substantially prevents or reduces chemical and / or physical instability of the protein of interest in lyophilized or liquid form. Exemplars’ stabilizers include sucrose, sorbitol, glycine, inositol, sodium chloride, methionine, arginine, and arginine hydrochloride.
[0052] Pharmaceutical compositions can be in in liquid form, for example in the form of a solution, emulsion, or in solid form, such as tablets, capsules, or semisolid. The formulation can be administered in a variety of ways including those particularly suitable for admixing with a foodstuff. The enzyme components can be active prior to or during ingestion, and may be treated, for example, by a suitable encapsulation, to control the timing of activity'.
[0053] For treating celiac disease, the pharmaceutical compositions may be formulated so as to release their activity in the gastric fluid. This ty pe of formulation can provide optimum activity in the right place, for example the release of the protease in the stomach.
[0054] Alternatively, a microorganism, such as a bacterial or yeast culture, capable of producing the active agents can be administered to a patient. Such a culture may be admixed with food preparations or formulated, for example, as a capsule.
[0055] Also included is a food supplement comprising a protease as described herein. The term “food supplement’" is interchangeable with the terms food additive, a dietary supplement and nutritional supplement.
[0056] Also included is a food supplement comprising a protease as described herein. The term “food supplement” is interchangeable with food additive, a dietary supplement, and nutritional supplement. As an example, the food supplement may be a granulated enzy me coated or uncoated product which may readily be mixed with food components, alternatively, food supplements can form a component of a pre-mix. Alternatively, the food supplements may be a stabilized liquid, an aqueous or oil-based slurry.
[0057] The pharmaceutical composition or the food supplement can be provided prior to meals, immediately before meals, with meals or immediately after meals, so that the proteases is released or activated in the upper gastrointestinal lumen where the protease cancomplement gastric and pancreatic enzymes to detoxify ingested gluten and prevent harmful peptides to pass the enterocytes layer.
[0058] Also included are food products comprising the proteases described herein.
[0059] The proteases described herein have numerous applications in the food processing industry', in particular they can be used in the manufacture of food supplements.
[0060] Also described herein is a method for degrading gluten oligopeptides which are resistant to cleavage by gastric and pancreatic enzymes and whose presence in the internal lumen results in toxic effects which comprises contacting said gluten oligopeptides with a protease as described herein.
[0061] In particular, one aspect of said method consists in the treatment or prevention of celiac disease (also known as celiac sprue), non-celiac gluten sensitivity, wheat allergy, gluten ataxia, dermatitis herpetiformis, and / or any other disorder associated with gluten intolerance which comprises administering to a patient in need thereof an effective amount of an enzy me composition or of at least one isolated endopeptidase of this invention, preferably, incorporated into a pharmaceutical formulation, food supplement, drink or beverage.
[0062] In an aspect, method of a treating a disorder associated with gluten intolerance comprises administering to a subject in need thereof an effective amount of the protease or the composition as described herein.
[0063] In another aspect, the use of a protease as described herein for the manufacture of a medicament for treating gluten intolerance in a subject, wherein the method comprises administering to the subject an effective amount of the protease or the composition as described herein.
[0064] In another aspect, included is a protease for use in a method of treating gluten intolerance in a subject, the method comprising administering to the subject an effective amount of the protease or the composition described herein.
[0065] Celiac disease is a highly prevalent disease in which dietary proteins found in wheat, barley, and rye products known as ‘glutens’ evoke an immune response in the small intestine of genetically predisposed individuals. The resulting inflammation can lead to the degradation of the villi of the small intestine, impeding the absorption of nutrients. Symptoms can appear in early childhood or later in life, and range widely in severity, from diarrhea, fatigue, weight loss, abdominal pain, bloating, excessive gas, indigestion, constipation, abdominal distension, nausea / vomiting, anemia, bruising easily, depression, anxiety, growth delay in children, hair loss, dermatitis, missed menstrual periods, mouth ulcers, muscle cramps, joint pain, nosebleeds, seizures, tingling or numbness in hands or feet,delayed puberty', defects in tooth enamel, and neurological symptoms such as ataxia or paresthesia. There are currently no effective therapies for this lifelong disease except the total elimination of glutens from the diet.
[0066] Treating a disorder associated with gluten intolerance can include (a) reducing the severity of disease; (b) limiting or preventing development of symptoms characteristic of disease; (c) inhibiting worsening of symptoms characteristic of disease; (d) limiting or preventing recurrence of disease in patients that have previously had the disorder; (e) limiting or preventing recurrence of symptoms in patients that were previously symptomatic fordisease; and (f) limiting development of disease in a subject at risk of developing disease, or not yet showing the clinical effects of disease.
[0067] As used herein, an "‘amount effective” refers to an amount of the polypeptide that is effective for treating a disorder associated with gluten intolerance.
[0068] Dosage regimens can be adjusted to provide the optimum desired response (e.g., a therapeutic or prophylactic response). A suitable dosage range may, for instance, be 0.1 ug / kg-100 mg / kg body weight; alternatively, it may be 0.5 ug / kg to 50 mg / kg; 1 ug / kg to 25 mg / kg, or 5 ug / kg to 10 mg / kg body weight. The polypeptides can be delivered in a single bolus, or may be administered more than once (e.g., 2, 3, 4, 5, or more times) as determined by an attending physician.
[0069] Also included is a method of testing a subject suspected of having gluten intolerance or gluten sensitivity, comprising administering to the subject a composition comprising an N21, H3 or 2C11 protease or variant thereof as described herein, waiting a period of time, and after the period of time, assessing one or more symptoms of gluten intolerance or gluten sensitivity in the subject, wherein a reduction in the one or more symptoms indicates the subject has gluten intolerance or gluten sensitivity.
[0070] In an aspect, the subject suspected of having gluten intolerance or gluten sensitivity is suffering from gastrointestinal distress (e.g., bloating, gas, diarrhea, constipation, abdominal pain, nausea, vomiting), fatigue, headache, joint pain, muscle pain, difficulty concentrating, depression, anxiety, anemia, unexplained weight loss, skin reactions, and / or injury due to inflammation. In an aspect, the one or more symptoms of gluten intolerance or gluten sensitivity comprises gastrointestinal distress (e.g., bloating, gas, diarrhea, constipation, abdominal pain, nausea, vomiting), fatigue, headache, j oint pain, muscle pain, difficulty concentrating, depression, anxiety, anemia, unexplained weight loss, skin reactions, and / or injury due to inflammation.
[0071] In an aspect, the period of time is 2 days, 1 week or longer, such as 1 week or 2 weeks.
[0072] In an aspect, a method of identifying an ancestral protease comprises identify ing a modem protease sequence, wherein the modem protease is neprosin; querying the modem protease sequence in a sequence database to provide a list of similar sequences,; selecting a subset of the list of similar sequences using a sequence similarity cut-off; aligning the subset of the list of similar sequences using a multiple sequence alignment tool to provide aligned sequences; preparing a phylogenetic tree based on the aligned sequences; selecting a node from the phylogenetic tree and selecting the ancestral proteases from the node, and performing a fluorescence assay with the fluorogenic peptide substrate described herein to verify the activity of the ancestral protease to degrade gluten peptides.
[0073] In an aspect, the sequence database is a database of bacterial kingdom sequences. In another aspect, the sequence similarity cutoff is 90% sequence similarity. In a further aspect, the phylogenetic tree is a Maximum Likelihood tree. In a still further aspect, the verifying is performed on crude cell extracts expressing the ancestral proteases in a multiwell assay.
[0074] In another aspect, the method further comprises further evolving the ancestral protease by mutating one or more amino acids of the ancestral protease, and selecting ancestral protease variants that have increased cleavage of the fluorogenic peptide substrate of claim 14 or claim 15 compared to the ancestral protease.
[0075] The invention is further illustrated by the following non-limiting examples.ExamplesMethods
[0076] Growth and purification protocol for ancestor N21Growth:Transform pET28a-AncNEP90-N21SMP-CO (codon-optimized N21) to BL21(DE3) Grow' 37°C overnight starter culture (LB with 50mg / L kanamycin)Inoculate 1 / 100 to IL LB, grow at 37°C to OD 0.6, induce w / 0.5mM IPTG, express 24°C for 16hPurify:Resuspend pellet in 50mM NazHPCL pH7.5, 300mM NaCl, DNAse I, 2.5mM MgCh Sonicate 15 min @25W on ice, spin down 16000xg 30min at 4°CSoluble fraction 0.2pm filtered and load onto lOmL TALON® Cobalt affinity column Washed w / 10CV buffer (no imidazole), eluted w / 2CV 50mM and 125mM imidazole 50mM elution desalted w / PD-10 column at RT (to 50mM Na2HPO4 pH7.5, 300mM NaCl) Protein concentration determined by 280nm absorbance
[0077] AncNEP N21 Lysate / Purified Activity AssaysLysate Prep: Wash pellet: resuspend in Y original media volume of PBS pH 7.4, spin down 2000g 5min (twice)Normalize to lOx final OD, take ImL of cell resuspension with known ODA) Spin down then resuspend in lx BugBuster®, rotate RT IhrB) Sonicate in lx PBS on ice (~20W, Imin)Activity Assay for either purified protein or cell lysate:10% !M H3PO4pH 0.61.5% 6M NaOH6% 5M NaCl62.5% ddH2O10% lOx Cell Lysate or purified protease10% lOx substrate in 10% DMSOFinal buffer condition: lOOmM NaxHxPO4pH 2.0, 300mM NaCl, 1% DMSOExcitation: 488nm, 9nm slit; emission: 530nm 25nm slit.Reactions incubated at 37C in black flat bottom 96well plates, 50pL reactions0 / 16 / 24 hr endpoints for protease activity measurement directly in cell lysates:Cells lysed either chemically by BugBuster®, or mechanically by sonicationReaction mix composition: lOOmM H3PO4NaOH pH 2.0, 300mMNaCl, 1% DMSO, O.lx BugBuster®0.1 OD equiv. Cell Lysate20pM DNE2 Fluorogenic SubstrateImL reactions, incubated 37C in foil covered Eppendorf® tubes, agitated at 500rpm.Acquisition: SpectraMax® i3xEx: 488nm, 9nm slit; Em: 530nm 25nm slit.37C, in Coming 3993 black flat bottom 96well plates, 50pL samplesNegative control is plasmid with super TEV, a protease that is not active at pH 2Example 1 : Identification and characterization of Ancestral NEProsin Node 21 SMP(Single Most Probable sequence) Full Length (AncNEP N21FL)
[0078] The fluorogenic peptide substrate of FIG. 1 and ASR (FIG. 3) were used to identify neprosin homologues.
[0079] As a first step, homologues of the ancestral protease in the bacterial kingdom were identified. FIG. 4 shows a method of preparing a neprosin-like tree and an exemplary tree.
[0080] Multiple Sequence Alignment (MSA) is generally the alignment of three or more bi ological sequences (protein or nucleic acid) of similar length. From the output, homology can be inferred and the evolutionary relationships between the sequences studied. MSA thus plays an important role in evolutionary analyses of biological sequences. MAFFT(Multiple Alignment using Fast Fourier Transform) is a high-speed multiple sequence alignment program.
[0081] Many of the sequences used have unknown functions. A phylogenetic tree was calculated, then ASR was performed to calculate ancestral sequences. An exemplary phylogenetic tree is a maximum likelihood tree.
[0082] A maximum likelihood phylogenetic tree, based on modem neprosin-like proteins of bacterial origin, was created. 258 modem sequences were identified at a 90% identity cutoff. Node relations (topology) within the phylogenetic tree can be calculated with IQ-TREE2. IQ-TREE2 takes as input a MSA and an evolutionary model and will reconstruct an evolutionary' tree that is best explained by the input data. Nodes represent branching points from the ancestral population. Node 21 is the last common ancestor between MaNEP, EaNEP and KaNEP (FIG. 5). While node 21 was selected in the current analysis, other nodes can also be made for this ASR in the future.
[0083] Within the identified node, high throughput screening of protease variants for the degradation of immunogenic peptides from gluten can be performed. As shown in FIG. 1, a sequence was designed that best mimics the sequence of the gluten peptides to be cleaved. Two fluorophores were added at the end so that if a protease cleaves, it releases the fluorescence quenching by the quencher. Advantageously, the fluorophores shown in FIG. 1 work at pH 2-4.
[0084] Using the selected protease variants and the designed substrate, a fluorescence assays can be performed in crude cell extract such as in 96 well plates. This screening of millions of variants in directed evolution. FIG. 6 shows BL21(DE3) cells expressing N21 vs an inactive protein.
[0085] Of note, Ancestral NEProsin Node 21 SMP (Single Most Probable sequence) Full Length (AncNEP N21FL) does not exist in nature. Characteristics of AncNEP N21FL include:• 47.3kDa, pl = 5.5, 436 amino acids• 2 domain neprosin-like protein with C-terminal scissile affinity tag• No tryptophan in 19kDa pro-peptide domain• Activity measured at pH 2, not observable at pH 4• AF2 predicted neprosin-like structure• Expressed solubly in BL21(DE3) E. coli• Yield: lOmg / L media or 2mg
[0086] FIG. 7 is a schematic of AncNEP N21FL, referred to as N21 herein. The sequence of N21 including the propeptide and fusion tag is SEQ ID NO: 6.
[0087] FIG. 8 is the Alphafold prediction of the structure of N21.
[0088] FIG. 9 shows that purified N21 is active at pH 2, not pH 4. Conditions were: 20 pM HiLyte™ Fluor 488 -SEQ ID NO: 4-(QXL®-520)-NH2; 100 mM sodium citrate, pH 2; IX BugBuster® lysis agent: 1% DMSO; 1 pM purified AncNEP N21FL; 37°C, 2 hr, 1 min intervals; Fluorescence normalized vs. free HiLyte™ Fluor 488. Detection was Spectramax® 13X w / coming 3993 plate; Excitation: Center @ 488 nm, 9 nm slit; Emission: Center 2 530 nm, 25 nm slit. Error bars: std. deviation from triplicates.
[0089] FIG. 10 shows a purified N21 band shift corresponds to activity delay at pH 2. Preliminary kinetics of N21 at pH 2, 100 nM enzyme concentration were performed. In 1% DMSO, the koat was 6.1 ± 0.9 X 10'3 / s, and the Km was 35 ± 16 pM. In 10% DMSO, the kcat was 13.5 ± 0.8 X 10'3 / s, and the Km was 40 ± 6 pM.
[0090] SEQ ID NOs. 15 and 16 are DNA sequences for N21 of SEQ ID NO: 6.DNA sequence forN21 (SEQ ID NO: 15)ATGGCACGTGCGCCGAAAAAGCTCACCCCGTTCAGCGAATTTATTGAATCAGTGA AGGCGGCGAAACATGAAGAGTTCAAGGCACGTCCAGGCGCAAAAGTGAAAGAT GCAGAAGAGTTTGAAGAAATGCGTCAGCATTTATTGAATCTGTACGAAGGCGTC GAAGTTCAACATAGCTTTGTTGACGAAGATGGCCAGATTTTCGATTGCATTCCGA TTGAACAGCAGCCTTCGCTGCGCGGTAGCGGTGCCAAAGTTATCGCTACCCCGCC GGATCTGCCTCCGGCTGCGGGCGCTTCGGCGAAAGAAGCGGAGGAGAGCGCCAA AGCGGTCCAACCGCCGTTATCTCCGGACCGCACCGATCGTTTTGGCAATGCGATG TCATGTCCGGATGGTACCATCCCTATGCGTCGCGTAACCCTGGAAGAACTCGCCC GCTTCGAGACCTTGGAGGACTTTTTCCGCAAGGGTCCGAACGGCGCGGGTAAAC GTCCTCCACGTGAAGCATCTGCGGCACCGCCAGCGGTTGCCGCGACCCATAAATA TGCCCACGCATATCAGAACGTGGATAATTTAGGTGGTCACAGCTTTCTTAACGTT TGGAATCCAGCTGTGGGTGCGAATCAAATTTTCAGCTTGAGTCAGCACTGGTACG TGGGCGGCTCGGGTGCGGGTCTGCAGACCGTGGAATGTGGCTGGCAGGTATACC CTGGCAAATACGGTAATAACAAACCGGTACTGTTCATCTATTGGACCGCCGATAA CTATAATAAAACCGGTTGTTATAACCTGGACTGCAGTGCCTTTGTCCAGACCAAT TCCAGCTGGGCCCTTGGCGGCGCGCTTTCGCCGGTCTCAACCAGTGGTGGTGCAC AGTATGAGATTGAACTGGCGTATTATCTGAGCGGTGGCAACTGGTGGCTGTACCT GAACGGTACCAGCGCCAGTGATGCCATTGGTTACTACCCTGCTACCCTGTTTGGC GGCGGCCAGTTAGCCACCAACGCGACCGAAATCGATTATGGTGGCGAGACCGTG GGTACCACCTCGTGGCCGCCGATGGGCTCCGGTGCCTTTCCGAGCGAAGGTTACC GTCACGCGGCATACCAGCGCGACATCTATTACTATCCGCCGTCCGGTGGTTCTCA ATCTGCCTCACTCACCCCGAGCCAACCGTCACCAAGCTGCTACACCATCGACGTG ACCAATGCCTCGGCAAGTTGGAACGAATATTTCTTTTTTGGCGGTCCGGGCGGCT CCAACTGCGAAAACCTGTATTTCCAGGGCAGTCACCATCACCATCACCATTAAFull plasmid sequence (pET28 (+) N21): (SEQ ID NO: 16)CACGTGCGCCGAAAAAGCTCACCCCGTTCAGCGAATTTATTGAATCAGTGAAGGCGGCGAAACATGAAGAGTTCAAGGCACGTCCAGGCGCAAAAGTGAAAGATGCAGAAGAGTTTGAAGAAATGCGTCAGCATTTATTGAATCTGTACGAAGGCGTCGAAGTTCAACATAGCTTTGTTGACGAAGATGGCCAGATTTTCGATTGCATTCCGATTGAACAGCAGCCTTCGCTGCGCGGTAGCGGTGCCAAAGTTATCGCTACCCCGCCGGATCTGCCTCCGGCTGCGGGCGCTTCGGCGAAAGAAGCGGAGGAGAGCGCCAAAGCGGTCCAACCGCCGTTATCTCCGGACCGCACCGATCGTTTTGGCAATGCGATGTCATGTCCGGATGGTACCATCCCTATGCGTCGCGTAACCCTGGAAGAACTCGCCCGCTTCGAGACCTTGGAGGACTTTTTCCGCAAGGGTCCGAACGGCGCGGGTAAACGTCCTCCACGTGAAGCATCTGCGGCACCGCCAGCGGTTGCCGCGACCCATAAATATGCCCACGCATATCAGAACGTGGATAATTTAGGTGGTCACAGCTTTCTTAACGTTTGGAATCCAGCTGTGGGTGCGAATCAAATTTTCAGCTTGAGTCAGCACTGGTACGTGGGCGGCTCGGGTGCGGGTCTGCAGACCGTGGAATGTGGCTGGCAGGTATACCCTGGCAAATACGGTAATAACAAACCGGTACTGTTCATCTATTGGACCGCCGATAACTATAATAAAACCGGTTGTTATAACCTGGACTGCAGTGCCTTTGTCCAGACCAATTCCAGCTGGGCCCTTGGCGGCGCGCTTTCGCCGGTCTCAACCAGTGGTGGTGCACAGTATGAGATTGAACTGGCGTATTATCTGAGCGGTGGCAACTGGTGGCTGTACCTGAACGGTACCAGCGCCAGTGATGCCATTGGTTACTACCCTGCTACCCTGTTTGGCGGCGGCCAGTTAGCCACCAACGCGACCGAAATCGATTATGGTGGCGAGACCGTGGGTACCACCTCGTGGCCGCCGATGGGCTCCGGTGCCTTTCCGAGCGAAGGTTACCGTCACGCGGCATACCAGCGCGACATCTATTACTATCCGCCGTCCGGTGGTTCTCAATCTGCCTCACTCACCCCGAGCCAACCGTCACCAAGCTGCTACACCATCGACGTGACCAATGCCTCGGCAAGTTGGAACGAATATTTCTTTTTTGGCGGTCCGGGCGGCTCCAACTGCGAAAACCTGTATTTCCAGGGCAGTCACCATCACCATCACCATTAACTCGAGCACCACCACCACCACCACTGAGATCCGGCTGCTAACAAAGCCCGAAAGGAAGCTGAGTTGGCTGCTGCCACCGCTGAGCAATAACTAGCATAACCCCTTGGGGCCTCTAAACGGGTCTTGAGGGGTTTTTTGCTGAAAGGAGGAACTATATCCGGATTGGCGAATGGGACGCGCCCTGTAGCGGCGCATTAAGCGCGGCGGGTGTGGTGGTTACGCGCAGCGTGACCGCTACACTTGCCAGCGCCCTAGCGCCCGCTCCTTTCGCTTTCTTCCCTTCCTTTCTCGCCACGTTCGCCGGCTTTCCCCGTCAAGCTCTAAATCGGGGGCTCCCTTTAGGGTTCCGATTTAGTGCTTTACGGCACCTCGACCCCAAAAAACTTGATTAGGGTGATGGTTCACGTAGTGGGCCATCGCCCTGATAGACGGTTTTTCGCCCTTTGACGTTGGAGTCCACGTTCTTTAATAGTGGACTCTTGTTCCAAACTGGAACAACACTCAACCCTATCTCGGTCTATTCTTTTGATTTATAAGGGATTTTGCCGATTTCGGCCTATTGGTTAAAAAATGAGCTGATTTAACAAAAATTTAACGCGAATTTTAACAAAATATTAACGCTTACAATTTAGGTGGCACTTTTCGGGGAAATGTGCGCGGAACCCCTATTTGTTTATTTTTCTAAATACATTCAAATATGTATCCGCTCATGAATTAATTCTTAGAAAAACTCATCGAGCATCAAATGAAACTGCAATTTATTCATATCAGGATTATCAATACCATATTTTTGAAAAAGCCGTTTCTGTAATGAAGGAGAAAACTCACCGAGGCAGTTCCATAGGATGGCAAGATCCTGGTATCGGTCTGCGATTCCGACTCGTCCAACATCAATACAACCTATTAATTTCCCCTCGTCAAAAATAAGGTTATCAAGTGAGAAATCACCATGAGTGACGACTGAATCCGGTGAGAATGGCAAAAGTTTATGCATTTCTTTCCAGACTTGTTCAACAGGCCAGCCATTACGCTCGTCATCAAAATCACTCGCATCAACCAAACCGTTATTCATTCGTGATTGCGCCTGAGCGAGACGAAATACGCGATCGCTGTTAAAAGGACAATTACAAACAGGAATCGAATGCAACCGGCGCAGGAACACTGCCAGCGCATCAACAATATTTTCACCTGAATCAGGATATTCTTCTAATACCTGGAATGCTGTTTTCCCGGGGATCGCAGTGGTGAGTAACCATGCATCATCAGGAGTACGGATAAAATGCTTGATGGTCGGAAGAGGCATAAATTCCGTCAGCCAGTTTAGTCTGACCATCTCATCTGTAACATCATTGGCAACGCTACCTTTGCCATGTTTCAGAAACAACTCTGGCGCATCGGGCTTCCCATACAATCGATAGATTGTCGCACCTGATTGCCCGACATTATCGCGAGCCCATTTATACCCATATAAATCAGCATCCATGTTGGAATTTAATCGCGGCCTAGAGCAAGACGTTTCCCGTTGAATATGGCTCATAACACCCCTTGTATTACTGTTTATGTAAGCAGACAGTTTTATTGTTCATGACCAAAATCCCTTAACGTGAGTTTTCGTTCCACTGAGCGTCAGACCCCGTAGAAAAGATCAAAGGATCTTCTTGAGATCCTTTTTTTCTGCGCGTAATCTGCTGCTTGCAAACAAAAAAACCACCGCTACCAGCGGTGGTTTGTTTGCCGGATCAAGAGCTACCAACTCTTTTTCCGAAGGTAACTGGCTTCAGCAGAGCGCAGATACCAAATACTGTCCTTCTAGTGTAGCCGTAGTTAGGCCACCACTTCAAGAACTCTGTAGCACCGCCTACATACCTCGCTCTGCTAATCCTGTTACCAGTGGCTGCTGCCAGTGGCGATAAGTCGTGTCTTACCGGGTTGGACTCAAGACGATAGTTACCGGATAAGGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGGAGCGAACGACCTACACCGAACTGAGATACCTACAGCGTGAGCTATGAGAAAGCGCCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCGGAACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCGGGTTTCGCCACCTCTGACTTGAGCGTCGATTTTTGTGATGCTCGTCAGGGGGGCGGAGCCTATGGAAAAACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTTCTTTCCTGCGTTATCCCCTGATTCTGTGGATAACCGTATTACCGCCTTTGAGTGAGCTGATACCGCTCGCCGCAGCCGAACGACCGAGCGCAGCGAGTCAGTGAGCGAGGAAGCGGAAGAGCGCCTGATGCGGTATTTTCTCCTTACGCATCTGTGCGGTATTTCACACCGCAATGGTGCACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAAGCCAGTATACACTCCGCTATCGCTACGTGACTGGGTCATGGCTGCGCCCCGACACCCGCCAACACCCGCTGACGCGCCCTGACGGGCTTGTCTGCTCCCGGCATCCGCTTACAGACAAGCTGTGACCGTCTCCGGGAGCTGCATGTGTCAGAGGTTTTCACCGTCATCACCGAAACGCGCGAGGCAGCTGCGGTAAAGCTCATCAGCGTGGTCGTGAAGCGATTCACAGATGTCTGCCTGTTCATCCGCGTCCAGCTCGTTGAGTTTCTCCAGAAGCGTTAATGTCTGGCTTCTGATAAAGCGGGCCATGTTAAGGGCGGTTTTTTCCTGTTTGGTCACTGATGCCTCCGTGTAAGGGGGATTTCTGTTCATGGGGGTAATGATACCGATGAAACGAGAGAGGATGCTCACGATACGGGTTACTGATGATGAACATGCCCGGTTACTGGAACGTTGTGAGGGTAAACAACTGGCGGTATGGATGCGGCGGGACCAGAGAAAAATCACTCAGGGTCAATGCCAGCGCTTCGTTAATACAGATGTAGGTGTTCCACAGGGTAGCCAGCAGCATCCTGCGATGCAGATCCGGAACATAATGGTGCAGGGCGCTGACTTCCGCGTTTCCAGACTTTACGAAACACGGAAACCGAAGACCATTCATGTTGTTGCTCAGGTCGCAGACGTTTTGCAGCAGCAGTCGCTTCACGTTCGCTCGCGTATCGGTGATTCATTCTGCTAACCAGTAAGGCAACCCCGCCAGCCTAGCCGGGTCCTCAACGACAGGAGCACGATCATGCGCACCCGTGGGGCCGCCATGCCGGCGATAATGGCCTGCTTCTCGCCGAAACGTTTGGTGGCGGGACCAGTGACGAAGGCTTGAGCGAGGGCGTGCAAGATTCCGAATACCGCAAGCGACAGGCCGATCATCGTCGCGCTCCAGCGAAAGCGGTCCTCGCCGAAAATGACCCAGAGCGCTGCCGGCACCTGTCCTACGAGTTGCATGATAAAGAAGACAGTCATAAGTGCGGCGACGATAGTCATGCCCCGCGCCCACCGGAAGGAGCTGACTGGGTTGAAGGCTCTCAAGGGCATCGGTCGAGATCCCGGTGCCTAATGAGTGAGCTAACTTACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCCAGGGTGGTTTTTCTTTTCACCAGTGAGACGGGCAACAGCTGATTGCCCTTCACCGCCTGGCCCTGAGAGAGTTGCAGCAAGCGGTCCACGCTGGTTTGCCCCAGCAGGCGAAAATCCTGTTTGATGGTGGTTAACGGCGGGATATAACATGAGCTGTCTTCGGTATCGTCGTATCCCACTACCGAGATATCCGCACCAACGCGCAGCCCGGACTCGGTAATGGCGCGCATTGCGCCCAGCGCCATCTGATCGTTGGCAACCAGCATCGCAGTGGGAACGATGCCCTCATTCAGCATTTGCATGGTTTGTTGAAAACCGGACATGGCACTCCAGTCGCCTTCCCGTTCCGCTATCGGCTGAATTTGATTGCGAGTGAGATATTTATGCCAGCCAGCCAGACGCA GACGCGCCGAGACAGAACTTAATGGGCCCGCTAACAGCGCGATTTGCTGGTGAC CCAATGCGACCAGATGCTCCACGCCCAGTCGCGTACCGTCTTCATGGGAGAAAAT AATACTGTTGATGGGTGTCTGGTCAGAGACATCAAGAAATAACGCCGGAACATTAGTGCAGGCAGCTTCCACAGCAATGGCATCCTGGTCATCCAGCGGATAGTTAATG ATCAGCCCACTGACGCGTTGCGCGAGAAGATTGTGCACCGCCGCTTTACAGGCTT CGACGCCGCTTCGTTCTACCATCGACACCACCACGCTGGCACCCAGTTGATCGGC GCGAGATTTAATCGCCGCGACAATTTGCGACGGCGCGTGCAGGGCCAGACTGGAGGTGGCAACGCCAATCAGCAACGACTGTTTGCCCGCCAGTTGTTGTGCCACGCGG TTGGGAATGTAATTCAGCTCCGCCATCGCCGCTTCCACTTTTTCCCGCGTTTTCGC AGAAACGTGGCTGGCCTGGTTCACCACGCGGGAAACGGTCTGATAAGAGACACC GGCATACTCTGCGACATCGTATAACGTTACTGGTTTCACATTCACCACCCTGAATTGACTCTCTTCCGGGCGCTATCATGCCATACCGCGAAAGGTTTTGCGCCATTCGA TGGTGTCCGGGATCTCGACGCTCTCCCTTATGCGACTCCTGCATTAGGAAGCAGC CCAGTAGTAGGTTGAGGCCGTTGAGCACCGCCGCCGCAAGGAATGGTGCATGCAAGGAGATGGCGCCCAACAGTCCCCCGGCCACGGGGCCTGCCACCATACCCACGC CGAAACAAGCGCTCATGAGCCCGAAGTGGCGAGCCCGATCTTCCCCATCGGTGA TGTCGGCGATATAGGCGCCAGCAACCGCACCTGTGGCGCCGGTGATGCCGGCCA CG ATGC GTCC GGC GTAGAGGATC GAGATCTC GATC CC GC GAA ATT AAT AC GAC TCACTATAGGGGAATTGTGAGCGGATAACAATTCCCCTCTAGAAATAATTTTGTTT AACTTTAAGAAGGAGATATACCATGGExample 2: Further evolution of N21
[0091] As shown in FIG. 11, a protocol was developed to further evolve the identifiedN21 sequence. Growth / prep / assay protocols were:Starter: Column6= WT; Column7= Catalytically dead mutant (See FIG. 12, plate activity heat map) Inoculate plate colonies in 200 pL LB-Kan50 in 96wp flat bottom; Shake at 450 RPM, 37°C O / N in tabletop shakerExpression: Inoculate 2 pL into 200 pL LB-Kan50 in Nunc™ Edge™ 96w flat bottom plate (Thermo); Shake at 450 RPM, 37°C for 3 hours in tabletop shaker; Induce w / 0.5 mM IPTG final (1:200 100 mM IPTG stock); Shake at 450 RPM, 24°C for additional 16 hours in tabletop shaker; Transfer 200 pL to 96WP round bottom, Centrifuge 1000g, 10 min, room temperature (RT); Store pellet at -80C.Preparation: Add 10 pL lx BugBuster® per 200 pL pellet, resuspend; Shake at 450 RPM, RT (25°C) for 1 hour in tabletop 96 wp shaker; Transfer 5 pL to assay plate.Reaction mix composition: 100 mM H3PO4 NaOH pH 2.0, 300 mM NaCl; 10% cell; Lysate from BugBuster® prep (O.lx final BugBuster®); 20 pM DNE2 Fluorogenic Substrate (FIG. 1), 1% DMSO.Acquisition: SpectraMax® i3x; Excitation: 488nm, 9nm slit; Emission: 530nm 25nm slit; 37°C 4 hours, in Coming 3993 black flat bottom 96 well plates, 50uL final volume.
[0092] FIG. 13 shows the N21 domain boundaries as assessed by LC / MS (data not shown). This determines the domain boundaries of SEQ ID NOs: 6, 7, 9, 10, 12, and 13; N- terminus of SEQ ID NOs 8, 11, and 14. This domain boundary is important, only after autoactivation (i.e., cleavage of the propeptide) is the enzyme active!
[0093] Using the method of FIG. 10, the H3 and 2C11 variants were identified.
[0094] H3 (SEQ ID NO: 11) has mutations F40^C. V61^A, AI7I^E, A29 I ^T.
[0095] 2C11 (SEQ ID NO: 10) has mutations F40^C, V61^A, A171->E, A291— >T; A145^T. F153^L, G290^A, S395^C.
[0096] The substrate cleavage activity of N21, H3 and 2C11 were compared using HiLyte™ Fluor 488 -SEQ ID NO: 5-(QXL®-520)-NH2(DNE2). Conditions were as follows:Reaction mix composition: 100 mM H3PO4 NaOH pH 2.0, 300 mM NaCl; 10% purified protease in 50 mM Tris-HCl, 300 mM NaCl, pH7.5; 20 pM DNE2 Fluorogenic Substrate, 1% DMSO; Final pH ~2.0.Acquisition: SpectraMax® i3x; Excitation: 488nm, 9nm slit; Emission: 530nm 25nm slit; 37°C 2 hours <7, 1 min interval, in Coming 3993 black flat bottom 96well plates, 50 pL final volume.
[0097] As shown in FIG. 14, both H3 and Cl 1 had higher cleavage activity compared to N21. At pH 2, N21 had an activity of 5.6 X 10'5± 7.6 X 10'7, H3 had an activity of 1.7 X IO’4± 2.4 X IO'6(3 X N21), and 2C11 had an activity of 4.1 X IO'4± 1 X IO’5(7.3 X N21).
[0098] As shown in FIG. 15, for H3, the peak fluorescence was observed at pH 4-5 (0. 135 / s at 20 pM DNE2 substrate. In the experiment, all datapoints are triplicates measured at 37C using SpectraMax® i3x. Conditions: 100 mM H3PO4 NaOH pH 1.0-7.0, 300 mM NaCl, 1% DMSO, 1 pM mature H3 by A280, 20pM DNE2 Substrate, 50pL volume in Coming 3993 black fluorescence plate. Excitation: 488 nm@9nm slit width, Emission: 530 nm@25nm slit width.
[0099] In FIG. 16, the pH-dependence of the activity of H3 was compared to Cl 1. All datapoints are single measurement at 37°C using a SpectraMax® i3x. Conditions: 100 mM H3PO4 NaOH pH 1.0-7.0 (final), 300 mM NaCl, 1% DMSO, 1 pM mature H3CD / 2C11 activated by A280, 20pM DNE2 Substrate, 50 pL volume in Coming 3993 black fluorescence plate. Excitation:488nm@9nm slit width, Emission: 530nm@25nm slit width;Linear phase (0-1000s) fited for slope. The pH of the activated 2C11 reaction was corrected retrospectively due to weak buffering. The pH profile of H3 and 2C11 show similar trends. Notably, H3 CD has maximal activity at the desired pH 4.
[0100] The KM curves for N21, H3 and 2C11 were compared in FIG. 17. Each substrate concentration is normalized against standard slope to remove the effect of inner filter effect of substrate and product. The data is summarized in Table 1.Table 1 : M curves at pH 2
[0101] The activation and stability of H3 was studied. H3 was activated in 100 mM NaH2PC>4 pH2, 0.3 M NaCl. N21 0429-H3: 5 pM final concentration at 25 mL (~6 mg total protein), incubated at 37°C in 50 mL conical tubes, 4 hours. Activated H3 was dialyzed to 50 mM NazHPCL pH 7.5, 0.3 M NaCl. Dialysis aggregates were spun down 3000g, 15 min, 10°C. A pH 2 SDS-PAGE sample quenched w / 9:30 volume of IM NaOH to pH >=7, added 4X SDS PAGE reducing loading buffer (BME) + 50mM TCEP pH7.5, 20 pL each lane. Gels: 4x12% Bis Tris Acrylamide in MES-SDS buffer, 180V, <120mA, 30min. The data is shown in FIG. 18. H3 survived 4 hours at 37°C, pH 2 without significant degradation. Activation neared completion at the 4 hour timepoint.
[0102] In FIG. 19, the stability of H3 was compared to TAK-062 (Pultz et al., “Gluten Degradation, Pharmacokinetics, Safety, and Tolerabilty of TAK-062, an Engineered Enzyme to treat Celiac Disease”, Gastroenterology, 2021; 161 : 81-93). Conditions were the same as in FIG. 18. H3 survived 4 hours at 37°C, pH 2 without significant degradation. Activation neared completion at the 4 hr timepoint. TAK-062 is greatly degraded after just 1 hr at pH 3 and >90% degraded at 3 hr.
[0103] Mass spectrometry w as used to determine the activity of H3 to digest a 33-mer gliaden peptide of SEQ ID NO: 17. Conditions: 100 mM Glycine pH2.5, IpM H3 / 10pM porcine pepsin by A280, lOmg / mL (-250 pM) 33-mer peptide, 50 pL samples at 37°C, 0 / 1 / 24 hour timepoints. As previously reported, pepsin is unable to further process thepeptide. The presence of pepsin does not affect H3 activity. After 1 hour, H3 starts degrading the 33-mer peptide (data not shown), and after 25 hours 100% of the 33-mer was cleaved (data not shown). H3 exclusively cleaves at QP|QL. H3 cleaves the 33-mer at multiple sites. The cleavage products are SEQ ID NO 18-24 provided in Table 2.Table 2: 33-mer peptide and cleavage fragmentsSEQ ID NO: 25 DNA sequence for H3ATGGCACGTGCGCCGAAAAAGCTCACCCCGTTCAGCGAATTTATTGAATCAGTGAAGGCGGCGAAACATGAAGAGTTCAAGGCACGTCCAGGCGCAAAAGTGAAAGATGCAGAAGAGTGTGAAGAAATGCGTCAGCATTTATTGAATCTGTACGAAGGCGTCGAAGTTCAACATAGCTTTGCTGACGAAGATGGCCAGATTTTCGATTGCATTCCGATTGAACAGCAGCCTTCGCTGCGCGGTAGCGGTGCCAAAGTTATCGCTACCCCGCCGGATCTGCCTCCGGCTGCGGGCGCTTCGGCGAAAGAAGCGGAGGAGAGCGCCAAAGCGGTCCAACCGCCGTTATCTCCGGACCGCACCGATCGTTTTGGCAATGCGATGTCATGTCCGGATGGTACCATCCCTATGCGTCGCGTAACCCTGGAAGAACTCGCCCGCTTCGAGACCTTGGAGGACTTTTTCCGCAAGGGTCCGAACGGCGCGGGTAAACGTCCTCCACGTGAAGCATCTGAGGCACCGCCAGCGGTTGCCGCGACCCATAAATATGCCCACGCATATCAGAACGTGGATAATTTAGGTGGTCACAGCTTTCTTAACGTTTGGAATCCAGCTGTGGGTGCGAATCAAATTTTCAGCTTGAGTCAGCACTGGTACGTGGGCGGCTCGGGTGCGGGTCTGCAGACCGTGGAATGTGGCTGGCAGGTATACCCTGGCAAATACGGTAATAACAAACCGGTACTGTTCATCTATTGGACCGCCGATAACTATAATAAAACCGGTTGTTATAACCTGGACTGCAGTGCCTTTGTCCAGACCAATTCCAGCTGGGCCCTTGGCGGCGCGCTTTCGCCGGTCTCAACCAGTGGTGGTACACAGTATGAGATTGAACTGGCGTATTATCTGAGCGGTGGCAACTGGTGGCTGTACCTGAACGGTACCAGCGCCAGTGATGCCATTGGTTACTACCCTGCTACCCTGTTTGGCGGTGGCCAGTTAGCCACCAACGCGACCGAAATCGATTATGGTGGCGAGACCGTGGGTACCACCTCGTGGCCGCCGATGGGCTCCGGTGCCTTTCCGAGCGAAGGTTACCGTCACGCGGCATACCAGCGCGACATCTATTACTATCCGCCGTCCGGTGGTTCTCAATCTGCCTCACTCACCCCGAGCCAACCGTCACCAAGCTGCTACACCATCGACGTGACCAATGCCTCGGCAAGTTGGAACGAATATTTCTTTTTTGGCGGTCCGGGCGGCTCCAACTGCGAAAACCTGTATTTCCAGGGCAGTCACCATCACCATCACCATTAASEQ ID NO: 26 Full plasmid sequence for H3CACGTGCGCCGAAAAAGCTCACCCCGTTCAGCGAATTTATTGAATCAGTGAAGGCGGCGAAACATGAAGAGTTCAAGGCACGTCCAGGCGCAAAAGTGAAAGATGCAGAAGAGTGTGAAGAAATGCGTCAGCATTTATTGAATCTGTACGAAGGCGTCGAAGTTCAACATAGCTTTGCTGACGAAGATGGCCAGATTTTCGATTGCATTCCGATTG AACAGCAGCCTTCGCTGCGCGGTAGCGGTGCCAAAGTTATCGCTACCCCGCCGG ATCTGCCTCCGGCTGCGGGCGCTTCGGCGAAAGAAGCGGAGGAGAGCGCCAAAG CGGTCCAACCGCCGTTATCTCCGGACCGCACCGATCGTTTTGGCAATGCGATGTC ATGTCCGGATGGTACCATCCCTATGCGTCGCGTAACCCTGGAAGAACTCGCCCGC TTCGAGACCTTGGAGGACTTTTTCCGCAAGGGTCCGAACGGCGCGGGTAAACGTC CTCCACGTGAAGCATCTGAGGCACCGCCAGCGGTTGCCGCGACCCATAAATATG CCCACGCATATCAGAACGTGGATAATTTAGGTGGTCACAGCTTTCTTAACGTTTG GAATCCAGCTGTGGGTGCGAATCAAATTTTCAGCTTGAGTCAGCACTGGTACGTG GGCGGCTCGGGTGCGGGTCTGCAGACCGTGGAATGTGGCTGGCAGGTATACCCT GGCAAATACGGTAATAACAAACCGGTACTGTTCATCTATTGGACCGCCGATAACT ATAATAAAACCGGTTGTTATAACCTGGACTGCAGTGCCTTTGTCCAGACCAATTC CAGCTGGGCCCTTGGCGGCGCGCTTTCGCCGGTCTCAACCAGTGGTGGTACACAG TATGAGATTGAACTGGCGTATTATCTGAGCGGTGGCAACTGGTGGCTGTACCTGA ACGGTAC C AGCGC C AGTGATGC C ATTGGTTACTACC CTGCT AC CCTGTTTGGCGG TGGCCAGTTAGCCACCAACGCGACCGAAATCGATTATGGTGGCGAGACCGTGGG TACCACCTCGTGGCCGCCGATGGGCTCCGGTGCCTTTCCGAGCGAAGGTTACCGT CACGCGGCATACCAGCGCGACATCTATTACTATCCGCCGTCCGGTGGTTCTCAAT CTGCCTCACTCACCCCGAGCCAACCGTCACCAAGCTGCTACACCATCGACGTGAC CAATGCCTCGGCAAGTTGGAACGAATATTTCTTTTTTGGCGGTCCGGGCGGCTCC AACTGCGAAAACCTGTATTTCCAGGGCAGTCACCATCACCATCACCATTAACTCG AGCACCACCACCACCACCACTGAGATCCGGCTGCTAACAAAGCCCGAAAGGAAG CTGAGTTGGCTGCTGCCACCGCTGAGCAATAACTAGCATAACCCCTTGGGGCCTC TAAACGGGTCTTGAGGGGTTTTTTGCTGAAAGGAGGAACTATATCCGGATTGGCG AATGGGACGCGCCCTGTAGCGGCGCATTAAGCGCGGCGGGTGTGGTGGTTACGC GCAGCGTGACCGCTACACTTGCCAGCGCCCTAGCGCCCGCTCCTTTCGCTTTCTTC CCTTCCTTTCTCGCCACGTTCGCCGGCTTTCCCCGTCAAGCTCTAAATCGGGGGCT CCCTTTAGGGTTCCGATTTAGTGCTTTACGGCACCTCGACCCCAAAAAACTTGAT TAGGGTGATGGTTCACGTAGTGGGCCATCGCCCTGATAGACGGTTTTTCGCCCTT TGACGTTGGAGTCCACGTTCTTTAATAGTGGACTCTTGTTCCAAACTGGAACAAC ACTCAACCCTATCTCGGTCTATTCTTTTGATTTATAAGGGATTTTGCCGATTTCGG CCTATTGGTTAAAAAATGAGCTGATTTAACAAAAATTTAACGCGAATTTTAACAA AATATTAACGCTTACAATTTAGGTGGCACTTTTCGGGGAAATGTGCGCGGAACCC CTATTTGTTTATTTTTCTAAATACATTCAAATATGTATCCGCTCATGAATTAATTC TTAGAAAAACTCATCGAGCATCAAATGAAACTGCAATTTATTCATATCAGGATTA TCAATACCATATTTTTGAAAAAGCCGTTTCTGTAATGAAGGAGAAAACTCACCGA GGCAGTTCCATAGGATGGCAAGATCCTGGTATCGGTCTGCGATTCCGACTCGTCC AACATCAATACAACCTATTAATTTCCCCTCGTCAAAAATAAGGTTATCAAGTGAG AAATCACCATGAGTGACGACTGAATCCGGTGAGAATGGCAAAAGTTTATGCATTTCTTTCCAGACTTGTTCAACAGGCCAGCCATTACGCTCGTCATCAAAATCACTCG CATCAACCAAACCGTTATTCATTCGTGATTGCGCCTGAGCGAGACGAAATACGCG ATCGCTGTTAAAAGGACAATTACAAACAGGAATCGAATGCAACCGGCGCAGGAA CACTGCCAGCGCATCAACAATATTTTCACCTGAATCAGGATATTCTTCTAATACC TGGAATGCTGTTTTCCCGGGGATCGCAGTGGTGAGTAACCATGCATCATCAGGAG TACGGATAAAATGCTTGATGGTCGGAAGAGGCATAAATTCCGTCAGCCAGTTTA GTCTGACCATCTCATCTGTAACATCATTGGCAACGCTACCTTTGCCATGTTTCAGA AACAACTCTGGCGCATCGGGCTTCCCATACAATCGATAGATTGTCGCACCTGATT GCCCGACATTATCGCGAGCCCATTTATACCCATATAAATCAGCATCCATGTTGGA ATTTAATCGCGGCCTAGAGCAAGACGTTTCCCGTTGAATATGGCTCATAACACCC CTTGTATTACTGTTTATGTAAGCAGACAGTTTTATTGTTCATGACCAAAATCCCTTAACGTGAGTTTTCGTTCCACTGAGCGTCAGACCCCGTAGAAAAGATCAAAGGATCTTCTTGAGATCCTTTTTTTCTGCGCGTAATCTGCTGCTTGCAAACAAAAAAACCACCGCTACCAGCGGTGGTTTGTTTGCCGGATCAAGAGCTACCAACTCTTTTTCCGAAGGTAACTGGCTTCAGCAGAGCGCAGATACCAAATACTGTCCTTCTAGTGTAGCCGTAGTTAGGCCACCACTTCAAGAACTCTGTAGCACCGCCTACATACCTCGCTCTGCTAATCCTGTTACCAGTGGCTGCTGCCAGTGGCGATAAGTCGTGTCTTACCGGGTTGGACTCAAGACGATAGTTACCGGATAAGGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGGAGCGAACGACCTACACCGAACTGAGATACCTACAGCGTGAGCTATGAGAAAGCGCCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCGGAACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCGGGTTTCGCCACCTCTGACTTGAGCGTCGATTTTTGTGATGCTCGTCAGGGGGGCGGAGCCTATGGAAAAACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTTCTTTCCTGCGTTATCCCCTGATTCTGTGGATAACCGTATTACCGCCTTTGAGTGAGCTGATACCGCTCGCCGCAGCCGAACGACCGAGCGCAGCGAGTCAGTGAGCGAGGAAGCGGAAGAGCGCCTGATGCGGTATTTTCTCCTTACGCATCTGTGCGGTATTTCACACCGCAATGGTGCACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAAGCCAGTATACACTCCGCTATCGCTACGTGACTGGGTCATGGCTGCGCCCCGACACCCGCCAACACCCGCTGACGCGCCCTGACGGGCTTGTCTGCTCCCGGCATCCGCTTACAGACAAGCTGTGACCGTCTCCGGGAGCTGCATGTGTCAGAGGTTTTCACCGTCATCACCGAAACGCGCGAGGCAGCTGCGGTAAAGCTCATCAGCGTGGTCGTGAAGCGATTCACAGATGTCTGCCTGTTCATCCGCGTCCAGCTCGTTGAGTTTCTCCAGAAGCGTTAATGTCTGGCTTCTGATAAAGCGGGCCATGTTAAGGGCGGTTTTTTCCTGTTTGGTCACTGATGCCTCCGTGTAAGGGGGATTTCTGTTCATGGGGGTAATGATACCGATGAAACGAGAGAGGATGCTCACGATACGGGTTACTGATGATGAACATGCCCGGTTACTGGAACGTTGTGAGGGTAAACAACTGGCGGTATGGATGCGGCGGGACCAGAGAAAAATCACTCAGGGTCAATGCCAGCGCTTCGTTAATACAGATGTAGGTGTTCCACAGGGTAGCCAGCAGCATCCTGCGATGCAGATCCGGAACATAATGGTGCAGGGCGCTGACTTCCGCGTTTCCAGACTTTACGAAACACGGAAACCGAAGACCATTCATGTTGTTGCTCAGGTCGCAGACGTTTTGCAGCAGCAGTCGCTTCACGTTCGCTCGCGTATCGGTGATTCATTCTGCTAACCAGTAAGGCAACCCCGCCAGCCTAGCCGGGTCCTCAACGACAGGAGCACGATCATGCGCACCCGTGGGGCCGCCATGCCGGCGATAATGGCCTGCTTCTCGCCGAAACGTTTGGTGGCGGGACCAGTGACGAAGGCTTGAGCGAGGGCGTGCAAGATTCCGAATACCGCAAGCGACAGGCCGATCATCGTCGCGCTCCAGCGAAAGCGGTCCTCGCCGAAAATGACCCAGAGCGCTGCCGGCACCTGTCCTACGAGTTGCATGATAAAGAAGACAGTCATAAGTGCGGCGACGATAGTCATGCCCCGCGCCCACCGGAAGGAGCTGACTGGGTTGAAGGCTCTCAAGGGCATCGGTCGAGATCCCGGTGCCTAATGAGTGAGCTAACTTACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCCAGGGTGGTTTTTCTTTTCACCAGTGAGACGGGCAACAGCTGATTGCCCTTCACCGCCTGGCCCTGAGAGAGTTGCAGCAAGCGGTCCACGCTGGTTTGCCCCAGCAGGCGAAAATCCTGTTTGATGGTGGTTAACGGCGGGATATAACATGAGCTGTCTTCGGTATCGTCGTATCCCACTACCGAGATATCCGCACCAACGCGCAGCCCGGACTCGGTAATGGCGCGCATTGCGCCCAGCGCCATCTGATCGTTGGCAACCAGCATCGCAGTGGGAACGATGCCCTCATTCAGCATTTGCATGGTTTGTTGAAAACCGGACATGGCACTCCAGTCGCCTTCCCGTTCCGCTATCGGCTGAATTTGATTGCGAGTGAGATATTTATGCCAGCCAGCCAGACGCAGACGCGCCGAGACAGAACTTAATGGGCCCGCTAACAGCGCGATTTGCTGGTGACCCAATGCGACCAGATGCTCCACGCCCAGTCGCGTACCGTCTTCATGGGAGAAAATAATACTGTTGATGGGTGTCTGGTCAGAGACATCAAGAAATAACGCCGGAACATTAGTGCAGGCAGCTTCCACAGCAATGGCATCCTGGTCATCCAGCGGATAGTTAATG ATCAGCCCACTGACGCGTTGCGCGAGAAGATTGTGCACCGCCGCTTTACAGGCTT CGACGCCGCTTCGTTCTACCATCGACACCACCACGCTGGCACCCAGTTGATCGGC GCGAGATTTAATCGCCGCGACAATTTGCGACGGCGCGTGCAGGGCCAGACTGGA GGTGGCAACGCCAATCAGCAACGACTGTTTGCCCGCCAGTTGTTGTGCCACGCGG TTGGGAATGTAATTCAGCTCCGCCATCGCCGCTTCCACTTTTTCCCGCGTTTTCGC AGAAACGTGGCTGGCCTGGTTCACCACGCGGGAAACGGTCTGATAAGAGACACC GGCATACTCTGCGACATCGTATAACGTTACTGGTTTCACATTCACCACCCTGAAT TGACTCTCTTCCGGGCGCTATCATGCCATACCGCGAAAGGTTTTGCGCCATTCGA TGGTGTCCGGGATCTCGACGCTCTCCCTTATGCGACTCCTGCATTAGGAAGCAGC CCAGTAGTAGGTTGAGGCCGTTGAGCACCGCCGCCGCAAGGAATGGTGCATGCA AGGAGATGGCGCCCAACAGTCCCCCGGCCACGGGGCCTGCCACCATACCCACGC CGAAACAAGCGCTCATGAGCCCGAAGTGGCGAGCCCGATCTTCCCCATCGGTGA TGTCGGCGATATAGGCGCCAGCAACCGCACCTGTGGCGCCGGTGATGCCGGCCA CGATGCGTCCGGCGTAGAGGATCGAGATCTCGATCCCGCGAAATTAATACGACT CACTATAGGGGAATTGTGAGCGGATAACAATTCCCCTCTAGAAATAATTTTGTTT AACTTTAAGAAGGAGATATACCATGGSEQ ID NO: 27 DNA sequence for 2C 11ATGGCACGTGCGCCGAAAAAGCTCACCCCGTTCAGCGAATTTATTGAATCAGTGA AGGC GGC GAAAC ATGA AGAGTTC A AGGC AC GTC C AGGCGC AAAAGTGAAAGAT GCAGAAGAGTGTGAAGAAATGCGTCAGCATTTATTGAATCTGTACGAAGGCGTC GAAGTTCAACATAGCTTTGCTGACGAAGATGGCCAGATTTTCGATTGCATTCCGA TTGAACAGCAGCCTTCGCTGCGCGGTAGCGGTGCCAAAGTTATCGCTACCCCGCC GGATCTGCCTCCGGCTGCGGGCGCTTCGGCGAAAGAAGCGGAGGAGAGCGCCAA AGCGGTCCAACCGCCGTTATCTCCGGACCGCACCGATCGTTTTGGCAATGCGATG TCATGTCCGGATGGTACCATCCCTATGCGTCGCGTAACCCTGGAAGAACTCACCC GCTTCGAGACCTTGGAGGACCTTTTCCGCAAGGGTCCGAACGGCGCGGGTAAAC GTCCTCCACGTGAAGCATCTGAGGCACCGCCAGCGGTTGCCGCGACCCATAAAT ATGCCCACGCATATCAGAACGTGGATAATTTAGGTGGTCACAGCTTTCTTAACGT TTGGAATCCAGCTGTGGGTGCGAATCAAATTTTCAGCTTGAGTCAGCACTGGTAC GTGGGCGGCTCGGGTGCGGGTCTGCAGACCGTGGAATGTGGCTGGCAGGTATAC CCTGGCAAATACGGTAATAACAAACCGGTACTGTTCATCTATTGGACCGCCGATA ACTATAATAAAACCGGTTGTTATAACCTGGACTGCAGTGCCTTTGTCCAGACCAA TTCCAGCTGGGCCCTTGGCGGCGCGCTTTCGCCGGTCTCAACCAGTGGTGCTACA CAGTATGAGATTGAACTGGCGTATTATCTGAGCGGTGGCAACTGGTGGCTGTACC TGAACGGTACCAGCGCCAGTGATGCCATTGGTTACTACCCTGCTACCCTGTTTGG CGGTGGCCAGTTAGCCACCAACGCGACCGAAATCGATTATGGTGGCGAGACCGT GGGTACCACCTCGTGGCCGCCGATGGGCTCCGGTGCCTTTCCGAGCGAAGGTTAC CGTCACGCGGCATACCAGCGCGACATCTATTACTATCCGCCGTCCGGTGGTTCTC AATCTGCCTCACTCACCCCGAGCCAACCGTCACCATGCTGCTACACCATCGACGT GACCAATGCCTCGGCAAGTTGGAACGAATATTTCTTTTTTGGCGGTCCGGGCGGC TCCAACTGCGAAAACCTGTATTTCCAGGGCAGTCACCATCACCATCACCATTAASEQ ID NO: 28 Full plasmid sequence for 2C11GCATTAAGCGCGGCGGGTGTGGTGGTTACGCGCAGCGTGACCGCTACACTTGCC AGCGCCCTAGCGCCCGCTCCTTTCGCTTTCTTCCCTTCCTTTCTCGCCACGTTCGC CGGCTTTCCCCGTCAAGCTCTAAATCGGGGGCTCCCTTTAGGGTTCCGATTTAGT GCTTTACGGCACCTCGACCCCAAAAAACTTGATTAGGGTGATGGTTCACGTAGTG GGCCATCGCCCTGATAGACGGTTTTTCGCCCTTTGACGTTGGAGTCCACGTTCTTTAATAGTGGACTCTTGTTCCAAACTGGAACAACACTCAACCCTATCTCGGTCTATT CTTTTGATTTATAAGGGATTTTGCCGATTTCGGCCTATTGGTTAAAAAATGAGCTG ATTTAACAAAAATTTAACGCGAATTTTAACAAAATATTAACGCTTACAATTTAGG TGGCACTTTTCGGGGAAATGTGCGCGGAACCCCTATTTGTTTATTTTTCTAAATAC ATTCAAATATGTATCCGCTCATGAATTAATTCTTAGAAAAACTCATCGAGCATCA AATGAAACTGCAATTTATTCATATCAGGATTATCAATACCATATTTTTGAAAAAG CCGTTTCTGTAATGAAGGAGAAAACTCACCGAGGCAGTTCCATAGGATGGCAAG ATCCTGGTATCGGTCTGCGATTCCGACTCGTCCAACATCAATACAACCTATTAAT TTCCCCTCGTCAAAAATAAGGTTATCAAGTGAGAAATCACCATGAGTGACGACTG AATCCGGTGAGAATGGCAAAAGTTTATGCATTTCTTTCCAGACTTGTTCAACAGG CCAGCCATTACGCTCGTCATCAAAATCACTCGCATCAACCAAACCGTTATTCATT CGTGATTGCGCCTGAGCGAGACGAAATACGCGATCGCTGTTAAAAGGACAATTA CAAACAGGAATCGAATGCAACCGGCGCAGGAACACTGCCAGCGCATCAACAATA TTTTCACCTGAATCAGGATATTCTTCTAATACCTGGAATGCTGTTTTCCCGGGGAT CGCAGTGGTGAGTAACCATGCATCATCAGGAGTACGGATAAAATGCTTGATGGT CGGAAGAGGCATAAATTCCGTCAGCCAGTTTAGTCTGACCATCTCATCTGTAACA TCATTGGCAACGCTACCTTTGCCATGTTTCAGAAACAACTCTGGCGCATCGGGCT TCCCATACAATCGATAGATTGTCGCACCTGATTGCCCGACATTATCGCGAGCCCA TTTATACCCATATAAATCAGCATCCATGTTGGAATTTAATCGCGGCCTAGAGCAA GACGTTTCCCGTTGAATATGGCTCATAACACCCCTTGTATTACTGTTTATGTAAGC AGACAGTTTTATTGTTCATGACCAAAATCCCTTAACGTGAGTTTTCGTTCCACTGA GCGTCAGACCCCGTAGAAAAGATCAAAGGATCTTCTTGAGATCCTTTTTTTCTGC GCGTAATCTGCTGCTTGCAAACAAAAAAACCACCGCTACCAGCGGTGGTTTGTTT GCCGGATCAAGAGCTACCAACTCTTTTTCCGAAGGTAACTGGCTTCAGCAGAGCG CAGATACCAAATACTGTCCTTCTAGTGTAGCCGTAGTTAGGCCACCACTTCAAGA ACTCTGTAGCACCGCCTACATACCTCGCTCTGCTAATCCTGTTACCAGTGGCTGCT GCCAGTGGCGATAAGTCGTGTCTTACCGGGTTGGACTCAAGACGATAGTTACCGG ATAAGGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGG AGCGAACGACCTACACCGAACTGAGATACCTACAGCGTGAGCTATGAGAAAGCG CCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCG GAACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATA GTCCTGTCGGGTTTCGCCACCTCTGACTTGAGCGTCGATTTTTGTGATGCTCGTCA GGGGGGCGGAGCCTATGGAAAAACGCCAGCAACGCGGCCTTTTTACGGTTCCTG GCCTTTTGCTGGCCTTTTGCTCACATGTTCTTTCCTGCGTTATCCCCTGATTCTGTG GATAACCGTATTACCGCCTTTGAGTGAGCTGATACCGCTCGCCGCAGCCGAACGA CCGAGCGCAGCGAGTCAGTGAGCGAGGAAGCGGAAGAGCGCCTGATGCGGTATT TTCTCCTTACGCATCTGTGCGGTATTTCACACCGCAATGGTGCACTCTCAGTACAA TCTGCTCTGATGCCGCATAGTTAAGCCAGTATACACTCCGCTATCGCTACGTGAC TGGGTCATGGCTGCGCCCCGACACCCGCCAACACCCGCTGACGCGCCCTGACGG GCTTGTCTGCTCCCGGCATCCGCTTACAGACAAGCTGTGACCGTCTCCGGGAGCT GCATGTGTCAGAGGTTTTCACCGTCATCACCGAAACGCGCGAGGCAGCTGCGGTAAAGCTCATCAGCGTGGTCGTGAAGCGATTCACAGATGTCTGCCTGTTCATCCGC GTCCAGCTCGTTGAGTTTCTCCAGAAGCGTTAATGTCTGGCTTCTGATAAAGCGG GCCATGTTAAGGGCGGTTTTTTCCTGTTTGGTCACTGATGCCTCCGTGTAAGGGG GATTTCTGTTCATGGGGGTAATGATACCGATGAAACGAGAGAGGATGCTCACGA TACGGGTTACTGATGATGAACATGCCCGGTTACTGGAACGTTGTGAGGGTAAAC AACTGGCGGTATGGATGCGGCGGGACCAGAGAAAAATCACTCAGGGTCAATGCC AGCGCTTCGTTAATACAGATGTAGGTGTTCCACAGGGTAGCCAGCAGCATCCTGC GATGCAGATCCGGAACATAATGGTGCAGGGCGCTGACTTCCGCGTTTCCAGACTT TACGAAACACGGAAACCGAAGACCATTCATGTTGTTGCTCAGGTCGCAGACGTTTTGCAGCAGCAGTCGCTTCACGTTCGCTCGCGTATCGGTGATTCATTCTGCTAACCAGTAAGGCAACCCCGCCAGCCTAGCCGGGTCCTCAACGACAGGAGCACGATCATGCGCACCCGTGGGGCCGCCATGCCGGCGATAATGGCCTGCTTCTCGCCGAAACGTTTGGTGGCGGGACCAGTGACGAAGGCTTGAGCGAGGGCGTGCAAGATTCCGAATACCGCAAGCGACAGGCCGATCATCGTCGCGCTCCAGCGAAAGCGGTCCTCGCCGAAAATGACCCAGAGCGCTGCCGGCACCTGTCCTACGAGTTGCATGATAAAGAAGACAGTCATAAGTGCGGCGACGATAGTCATGCCCCGCGCCCACCGGAAGGAGCTGACTGGGTTGAAGGCTCTCAAGGGCATCGGTCGAGATCCCGGTGCCTAATGAGTGAGCTAACTTACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCCAGGGTGGTTTTTCTTTTCACCAGTGAGACGGGCAACAGCTGATTGCCCTTCACCGCCTGGCCCTGAGAGAGTTGCAGCAAGCGGTCCACGCTGGTTTGCCCCAGCAGGCGAAAATCCTGTTTGATGGTGGTTAACGGCGGGATATAACATGAGCTGTCTTCGGTATCGTCGTATCCCACTACCGAGATATCCGCACCAACGCGCAGCCCGGACTCGGTAATGGCGCGCATTGCGCCCAGCGCCATCTGATCGTTGGCAACCAGCATCGCAGTGGGAACGATGCCCTCATTCAGCATTTGCATGGTTTGTTGAAAACCGGACATGGCACTCCAGTCGCCTTCCCGTTCCGCTATCGGCTGAATTTGATTGCGAGTGAGATATTTATGCCAGCCAGCCAGACGCAGACGCGCCGAGACAGAACTTAATGGGCCCGCTAACAGCGCGATTTGCTGGTGACCCAATGCGACCAGATGCTCCACGCCCAGTCGCGTACCGTCTTCATGGGAGAAAATAATACTGTTGATGGGTGTCTGGTCAGAGACATCAAGAAATAACGCCGGAACATTAGTGCAGGCAGCTTCCACAGCAATGGCATCCTGGTCATCCAGCGGATAGTTAATGATCAGCCCACTGACGCGTTGCGCGAGAAGATTGTGCACCGCCGCTTTACAGGCTTCGACGCCGCTTCGTTCTACCATCGACACCACCACGCTGGCACCCAGTTGATCGGCGCGAGATTTAATCGCCGCGACAATTTGCGACGGCGCGTGCAGGGCCAGACTGGAGGTGGCAACGCCAATCAGCAACGACTGTTTGCCCGCCAGTTGTTGTGCCACGCGGTTGGGAATGTAATTCAGCTCCGCCATCGCCGCTTCCACTTTTTCCCGCGTTTTCGCAGAAACGTGGCTGGCCTGGTTCACCACGCGGGAAACGGTCTGATAAGAGACACCGGCATACTCTGCGACATCGTATAACGTTACTGGTTTCACATTCACCACCCTGAATTGACTCTCTTCCGGGCGCTATCATGCCATACCGCGAAAGGTTTTGCGCCATTCGATGGTGTCCGGGATCTCGACGCTCTCCCTTATGCGACTCCTGCATTAGGAAGCAGCCCAGTAGTAGGTTGAGGCCGTTGAGCACCGCCGCCGCAAGGAATGGTGCATGCAAGGAGATGGCGCCCAACAGTCCCCCGGCCACGGGGCCTGCCACCATACCCACGCCGAAACAAGCGCTCATGAGCCCGAAGTGGCGAGCCCGATCTTCCCCATCGGTGATGTCGGCGATATAGGCGCCAGCAACCGCACCTGTGGCGCCGGTGATGCCGGCCACGATGCGTCCGGCGTAGAGGATCGAGATCTCGATCCCGCGAAATTAATACGACTCACTATAGGGGAATTGTGAGCGGATAACAATTCCCCTCTAGAAATAATTTTGTTTAACTTTAAGAAGGAGATATACCATGGCACGTGCGCCGAAAAAGCTCACCCCGTTCAGCGAATTTATTGAATCAGTGAAGGCGGCGAAACATGAAGAGTTCAAGGCACGTCCAGGCGCAAAAGTGAAAGATGCAGAAGAGTGTGAAGAAATGCGTCAGCATTTATTGAATCTGTACGAAGGCGTCGAAGTTCAACATAGCTTTGCTGACGAAGATGGCCAGATTTTCGATTGCATTCCGATTGAACAGCAGCCTTCGCTGCGCGGTAGCGGTGCCAAAGTTATCGCTACCCCGCCGGATCTGCCTCCGGCTGCGGGCGCTTCGGCGAAAGAAGCGGAGGAGAGCGCCAAAGCGGTCCAACCGCCGTTATCTCCGGACCGCACCGATCGTTTTGGCAATGCGATGTCATGTCCGGATGGTACCATCCCTATGCGTCGCGTAACCCTGGAAGAACTCACCCGCTTCGAGACCTTGGAGGACCTTTTCCGCAAGGGTCCGAACGGCGCGGGTAAACGTCCTCCACGTGAAGCATCTGAGGCACCGCCAGCGGTTGCCGCGACCCATAAATATGCCCACGCATATCAGAACGTGGATAATTTAGGTGGTCACAGCTTTCTTAACGTTTGGAATCCAGCTGTGGGTGCGAATCAAATTTTCAGCTTGAGTCAGCACTGGTACGTGGGCGGCTCGGGTGCGGGTCTGCAGACCGTGGAATGTGGCTGGCAGGTATACCCTGGCAAATACGGTAATAACAAACCGGTACTGTTCATCTATTGGACCGCCGATAACTATAATAAAACCGGTTGTTATAACCTGGACTGCAGTGCCTTTGTCCAGACCAATTCCAGCTGGGCCCTTGGCGGCGCGCTTTCGCCGGTCTCAACCAGTGGTGCTACACAGTATGAGATTGAACTGGCGTATTATCTGAGCGGTGGCAACTGGTGGCTGTACCTGAACGGTACCAGCGCCAGTGATGCCATTGGTTACTACCCTGCTACCCTGTTTGGCGGTGGCCAGTTAGCCACCAACGCGACCGAAATCGATTATGGTGGCGAGACCGTGGGTACCACCTCGTGGCCGCCGATGGGCTCCGGTGCCTTTCCGAGCGAAGGTTACCGTCACGCGGCATACCAGCGCGACATCTATTACTATCCGCCGTCCGGTGGTTCTCAATCTGCCTCACTCACCCCGAGCCAACCGTCACCATGCTGCTACACCATCGACGTGACCAATGCCTCGGCAAGTTGGAACGAATATTTCTTTTTTGGCGGTCCGGGCGGCTCCAACTGCGAAAACCTGTATTTCCAGGGCAGTCACCATCACCATCACCATTAACTCGAGCACCACCACCACCACCACTGAGATCCGGCTGCTAACAAAGCCCGAAAGGAAGCTGAGTTGGCTGCTGCCACCGCTGAGCAATAACTAGCATAACCCCTTGGGGCCTCTAAACGGGTCTTGAGGGGTTTTTTGCTGAAAGGAGGAACTATATCCGGATTGGCGAATGGGACGCGCCCTGTAGCGGC
[0104] The use of the terms “a” and ’‘an” and "the" and similar referents (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms first, second etc. as used herein are not meant to denote any particular ordering, but simply for convenience to denote a plurality of, for example, layers. The terms “comprising’’, “having”, “including”, and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to”) unless otherwise noted. Recitation of ranges of values are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it w ere individually recited herein. The endpoints of all ranges are included w ithin the range and independently combinable. All methods described herein can be performed in a suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”), is intended merely to better illustrate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention as used herein.
[0105] While the invention has been described with reference to an exemplary embodiment, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof without departing from the scope of the invention. In addition, many modifications may be made to adapt a particular situation or matenal to the teachings of the invention without departing from the essential scope thereof. Therefore, it is intended that the invention not be limited to the particularembodiment disclosed as the best mode contemplated for carrying out this invention, but that the invention will include all embodiments falling within the scope of the appended claims. Any combination of the above-described elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or otherwise clearly contradicted by context.
Claims
Claims1. A protease of any of SEQ ID NOs: 6-14 or a proteolytically active variant thereof with greater than 85%, 90%, 95% or 98% sequence identity thereto, wherein the proteolytically active variant cleaves a peptide of SEQ ID NO: 17 at QP|QL to provide peptides of any of SEQ ID NOs. 18-24.
2. The protease of claim 1, wherein the protease cleaves the peptide of SEQ ID NO: 17 at a pH of 2 to 4.
3. The protease of claim 1 or 2, wherein 1 pM of the protease and 1 mg / mL SEQ ID NO: 17 peptide, when incubated at a pH of 2.5 at 37°C for 24 hours, provides cleavage of 80% or more of the peptide of SEQ ID NO: 17 at QP|QL to provide the peptides of any of SEQ ID NOs. 18-24.
4. A composition comprising the protease of any of claims 1 to 3 and an excipient.
5. The composition of claim 4 in the form of a pharmaceutical composition, a food supplement, or a food product.
6. A polynucleotide encoding the protease of any of claims 1-3.
7. An expression cassette comprising the polynucleotide of claim 6.
8. A cell comprising the polynucleotide of claim 6, or the expression cassette of claim 7.
9. A method for degrading gluten oligopeptides, comprising contacting the gluten oligopeptides with the protease of any of claims 1-3 or the composition of claims 4-5.
10. A method of treating a disorder associated with gluten intolerance, comprising administering to a subject in need thereof an effective amount of the protease of any of claims 1-3 or the composition of claims 4-5.
11. The method of claim 10, wherein the disorder associated with gluten intolerance comprises celiac disease, non-celiac gluten sensitivity, wheat allergy, gluten ataxia, or dermatitis herpetiformis.
12. A protease for use in a method of treating gluten intolerance in a subject, wherein the method comprises administering to the subject an effective amount of the protease of any of claims 1-3 or the composition of claims 4-5.
13. Use of a protease of any of claims 1-3 for the manufacture of a medicament for treating gluten intolerance in a subject, the method comprising administering to the subject an effective amount of the protease.
14. A method of testing a subject suspected of having gluten intolerance or gluten sensitivity, comprising administering to the subject a composition comprising the protease of any of claims 1-3, waiting a period of time, and after the period of time, assessing one or more symptoms of gluten intolerance or gluten sensitivity in the subject, wherein a reduction in the one or more symptoms indicates the subject has gluten intolerance or gluten sensitivity.
15. The method of claim 14, wherein the subject suspected of having gluten intolerance or gluten sensitivity is suffering from gastrointestinal distress fatigue, headache, joint pain, muscle pain, difficulty concentrating, depression, anxiety, anemia, unexplained weight loss, skin reactions, and / or injury' due to inflammation.
16. The method of claim 14, wherein the one or more symptoms is gastrointestinal distress fatigue, headache, joint pain, muscle pain, difficulty concentrating, depression, anxiety, anemia, unexplained weight loss, skin reactions, and / or injury due to inflammation.
17. The method of claim 14, wherein the period of time is 1 week or longer.
18. A fluorogenic peptide substrate, comprisingFluorophorel - SEQ ID NO: 4 - K - Fluorophore 2, wherein Fluorophore 1 and Fluorophore 2 are active at pH 2-4.
19. The fluorogenic peptide substrate of claim 18, comprisingHiLyte™ Fluor 488-SEQ ID NO: 4-QXLTM520-NH2.
20. A method of identifying an ancestral protease comprising identity ing a modem protease sequence, wherein the modem protease is neprosin, query ing the modem protease sequence in a sequence database to provide a list of similar sequences, selecting a subset of the list of similar sequences using a sequence similarity cut-off. aligning the subset of the list of similar sequences using a multiple sequence alignment tool to provide aligned sequences, preparing a phylogenetic tree based on the aligned sequences, selecting a node from the phylogenetic tree and selecting the ancestral proteases from the node, and performing a fluorescence assay with the fluorogenic peptide substrate of claim 14 or claim 15 to verity' the activity' of the ancestral protease to degrade gluten peptides.
21. The method of claim 20, wherein the sequence database is a database of bacterial kingdom sequences.
22. The method of claim 20, wherein the sequence similarity cutoff is 90% sequence similarity.
23. The method of claim 20, wherein the phylogenetic tree is a Maximum Likelihood tree.
24. The method of claim 20, wherein the verifying is performed on crude cell extracts expressing the ancestral proteases in a multi-well assay.
25. The method of claim 20, further comprising further evolving the ancestral protease by mutating one or more amino acids of the ancestral protease, and selecting ancestral protease variants that have increased cleavage of the fluorogenic peptide substrate of claim 14 or claim 15 compared to the ancestral protease.
Citation Information
Patent Citations
Rothia species gluten-degrading enzymes and uses thereof
US20130171109A1
Enzyme treatment of foodstuffs for celiac sprue
US20170119860A1