Polypeptides from the nesting materials of solitary colletid bees
Colletid bee-derived polypeptides offer eco-friendly alternatives to synthetic polymers, addressing environmental and health concerns by providing sustainable materials with desirable properties.
Patent Information
- Application Number
- PCT/IB2025/055575
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-04
AI Technical Summary
Current synthetic polymers used in textiles, biomedical devices, and cosmetics are not eco-friendly, posing environmental hazards and health risks, and their alternatives face challenges in cost-efficiency and material availability.
Utilization of polypeptides derived from the nesting materials of solitary Colletid bees, such as Hylaeus, Euryglossa, and Meroglossa, which exhibit eco-friendly properties and can be recombinantly produced for applications similar to synthetic polymers.
The Colletid bee-derived polypeptides provide sustainable alternatives with desirable properties like hydrophilicity, mechanical strength, and biocompatibility, reducing environmental impact and health risks while maintaining performance.
Smart Images

Figure IMGF000062_0001 
Figure IMGF000063_0001 
Figure IMGF000064_0001
Abstract
Description
[0001] POLYPEPTIDES FROM THE NESTING MATERIALS OF SOLITARY COLLETID BEES
[0002] FIELD OF THE INVENTION
[0003] This invention generally relates to polypeptides from Colletid bees. The invention also relates to recombinant production of such polypeptides, to methods of making the polypeptides and to the use of the polypeptides to make various articles of manufacture comprising desirable properties.
[0004] BACKGROUND
[0005] Synthetic polymers (components of "plastics") are an exceptionally useful group of materials. Synthetic polymers are comprised in many of the everyday products we dress in, sleep on, and create with. In many instances, synthetic polymers are used in the manufacture of various materials to impart desirable properties. Examples of such properties include imparting mechanical strength, heat resistance, wicking or the ability to repel water or resist wetting.
[0006] Industries where the use of synthetic polymers is prevalent include textiles, biomedical devices, and cosmetics. In one example, mass produced synthetic clothing is composed of various polymers that have been spun into synthetic fibers including polyester, nylon, vinyl, and acrylics. In some cases, these synthetic fibers are hydrophobic and are poor absorbents of water or sweat, impacting their wear comfort. In others, the fibers are hydrophilic and overly absorbent, again impacting wearability.
[0007] In other examples, synthetic polymers may be used to impart wicking, hygroscopic or hydrophilic properties to a textile or other material. Synthetic polymers currently being used for such applications are not eco-friendly. For example, polyether amines are used to form a hydrophilic coating on nylon clothing. Unfortunately, these polymers degrade over time, releasing harmful by-products that are toxic to aquatic life into the environment.
[0008] Environmental concerns associated with textile finishing chemicals have shifted the focus of major manufacturing companies toward green (bio-based) chemicals, which are eco- friendly. Green chemicals are produced using animal and plant fats / oils, making them eco-friendly and cost-efficient compared to their conventional counterparts. However, the additional weight of oil-based products, the need to reapply them, and the fluctuating availability and prices of raw materials pose a challenge for the market players to achieve profitability and economies of scale.
[0009] Another important use of such coatings is in the medical device industry. Examples of various biomedical devices currently coated with hygroscopic / hydrophilic coatings include catheters, implants, tubes, lenses, and disposable plastic slides. In many examples, these coatings provide the coated biomedical devices (particularly those used in situ) with excellent biocompatibility, hydrophilicity, hydrophobicity and / or friction resistance, allowing for effective performance.
[0010] Typically, these coatings are composed of polyurethane, silicone, and polyethylene terephthalate materials.
[0011] Synthetic polymers are also used in the production of cosmetics / personal care products to impart properties like lubricity and viscosity. Phthalates are a group of chemicals called "everywhere" chemicals, found in products like nail polish, perfumes, deodorants, hair gels, shampoos, soaps, hair sprays, and body lotions. Phthalates have been identified as endocrine-disrupting chemicals in humans, causing hormonal imbalances, and numerous reproductive health, and developmental problems. Phthalates also bioaccumulate in fish, proving toxic to aquatic ecosystems and introducing hazards to humans.
[0012] The increasing public awareness of the environmental concerns associated with synthetic polymer use has led to an urgent need in the art for industries to provide alternatives that are non-toxic and eco-friendly.
[0013] It is an object of the present invention to provide a sustainable polypeptide or derivative thereof that is that may be used as an eco-friendly alternative for at least some of the synthetic polymers currently used in the various industries described above and / or to at least impart beneficial properties to at least some of the synthetic polymers currently used in the various industries described above and / or to provide a method of making such a polypeptide or derivative thereof and / or to at least provide the public with a useful choice.
[0014] In this specification where reference has been made to patent specifications, other external documents, or other sources of information, this is generally for the purpose of providing a context for discussing the features of the invention. Unless specifically stated otherwise, reference to such external documents is not to be construed as an admission that such documents, or such sources of information, in any jurisdiction, are prior art, or form part of the common general knowledge in the art.
[0015] SUMMARY OF THE INVENTION
[0016] Disclosed herein are polypeptides from Colletid bees in the genera Hylaeus, Euryglossa and Meroglossa and polynucleotides encoding such polypeptides. Also disclosed are compositions comprising these polypeptides and / or polynucleotides, and methods of making and using these polypeptides, polynucleotides and compositions.
[0017] In one aspect the present invention relates to an isolated polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0018] In another aspect the invention relates to an isolated polypeptide or functional portion thereof comprising, consisting essentially of or consisting of at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0019] In another aspect, the invention relates to an isolated polynucleotide or functional portion thereof comprising, consisting essentially of or consisting of at least 70% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0020] In another aspect the invention relates to an isolated polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0021] In another aspect the invention relates to an isolated polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0022] In another aspect, the invention relates to an isolated polynucleotide or functional portion thereof comprising, consisting essentially of or consisting of a polynucleotide having at least 70% nucleic acid sequence identity to SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65. In another aspect the invention relates to a vector comprising an isolated polynucleotide as described herein.
[0023] In another aspect the invention relates to a vector that encodes an isolated polypeptide as described herein.
[0024] In another aspect the invention relates to an isolated host cell comprising an isolated polynucleotide, isolated polypeptide, and / or vector as described herein.
[0025] In another aspect the invention relates to a composition comprising an isolated polynucleotide, isolated polypeptide and / or vector as described herein, and a carrier, diluent or excipient.
[0026] In another aspect the invention relates to a method of making an isolated Colletid bee nesting material (BNMP) polypeptide or functional portion thereof as described herein, the method comprising heterologously expressing the BNMP polypeptide or functional portion thereof in an isolated host cell, and optionally purifying the BNMP polypeptide.
[0027] In another aspect the invention relates to a polypeptide or functional portion thereof as described herein made by a method as described herein.
[0028] Various embodiments of the different aspects of the invention as discussed above are also set out below in the detailed description of the invention, but the invention is not limited thereto.
[0029] Other aspects of the invention may become apparent from the following description which is given by way of example only and with reference to the accompanying drawings.
[0030] BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The invention will now be described by way of non-limiting example only and with reference to the drawings in which:
[0032] Figure 1: Muscle alignment of the native predicted BNMP amino acid sequences HnMl (SEQ ID NO: 2), EaMl (SEQ ID NO: 6), EsMl (SEQ ID NO: 10), HeMl (SEQ ID NO: 14) and MiMl (SEQ ID NO: 18). Regions are indicated with lines and labels above the aligned sequences. A consensus plot is positioned beneath the aligned sequences and an amino acid count is positioned to the right of the aligned sequences. Figure 2: Length and amino acid percentage in the construct sequences (HnMl-01 = HnBNMP (SEQ ID NO: 4), EaMl-01= EaBNMP (SEQ ID NO: 8), EsMl-01= EsBNMP (SEQ ID NO: 12), HeMl-01= HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl-01(SEQ ID NO: 20) ), wildtype sequences (HnMl (SEQ ID NO: 2), EaMl (SEQ ID NO: 6), EsMl (SEQ ID NO: 10), HeM l (SEQ ID NO: 14) and MiMl (SEQ ID NO: 18)) and the signal peptide, N- terminal region, signal peptide + N-terminal region, charged region, repeat region and C- terminal region as tabulated by Emboss pepstats.
[0033] Figure 3: Beta-Serpentine input parameters including Beta-arch and Beta-serpentine thresholds used for predicting the Beta-serpentine structures for the repeat region from HnMl-01 = HnBNMP (SEQ ID NO: 4), EaMl-01= EaBNMP (SEQ ID NO: 8), EsMl-01 = EsBNMP (SEQ ID NO: 12), HeMl-01 = HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl- 01(SEQ ID NO: 20), as well as the maximum number of beta-arches per residue and the total number of predicted beta-serpentine structures for each construct's repeat region. Also shown is a diagram of the highest-scoring Beta-serpentine structure, with amino acid residues noted at the sequence terminals, and the corresponding BSS score for each construct's repeat region.
[0034] Figure 4: Beta-serpentine structures predicted by Beta-Serpentine for the repeat region of the construct sequences HnMl-01= HnBNMP (SEQ ID NO: 4), EaM l-01= EaBNMP (SEQ ID NO: 8), EsMl-01 = EsBNMP (SEQ ID NO: 12), HeMl-01 = HeBNMP (SEQ ID NO: 16) and MiBNMP = MiM l-01(SEQ ID NO: 20). Amino acids with a light grey background are predicted to form a beta-strand, amino acids with a dark grey background are predicted to form a beta-arc and amino acids with a white background are undefined. Amino acids in bold typeface indicate that the structural prediction is statistically significant (p-value < 0.001) at the indicated position.
[0035] Figure 5: Muscle alignment of the repeats from the repeat region of BNMP constructs HnMl-01 = HnBNMP (SEQ ID NO: 4), EaMl-01= EaBNMP (SEQ ID NO: 8), EsMl-01 = EsBNMP (SEQ ID NO: 12), HeMl-01= HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl- 01(SEQ ID NO: 20) identified by RADAR. The consensus is indicated below the alignment using a bar plot with black indicating consensus at an amino acid position and lighter greys indicating decreasing consensus. The consensus sequence of the repeat is positioned below the consensus graph with consensus positions showing the majority amino acid or 'X' if there is no consensus at a specific amino acid position. Figure 6: Emboss pepwindow of each repeat region, with a peptide window length of 11 amino acids, implies alternating regions of hydrophilicity and hydrophobicity in constructs HnMl-01= HnBNMP (SEQ ID NO: 4), EaMl-01= EaBNMP (SEQ ID NO: 8), EsMl-01 = EsBNMP (SEQ ID NO: 12), HeMl-01 = HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl- 01(SEQ ID NO: 20). The more positive the hydropathy value the more hydrophobic the region the more negative the hydropathy value the more hydrophilic the region.
[0036] Figure 7: The percentage identity of the signal peptide, N-terminal region, signal peptide + N-terminal region, charged region, repeat region, C-terminal region, the native predicted BNMP amino acid sequences HnMl = HnBNMP (SEQ ID NO: 2), EaMl = EaBNMP (SEQ ID NO: 6), EsMl = EsBNMP (SEQ ID NO: 10), HeMl = HeBNMP (SEQ ID NO: 14) and MiBNMP = MiM l(SEQ ID NO: 18) and the construct sequences HnMl-01= HnBNMP (SEQ ID NO: 4), EaMl-01= EaBNMP (SEQ ID NO: 8), EsMl-01= EsBNMP (SEQ ID NO: 12), HeMl-01 = HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl-01(SEQ ID NO: 20).
[0037] Figure 8: Wettability of BNMPs (HnMl-01= HnBNMP (SEQ ID NO: 4), EaMl-01 = EaBNMP (SEQ ID NO: 8), EsMl-01= EsBNMP (SEQ ID NO: 12), HeMl-01= HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl-01(SEQ ID NO: 20)) upon contact and after 60 seconds. Comparison of water contact angle profile of BNMPs coated on a glass substrate with and without ethanol treatment and the negative control which did not contain protein.
[0038] Figure 9: Stabilisation of oil in water emulsion by BNMPs (HnM l-01= HnBNMP (SEQ ID NO: 4), EaMl-01 = EaBNMP (SEQ ID NO: 8), EsMl-01= EsBNMP (SEQ ID NO: 12), HeMl- 01= HeBNMP (SEQ ID NO: 16) and MiBNMP = MiMl-01(SEQ ID NO: 20)). 1% solution of BNMPs in 2M guanidinium hydrochloride with deionised water at 0 and 4 days.
[0039] Figure 10: Hydrogel formation. Photographs of inverted tubes demonstrating hydrogel formation by BNMPs (HnMl-01= HnBNMP (SEQ ID NO: 4), EaMl-01= EaBNMP (SEQ ID NO: 8), EsMl-01 = EsBNMP (SEQ ID NO: 12), HeMl-01= HeBNMP (SEQ ID NO: 16) and MiBNMP = MiM l-01(SEQ ID NO: 20)).
[0040] DETAILED DESCRIPTION OF THE INVENTION
[0041] DEFINITIONS
[0042] The following definitions are presented to better define the present invention and as a guide for those of ordinary skill in the art of the practice of the present invention. Unless otherwise specified, all technical and scientific terms used herein are to be understood as having the same meanings as is understood by one of ordinary skill in the relevant art to which this disclosure pertains.
[0043] The term "comprising" as used in this specification and claims means "consisting at least in part of;" that is to say when interpreting statements in this specification and claims which include "comprising", the features prefaced by this term in each statement all need to be present but other features can also be present. Related terms such as "comprise" and "comprised" are to be interpreted in a similar manner.
[0044] The term "consisting essentially of" as used herein means the specified materials or steps and those that do not materially affect the basic and novel characteristic(s) of the claimed invention.
[0045] The term "consisting of" as used herein means the specified materials or steps of the claimed invention, excluding any element, step, or ingredient not specified in the claim.
[0046] The term "vector" as used herein refers to any type of polynucleotide molecule that may be used to manipulate genetic material so that it can be amplified, replicated, manipulated, partially replicated, modified and / or expressed, but not limited thereto. In some embodiments a vector may be used to transport a polynucleotide comprised in that vector into a cell or organism.
[0047] The term "polynucleotide(s)," as used herein, means a single or double-stranded deoxyribonucleotide or ribonucleotide polymer of any length, and includes as non-limiting examples, coding and non-coding sequences of a gene, sense and antisense sequences, exons, introns, genomic DNA, cDNA, pre-mRNA, mRNA, rRNA, siRNA, miRNA, tRNA, ribozymes, recombinant polynucleotides, isolated and purified naturally occurring DNA or RNA sequences, synthetic RNA and DNA sequences, nucleic acid probes, primers, fragments, genetic constructs, vectors and modified polynucleotides. Reference to nucleic acids, nucleic acid molecules, nucleotide sequences and polynucleotide sequences is to be similarly understood.
[0048] The term "gene" as used herein refers to the biologic unit of heredity, self-reproducing and located at a definite position (locus) on a particular chromosome. In one embodiment the particular chromosome is a eukaryotic or bacterial chromosome. The term bacterial chromosome is used interchangeably herein with the term bacterial genome. The term "endogenous" as used herein refers to a constituent of a cell, tissue or organism that originates or is produced naturally within that cell, tissue or organism. An "endogenous" constituent may be any constituent including but not limited to a polynucleotide, a polypeptide including a non-ribosomal polypeptide, a fatty acid or a polyketide, but not limited thereto.
[0049] The term "exogenous" as used herein refers to any constituent of a cell, tissue or organism that does not originate or is not produced naturally within that cell, tissue or organism. An exogenous constituent may be, for example, a polynucleotide sequence that has been introduced into a cell, tissue or organism, or a polypeptide expressed in that cell, tissue or organism from that polynucleotide sequence.
[0050] "Naturally occurring" as used herein with reference to a polynucleotide sequence according to the invention refers to a primary polynucleotide sequence that is found in nature. A synthetic polynucleotide sequence that is identical to a wild polynucleotide sequence is, for the purposes of this disclosure, considered a naturally occurring sequence. What is important for a naturally occurring polynucleotide sequence is that the actual sequence of nucleotide bases that comprise the polynucleotide is found or known from nature.
[0051] For example, a wild-type polynucleotide sequence is a naturally occurring polynucleotide sequence, but not limited thereto. A naturally occurring polynucleotide sequence also refers to variant polynucleotide sequences as found in nature that differ from wild type. For example, allelic variants and naturally occurring recombinant polynucleotide sequences due to hybridization or horizontal gene transfer, but not limited thereto.
[0052] "Non-naturally occurring" as used herein with reference to a polynucleotide sequence according to the invention refers to a polynucleotide sequence that is not found in nature. Examples of non-naturally occurring polynucleotide sequences include artificially produced mutant and variant polynucleotide sequences, made for example by point mutation, insertion, or deletion, but not limited thereto. Non-naturally occurring polynucleotide sequences also include chemically evolved sequences. What is important for a non- naturally occurring polynucleotide sequence according to the invention is that the actual sequence of nucleotide bases that comprise the polynucleotide is not found or known from nature. The term, "wild type" when used herein with reference to a polynucleotide refers to a naturally occurring; non-mutant form of a polynucleotide. A mutant polynucleotide means a polynucleotide that has sustained a mutation as known in the art, such as point mutation, insertion, deletion, substitution, amplification or translocation, but not limited thereto.
[0053] The term, "wild type" when used herein with reference to a polypeptide refers to a naturally occurring, non-mutant form of a polypeptide. A wild-type polypeptide is a polypeptide that is capable of being expressed from a wild-type polynucleotide.
[0054] The term "coding sequence" or "open reading frame" (ORF) refers to the sense strand of a genomic DNA sequence or a cDNA sequence that is capable of producing a transcription product and / or a polypeptide under the control of appropriate regulatory sequences. The CDS is identified by the presence of a 5' translation start codon and a 3' translation stop codon. When inserted into a genetic construct or an expression cassette, a "coding sequence" (CDS) is capable of being expressed when it is operably linked to a promoter sequence and / or other regulatory elements.
[0055] "Operably-linked" means that the sequence to be expressed is placed under the control of regulatory elements.
[0056] "Regulatory elements" as used herein refers to any nucleic acid sequence element that controls or influences the expression of a polynucleotide insert from a vector, genetic construct or expression cassette and includes promoters, transcription control sequences, translation control sequences, origins of replication, tissue-specific regulatory elements, temporal regulatory elements, enhancers, polyadenylation signals, repressors and terminators. Regulatory elements can be "homologous" or "heterologous" to the polynucleotide insert to be expressed from a genetic construct, expression cassette or vector as described herein. When a genetic construct, expression cassette or vector as described herein is present in a cell, a regulatory element can be "endogenous", "exogenous", "naturally occurring" and / or "non-naturally occurring" with respect to cell.
[0057] The term "noncoding region" refers to untranslated sequences that are upstream of the translational start site and downstream of the translational stop site. These sequences are also referred to respectively as the 5' UTR and the 3' UTR. These regions include elements required for transcription initiation and termination and for regulation of translation efficiency. Terminators are sequences, which terminate transcription, and are found in the 3' untranslated ends of genes downstream of the translated sequence. Terminators are important determinants of mRNA stability and in some cases have been found to have spatial regulatory functions.
[0058] The term "promoter" refers to non-transcribed cis-regulatory elements upstream of the coding region that regulate the transcription of a polynucleotide sequence. Promoters comprise cis-initiator elements which specify the transcription initiation site and conserved boxes. In one non-limiting example, bacterial promoters may comprise a "Pribnow box" (also known as the -10 region), and other motifs that are bound by transcription factors and promote transcription. Promoters can be homologous or heterologous with respect to polynucleotide sequence to be expressed. When the polynucleotide sequence is to be expressed in a cell, a promoter may be an endogenous or exogenous promoter. Promoters can be constitutive promoters, inducible promoters or regulatable promoters as known in the art.
[0059] "Homologous" as used herein with reference to polynucleotide regulatory elements, means a polynucleotide regulatory element that is a native and naturally occurring polynucleotide regulatory element. A homologous polynucleotide regulatory element may be operably linked to a polynucleotide of interest such that the polynucleotide of interest can be expressed from a vector according to the invention.
[0060] "Homologous" as used herein with reference to polynucleotide or polypeptide in a host organism means that the polynucleotide or polypeptide is a native and naturally occurring polynucleotide or polynucleotide within that host organism. A homologous polynucleotide may be operably linked to a homologous or heterologous regulatory element so that a homologous polypeptide may be expressed from a vector comprising the homologous polynucleotide as described herein.
[0061] "Heterologous" as used herein with reference to polynucleotide regulatory elements, means a polynucleotide regulatory element that is not a native and naturally occurring polynucleotide regulatory element. A heterologous polynucleotide regulatory element is not normally associated with the CDS to which it is operably linked. A heterologous regulatory element may be operably linked to a polynucleotide of interest such that the polynucleotide of interest can be expressed from a polynucleotide or vector according to the invention. Such promoters may include promoters normally associated with other genes, ORFs or coding regions, and / or promoters isolated from any other bacterial, viral, eukaryotic, or mammalian cell.
[0062] "Heterologous" as used herein with reference to a polynucleotide or polypeptide in a host organism (i.e., a "heterologous polynucleotide" or "heterologous polypeptide") means a polynucleotide or polypeptide that is not a native and naturally occurring polynucleotide or polypeptide in that host organism. A heterologous polynucleotide may be operably linked to a heterologous or homologous regulatory element so that a heterologous polypeptide may be expressed from a vector comprising the heterologous polynucleotide as described herein.
[0063] The terms "heterologously expressing" and "heterologous expression" mean the expression of a heterologous polypeptide in a host cell.
[0064] The phrases "functional variant or fragment thereof" and "functional portion thereof" of a polypeptide (and other grammatical variations of these phrases) both refer to a subsequence of the polypeptide that performs a function that is required for the biological activity or binding of that polypeptide and / or provides the three-dimensional structure of the polypeptide. The term may refer to a polypeptide, an aggregate of a polypeptide such as a dimer or other multimer, a fusion polypeptide, a polypeptide fragment, a polypeptide variant, or functional polypeptide derivative thereof that is capable of performing the polypeptide activity.
[0065] "Isolated" as used herein with reference to polynucleotide or polypeptide sequences describes a sequence that has been removed from its natural cellular environment. An isolated molecule may be obtained by any method or combination of methods as known and used in the art, including biochemical, recombinant, and synthetic techniques. The polynucleotide or polypeptide sequences may be prepared by at least one purification step.
[0066] "Isolated" when used herein in reference to a cell or host cell describes a cell or host cell that has been obtained or removed from an organism or from its natural environment and is subsequently maintained in a laboratory environment as known in the art. The term encompasses single cells, per se, as well as cells or host cells that are comprised in a cell culture and can include a single cell or single host cell.
[0067] The term "recombinant" refers to a polynucleotide sequence that is removed from sequences that surround it in its natural context and / or is recombined with sequences that are not present in its natural context. A "recombinant" polypeptide sequence is produced by translation from a "recombinant" polynucleotide sequence.
[0068] As used herein, the term "variant" refers to polynucleotide or polypeptide sequences different from the specifically identified sequences, wherein one or more nucleotides or amino acid residues are deleted, substituted, or added. Variants may be naturally occurring allelic variants or non-naturally occurring variants. Variants may be from the same or from other species and may encompass homologues, paralogues and orthologues. In certain embodiments, variants of the polypeptides useful in the invention have biological activities that are the same or similar to those of a corresponding wild type molecule; i.e., the parent polypeptides or polynucleotides.
[0069] In certain embodiments, variants of the polypeptides described herein have biological activities that are similar, or that are substantially similar to their corresponding wild type molecules. In certain embodiments the similarities are similar activity and / or binding specificity.
[0070] In certain embodiments, variants of polypeptides described herein have biological activities that differ from their corresponding wild type molecules. In certain embodiments the differences are altered activity and / or binding specificity.
[0071] The term "variant" with reference to polynucleotides and polypeptides encompasses all forms of polynucleotides and polypeptides as defined herein.
[0072] Variant polynucleotide sequences preferably exhibit at least 50%, at least 60%, preferably at least 70%, preferably at least 71%, preferably at least 72%, preferably at least 73%, preferably at least 74%, preferably at least 75%, preferably at least 76%, preferably at least 77%, preferably at least 78%, preferably at least 79%, preferably at least 80%, preferably at least 81%, preferably at least 82%, preferably at least 83%, preferably at least 84%, preferably at least 85%, preferably at least 86%, preferably at least 87%, preferably at least 88%, preferably at least 89%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, and preferably at least 99% identity to a sequence of the present invention. Identity is found over a comparison window of at least 8 nucleotide positions, preferably at least 10 nucleotide positions, preferably at least 15 nucleotide positions, preferably at least 20 nucleotide positions, preferably at least 27 nucleotide positions, preferably at least 40 nucleotide positions, preferably at least 50 nucleotide positions, preferably at least 60 nucleotide positions, preferably at least 70 nucleotide positions, preferably at least 80 nucleotide positions, preferably over the entire length of a polynucleotide used in or identified according to a method of the invention.
[0073] Polynucleotide variants also encompass those which exhibit a similarity to one or more of the specifically identified sequences that is likely to preserve the functional equivalence of those sequences, and which could not reasonably be expected to have occurred by random chance.
[0074] Polynucleotide sequence identity and similarity can be determined readily by those of skill in the art.
[0075] Variant polynucleotides also encompass polynucleotides that differ from the polynucleotide sequences described herein but that, as a consequence of the degeneracy of the genetic code, encode a polypeptide having similar activity to a polypeptide encoded by a polynucleotide of the present invention. A sequence alteration that does not change the amino acid sequence of the polypeptide is a "silent variation". Except for ATG (methionine) and TGG (tryptophan), other codons for the same amino acid may be changed by art recognized techniques, e.g., to optimise codon expression in a particular host organism.
[0076] Polynucleotide sequence alterations resulting in conservative substitutions of one or several amino acids in the encoded polypeptide sequence without significantly altering its biological activity are also included in the invention. A skilled artisan will be aware of methods for making phenotypically silent amino acid substitutions (see, e.g., Bowie et al., 1990, Science 247, 1306).
[0077] The term "variant" with reference to polypeptides also encompasses naturally occurring, recombinantly and synthetically produced polypeptides. Variant polypeptide sequences preferably exhibit at least 35%, preferably at least 40%, preferably at least 50%, preferably at least 60%, preferably at least 70%, preferably at least 71%, preferably at least 72%, preferably at least 73%, preferably at least 74%, preferably at least 75%, preferably at least 76%, preferably at least 77%, preferably at least 78%, preferably at least 79%, preferably at least 80%, preferably at least 81%, preferably at least 82%, preferably at least 83%, preferably at least 84%, preferably at least 85%, preferably at least 86%, preferably at least 87%, preferably at least 88%, preferably at least 89%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, and preferably at least 99% identity to a sequence of the present invention. Identity is found over a comparison window of at least 2 amino acid positions, preferably at least 3 amino acid positions, preferably at least 4 amino acid positions, preferably at least 5 amino acid positions, preferably at least 7 amino acid positions, preferably at least 10 amino acid positions, preferably at least 15 amino acid positions, preferably at least 20 amino acid positions, preferably over the entire length of a polypeptide used in or identified according to a method of the invention.
[0078] Polypeptide variants also encompass those which exhibit a similarity to one or more of the specifically identified sequences that is likely to preserve the functional equivalence of those sequences, and which could not reasonably be expected to have occurred by random chance.
[0079] Polypeptide sequence identity and similarity can be determined readily by those of skill in the art.
[0080] A variant polypeptide includes a polypeptide wherein the amino acid sequence differs from a polypeptide herein by one or more conservative amino acid or non-conservative substitutions, deletions, additions, or insertions which do not affect the biological activity of the peptide.
[0081] Conservative substitutions typically include the substitution of one amino acid for another with similar characteristics as known and used in the art.
[0082] Analysis of evolved biological sequences has shown that not all sequence changes are equally likely, reflecting at least in part the differences in conservative versus nonconservative substitutions at a biological level. For example, certain amino acid substitutions may occur frequently, whereas others are very rare. Evolutionary changes or substitutions in amino acid residues can be modelled by a scoring matrix also referred to as a substitution matrix. Such matrices are used in bioinformatics analysis to identify relationships between sequences and are known to the skilled worker.
[0083] Other variants include peptides with modifications which influence peptide stability. Such analogs may contain, for example, one or more non-peptide bonds (which replace the peptide bonds) in the peptide sequence. Also included are analogs that include residues other than naturally occurring L-amino acids, e.g., D-amino acids or non-naturally occurring synthetic amino acids, e.g., beta or gamma amino acids and cyclic analogs.
[0084] Substitutions, deletions, additions or insertions may be made by mutagenesis methods known in the art. A skilled worker will be aware of methods for making phenotypically silent amino acid substitutions. See for example Bowie et al., 1990, Science 247, 1306.
[0085] A polypeptide as used herein can also refer to a polypeptide that has been modified during or after synthesis, for example, by biotinylation, benzylation, glycosylation, phosphorylation, amidation, by derivatization using blocking / protecting groups and the like. Such modifications may increase stability or activity of the polypeptide.
[0086] The terms "modulate(s) expression", "modulated expression" and "modulating expression" of a polynucleotide or polypeptide, are intended to encompass the situation where genomic DNA corresponding to a polynucleotide to be expressed according to the invention is modified thus leading to modulated expression of a polynucleotide or polypeptide of the invention. Modification of the genomic DNA may be through genetic transformation or other methods known in the art for inducing mutations. The "modulated expression" can be related to an increase or decrease in the amount of messenger RNA and / or polypeptide produced and may also result in an increase or decrease in the activity of a polypeptide due to alterations in the sequence of a polynucleotide and polypeptide produced.
[0087] The terms "modulate(s) activity", "modulated activity" and "modulating activity" of a polynucleotide or polypeptide, are intended to encompass the situation where genomic DNA corresponding to a polynucleotide to be expressed according to the invention is modified thus leading to modulated expression of a polynucleotide or modulated expression or activity of polypeptide of the invention. Modification of the genomic DNA may be through genetic transformation or other methods known in the art for inducing mutations. The "modulated activity" can be related to an increase or decrease in the amount of messenger RNA and / or polypeptide produced and may also result in an increase or decrease in the functional activity of a polypeptide due to alterations in the sequence of a polynucleotide and polypeptide produced.
[0088] The term "sub repeat region" and grammatical variations thereof as used herein refers to a part or subsequence of a repeat region as described herein. As used herein the terminology "similar or substantially similar" refers to repeat sub regions of a polypeptide that form beta sheets as detected by Rapid Automatic Detection and Alignment of Repeats (RADAR) (Heger and Holm, 2000).
[0089] It is intended that reference to a range of numbers disclosed herein (for example 1 to 10) also incorporates reference to all related numbers within that range (for example, 1, 1.1, 2, 3, 3.9, 4, 5, 6, 6.5, 7, 8, 9 and 10) and also any range of rational numbers within that range (for example 2 to 8, 1.5 to 5.5 and 3.1 to 4.7) and, therefore, all sub-ranges of all ranges expressly disclosed herein are expressly disclosed. These are only examples of what is specifically intended and all possible combinations of numerical values between the lowest value and the highest value enumerated are to be considered to be expressly stated in this application in a similar manner.
[0090] DETAILED DESCRIPTION
[0091] The present invention generally relates to polypeptides from the nesting materials of solitary Colletid bees, and to functional portions, functional analogs, functional variants and / or functional derivatives thereof. The invention also relates generally to polynucleotides encoding such polypeptides, compositions comprising such polynucleotides and / or polypeptides, portions, analogs, variants and / or derivatives, to methods of making such polynucleotides and polypeptides including by heterologous expression of such polynucleotides in an appropriate isolated host cell, and to methods of using such polypeptides.
[0092] The polynucleotides and polypeptides as described herein are shown in Table 1.
[0093] Bees in the family Colletidae (Hymenoptera: Colletidae) produce a nesting material described as 'cellophane-like' (Almeida, E.A.B. Colletidae nesting biology (Hymenoptera : Apoidea). Apidologie 39, 16-29 (2008)). Collectively these bees are referred to as the "polyester" bees or the "plasterer bees" due to their habit of lining their nest cells with secretions that dry to a cellophane-like material.
[0094] Previously published results on nesting lining materials (also called "nest material" and "nesting materials" herein) produced by Colletid bees suggest the nesting material is a unique composite of lipid polymer and protein biopolymers. To better understand the basis for the observed properties, the inventors have undertaken an extensive investigation of Colletid bee nesting material seeking to identify the constituents of this material that are responsible for its surprising properties. Molecular analyses of the Dufour's, labial, mandibular glands, and nesting material of H. nubilosus, have identified a single major protein component of the bee nesting material of H. nubilosus (>10%), termed HnBNMP herein.
[0095] The inventors have employed a molecular approach (including genomic and transcriptomic sequencing and proteomics) to identify and characterize the nucleic and amino acid sequences of a number of Colletid bee nesting material polypeptides. As described herein, molecular analyses from a representative sample of Colletid bees from the genera Euryglossa, Hylaeus and Meroglossa have identified a number of polypeptides that are homologues of a major protein component of the bee nesting material from the H. nubilosus (>10%). These polypeptides are termed "bee nesting material polypeptides (BNMP)" herein.
[0096] The H. nubilosus bee nesting material polypeptide (HnBNMP) is a silk-like protein but is not similar or homologous to the silks known to be produced by Hymenoptera (ants, wasps, and bees). The dissimilarity is due to the properties of the HnBNMP and the way that this polypeptide is produced.
[0097] For example, larval honeybees (Apis mellifera') produce silk rich in proteins in their labial glands for end-capping cells prior to pupation. Honeybee silk is made of proteins that fold into o-helices, which then assemble into larger super secondary structures, coiled-coils. The honeybee proteins are composed of four small fibroin subunits ~30 kDa long and comprise ~30% alanine.
[0098] In contrast, HnBNMP from H. nubilosus takes a beta-sheet structure which is the same structure as in silk fibres made by the silkworm (Bombyx mori) or dragline silk of araneomorph spiders, but different to the o-helica I structure that dominates larval honeybee silk. The HnBNMP is rich in glutamine and serine, which is not a feature of either honeybee silk, silkworm silk, or spider silk. Although one asparagine-rich silk protein (asparagine and glutamine are similar amino acids) has been reported from a distantly related parasitic wasp (Cotesia glomerata), and the ability of glutamine to form beta-sheets by hydrogen bonding is important for the mechanical properties of spider silks, silk proteins rich in glutamine are not known.
[0099] Using a transcriptomic sequencing, and proteomics approach the inventors have also determined the nucleic and amino acid sequences of the BNMPs from a representative sample of Colletid bees, including numerous sequences rich in glutamine and serine. Based on the work disclosed herein, the inventors believe that the bee nesting material polypeptides (BNMPs) or their isoforms from each bee as described herein are comprised in, or comprise, that bee's nesting material. Without wishing to be bound by theory, the inventors believe that the BNMPs from each bee as described herein or at least a portion thereof, impart important structural and functional properties to that bee's nest material.
[0100] The inventors believe that the BNMPs and functional variants, analogues and derivatives thereof as described herein, have numerous applications in materials in which known synthetic polymers and biopolymers (including silk proteins) are normally employed.
[0101] As described herein, the inventors are the first to provide a nesting material protein or part thereof from a representative sample of Colletid bees. To achieve this goal, the inventors carried out transcriptome analysis of RNA extracted from five Colletid bees, Hylaeus nubilosus, Hylaeus euxanthus Euryglossa adelaidae, Euryglossa subsericea and Meroglossa impressifrons penetrate. From this work they identified five predicted coding sequences (termed "native" herein) that are each comprised in a BNMP, the BNMP itself being comprised in or comprising the nest material of one of the above bee species. Without wishing to be bound by theory the inventors believe that the organization of the genes coding for Colletid bee nesting material proteins is modular and is expressed as various isoforms, with the ultimate nesting material polypeptides in each bee comprising various combinations of isoforms of their respective nest material proteins. Again, without wishing to be bound by theory the inventors believe that the BNMPs described herein are short isoforms of these proteins.
[0102] Described herein are the nucleic acid sequences (cDNAs) encoding the predicted amino acid sequences from H. nubilosus, H. euxanthus E. adelaidae, E. subsericea and M. impressifrons penetrate. Each of these native amino acid sequences is comprised in, or comprises, the nesting material from each one of the above bees respectively. These nucleic acid sequences are disclosed herein in Table 1 and below as:
[0103] SEQ ID NO: 1 - Nucleic acid sequence (cDNA) encoding the native amino acid sequence of SEQ ID NO: 2 from H. nubilosus. This amino acid sequence is referred to herein as an HnMl polypeptide.
[0104] SEQ ID NO: 5 - Nucleic acid sequence (cDNA) encoding the native amino acid sequence of SEQ ID NO: 6 from E. adelaidae. This amino acid sequence is referred to herein as an EaMl polypeptide. SEQ ID NO: 9 - Nucleic acid sequence (cDNA) encoding the native amino acid sequence of SEQ ID NO: 10 from E. subsericea. This amino acid sequence is referred to herein as an EsMl polypeptide.
[0105] SEQ ID NO: 13 - Nucleic acid sequence (cDNA) encoding the native amino acid sequence of SEQ ID NO: 14 from H. euxanthus. This amino acid sequence is referred to herein as an HeMl polypeptide.
[0106] SEQ ID NO: 17 - Nucleic acid sequence (cDNA) encoding the native amino acid sequence of SEQ ID NO: 18 from M. impressifrons penetrate. This amino acid sequence is referred to herein as an MiMl polypeptide.
[0107] Collectively, the above cDNAs are termed herein the "Module 1" coding sequences of each respective bee, e.g., SEQ ID NO: 1 encodes the H. nubilosus Module 1 polypeptide, SEQ ID NO: 5 encodes the E. adelaidae Module 1 polypeptide, etc. As noted previously, the inventors believe that the polypeptides expressed from each module 1 coding sequence is a short isoform of the BNMP from the bee from which it was obtained.
[0108] As described herein, the above cDNAs may be modified for recombinant expression in host cell, preferably an isolated host cell. These modified cDNA expression sequences are described herein as:
[0109] SEQ ID NO: 3 - Modified nucleic acid sequence (cDNA) from H. nubilosus encoding the recombinantly expressed amino acid sequence of SEQ ID NO: 4. This amino acid sequence is referred to herein as an HnMl-01 polypeptide.
[0110] SEQ ID NO: 7 - Modified nucleic acid sequence (cDNA) from E. adelaidae encoding the recombinantly expressed amino acid sequence of SEQ ID NO: 8. This amino acid sequence is referred to herein as an EaMl-01 polypeptide.
[0111] SEQ ID NO: 11 - Modified nucleic acid sequence (cDNA) from E. subsericea encoding the recombinantly expressed amino acid sequence of SEQ ID NO: 12. This amino acid sequence is referred to herein as an EsMl-01 polypeptide.
[0112] SEQ ID NO: 15 - Modified nucleic acid sequence (cDNA) from H. euxanthus encoding the recombinantly expressed amino acid sequence of SEQ ID NO: 16. This amino acid sequence is referred to herein as an HeMl-01 polypeptide. SEQ ID NO: 19 - Modified nucleic acid sequence (cDNA) from M. impressifrons penetrate encoding the recombinantly expressed amino acid sequence of SEQ ID NO: 20. This amino acid sequence is referred to herein as an MiMl-01 polypeptide.
[0113] Also disclosed herein is SEQ ID NO: 21, the nucleic acid sequence encoding a polypeptide consensus sequence of the amino acid sequences of SEQ ID Nos: 2, 6, 10, 14 and 18. This consensus amino acid sequence is disclosed herein as SEQ ID NO: 22 and termed CsBNMP.
[0114] Also disclosed herein is SEQ ID NO: 23, the nucleic acid sequence encoding a polypeptide comprising SEQ ID NO: 24, a consensus sequence of the N-terminal amino acid sequences of SEQ ID Nos: 2, 6, 10, 14 and 18. This consensus amino acid sequence is disclosed herein as SEQ ID NO: 22 and termed CsNTBNMP.
[0115] All of the nucleic acid and amino acid sequences disclosed herein may be found in Table 1.
[0116] Also disclosed herein are the nucleic acid coding sequences (cDNAs) and native amino acid sequences of the N-terminal region from "Module 1" of each respective bee, these cDNAs that have been modified for recombinant expression and the recombinantly expressed amino acid sequences (Table 1).
[0117] Also disclosed herein are the nucleic acid coding sequences (cDNAs) and the native amino acid sequences of the signal peptide + N-terminal region from "Module 1" of each respective bee, these cDNAs that have been modified for recombinant expression and the recombinantly expressed amino acid sequences (Table 1).
[0118] Polypeptide characterization
[0119] The applicants are also the first to provide methods of heterologous expression of a silklike protein from H. nubilosus, E. adelaidae, E. subsericea, H. euxanthus, and M. impressifrons penetrate in an isolated host cell.
[0120] The gene for the full-length Hylaeus nubilosus was first identified by the inventors as described in the examples herein and in Table 1 as SEQ ID NO: 25 (nucleic acid) and SEQ ID NO: 26 (amino acid). The coding sequence for the short isoform of BNMP, such as HnMl (SEQ ID NO: 2), from H. nubilosus was then identified by mapping transcripts from H. nubilosus onto the assembled scaffold containing the bee nest material gene, transcripts of various length accurately mapped to the bee nest material gene implying diverse splice variants which produce isoforms of the H. nubilosus bee nesting material gene (HnBNMG) which differ in length and amino acid composition. Several full-length RIMA and cDNA sequences from H. nubilosus mapped onto the first and second exon of the H. nubilosus nest material gene, revealing a short-expressed isoform of the HnBNMG named HnMl. The HnBNMP module 1 isoform contains the conserved N-terminal region, a glutamine and serine-rich repeat region and a C-terminal region, which is repetitive or enriched in these residues.
[0121] As shorter proteins are more amenable to expression and purification but may still be expected to have some properties of the native bee nest material. Similar-sized homologs were searched for in the assembled transcriptomes of several additional bee species. The coding sequences for the module 1 BNMPs from E. adelaidae, E. subsericea, H. euxanthus, and M. impressifrons penetrate disclosed herein were identified as homologues of the module 1 HnBNMP as described in the examples herein. Clear homology was present in the N-terminal region of the BNMPs from E. adelaidae, E. subsericea, H. euxanthus, and M. impressifrons penetrate as determined by a protein blast search using the first 144 amino acids of HnBNMP as a query sequence. Several BNMP isoforms were identified for each bee species, varying in length and amino acid composition but maintaining a glutamine and serine-rich repeat region and N-terminal homology. BNMP isoforms from E. adelaidae, E. subsericea, H. euxanthus, and M. impressifrons penetrate were selected based on their short length. The nesting material nucleic acid and amino acid sequences disclosed herein were selected to capture representative variations of Colletid bee nest material polypeptides. EaMl, EsMl, HeMl and MiMl for nest materials from E. adelaidae, E. subsericea, H. euxanthus, and M. impressifrons penetrate, respectively.
[0122] The bee nest material polypeptide constructs described herein are 216, 232, 257, 340 and 376 amino acids in length for HnMl-01 (SEQ ID NO:4), HeMl-01 (SEQ ID NO:16), EaMl-01 (SEQ ID NO:8), MiMl-01 (SEQ ID NO:20) and EsMl-01 (SEQ ID NO: 12), respectively (Figure 1, Figure 2). The wild-type sequences, including their signal peptides, are 232, 248, 273, 356 and 392 amino acids in length for HnMl (SEQ ID NO:2), HeMl (SEQ ID NO:14), EaMl (SEQ ID NO:6), MiMl (SEQ ID NO:18) and EsMl (SEQ ID NO:10), respectively. The native BNMPs are highly polar, with polar amino acids making up 62%, 65%, 69%, 69% and 76% in the wild-type sequences HeMl, EaMl, MiMl, HnMl and EsMl, respectively (Figure 2). The BNMP constructs are also highly polar, with polar amino acids making up between 66%, 68%, 72%, 74% and 78% of amino acid residues in the constructs HeMl-01, EaMl-01, MiMl-01, HnMl-01 and EsMl-01, respectively. (Figure 2).
[0123] Also, without wishing to be bound by theory, the inventors believe that the aggregation propensity of the bee nesting material polypeptides is likely the result of the length of the repeat region, the high proportion of serine and glutamine residues and the repeating pattern of higher and lower hydropath regions in the polypeptide resulting from the hydrophobic residues dispersed in between the repeat region at semi-regular intervals (Figure 2, 5 and 6). The BNMPs are amphipathic as is implied by the ability to stabilize an oil in water emulsion this property is also likely the result of the alternative regions of comparative high and low hydropathy in the repeat region of the BNMPs. BNMPs form beta-sheet like structures which are known to facilitate aggregation in silks and insoluble fibrils or plaques associated with diseases like Alzheimer's, Parkinson's and amyloidosis through the formation of intermolecular hydrogen bonds.
[0124] Also, without wishing to be bound by theory, the inventors believe that the aggregation propensity of the bee nesting material polypeptides is likely the result of aggregation in repeat region of the BNMPs as plausible sites of nucleation were identified using the BetaSerpentine software (Bondarev et al., 2018) which is used to identify Beta-serpentine structures which form intermolecular hydrogen bonds driving aggregation in BNMPs by forming super-pleated beta-structures. Predicted repeat regions and beta-sheet formation for EaMl-01, EsMl-01, HeMl-01, HnMl-01 and MiMl-01 are shown in (Figures 1, 3, 4 and 5). The repeat region comprises alternating regions of hydrophilicity and hydrophobicity in HnMl-01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01 (Figure 6).
[0125] Avoiding protein aggregation and maintaining protein stability homeostasis in a cell is energy-expensive; disordered regions and those prone to aggregation can play fundamental roles in receptor binding or dynamic protein structure. It is important to avoid spurious aggregations that result in disease. BNMPs must balance solubility with aggregation so that the bees that produce them can reproducibly and reliably express these structural proteins for nest building. The N-terminal regions of silks play key roles in controlling aggregation in silks from both Bombyx and spiders. Also, without wishing to be bound by theory, the inventors believe that the N-terminal region of the BNMPs is highly conserved as it is likely to play an important role in controlling the aggregation of the BNMPs. The ability of the conserved N-terminal region of the BNMPs to control aggregation decreases with the length of the repeat region, making longer polypeptides more likely to aggregate, resulting in molecules less amenable to long-term storage, which is often required in the process of making materials. BNMPs must be optimised to avoid spurious aggregation while being able to produce structural proteins reproducibly, which requires controlled and systematic aggregation. Expression, which relies on solubility and controlled aggregation, is useful for forming materials such as hydrogels and other silks in expression vectors and at scale.
[0126] Also, without wishing to be bound by theory, the inventors believe that the aggregation propensity of the bee nesting material polypeptides is likely the result of both the N- terminal region and the repeat region of the BNMPs which are fundamental to its properties as they form hydrogels in a controllable and reproducible manner as disclosed herein (Figure 10). The BNMPs have greater than 70%, and up to 90%, amino acid sequence identity in the combined Signal Peptide and N-terminal region as this region likely performs a function reliant on a conserved sequence order for signal peptide processing and modulation of aggregation. The percent amino acid identity of the N- terminal regions of the BNMPs ranges from 67% to 85%. This is the region with the second most identity after the signal peptide region which ranges from 71% to 100% identity. The N-terminal region is the most conserved region maintained in the mature protein, implying it has a conserved role in bee nest material formation (Figure 7).
[0127] The repeat regions of the BNMPs have a relatively low percentage identity, ranging from 19% to 62%. The repeat region of the native and construct BNMPs are identical in sequence and vary in length and are 110, 121, 146, 204 and 258 amino acids long for HnMl and HnMl-01, HeMl and HeMl-01, EaMl and EaMl-01, MiMl and MiMl-01 and EsMl and EsMl-01, respectively (Figure 2 8i 7). The repeat region of the BNMPs contains a high proportion of glutamine and serine residues with EaMl and EaMl-01, HeMl and HeMl-01, HnMl and HnMl-01, EsMl and EsMl-01 and MiMl and MiMl-01 composed of 14%, 22%, 23%, 34% and 36% glutamine residues and 23%, 18%, 28%, 31% and 29% serine, respectively. The repeat region of the BNMPs contains a high proportion of polar residues comprising 70%, 76%, 79%, 80% and 84% in EaMl and EaMl-01, HeMl and He-01, MiMl and Mi-01, HnMl and Hn-01 and EsMl and Es-01, respectively. The repeat region of the BNMPs comprise 30%, 24%, 21%, 20% and 16% non-polar residues in EaMl and Ea-01, HeMl and He-01, MiMl and Mi-01, HnMl and Hn-01 and EsMl and Es-01, respectively. The repeat regions of the BNMPs comprise 10%, 13%, 18%, 22% and 23% charged residues in MiMl and MiMl-01, EsMl and EsMl-01, EaMl and EaMl- 01, HnMl and HnMl-01 and HeMl and HeMl-01, respectively (Figure 2).
[0128] Protein expression To confirm the function and characterise the properties of the BNMPs described herein, the inventors constructed a series of expression vectors. These vectors were transformed into an appropriate host for heterologous production of the BNMPs.
[0129] The successful production of recombinant BNMPs as described herein confirms that heterologous expression of the "Module 1" BNMPs is viable using a recombinant biosynthetic system. The recombinant BNMPs produced as described herein are identified as SEQ ID NO: 4, 8, 12, 16 and 20.
[0130] Accordingly, in one aspect the present invention relates to an isolated polynucleotide that encodes a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0131] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0132] In one embodiment the polynucleotide encodes a polypeptide comprising, consisting essentially of or consisting of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0133] In one embodiment the polynucleotide encodes a polypeptide comprising a repeat region of about, preferably of, 110 to 258 amino acids in length.
[0134] In one embodiment the repeat region comprises at least two regions of alternating relative hydrophobicity and hydrophilicity.
[0135] In one embodiment the region of relative hydrophobicity comprises at least 20% to 30% hydrophobic amino acid residues.
[0136] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is from the nesting material of a Colletid bee.
[0137] In one embodiment the polypeptide is expressed in the Salivary or Dufour's gland of the Colletid bee.
[0138] In one embodiment the polypeptide is a component of the nesting material from a Colletid bee. In one embodiment the Colletid bee is selected from the following genera : Hylaeus spp., Euryglossa spp. and Meroglossa spp.
[0139] In one embodiment the Colletid bee is selected from Hylaeus nubilosus, Hylaeus euxanthus Euryglossa adelaidae, Euryglossa subsericea and Meroglossa impressifrons penetrate.
[0140] In one embodiment the Colletid bee is a Hylaeus spp. bee. In one embodiment the Hylaeus spp. bee is H. nubilosus or H. euxanthus.
[0141] In one embodiment the nesting material is from a Hylaeus spp. bee. In one embodiment the Hylaeus spp. bee is H. nubilosus or H. euxanthus.
[0142] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is an H. nubilosus bee nesting material polypeptide (HnBNMP).
[0143] In one embodiment the HnBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2 or SEQ ID NO: 4, preferably SEQ ID NO: 2, preferably SEQ ID NO: 4.
[0144] In one embodiment the HnBNMP comprises, consists essentially of or consists of SEQ ID NO: 2 or SEQ ID NO: 4, preferably SEQ ID NO: 2, preferably SEQ ID NO: 4.
[0145] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an H. euxanthus bee nesting material polypeptide (HeBNMP).
[0146] In one embodiment the HeBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 14 or SEQ ID NO: 16, preferably SEQ ID NO: 14, preferably SEQ ID NO: 16.
[0147] In one embodiment the HeBNMP comprises, consists essentially of or consists of SEQ ID NO: 14 or SEQ ID NO: 16, preferably SEQ ID NO: 14, preferably SEQ ID NO: 16.
[0148] In one embodiment the nesting material is from a Euryglossa spp. bee. In one embodiment the Euryglossa spp. bee is E. adelaidae or E. subsericea. In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an E. adelaidae bee nesting material polypeptide (EaBNMP).
[0149] In one embodiment the EaBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 6or SEQ ID NO: 8, preferably SEQ ID NO: 6, preferably SEQ ID NO: 8.
[0150] In one embodiment the EaBNMP comprises, consists essentially of or consists of SEQ ID NO: 6 or SEQ ID NO: 8, preferably SEQ ID NO: 6, preferably SEQ ID NO: 8.
[0151] In one embodiment the polynucleotide encodes a polypeptide that is an E. subsericea bee nesting material polypeptide (EsBNMP).
[0152] In one embodiment the EsBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 10 or SEQ ID NO: 12, preferably SEQ ID NO: 10, preferably SEQ ID NO: 12.
[0153] In one embodiment the EsBNMP comprises, consists essentially of or consists of SEQ ID NO: 10 or SEQ ID NO: 12, preferably SEQ ID NO: 10, preferably SEQ ID NO: 12.
[0154] In one embodiment the nesting material is from a Meroglossa spp. bee. In one embodiment the Meroglossa spp. bee is M. impressifrons penetrate.
[0155] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an M. impressifrons penetrate bee nesting material polypeptide (MiBNMP).
[0156] In one embodiment the MiBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 18 or SEQ ID NO: 20, preferably SEQ ID NO: 18, preferably SEQ ID NO: 20.
[0157] In one embodiment the MiBNMP comprises, consists essentially of or consists of SEQ ID NO: 18 or SEQ ID NO: 20, preferably SEQ ID NO: 18, preferably SEQ ID NO: 20.
[0158] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is a consensus sequence bee nesting material polypeptide (CsBNMP). In one embodiment the CsBNMP is an amino acid consensus sequence from the N- terminal region of at least one, preferably at least two, three, four, preferably at least five colletid bee nesting material polypeptides.
[0159] In one embodiment the CsBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 22 or SEQ ID NO: 24, preferably SEQ ID NO: 22, preferably SEQ ID NO: 24.
[0160] In one embodiment the CsBNMP comprises, consists essentially of or consists of SEQ ID NO: 22 or SEQ ID NO: 24, preferably SEQ ID NO: 22, preferably SEQ ID NO: 24.
[0161] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is at least partially capable of modifying the hydrophobicity of a material.
[0162] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is amphipathic.
[0163] In one embodiment the polypeptide or functional portion thereof is an isolated polypeptide or functional portion thereof.
[0164] In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0165] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0166] In one embodiment the polynucleotide is modified for expression in a heterologous host cell. In one embodiment the heterologous host cell is a prokaryotic host cell, preferably a bacterial host cell, preferably E. coli.
[0167] In one embodiment the polynucleotide is comprised in a vector or nucleic acid expression construct. In one embodiment the vector or nucleic acid expression construct further comprises a heterologous regulatory element.
[0168] The isolated polynucleotide molecules described herein can be isolated from a biological sample using a variety of techniques known to those of ordinary skill in the art. By way of example, such polynucleotides can be isolated through use of the polymerase chain reaction (PCR) as known in the art. The nucleic acid molecules described can be amplified using primers, as defined herein, derived from the polynucleotide sequences as described herein.
[0169] Further methods for isolating polynucleotides include use of all, or portions of, a polynucleotide as described herein as hybridization probes. The technique of hybridizing labelled polynucleotide probes to polynucleotides immobilized on solid supports such as nitrocellulose filters or nylon membranes, can be used to screen genomic or cDNA libraries. Similarly, probes may be coupled to beads and hybridized to the target sequence. Isolation can be affected using known art protocols such as magnetic separation. The choice of appropriately stringent hybridization and wash conditions is believed to be within the skill of those in the art.
[0170] Polynucleotide fragments may be produced by techniques well-known in the art such as restriction endonuclease digestion and oligonucleotide synthesis.
[0171] A partial polynucleotide sequence may be used as a probe, in methods well-known in the art to identify the corresponding full-length polynucleotide sequence in a sample. Such methods include PCR-based methods, 5'RACE and hybridization-based method, computer / database-based methods as known in the art. Detectable labels such as radioisotopes, fluorescent, chemiluminescent and bioluminescent labels may be used to facilitate detection. Inverse PCR also permits acquisition of unknown sequences, flanking the polynucleotide sequences disclosed herein, starting with primers based on a known region as known and used in the art. The method uses several restriction enzymes to generate a suitable fragment in the known region of a gene. The fragment is then circularized by intramolecular ligation and used as a PCR template. Divergent primers are designed from the known region. In order to physically assemble full-length clones, standard molecular biology approaches can be utilized as known in the art. Primers and primer pairs which allow amplification of polynucleotides of the invention, are also contemplated as embodiments disclosed herein.
[0172] Variants (including orthologues) may be identified by the methods described. Variant polynucleotides may be identified using PCR-based methods as known in the art. Typically, the polynucleotide sequence of a primer, useful to amplify variants of polynucleotide molecules by PCR, may be based on a sequence encoding a conserved region of the corresponding amino acid sequence. Further methods for identifying variant polynucleotides include use of all, or portions of the specified polynucleotides as hybridization probes to screen genomic or cDNA libraries as described above. Typically probes based on a sequence encoding a conserved region of the corresponding amino acid sequence may be used. Hybridization conditions may also be less stringent than those used when screening for sequences identical to the probe.
[0173] In another aspect the invention relates to an isolated polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0174] In one embodiment the polypeptide or functional portion thereof comprises consists essentially of or consists of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0175] In one embodiment the polypeptide comprises, consists essentially of or consists of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0176] In one embodiment the polypeptide is encoded by a polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0177] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0178] In one embodiment the polypeptide comprises a repeat region of about, preferably of, 110 to 258 amino acids in length.
[0179] In one embodiment the repeat region comprises at least two regions of alternating relative hydrophobicity and hydrophilicity.
[0180] In one embodiment the region of relative hydrophobicity comprises at least 20% to 30% hydrophobic amino acid residues.
[0181] In one embodiment the polypeptide or functional portion thereof is from the nesting material of a Colletid bee. In one embodiment the polypeptide is expressed in the Salivary or Dufour's gland of the Colletid bee.
[0182] In one embodiment the polypeptide is a component of the nesting material from a Colletid bee.
[0183] In one embodiment the Colletid bee is selected from the following genera: Hylaeus spp., Euryglossa spp. and Meroglossa spp.
[0184] In one embodiment the Colletid bee is selected from Hylaeus nubilosus, Hylaeus euxanthus Euryglossa adelaidae, Euryglossa subsericea and Meroglossa impressifrons penetrate.
[0185] In one embodiment the nesting material is from a Hylaeus spp. bee. In one embodiment the Hylaeus spp. bee is H. nubilosus or H. euxanthus.
[0186] In one embodiment the polypeptide or functional portion thereof is an H. nubilosus bee nesting material polypeptide (HnBNMP).
[0187] In one embodiment the HnBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2 or SEQ ID NO: 4, preferably SEQ ID NO: 2, preferably SEQ ID NO: 4.
[0188] In one embodiment the HnBNMP comprises, consists essentially of or consists of SEQ ID NO: 2 or SEQ ID NO: 4, preferably SEQ ID NO: 2, preferably SEQ ID NO: 4.
[0189] In one embodiment the polypeptide or functional portion thereof is an H. euxanthus bee nesting material polypeptide (HeBNMP).
[0190] In one embodiment the HeBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 14 or SEQ ID NO: 16, preferably SEQ ID NO: 14, preferably SEQ ID NO: 16.
[0191] In one embodiment the HeBNMP comprises, consists essentially of or consists of SEQ ID NO: 14 or SEQ ID NO: 16, preferably SEQ ID NO: 14, preferably SEQ ID NO: 16. In one embodiment the nesting material is from a Euryglossa spp. bee. In one embodiment the Euryglossa spp. bee is E. adelaidae or E. subsericea.
[0192] In one embodiment the polypeptide or functional portion thereof is an E. adelaidae bee nesting material polypeptide (EaBNMP).
[0193] In one embodiment the EaBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 6 or SEQ ID NO: 8, preferably SEQ ID NO: 6, preferably SEQ ID NO: 8.
[0194] In one embodiment the EaBNMP comprises, consists essentially of or consists of SEQ ID NO: 6 or SEQ ID NO: 8, preferably SEQ ID NO: 6, preferably SEQ ID NO: 8.
[0195] In one embodiment the polypeptide is an E. subsericea bee nesting material polypeptide (EsBNMP).
[0196] In one embodiment the EsBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 10 or SEQ ID NO: 12, preferably SEQ ID NO: 10, preferably SEQ ID NO: 12.
[0197] In one embodiment the EsBNMP comprises, consists essentially of or consists of SEQ ID NO: 10 or SEQ ID NO: 12, preferably SEQ ID NO: 10, preferably SEQ ID NO: 12.
[0198] In one embodiment the nesting material is from a Meroglossa spp. bee. In one embodiment the Meroglossa spp. bee is M. impressifrons penetrate.
[0199] In one embodiment the polypeptide or functional portion thereof is an M. impressifrons penetrate bee nesting material polypeptide (MiBNMP).
[0200] In one embodiment the MiBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 18 or SEQ ID NO: 20, preferably SEQ ID NO: 18, preferably SEQ ID NO: 20.
[0201] In one embodiment the MiBNMP comprises, consists essentially of or consists of SEQ ID NO: 18 or SEQ ID NO: 20, preferably SEQ ID NO: 18, preferably SEQ ID NO: 20. In one embodiment the polypeptide or functional portion thereof is a consensus sequence bee nesting material polypeptide (CsBNMP).
[0202] In one embodiment the polypeptide or functional portion thereof is a consensus sequence bee nesting material polypeptide from the N-terminal region (CsNTBNMP) of at least two, preferably at least three, four, preferably at least five colletid bee nesting material polypeptides.
[0203] In one embodiment the CsBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 22.
[0204] In one embodiment the CsBNMP comprises, consists essentially of or consists of SEQ ID NO: 22.
[0205] In one embodiment the CsNTBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 24.
[0206] In one embodiment the CsNTBNMP comprises, consists essentially of or consists of SEQ ID NO: 24.
[0207] In one embodiment the polypeptide or functional portion thereof is amphipathic.
[0208] In one embodiment the polypeptide or functional portion thereof is at least partially capable of modifying the hydrophobicity of a material.
[0209] Specifically contemplated as embodiments of this aspect of the invention are all of the polypeptide embodiments set forth above in the isolated polynucleotide aspect of the invention.
[0210] In another aspect, the invention relates to an isolated polynucleotide or functional portion thereof comprising at least 70% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0211] In one embodiment the polynucleotide or functional portion thereof comprises, consists essentially of or consists of a polynucleotide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23. In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
[0212] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof comprises, consists essentially of or consists of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24, respectively.
[0213] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that comprises, consists essentially of or consists of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24, respectively.
[0214] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is from the nesting material of a Colletid bee.
[0215] In one embodiment the polynucleotide is expressed in the Salivary or Dufour's gland of the Colletid bee.
[0216] In one embodiment the polynucleotide encodes a polypeptide that is a component of the nesting material from a Colletid bee.
[0217] In one embodiment the Colletid bee is selected from the following genera: Hylaeus spp., Euryglossa spp. and Meroglossa spp.
[0218] In one embodiment the Colletid bee is selected from Hylaeus nubilosus, Hylaeus euxanthus Euryglossa adelaidae, Euryglossa subsericea and Meroglossa impressifrons penetrate.
[0219] In one embodiment the nesting material is from a Hylaeus spp. bee. In one embodiment the Hylaeus spp. bee is H. nubilosus or H. euxanthus.
[0220] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is an H. nubilosus bee nesting material polypeptide (HnBNMP).
[0221] In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 1 or SEQ ID NO: 3, preferably SEQ ID NO: 1, preferably SEQ ID NO: 3. In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 1 or SEQ ID NO: 3, preferably SEQ ID NO: 1, preferably SEQ ID NO: 3.
[0222] In one embodiment the HnBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2 or SEQ ID NO: 4, preferably SEQ ID NO: 2, preferably SEQ ID NO: 4.
[0223] In one embodiment the HnBNMP comprises, consists essentially of or consists of SEQ ID NO: 2 or SEQ ID NO: 4, preferably SEQ ID NO: 2, preferably SEQ ID NO: 4.
[0224] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an H. euxanthus bee nesting material polypeptide (HeBNMP).
[0225] In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 13 or SEQ ID NO:
[0226] 15, preferably SEQ ID NO: 13, preferably SEQ ID NO: 15.
[0227] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 13 or SEQ ID NO: 15, preferably SEQ ID NO: 13, preferably SEQ ID NO: 15.
[0228] In one embodiment the HeBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 14 or SEQ ID NO:
[0229] 16, preferably SEQ ID NO: 14, preferably SEQ ID NO: 16.
[0230] In one embodiment the HeBNMP comprises, consists essentially of or consists of SEQ ID NO: 14 or SEQ ID NO: 16, preferably SEQ ID NO: 14, preferably SEQ ID NO: 16.
[0231] In one embodiment the nesting material is from a Euryglossa spp. bee. In one embodiment the Euryglossa spp. bee is E. adelaidae or E. subsericea.
[0232] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an E. adelaidae bee nesting material polypeptide (EaBNMP).
[0233] In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 5 or SEQ ID NO: 7, preferably SEQ ID NO: 5, preferably SEQ ID NO: 7.
[0234] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 5 or SEQ ID NO: 7, preferably SEQ ID NO: 5, preferably SEQ ID NO: 7.
[0235] In one embodiment the EaBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 6 or SEQ ID NO: 8, preferably SEQ ID NO: 6, preferably SEQ ID NO: 8.
[0236] In one embodiment the EaBNMP comprises, consists essentially of or consists of SEQ ID NO: 6 or SEQ ID NO: 8, preferably SEQ ID NO: 6, preferably SEQ ID NO: 8.
[0237] In one embodiment the polynucleotide encodes a polypeptide that is an E. subsericea bee nesting material polypeptide (EsBNMP).
[0238] In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 9 or SEQ ID NO: 11, preferably SEQ ID NO: 9, preferably SEQ ID NO: 11.
[0239] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 9 or SEQ ID NO: 11, preferably SEQ ID NO: 9, preferably SEQ ID NO: 11
[0240] In one embodiment the EsBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 10 or SEQ ID NO: 12, preferably SEQ ID NO: 10, preferably SEQ ID NO: 12.
[0241] In one embodiment the EsBNMP comprises, consists essentially of or consists of SEQ ID NO: 10 or SEQ ID NO: 12, preferably SEQ ID NO: 10, preferably SEQ ID NO: 12.
[0242] In one embodiment the nesting material is from a Meroglossa spp. bee. In one embodiment the Meroglossa spp. bee is M. impressifrons penetrate.
[0243] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an M. impressifrons penetrate bee nesting material polypeptide (MiBNMP). In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 17 or SEQ ID NO: 19, preferably SEQ ID NO: 17, preferably SEQ ID NO: 19.
[0244] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 17 or SEQ ID NO: 19, preferably SEQ ID NO: 17, preferably SEQ ID NO: 19.
[0245] In one embodiment the MiBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 18 or SEQ ID NO: 20, preferably SEQ ID NO: 18, preferably SEQ ID NO: 20.
[0246] In one embodiment the MiBNMP comprises, consists essentially of or consists of SEQ ID NO: 18 or SEQ ID NO: 20, preferably SEQ ID NO: 18, preferably SEQ ID NO: 20.
[0247] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is a consensus sequence bee nesting material polypeptide (CsBNMP).
[0248] In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 21.
[0249] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 21.
[0250] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is an amino acid consensus sequence from the N-terminal region of at least two, preferably at least three, four, preferably at least five colletid bee nesting material polypeptides.
[0251] In one embodiment the CsNTBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 24.
[0252] In one embodiment the CsNTBNMP comprises, consists essentially of or consists of SEQ ID NO: 24. In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is hydrophobic.
[0253] In one embodiment the polynucleotide is modified for expression in a heterologous host cell. In one embodiment the heterologous host cell is a prokaryotic host cell, preferably a bacterial host cell, preferably E. coli.
[0254] In one embodiment the polynucleotide is comprised in a vector or nucleic acid expression construct. In one embodiment the vector or nucleic acid expression construct further comprises a heterologous regulatory element.
[0255] Further, specifically contemplated as embodiments of the isolated polynucleotide aspects of the invention that relate to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23 and the isolated polypeptide aspects of the invention that relate to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24 as set forth above, are all of the embodiments set forth below in the isolated polynucleotide aspects of the invention that relate to SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65 and the isolated polypeptide aspects that relate to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66 including all embodiments related to % amino acid sequence identity, % nucleic acid sequence identity, nesting material, Colletid bees, repeat regions, repeat sub regions, regions of alternating relative hydrophobicity and hydrophilicity, HnBNMP, HeBNMP, EaBNMP, EsBNMP, MiBNMP, % amino acid residues, hydrogels, recombinant expression, isolation and heterologous regulatory elements and subsequences.
[0256] In another aspect the invention relates to an isolated polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0257] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66. In one embodiment the polynucleotide encodes a polypeptide comprising, consisting essentially of or consisting of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0258] In one embodiment the polynucleotide encodes a polypeptide comprising a repeat region of about, preferably of, 110 to 258 amino acids in length.
[0259] In one embodiment the repeat region comprises at least two regions of alternating relative hydrophobicity and hydrophilicity.
[0260] In one embodiment the region of relative hydrophobicity comprises at least 20% to 30% hydrophobic amino acid residues.
[0261] In one embodiment the polypeptide is from the nesting material of a Colletid bee.
[0262] In one embodiment the polypeptide is expressed in the Salivary or Dufour's gland of the Colletid bee.
[0263] In one embodiment the polypeptide is a component of the nesting material from a Colletid bee.
[0264] In one embodiment the Colletid bee is selected from the following genera: Hylaeus spp., Euryglossa spp. and Meroglossa spp.
[0265] In one embodiment the Colletid bee is selected from Hylaeus nubilosus, Hylaeus euxanthus Euryglossa adelaidae, Euryglossa subsericea and Meroglossa impressifrons penetrate.
[0266] In one embodiment the Colletid bee is a Hylaeus spp. bee. In one embodiment the Hylaeus spp. bee is H. nubilosus or H. euxanthus.
[0267] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof that is an H. nubilosus bee nesting material polypeptide (HnBNMP).
[0268] In one embodiment the HnBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 28, SEQ ID NO: 30, SEQ ID NO: 32, or SEQ ID NO: 34, preferably SEQ ID NO: 28, preferably SEQ ID NO: 30, preferably SEQ ID NO: 32, preferably SEQ ID NO: 34. In one embodiment the HnBNMP comprises, consists essentially of or consists of SEQ ID NO: 28, SEQ ID NO: 30, SEQ ID NO: 32, or SEQ ID NO: 34, preferably SEQ ID NO: 28, preferably SEQ ID NO: 30, preferably SEQ ID NO: 32, preferably SEQ ID NO: 34.
[0269] In one embodiment the HnBNMP comprises an amino acid repeat region of about, preferably of, 110 amino acids in length.
[0270] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an H. euxanthus bee nesting material polypeptide (HeBNMP).
[0271] In one embodiment the HeBNMP comprises, consists essentially of or consists of a polypeptide having at least at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56, or SEQ ID NO: 58, preferably SEQ ID NO: 52 preferably SEQ ID NO: 54, preferably SEQ ID NO: 56, preferably SEQ ID NO: 58.
[0272] In one embodiment the HeBNMP comprises, consists essentially of or consists of SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56, or SEQ ID NO: 58, preferably SEQ ID NO: 52 preferably SEQ ID NO: 54, preferably SEQ ID NO: 56, preferably SEQ ID NO: 58.
[0273] In one embodiment the HeBNMP comprises an amino acid repeat region of about, preferably of, 121 amino acids in length.
[0274] In one embodiment the Colletid bee is a Euryglossa spp. bee. In one embodiment the Euryglossa spp. bee is E. adelaidae or E. subsericea.
[0275] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an E. adelaidae bee nesting material polypeptide (EaBNMP).
[0276] In one embodiment the EaBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 36, SEQ ID NO: 38, SEQ ID NO: 40, or SEQ ID NO: 42, preferably SEQ ID NO: 36 preferably SEQ ID NO: 38, preferably SEQ ID NO: 40, preferably SEQ ID NO: 42.
[0277] In one embodiment the EaBNMP comprises, consists essentially of or consists of SEQ ID NO: 36, SEQ ID NO: 38, SEQ ID NO: 40, or SEQ ID NO: 42, preferably SEQ ID NO: 36 preferably SEQ ID NO: 38, preferably SEQ ID NO: 40, preferably SEQ ID NO: 42. In one embodiment the EaBNMP comprises an amino acid repeat region of about, preferably of, 146 amino acids in length.
[0278] In one embodiment the polynucleotide encodes a polypeptide that is an E. subsericea bee nesting material polypeptide (EsBNMP).
[0279] In one embodiment the EsBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 44, SEQ ID NO: 46, SEQ ID NO: 48, or SEQ ID NO: 50, preferably SEQ ID NO: 44 preferably SEQ ID NO: 46, preferably SEQ ID NO: 48, preferably SEQ ID NO: 50.
[0280] In one embodiment the EsBNMP comprises, consists essentially of or consists of SEQ ID NO: 44, SEQ ID NO: 46, SEQ ID NO: 48, or SEQ ID NO: 50, preferably SEQ ID NO: 44 preferably SEQ ID NO: 46, preferably SEQ ID NO: 48, preferably SEQ ID NO: 50.
[0281] In one embodiment the EsBNMP comprises an amino acid repeat region of about, preferably of, 258 amino acids in length.
[0282] In one embodiment the Colletid bee is a Meroglossa spp. bee. In one embodiment the Meroglossa spp. bee is M. impressifrons penetrate.
[0283] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof is an M. impressifrons penetrate bee nesting material polypeptide (MiBNMP).
[0284] In one embodiment the MiBNMP comprises, consists essentially of or consists of a polypeptide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 64, or SEQ ID NO: 66, preferably SEQ ID NO: 60, preferably SEQ ID NO: 62, preferably SEQ ID NO: 64, preferably SEQ ID NO: 66.
[0285] In one embodiment the MiBNMP comprises, consists essentially of or consists of SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 64, or SEQ ID NO: 66, preferably SEQ ID NO: 60, preferably SEQ ID NO: 62, preferably SEQ ID NO: 64, preferably SEQ ID NO: 66.
[0286] In one embodiment the MiBNMP comprises an amino acid repeat region of about, preferably of, 146 amino acids in length. In one embodiment the repeat region of each of HnBNMP, HeBNMP, EaBNMP, EsBNMP and MiBNMP as set forth above comprises at least two, preferably at least three repeat regions.
[0287] In one embodiment the repeat regions are similar or substantially similar repeat sub regions that form beta sheets.
[0288] In one embodiment the repeat sub regions that form beta sheets are identified by rapid automatic detection and alignment of repeats (RADAR).
[0289] In one embodiment the repeat regions comprise at least 14% to 36% glutamine residues and at least 18% to 31% serine residues.
[0290] In one embodiment the HnBNMP comprises a repeat region comprising 23% glutamine residues and 28% serine residues.
[0291] In one embodiment the repeat region comprises at least three regions of alternating relative hydrophobicity and hydrophilicity.
[0292] In one embodiment the region of relative hydrophobicity comprises about 20%, preferably 20% hydrophobic amino acid residues.
[0293] In one embodiment the region of relative hydrophilicity comprises about 80%, preferably 80% hydrophobic amino acid residues.
[0294] In one embodiment the HeBNMP comprises a repeat region comprising 22% glutamine residues and 18% serine residues.
[0295] In one embodiment the repeat region comprises at least two regions of alternating relative hydrophobicity and hydrophilicity.
[0296] In one embodiment the region of relative hydrophobicity comprises about 24%, preferably 24% hydrophobic amino acid residues.
[0297] In one embodiment the region of relative hydrophilicity comprises about 76%, preferably 76% hydrophobic amino acid residues.
[0298] In one embodiment the EaBNMP comprises a repeat region comprising 14% glutamine residues and 23% serine residues. In one embodiment the EaBNMP repeat region comprises at least three regions of alternating relative hydrophobicity and hydrophilicity.
[0299] In one embodiment the region of relative hydrophobicity comprises about 30%, preferably 30% hydrophobic amino acid residues.
[0300] In one embodiment the region of relative hydrophilicity comprises about 70%, preferably 70% hydrophobic amino acid residues.
[0301] In one embodiment the EsBNMP comprises a repeat region comprising 34% glutamine residues and 31% serine residues.
[0302] In one embodiment the EsBNMP repeat region comprises at least four regions of alternating relative hydrophobicity and hydrophilicity.
[0303] In one embodiment the region of relative hydrophobicity comprises about 16%, preferably 16% hydrophobic amino acid residues.
[0304] In one embodiment the region of relative hydrophilicity comprises about 84%, preferably 84% hydrophobic amino acid residues.
[0305] In one embodiment the MiBNMP comprises a repeat region comprising 36% glutamine residues and 29% serine residues.
[0306] In one embodiment the MiBNMP repeat region comprises at least four regions of alternating relative hydrophobicity and hydrophilicity.
[0307] In one embodiment the region of relative hydrophobicity comprises about 21%, preferably 21% hydrophobic amino acid residues.
[0308] In one embodiment the region of relative hydrophilicity comprises about 79%, preferably 79% hydrophobic amino acid residues.
[0309] In one embodiment the polypeptide forms a hydrogel in about 6M guanidine thiocyanate, preferably wherein the hydrogel comprises at least lOmg / ml of the polypeptide.
[0310] In one embodiment the polypeptide or functional portion thereof is an isolated polypeptide or functional portion thereof. In one embodiment the polynucleotide comprises, consists essentially of or consists of a polynucleotide having at least 70%, preferably at least 75%, 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65.
[0311] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65.
[0312] In one embodiment the polynucleotide is modified for expression in a heterologous host cell. In one embodiment the heterologous host cell is a prokaryotic host cell, preferably a bacterial host cell, preferably E. coli.
[0313] In one embodiment the polynucleotide is comprised in a vector or nucleic acid expression construct. In one embodiment the vector or nucleic acid expression construct further comprises a heterologous regulatory element.
[0314] In one embodiment the polynucleotide is a subsequence of an polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0315] In one embodiment the polynucleotide is a subsequence of an polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0316] In one embodiment the polynucleotide is a subsequence of an polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0317] In one embodiment the polynucleotide is a subsequence of an polynucleotide encoding a polypeptide comprising, consisting essentially of or consisting of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
[0318] In another aspect the invention relates to an isolated polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0319] In one embodiment the polypeptide or functional portion thereof comprises, consists essentially of or consists of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0320] In one embodiment the polypeptide comprises, consists essentially of or consists of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
[0321] Specifically contemplated as embodiments of each of the isolated polypeptides of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66 as set forth above are all of the polypeptide embodiments set forth in the previous aspect of the invention that relates to isolated polynucleotides that encode polypeptides comprising, consisting essentially of or consisting of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66 including all embodiments related to % amino acid sequence identity, % nucleic acid sequence identity, nesting material, Colletid bees, repeat regions, repeat sub regions, regions of alternating relative hydrophobicity and hydrophilicity, HnBNMP, HeBNMP, EaBNMP, EsBNMP, MiBNMP, % amino acid residues, hydrogels, recombinant expression, isolation and heterologous regulatory elements and subsequences.
[0322] In another aspect, the invention relates to an isolated polynucleotide or functional portion thereof comprising, consisting essentially of or consisting of a polynucleotide having at least 70% nucleic acid sequence identity to SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65.
[0323] In one embodiment the polynucleotide or functional portion thereof comprises, consists essentially of or consists of a polynucleotide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% nucleic acid sequence identity to SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65.
[0324] In one embodiment the polynucleotide comprises, consists essentially of or consists of SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65. In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66, respectively.
[0325] In one embodiment the polynucleotide encodes a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 75%, preferably at least 80%, 85%, 90%, 95%, preferably at least 99% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66, respectively.
[0326] In one embodiment the polynucleotide encodes a polypeptide comprising, consisting essentially of or consisting of the amino acid sequence of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66, respectively.
[0327] Specifically contemplated as embodiments of each of the isolated polynucleotides of SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65 as set forth above are all of the polynucleotide embodiments set forth in the previous aspect of the invention that relates to isolated polynucleotides that encode polypeptides comprising, consisting essentially of or consisting of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66 including all embodiments related to % amino acid sequence identity, % nucleic acid sequence identity, nesting material, Colletid bees, repeat regions, repeat sub regions, regions of alternating relative hydrophobicity and hydrophilicity, HnBNMP, HeBNMP, EaBNMP, EsBNMP, MiBNMP, % amino acid residues, hydrogels, recombinant expression, isolation, vectors, heterologous regulatory elements and subsequences.
[0328] In another aspect the invention relates to a vector comprising an isolated polynucleotide according to the invention.
[0329] In another aspect the invention relates to a vector that encodes an isolated polypeptide according to the invention.
[0330] In one embodiment the vector is selected from the group consisting of plasmids, BACs, (PACs), YACs, bacteriophage, phagemids, and cosmids. In one embodiment the vector is a plasmid. In one embodiment the vector is selected from the group consisting of plasmids, BACs, PACs, YACs, bacteriophage, phagemids, and cosmids. Preferably the vector is a plasmid.
[0331] In one embodiment the vector is an expression vector.
[0332] Examples of suitable expression vectors include, but are not limited to, plasmid DNA vectors, viral DNA vectors (such as adenovirus and adeno-associated virus), or viral RNA vectors (such as retroviral vectors). In some embodiments the plasmid and / or phage vectors may be selected from the following vectors or variants thereof including pET, pFastBac, pUC18, pU19, Mpl8, Mpl9, ColEl, PCR1 and pKRC; lambda gtlO and M13 plasmids such as pBR322, pACYC184, pT127, RP4, plJlOl, SV40 and BPV. Additional non-limiting examples of vectors include cosmids, YACS, BACs, shuttle vectors such as pSA3, and PAT28 transposons.
[0333] Suitable viral vectors include but are not limited to vectors derived from adenovirus (AV); adeno-associated virus (AAV); retroviruses (e.g., lentiviruses (LV), Rhabdoviruses, murine leukaemia virus); herpes virus, and the like. Viral vectors employed herein can be appropriately modified by pseudo-typing with envelope proteins or other surface antigens from other viruses, or by substituting different viral capsid proteins, as known, and used in the art.
[0334] The vector can be constructed to drive expression of a polypeptide as described herein, either in vitro or in vivo. In one embodiment, the vector comprises a polynucleotide of the invention operatively linked to 5' or 3' untranslated regulatory sequences. The design of a vector will depend on various factors including the host cells in which the operatively linked polynucleotide is to be expressed and the desired level of polynucleotide expression.
[0335] Likewise, the selection of various promoters, enhancers and / or other genetic elements for the vector will depend on various factors including the host cells and expression levels discussed above. In one embodiment, the vector comprises a homologous promoter operatively linked to a polynucleotide of the invention. In another embodiment, the vector comprises a heterologous promoter operatively linked to a polynucleotide of the invention. In one embodiment, the homologous or heterologous promoter is an inducible, repressible, or regulatable promoter. A suitable promoter may be chosen and used under the appropriate conditions to direct high-level expression of a polynucleotide of the invention. Many such elements are described in the literature and are available through commercial suppliers. By way of example only, promoters useful in the vector can be any suitable eukaryotic or prokaryotic promoter. In one embodiment, the eukaryotic promoter can be a eukaryotic RIMA polymerase I (pol I), RNA polymerase II (pol II), or RNA polymerase III (pol III). Expression levels of an operably linked polynucleotide in a particular cell type will be determined by the nearby presence (or absence) of specific gene regulatory sequences (e.g., enhancers, silencers and the like). Any suitable promoter / enhancer combination (see: Eukaryotic Promoter Data Base EPDB) can be used to drive expression of a polynucleotide of the invention.
[0336] Additional promoters useful in expression cassettes include beta-lactamase, alkaline phosphatase, tryptophan, and tac promoter systems which are all well known in the art. Yeast promoters include 3-phosphoglycerate kinase, enolase, hexokinase, pyruvate decarboxylase, glucokinase, and glyceraldehydrate-3-phosphanate dehydrogenase but are not limited thereto.
[0337] Prokaryotic promoters useful in expression cassettes include constitutive promoters as known in the art (such as the int promoter of bacteriophage lamda and the bla promoter of the beta -lactamase gene sequence of pBR322) and regulatable promoters (such as lacZ, recA and gal). A ribosome binding site upstream of the CDS may also be required for expression.
[0338] Enhancers useful in a vector as described herein include SV40 enhancer, cytomegalovirus early promoter enhancer, globin, albumin, insulin and the like.
[0339] In one embodiment, a vector may be driven by a T3, T7 or SP6 cytoplasmic expression system.
[0340] In another aspect the invention relates to an isolated host cell comprising an isolated polynucleotide, isolated polypeptide, and / or vector according to the invention.
[0341] In one embodiment the isolated host cell is prokaryotic or eukaryotic.
[0342] In one embodiment the prokaryotic cell is a bacterial cell, preferably wherein the bacterial cell is an Escherichia coli (E. coli), Pseudomonas, Bacillus, Serratia, Klebsiella, Streptomyces, Listeria, Salmonella or Mycobacteria cell.
[0343] In one embodiment the bacterial cell is an E. coli cell. In one embodiment the eukaryotic cell is an animal cell, a plant cell, a fungal cell, or a protist cell.
[0344] In one embodiment the eukaryotic cell is a fungal cell. In one embodiment the fungal cell is a yeast cell. In one embodiment the yeast cell is a Pichia pastoris or Saccharomyces spp cell. In one embodiment the fungal cell is an Aspergillus spp. cell. In one embodiment the Aspergillus spp is Aspergillus niger.
[0345] In one embodiment the animal cell is an insect cell or a mammalian cell. In one embodiment the animal cell is a non-human animal cell. In one embodiment the mammalian cell is a non-human mammalian cell.
[0346] In one embodiment the insect cell comprises a polynucleotide as described herein in a viral vector, preferably a baculovirus. In one embodiment the insect cell is an Sf9 or High Five cell.
[0347] In another aspect the invention relates to a composition comprising at least one isolated polynucleotide, isolated polypeptide, and / or vector as described herein, and a carrier, diluent, or excipient.
[0348] In one embodiment, the composition consists essentially of the isolated polynucleotide, isolated polypeptide, and / or vector as described herein.
[0349] In one embodiment the composition is a cosmetic composition. In one embodiment the cosmetic composition is a hair or skin care composition.
[0350] In one embodiment the carrier, diluent, or excipient is a cosmetically acceptable carrier, diluent, or excipient.
[0351] In one embodiment the carrier, diluent, or excipient is a surfactant or dispersant.
[0352] In another aspect the invention relates to a method of making an isolated BNMP polypeptide or functional portion thereof, wherein the BNMP polypeptide or functional portion comprises at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64 or 66, the method comprising heterologously expressing the BNMP polypeptide or functional portion thereof in an isolated host cell, and optionally purifying the BNMP polypeptide. Specifically contemplated as embodiments of this aspect of the invention are all of the embodiments above relating to an isolated polypeptide or functional portion thereof comprising at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64 or 66 including % amino acid sequence identity, % nucleic acid sequence identity, nesting material, Colletid bees, repeat regions, repeat sub regions, regions of alternating relative hydrophobicity and hydrophilicity, HnBNMP, HeBNMP, EaBNMP, EsBNMP, MiBNMP, % amino acid residues, hydrogels, recombinant expression, isolation, vectors, heterologous regulatory elements and subsequences.
[0353] In one embodiment expression is from an isolated polynucleotide encoding the BNMP or functional portion thereof.
[0354] In one embodiment the polynucleotide comprises at least 70% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65.
[0355] Specifically contemplated as further embodiments of this aspect of the invention are all of the embodiments set forth above relating to an isolated polynucleotide encoding a polypeptide or functional portion thereof comprising at least 70% amino acid sequence identity to SEQ ID NO: 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64 or 66 and to an isolated polynucleotide or functional portion thereof comprising at least 70% nucleic acid sequence identity to SEQ ID NO: SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65, including, but not limited to, isolated polynucleotides that encode HnBNMP, HeBNMP, EaBNMP, EsBNMP and / or MiBNMP polypeptides or functional portions thereof and all embodiments relating to % nucleic acid sequence identity of the HnBNMP, HeBNMP, EaBNMP, EsBNMP and MiBNMP polynucleotides or functional portions thereof
[0356] Likewise, specifically contemplated as embodiments of this aspect of the invention are all the embodiments set forth in previous aspects of the invention relating to polynucleotide expression including vectors, heterologous regulatory elements, and isolated host cells, but not limited thereto.
[0357] In one embodiment the method comprises purifying the BNMP polypeptide after it is secreted from an isolated host cell. In one embodiment the method comprises purifying the BNMP polypeptide from an isolated host cell.
[0358] In one embodiment the isolated host cell is a bacterial cell. In one embodiment the bacterial cell is E. coli.
[0359] In another aspect the invention relates to an isolated polypeptide or BNMP polypeptide or functional portion of either when made by a method of the invention.
[0360] In another aspect the invention relates to the use of an isolated polypeptide or functional portion thereof as described herein to coat or to form a coating on an article of manufacture. In one embodiment the coating is hydrophilic, hydrophobic or amphipathic.
[0361] In another aspect the invention relates to the use of an isolated polypeptide or functional portion thereof as described herein to make a film on an article of manufacture.
[0362] In one embodiment the film is hydrophobic.
[0363] In one embodiment the article of manufacture is selected from the group consisting of textiles or textile components or parts thereof and biomedical devices or components or parts thereof.
[0364] In one embodiment the article of manufacture or component or part thereof is a synthetic fiber. In one embodiment the component or part thereof is a synthetic polymer. In one embodiment the synthetic fiber or synthetic polymer is selected from the group consisting of polyester, spandex, rayon, nylon, acrylic, microfiber, neoprene, polyamide, acetate, polyvinyl chloride (PVC) and synthetic or "faux" leather or fur fibers and polymers.
[0365] In one embodiment the article of manufacture or component or part thereof is a natural fiber. In one embodiment the natural fiber is selected from the group consisting of cotton, wool, silk, coir, alpaca, flax, hemp, bamboo, sisal, and jute.
[0366] In one embodiment article of manufacture is a textile, preferably a natural textile or a synthetic textile.
[0367] In one embodiment the article of manufacture or component or part thereof is or is comprised in or on a biomedical device. In one embodiment the biomedical device is an implantable biomedical device. In one embodiment the implantable biomedical device is selected from the group consisting of cardiovascular devices including cardioverter defibrillators, pacemakers and left ventricular assist devices, breast implants, cochlear implants, intraocular lenses, joint replacements including hip implants, catheters, dialysis tubing, contraceptive intrauterine devices, stents, sutures, staples, bandages, and wound dressings.
[0368] In one embodiment the article of manufacture is an air filtration device, component, or part thereof. In one embodiment the component or part thereof is an air filter. In one embodiment the component or part thereof is a synthetic or natural fiber and / or polymer. In one embodiment the synthetic or natural fiber and / or polymer is comprised in the air filter. In one embodiment the synthetic or natural fiber and / or polymer is in the form of a nanofiber.
[0369] Specifically contemplated as embodiments within the method and use aspects of this invention are the embodiments set out above in the previous aspects of the invention that relate to isolated polypeptides and isolated polynucleotides and functional portions of either including but not limited to polynucleotide expression including vectors, heterologous regulatory elements, and isolated host cells.
[0370] The invention will now be illustrated in a non-limiting way by reference to the following examples.
[0371] EXAMPLES
[0372] Example 1 - Module 1 BNMPs from a representative sample of Colletid bees
[0373] Identification of Module 1
[0374] The coding sequence for the HnMl BNMP from H. nubilosus (SEQ ID NO: 2) was identified by mapping transcripts from H. nubilosus onto the scaffold containing the nest material gene using Minimap2 (Le., 2018), implying diverse splice variants which produce isoforms of the HnBNMG which differ in length and amino acid composition. Several full-length RNA and cDNA sequences from H. nubilosus mapped onto the first and second exon of the H. nubilosus nest material gene, revealing a short-expressed isoform of the HnBNMG named HnMl. The HnBNMP module 1 isoform contains the conserved N-terminal region, a glutamine and serine-rich repeat region and a C-terminal region, which is not repetitive or enriched in these residues (Figure 1). Sourcing and Collection of Colletid bees
[0375] Three E. adelaidae, 12 E. subsericea, 11 H. euxanthus, and 2 M. impressifrons penetrata were collected by trained entomologists while foraging in Queensland, Australia. Bees were incised, collected in RNAIater, and stored at 4°C to slow RIMA degradation. A single female of each bee species (E. adelaidae, E. subsericea, H. euxanthus and M. impressifrons penetrata) was selected for RNA extraction.
[0376] RNA extraction
[0377] First, the cuticle was removed from the bee from which RNA was extracted. The remaining tissue was transferred into a DNA Lobind 1.5ml tube. 50 pL of TRIzol reagent was transferred into the tube containing the samples. The tissue was then disrupted by crushing with a DNAse-free pestle. An additional 200 pL of TRIzol reagent was used to rinse the pestle, washing the remaining tissue into the tube. The sample was then incubated at room temperature for 2 minutes. 50 pL of chloroform was added to the tube containing the sample. The sample tube was vigorously shaken for 15 seconds, incubated for 5 minutes at room temperature, and then centrifuged at 12,000g for 15 minutes at 4°C. The upper aqueous phase, containing the RNA, was transferred to a new DNA Lobind 1.5 ml tube. 0.125 mL of isopropanol was added to the aqueous phase tube. The sample was mixed through inversion, resulting in RNA precipitation. The sample was then incubated at room temperature for 2 minutes. The sample was subsequently centrifuged at 12,000 g for 10 minutes at 4°C. The supernatant was then removed and discarded. The sample was washed with 250 pL of 75% ethanol, centrifuged at 7500 g for 5 minutes at 4 °C, and the supernatant removed. The sample was then air dried until all droplets had evaporated from the RNA pellet. The RNA pellet was then gently resuspended in 32 pL of RNAse-free water by flicking the tube. Using a Nanodrop, the RNA concentration was assessed by analysing the A 260 / 280 and A 260 / 230 ratios. The sample was then flash frozen in liquid nitrogen and stored at - 80°C till. The RNA extraction then underwent poly-A enrichment followed by reverse transcription into cDNA. cDNA sequencing
[0378] Garvan Institute of Medical Research (Darlinghurst NSW, Australia) performed poly-A enrichment on the extracted RNA from each bee and then reverse-transcribed this to cDNA. The cDNA was then screened for sequence degradation. The samples were multiplexed and sequenced on a Nanopore PromethlON flow cell. Base-calling was performed using the Oxford Nanopore Technologies Guppy version 6.1.5. Transcriptome assembly
[0379] The cDNA reads were processed using porechop ABI to trim the adapters from the reads (reads with middle adapters were discarded)(Bonenfant and Touzet, 2023). Reads shorter than 1000 amino acids or with a quality score of under 10 were removed by nanoq (Steinig and Coin, 2022). Read quality was evaluated using fastqc and nanoq stats. The cDNA assemblers RNAbloom2 and RATTLE were used to assemble the processed reads using Kmer values of 25, 35, 45 and 55 for RATTLE and similarity thresholds of 90, 95 and 99 for RNAbloom2 (Nip et al., 2023 and de la Rubia et al., 2022). Coding strand prediction was performed using CodAn to generate protein predictions (Nachtigall et al., 2020).
[0380] BN MG identification
[0381] The first 1-144 amino acids of the H. nubilosus full-length polypeptide (signal peptide, N- terminal region, and charged region) were used to identify bee nest material gene homologs for E. adelaidae, E. subsericea, H. euxanthus and Meroglossa impressifrons penetrata. The H. nubilosus N-terminal region includes the signal peptide, N-terminal region and charged region. The N-terminal amino acid sequence was used for a blast search (tblastn) to identify nest material gene transcripts in each transcriptome (Camacho et al., 2009). The corresponding amino acid sequence was acquired using CoDan and a custom translation script (Nachtigall et al., 2020).
[0382] For further physical and in-silico characterization, isoforms similar in size to the bee nest material gene isoform 'Module 1' from H. nubilosus were selected. The selected isoforms were called the Module 1 BNMPs (HnMl, EaM l, EsMl, HeMl and MiMl), and the first expression constructs are HnMl-01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01 respectively.
[0383] Bioinformatic characterization
[0384] Regions of amino acid enrichment were visually determined in the BNMPs through alignment with muscle (Edgar, 2004; Figure 1). Regional enrichment was confirmed using Composition Profiler (Vacic et al., 2007).
[0385] The BNMPs range in size with HnMl (SEQ ID NO: 2), EaM l (SEQ ID NO: 6), EsMl (SEQ ID NO: 10), HeMl (SEQ ID NO: 14) and MiMl (SEQ ID NO: 18) consisting of 232, 273, 392, 248 and 356 amino acids respectively. The BNMPs comprise 10% to 25% glutamine and 11% to 23% serine. (Figure 2) The signal peptides of the BNMPs are 17 amino acids long and the N-terminal region is 39-41 amino acids in length in the BNMPs. The signal peptides of the BNMPs are between 72% and 100% identical (calculated using esl-alipid in the HMMER suite of packages; Finn et al., 2011), showing a high level of homology. (Figures 2 and 7)
[0386] The N-terminal region is the most homologous region retained in the mature native BNM, ranging from 66% to 84% identical. The signal peptides and N-terminal regions of the BNMPs are between 70% and 90% identical (Figure 7). The N-terminal region of the BNMPs is enriched in histidine, proline and lysine and reduced in glutamine and serine residues when compared to the Swissprot Database protein using Composition Profiler (Bairoch and Apweiler, 2000; Vacic et al., 2007). (Figure 2)
[0387] The N-terminal region is flanked by the charged region, which is enriched with charged amino acids with MiMl, EsMl, EaMl, HeMl and HnMl being comprised of 27%, 39%, 48%, 49% and 50% charged amino acids respectively. The charged region is enriched in both acidic such as lysine and basic amino acids such as aspartate and glutamate, with MiMl, EsMl, EaMl, HeMl and HnMl being comprised of 14%, 20%, 20%, 28%, 28% acidic and 12%, 18%, 27%, 21%, 22% basic amino acids. (Figure 2)
[0388] The repeat region of the BNMPs comprises 2, 3, 3, 4 and 4 repeats with variable fidelity in HeMl, HnMl, EaMl, MiMl and EsMl, respectively, as detected by RADAR (Figure 5;
[0389] Heger and Holm., 2000). The repeat region is enriched in glutamine and serine residues when compared to the Swissprot protein using Composition Profiler (Bairoch and Apweiler, 2000; Vacic et al., 2007). The repeat regions of the BNMPs have a low percentage identity when compared ranging from 19% to 62% (Figure 7). The repeat region of the BNMPs vary in length and are 110, 121, 146, 204 and 258 amino acids long for HnMl, HeMl, EaMl, MiMl and EsMl, respectively. Each of these repeat regions has 100% amino acid sequence identity to the repeat regions in the expressed construct polypeptides HnMl-01, HeMl-01, EaMl-01, MiMl-01, and EsMl-01. The repeat region of the BNMPs contains a high proportion of glutamine and serine residues with EaMl, HeMl, HnMl, EsMl and MiMl composed of 14%, 22%, 23%, 34% and 36% glutamine residues and 23%, 18%, 28%, 31% and 29% serine, respectively. The repeat region of the BNMPs contains a high proportion of polar residues comprising 70%, 76%, 79%, 80% and 84% in EaMl, HeMl, MiMl, HnMl and EsMl, respectively. The repeat region of the BNMPs comprise 30%, 24%, 21%, 20% and 16% non-polar residues in EaMl, HeMl, MiMl, HnMl and EsMl, respectively. The repeat regions of the BNMPs comprise 10%, 13%, 18%, 22% and 23% charged residues in MiMl, EsMl, EaMl, HnMl and HeMl, respectively (Figure 2). Emboss pepwindow, with a window length of 11 amino acids, of the repeat region for HeMl, HnMl, EaMl, MiMl and EsMl implies alternating regions of relative hydrophilicity and hydrophibicity (Figure 6; Rice et al., 2000).
[0390] The C-terminal region of HnMl, HeMl, EaMl, EsMl and MiMl is enriched with cysteine, proline, serine and threonine residues when compared to the Swissprot protein database using Composition Profiler (Bairoch and Apweiler, 2000; Vacic et al., 2007).
[0391] Beta-Serpentine (Bondarev et al., 2018), which predicts a propensity for beta-sheet- driven aggregation by the formation of beta-serpentine structures, was used to identify regions of aggregation in the bee nest material polypeptides and identified nucleation sites in HnMl, HeMl, EaMl, EsMl and MiMl constructs and wildtype BNMPs (Figures 3 and 4).
[0392] Example 2 - Expression and purification of Module 1 BNMPs from a representative sample of Colletid bees
[0393] Expression
[0394] This example describes the production of a polypeptide in E.coli BL21 (DE3) using a kanamycin-resistant plasmid and the resulting production of purified polypeptide HnMl- 01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01.
[0395] Nucleic acid sequences coding for synthesis of HnMl-01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01 protein were synthesized using non-template PCR. In short, virtual nucleic acid sequences were converted into oligonucleotide sequences using software suite LIMS (DNA TwoPointO, Inc., Newark, CA, USA). Full-length nucleic acid sequences were synthesized by assembling oligonucleotides using template-free PCR. An amplicon was purified and cloned using standard cloning methods (Molecular Cloning. A Laboratory Manual. 2012. Green and Sambrook).
[0396] A gene coding for synthesis of the HnMl-01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01 proteins were cloned into expression vector pD451-SR, containing T7 inducible promoter (DNA TwoPointO, Inc., Newark, CA, USA). The purified plasmid containing the gene was transformed into chemically competent E. coli BL21(DE3) cells via heat shock and plated on non-inducing agar with 0.1 mg / mL kanamycin for HnMl-01 and HeMl-Oland 0.05 mg / mL kanamycin for EaMl-01, EsMl-01 and MiMl-01. Plates were incubated overnight at 37°C. Glycerol stocks were prepared by selecting and growing a single colony from a transformation plate in non-inducing media, followed by suspension of cells in media containing glycerol and preservation by storage at -80°C.
[0397] E. coli BL21(DE3) containing plasmid capable of expressing HnMl-01, EsMl-01, HeMl-01 and MiMl-01 protein were grown in 100L fermenters. Media was prepared and autoclaved in the fermenter. Media components and concentrations were as follows: casein hydrolysates, 12 g / L; yeast extract, 24 g / L; NaCI, 10 g / L; K2HPO4, 8 g / L; glycerol, 30 g / L. Kanamycin, 50 mg / L was added when the media had cooled. Preculture 1 flasks were grown at 37°C for approximately 6 hours. Pre-culture 2 flasks were inoculated from pre-culture 1 and grown at 25°C for approximately 12 hours for HnMl- 01. Pre-culture 2 flasks were inoculated from pre-culture 1 and grown at 28°C for approximately 12 hours for EsMl-01 and HeMl-01. Pre-culture 2 flasks were inoculated from pre-culture 1 and grown at 28°C for approximately 11 hours for MiMl-01. The fermenter was inoculated from pre-culture 2 flasks and temperature controlled at 37°C for the initial growth phase. Dissolved oxygen was controlled to 30% air saturation and the fermenter was maintained at pH 7. Pluronic antifoam, 5 g / L was added to control foaming.
[0398] E. coli BL21(DE3) containing plasmid capable of expressing EaMl-01 protein was grown in 10L fermenters. Media was prepared and autoclaved in the fermenter. Media components and concentrations were as follows: casein hydrolysates, 12 g / L; yeast extract, 24 g / L; NaCI, 10 g / L; K2HPO4, 8 g / L; glycerol, 30 g / L. Kanamycin, 50 mg / L was added when the media had cooled. Pre-culture flasks were grown at 22°C for approximately 16 hours. The fermenter was inoculated from pre-culture flasks and temperature controlled at 37°C for the initial growth phase. Dissolved oxygen was controlled to 30% air saturation and the fermenter was maintained at pH 7. Pluronic antifoam, 5 g / L was added to control foaming.
[0399] Immediately before induction, the fermenter was cooled down to 20°C and expression of HnMl-01, EaMl-01, EsMl-01 and HeMl-01 protein was induced at approximately OD600 2.0 using 0.2mM IPTG. For expression of MiMl-01 protein, the fermenter was cooled down to 28°C before induction. The biomass was concentrated by tangential flow filtration (TFF) approximately 20-22 hours post-induction and harvested by centrifugation for HnMl-01, EsMl-01, HeMl-01 and MiMl-01. EaMl-01 biomass was harvested by centrifugation approximately 20 hours post-induction. Biomass was frozen at -20°C until further processing. Purification
[0400] Biomass was thawed overnight and resuspended in lysis buffer (25mM Tris, 2mM MgCI2 0.5% (w / v) TritonX-100 pH 8.0) using a Miccra D-9 rotor-stator. Lysis was performed at room temperature for 40 minutes using 2mg lysozyme per gram of biomass. DNA was degraded with 25 units of benzonase per gram of biomass. Insoluble material was collected by centrifugation at 17000 g for 20 minutes for HnMl-01 and MiMl-01, 17500g for 40 minutes for EaMl-01 and HeMl-01 and 17500 g for 20 minutes for EsMl-01. The lysate pellet (i.e., 'insoluble' fraction) was subjected to washing in 25 mM Tris, 2 mM MgCI2, 0.5% (w / v) TritonX-100 pH 8.0 for 40 minutes. Insoluble material was collected by centrifugation at 17000 g for 40 minutes for HnMl-01 and MiMl-01 and 17500g for 40 minutes for EaMl-01, EsMl-01 and HeMl-01. The washed pellet was further washed in 0.05M sodium phosphate, pH 11.5 for HnMl-01, EsMl-01, HeMl-01, and MiMl-01 and 0.05M sodium phosphate, pH 11.3 for EaMl-01. Insoluble material was collected by centrifugation at 17,000g for 40 minutes for HnMl-01 and MiMl-01 and 17500 g for 40 minutes for EaMl-01, EsMl-01, and HeMl-01. The washed pellet was subjected to extraction in 10 mM Tris 4M guanidine at pH 8.0 for 40 minutes at room temperature HnMl-01, EaMl-01 and MiMl-01. The washed pellet was extracted for HeMl-01 in 10 mM Tris and 8 M guanidine at pH 8.0 for 90 minutes at room temperature. The washed pellet was extracted in 10 mM Tris, 6 M guanidine, pH 8.0, for 40 minutes at room temperature for EsMl-01. The extracted protein fraction was centrifuged at 17,000g for 20 mins to remove debris, and the supernatant was filtered for HnMl-01 and MiMl-01. The extracted protein fraction was centrifuged at 17500 g for 40 mins to remove debris, and the supernatant was filtered for EaMl-01 and HeMl-01. The extracted protein fraction was centrifuged at 17500 g for 20 mins to remove debris, and the supernatant was filtered for EsMl-01. The extracted fraction was diluted into immobilized metal affinity chromatography (IMAC) loading conditions (lOmM Tris, 3 M guanidine 500 mM NaCI, 20 mM Imidazole, pH 8.0). The diluted material was loaded onto a HiScale column packed with IMAC Sepharose 6 Fast Flow resin (Cytiva) charged with Nickel. The IMAC column was washed with loading buffer (10 mM Tris, 2 M guanidine, 0.5 M NaCI, 20 mM imidazole, pH 8.0), and the proteins recovered with elution buffer (10 mM tris, 2 M guanidine, 0.5M NaCI, 500 mM imidazole pH 8.0).
[0401] The proteins were precipitated from elution fractions using 1.5 M ammonium sulphate for HnMl-01, EaMl-01, HeMl-01, and MiMl-01 and 2 M ammonium sulphate for EsMl-01. Precipitated proteins were recovered by centrifugation at 17500 g for 40 minutes. Precipitated protein pellets were resuspended in ultrapure water and washed, and the protein precipitate was collected by centrifugation at 12000 g for 10 minutes.
[0402] Example 3 - Water contact angle of Module 1 BNMPs from a representative sample of Colletid bees
[0403] Method:
[0404] Protein solutions of Module 1 BNMPs (HnMl-01 (SEQ ID NO: 4), EaMl-01 (SEQ ID NO: 8), HeM l-01 (SEQ ID NO: 16) and MiMl-01 (SEQ ID NO: 20) were prepared at a concentration of 1 mg / ml dissolved in 98% formic acid. Protein solutions were drop-cast on glass slides and then dried at room temperature. The coated glass slides were fixed in a Theta Flow Tensiometer (Biolin Scientific, UK). The water contact angles (WCA) were recorded continuously over 60 s.
[0405] Result:
[0406] Images from WCA measurements are shown in Figure 8. The surface of glass treated with formic acid only (control condition) was hydrophilic, showing a water-contact angle of 37- 38° which reduced to 30-31° after 60 seconds of equilibration. Glass coated with HnMl-01 showed contact angles of 77-81° on contact, which reduced to 45° after 60 seconds equilibration. Glass coated with EaMl-01 showed contact angles of 84° on contact, which stayed stable at 82-83° after 60 seconds of equilibration. Glass coated with HeMl-01 showed contact angles of 82-85° on contact, which reduced to 45-54° after 60 seconds equilibration. Glass coated with MiM l-01 showed contact angles of 86-88° on contact, which stayed stable at 85-87° after 60 seconds of equilibration.
[0407] Without wishing to be bound by theory, the inventors hereby demonstrate that Module 1 proteins are capable of modifying the relative hydrophobicity of materials due to their amphipathic nature, such as may be desirable when used as a film or coating. The inventors also believe that EsMl-01 is reasonably expected to modify the relative hydrophobicity of various materials due to the structural similarities observed between this polypeptide and the other related Colletid bee polypeptides described herein.
[0408] Example 4 - Protein stabilisation of an oil in water emulsion of Module 1 BNMPs from a representative sample of Colletid bees
[0409] Method:
[0410] Protein solutions of Module 1 BNMPs (HnMl-01 (SEQ ID NO: 4), EaMl-01 (SEQ ID NO: 8), EsMl-01 (SEQ ID NO: 12), HeMl-01 (SEQ ID NO: 16) and MiMl-01 (SEQ ID NO: 20) were dissolved in 8 M guanidine hydrochloride to make a 4% (40 mg / mL) solution. The dissolved protein was then diluted with deionised water creating a 1% protein solution in 2M guanidine hydrochloride. Emulsions were created at a 7:3 ratio of aqueous protein solutiomMCT oil (Medium-chain triglycerides; New Directions Australia) containing 5 pg / mL Nile Red for HnMl-01, EaMl-01, and MiMl-01. Emulsions were created at a 1 : 1 ratio of aqueous protein solution: MCT oil containing 5 pg / mL Nile Red for EsMl-01 and HeMl-01. The sample was emulsified by ultrasonication (Omni Sonic Ruptor 400 Ultrasonic Homogeniser) using 15 x 1-second pulses with a high-intensity tip. The oil in water emulsion was incubated at room temperature and monitored for 4 days. Controls without protein were prepared with 2M guanidine hydrochloride similarly. The tubes were photographed to assess bulk phase separation.
[0411] Result:
[0412] Images of the water:oil emulsions are shown in Figure 9. The absence of an emulsifying agent would be expected to result in the separation of the oil and water phases, with the Nile red dye remaining in the top (oil) phase. This separation is visible in the control sample tubes on day 0 and day 4. The presence of an emulsifying agent would show a lack of phase separation with a turbid appearance. Tubes containing HnMl-01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01 protein show an emulsion is maintained for at least four days.
[0413] On the basis of these observations, the inventors have shown that the Module 1 proteins as described herein possess surface activity. Without wishing to be bound by theory, the inventors believe this surface activity is likely due to their amphipathic structure. Therefore, in an oil / Module 1 protein solution system, protein molecules are thought to be adsorbed at the oil / buffer interface to minimise the surface tension. Module 1 proteins at the interface are likely rearranged to expose hydrophilic chains toward the buffer phase and hydrophobic chains toward the oil phase, which consequently facilitates the stabilisation of oil / buffer emulsions in this mixture.
[0414] Example 5 - Hydrogel formation of Module 1 BNMPs from a representative sample of Colletid bees
[0415] Method:
[0416] Module 1 BNMPs were dissolved in 6 M guanidine thiocyanate at 1 mL per lOmg of EsMl- 01 and MiMl-01 and 1 mL per lOOmg of HnMl-01, EaMl-01 and HeMl-01. The protein was dispersed by vortexing and homogenisation (Miccra minibatch D-9 with DS-8-P attachment). The protein dispersion was subjected to inversion mixing for 20 minutes at room temperature. The resulting dissolved sample was centrifuged at 12,000g for 10 mins at 20°C. The protein solution was incubated at room temperature until hydrogel formation. Result:
[0417] Images of hydrogels formed by this method are shown in Figure 10. A hydrogel would be expected to remain intact in the base of a tube after inversion for at least 30 seconds ('inversion test')- After overnight incubation, the HnMl-01, EaMl-01, EsMl-01, HeMl-01 and MiMl-01 protein solutions in 6M guanidine thiocyanate formed hydrogels which passed the inversion test.
[0418] On the basis of these observations, the inventors have shown that the module Module 1 proteins as described herein form hydrogels, which have applications in many fields including biomedical, industrial materials and consumer products. Without wishing to be bound by theory, the inventors believe this is due to the common structural features of these proteins as described herein.
[0419] Table 1 - Nucleic acid and Amino acid sequences
[0420]
[0421]
[0422]
[0423]
[0424]
[0425]
[0426]
[0427]
[0428]
[0429]
[0430]
[0431]
[0432]
[0433]
[0434]
[0435]
[0436]
[0437]
[0438]
[0439]
[0440]
[0441] References
[0442] 1. Bairoch, A. (2000). The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000. Nucleic Acids Research, 28(1), 45-48. https: / / doi.Org / 10.1093 / nar / 28. l.45
[0443] 2. Bondarev, S. A., Bondareva, O. v, Zhouravleva, G. A., & Kajava, A. v. (2018). BetaSerpentine: a bioinformatics tool for reconstruction of amyloid structures. Bioinformatics, 34(4), 599-608. https: / / doi.org / 10.1093 / bioinformatics / btx629
[0444] 3. Bonenfant, Q., Noe, L., 8i Touzet, H. (2023). Porechop_ABI: discovering unknown adapters in Oxford Nanopore Technology sequencing reads for downstream trimming. Bioinformatics Advances, 3(1). https: / / doi.org / 10.1093 / bioadv / vbac085
[0445] 4. Camacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., & Madden, T. L. (2009). BLAST+ : architecture and applications. BMC Bioinformatics, 10(1), 421. https: / / doi.org / 10.1186 / 1471-2105-10-421
[0446] 5. de la Rubia, I., Srivastava, A., Xue, W., Indi, J. A., Carbonell-Sala, S., Lagarde, J., Alba, M. M., & Eyras, E. (2022). RATTLE: reference-free reconstruction and quantification of transcriptomes from Nanopore sequencing. Genome Biology, 23(1), 153. https: / / doi.org / 10.1186 / sl3059-022-02715-w
[0447] 6. Edgar, R. C. (2004). MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research, 32(5), 1792-1797. https: / / doi.org / 10.1093 / nar / gkh340
[0448] 7. Finn, R. D., Clements, J., & Eddy, S. R. (2011). HMMER web server: interactive sequence similarity searching. Nucleic Acids Research, 39(suppl), W29-W37. https: / / doi.org / 10.1093 / nar / gkr367
[0449] 8. Heger, A., & Holm, L. (2000). Rapid automatic detection and alignment of repeats in protein sequences. Proteins: Structure, Function, and Genetics, 41 (2), 224-237. https: / / doi.org / 10.1002 / 1097-0134(20001101)41 :2<224: :AID- PROT70>3.0.CO;2-Z
[0450] 9. Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34(18), 3094-3100. https: / / doi.org / 10.1093 / bioinformatics / btyl91 Na chtiga 11, P. G., Kashiwabara, A. Y., & Durham, A. M. (2021). CodAn : predictive models for precise identification of coding regions in eukaryotic transcripts. Briefings in Bioinformatics, 22(3'). https: / / doi.org / 10.1093 / bib / bbaa045 Nip, K. M., Hafezqorani, S., Gagalova, K. K., Chiu, R., Yang, C., Warren, R. L., & Birol, I. (2023). Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2. Nature Communications, 14(1), 2940. https: / / doi.org / 10.1038 / s41467-023-38553-y Rice, P., Longden, I., & Bleasby, A. (2000). EMBOSS: the European molecular biology open software suite. Trends in Genetics, 16(6), 276-277. Steinig, E., 8i Coin, L. (2022). Nanoq : ultra-fast quality control for nanopore reads. Journal of Open Source Software, 7(69), 2991. https: / / doi.org / 10.21105 / joss.02991 Vacic, V., Uversky, V. N., Dunker, A. K., & Lonardi, S. (2007). Composition Profiler: a tool for discovery and visualization of amino acid composition differences. BMC Bioinformatics, 8(1), 211. https: / / doi.org / 10.1186 / 1471-2105-8- 211
Claims
What we claim is:
1. An isolated polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
2. An isolated polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 or 24.
3. An isolated polynucleotide or functional portion thereof comprising at least 70% nucleic acid sequence identity to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 or 23.
4. An isolated polynucleotide encoding a polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
5. An isolated polypeptide or functional portion thereof comprising, consisting essentially of or consisting of a polypeptide having at least 70% amino acid sequence identity to SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, or 66.
6. An isolated polynucleotide or functional portion thereof comprising, consisting essentially of or consisting of a polynucleotide having at least 70% nucleic acid sequence identity to SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63 or 65.
7. A polypeptide of claim 1, 2, 4 or 5 that is from the nesting material of a Colletid bee.
8. A polypeptide of claim 1, 2, 4, 5 or 7 that is expressed in the Salivary or Dufour's gland of the Colletid bee.
9. A polypeptide of claim 1, 2, 4, 5, 7 or 8 that is a component of the nesting material from a Colletid bee.
10. A polypeptide of claim 7, 8 or 9 wherein the Colletid bee is selected from the following genera: Hylaeus spp., Euryglossa spp., and Meroglossa spp.
11. A polypeptide of claim 7, 8, 9 or 10 wherein the Colletid bee is selected from Hylaeus nubilosus, Hylaeus euxanthus Euryglossa adelaidae, Euryglossa subsericea and Meroglossa impressifrons penetrate.
12. A polypeptide of claim 1 or 2 that comprises, or a polypeptide of claim 4 or 5 that comprises a repeat region of about, preferably of, 110 to 258 amino acids in length.
13. A polypeptide of claim 12 wherein the repeat region comprises at least two regions of alternating relative hydrophobicity and hydrophilicity.
14. A polypeptide of claim 13 wherein the region of relative hydrophobicity comprises at least 20% to 30% hydrophobic amino acid residues.
15. A polypeptide of claim 13 or 14 wherein the repeat regions are similar or substantially similar repeat sub regions that form beta sheets.
16. A polypeptide of claim 15 wherein the repeat sub regions that form beta sheets are identified by rapid automatic detection and alignment of repeats (RADAR).
17. A polypeptide of any one of claims 12 to 16 wherein the repeat regions comprise at least 14% to 36% glutamine residues and at least 18% to 31% serine residues.
18. A polypeptide of any one of claims 1,2, 4, 5, or 7 to 17 wherein the polypeptide forms a hydrogel in about 6M guanidine thiocyanate, preferably wherein the hydrogel comprises at least lOmg / ml of the polypeptide.
19. A polynucleotide of claim 3 or claim 6 that is modified for expression in a heterologous host cell, preferably an isolated heterologous host cell, preferably a prokaryotic host cell, preferably a bacterial host cell, preferably E. coli.
20. A polynucleotide of claim 3, 6 or 19 that is comprised in a vector or nucleic acid expression construct, preferably wherein the vector or nucleic acid expression construct further comprises a heterologous regulatory element.
Citation Information
Patent Citations
Silk proteins
WO2007038837A1
Processes for producing silk dope
WO2011022771A1
Recombinant polypeptide
WO2024116130A1