Recombinant spider silk proteins
Recombinant spider silk proteins with extended REP domains from Larinioides sclopetarius address the spinnability and mechanical property limitations of existing proteins, achieving silk fibers with enhanced performance.
Patent Information
- Application Number
- PCT/SE2025/050253
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2025-03-21
- Publication Date
- 2025-09-25
AI Technical Summary
Existing recombinant spider silk proteins lack long REP domains necessary for producing silk fibers with mechanical properties similar to naturally occurring spider silk, limiting their spinnability and effectiveness.
Development of recombinant spider silk proteins with REP domains derived from Larinioides sclopetarius, comprising at least 100 amino acids, which are spinnable into silk fibers with excellent mechanical properties, mimicking naturally occurring silk proteins by having REP domains that are 3-5 times longer than existing proteins like NT2RepCT.
The recombinant spider silk proteins produce silk fibers with improved mechanical properties, mirroring those of naturally occurring silk fibers, despite being larger in amino acid length, enabling scalable production and spinnability.
Smart Images

Figure SE2025050253_25092025_PF_FP_ABST
Abstract
Description
[0001]RECOMBINANT SPIDER SILK PROTEINS TECHNICAL FIELD The present invention generally relates to recombinant spider silk proteins, and in particular such recombinant spider silk proteins comprising a REP domain derived from the repetitive region of a silk protein of Larinioides sclopetarius. BACKGROUND Spiders can spin seven types of silk, each with unique mechanical properties, produced in different glands; major ampullate, minor ampullate, flagelliform, tubuliform, aciniform and aggegate. These silksare made up of silk proteins that are named according to their primary gland of expression, majorampullate spider silk protein (MaSp), minor ampullate spider silk protein (MiSp), flagelliform spider silk protein (FlSp), tubuliform spider silk protein (TuSp), aciniform spider silk protein (AcSp), aggregate spider silk protein (AgSp) and pyriform spider silk protein (PySp), respectively. Spider silk proteins, also referredto as spidroins, have an N-terminal (NT) domain, an extensive repetitive region (REP), and a C-terminal(CT) domain. The mechanical properties are believed to be dictated by the REP domain. The most extensible fiber, the flagelliform silk, is mainly made from spider silk proteins (FlSps) that carry a Pro-rich REP region, which is predicted to form spring-like structures. The strongest fiber, the major ampullate silk, also referred to as dragline, is mainly composed of spider silk proteins (MaSps) that carry a repeat region of iterated Gly-rich and poly-Ala repeats. The tensile strength of the major ampullate silk is derived from the MaSp poly-Ala blocks that form beta-sheet crystals in the silk fiber, while the Gly-rich parts mediate the fiber’s extensibility. The MaSp silk is the toughest natural fiber known (around 150 MJ / m3). A recombinant spider silk protein having improved solubility in water and thereby allowing scalable production at high yields is known in the art as NT2RepCT (WO 2018 / 002216); Andersson et al., NatureChemical Biology 11: 309-315 (2017)). NT2RepCT contains a His6-tag, a NT domain from Euprosthenopsaustralis MaSp1, two Gly-rich and poly-Ala tandem repeats from E. australis (2Rep) and a CT domainfrom Araneus ventricosus MiSp.WO 2020 / 092769 discloses an engineered polypeptide including at least two units, wherein each unit includes a MaSp4 repeat unit of Caerostris darwini. The C. darwini MaSp4 differs from other MaSps by being largely comprised of GPGPQ amino acid motifs, which give the spider fiber extensibility. This high extensibility is said to lead to the high toughness of the C. darwini silk fibers.WO 2023 / 167628 discloses a recombinant spider silk protein comprising an NT domain, a REP domainand a CT domain. The REP domain comprises a set of domains according to the formula pA1-pG-pA2. pG represents a glycine-rich domain and pA1 and pA2 represent alanine-rich domains. One of pA1 and pA2 is a poly-alanine domain and the other of pA1 and pA2 is a poly-alanine domain having every third or fourth alanine residue replaced by an isoleucine residue or a valine residue. There is, though, still a need for a recombinant spider protein with long REP domains that can be expressed and are spinnable into silk fibers. SUMMARY It is a general objective to provide a recombinant spider silk protein with long REP domains that is capable of producing silk fibers. This and other objectives are met by embodiments of the present invention. The present invention is defined in the independent claims. Further embodiments of the invention are defined in the dependent claims.An aspect of the invention relates to a recombinant spider silk protein comprising an N-terminal (NT)domain, a repetitive region (REP) domain and a C-terminal (CT) domain. The REP domain comprises atleast 100 amino acids derived from the repetitive region of a spider silk protein of Larinioides sclopetarius.The NT domain is derived from the NT domain of a spider silk protein other than the spider silk protein, from which the REP domain is derived and / or the CT domain is derived from the CT domain of a spidersilk protein other than the spider silk protein, from which the REP domain is derived.Further aspects of the invention relate to a silk fiber made of a recombinant spider silk protein accordingto above, a synthetic material comprising a silk fiber according to above, a nucleic acid molecule encoding a recombinant spider silk protein according to above, an expression vector comprising a nucleic acid molecule according to above, a hos cell comprising the expression vector according to above.An additional aspect of the invention relates to a method for producing a silk fiber. The method comprisesextruding a spinning dope comprising a recombinant spider silk protein according to above into an aqueous buffer having an acidic pH to induce polymerization of the recombinant spider silk protein into a silk fiber. The method also comprises isolating the silk fiber from the aqueous buffer. The recombinant spider silk proteins of the invention can be expressed, purified and spun into silk fibershaving excellent mechanical properties with comparatively long REP domains, including at least 4-5 timeslonger REP domains as compared to NT2RepCT. The recombinant spider silk proteins of the invention thereby closer mimics the naturally occurring silk proteins in terms of REP domain lengths. BRIEF DESCRIPTION OF THE DRAWINGS The embodiments, together with further objects and advantages thereof, may best be understood by making reference to the following description taken together with the accompanying drawings, in which:Fig.1. Swedish bridge spider, L. sclopetarius and stress-strain curves of its dragline silk. (1a) A femaleL. sclopetarius spider. (1b) Stress-strain curves of the dragline silk obtained from forcefully silked spiders.Fig. 2. Spidroin catalogue from L. sclopetarius. (2a) Schematic view of an orb-weaving spider with oneset of each type of silk gland indicated. (2b) Schematic figure of the major ampullate gland. The glandhas three anatomical parts: the tail, the sac, and the duct. The tail and sac are composed of a singlelayered epithelium, in which three morphologically distinct cell types are found, each localized to one ofthree zones (A–C). (2c) Phylogenetic tree of the N-terminal domain from the 35 spidroins identified in L.sclopetarius. Numbers on the branches indicate bootstrap values. Spidroins in bold were identified byproteomics analysis of the major ampullate gland and silk. (2d) Heat map with the expression of allspidroins in different tissues as determined by bulk RNA sequencing. The grayscales correspond to thenormalized counts shown in the scale (bottom left inset). (2e) Schematic illustration of the spidroin genes.All spidroin genes encoded proteins with a signal peptide (not shown) and an N-terminal domain. Mostspidroin genes encoded a canonical C-terminal domain, except FlSp-like, which was found to have anon-canonical C-terminal domain, and AmSp-like1 and AmSp-like2 and AgSp-like spidroins, whichcompletely lacked the C-terminal domain. The repetitive motifs in each spidroin are represented asblocks. (2f) Table showing the number of typical MaSp repeat motifs found in each of the MaSps. (A)nrefers to poly-alanine motifs, X in GGX and GPGXX represents any amino acid residue.Fig. 3. Expression of the 17 silk genes in the major ampullate tail, sac and duct. (3a) Relativequantification of the proteins identified in the major ampullate gland and in dissolved silk fibers using LC-MS / MS proteomics, marked according to protein classes (MaSp1, MaSp2, MaSp3, MaSp4, AmSp-like,and SpiCE-LMa). (3b) Schematic figure of the gland showing the three parts (tail, sac, and duct), thatwere separated for RNA sequencing experiments. (3c) PLS analysis of the bulk RNA data separates the three different parts (tail, sac and duct). Diamonds represent samples, small circles represent the overlayof the 17 silk genes. (3d) Heatmap showing the relative expression levels for the 17 silk genes in the tail,sac and duct samples. Gene names are shown in the right. This analysis separates the genes into six clusters based on their expression profiles shown as a dendrogram on the y-axis. The clusters areindicated based on whether the highest gene expressions were in the tail or sac parts, respectively. Thebar on the right indicates protein length. The scale ranges from 0 (white) to 10000 aa (black). (3e) Fractionof different categories of amino acid residues in each of the 17 silk proteins (%). The figure indicatesdifferent categories (small nonpolar: A, G, P, S, T; hydrophobic: I, L, M, V; polar: D, E, H, K, N, Q, R; aromatic and cysteine: C, F, W, Y).Fig. 4. Spatial transcriptomics of silk glands. (4a) An H&E-stained section of the spider abdomen. Theinset shows a lateral view of the abdomen with the approximate plane where the section was made. (4b) The spots were annotated as different silk glands based on morphology of the tissue and the spatiallocation. In (4a) and (4b), sp and pd indicate the location of the spinnerets and the pedicel, respectively.(4c) UMAP analysis of all spots, from eight sections, that were manually annotated as silk glands. Eachdot represents a spot in the spatial sections.Fig.5. Spatial resolution of silk protein expression in the zone A, B and C. (5a) One of the major ampullateglands in section 1 contains several cross sections of the tail and a cross sectioned sac (H&E staining).Inset shows the original image from which the region was magnified (white square). (5b) The spotscorresponding to the major ampullate gland were annotated as zone A, B or C based on the morphologyand staining of the epithelium overlaying each spot. (5c) Heatmap showing the expression of the 17 silkgenes in the 847 spots annotated as zone A, B and C, respectively. Each bar on the heatmap representsa spot on the spatial section and black boxes indicate marker genes in different zones. (5d) Expressionprofiles of MaSp1a, MaSp2b, MaSp3a, SpiCE-LMa1, SpiCE-LMa2 and SpiCE-LMa3 (in order) in thethree zones of major ampullate gland shown in (5a) and (5b).Fig 6. Single cell RNA sequencing analysis of major ampullate gland reveals eight cell types. (6a) UMAPreveal eight cell types in the major ampullate gland. Each dot represents a cell. The cell types were annotated by correlating the gene expression with the bulk RNA and spatial transcriptomic data. This revealed that six cell types make up the secretory epithelium of the tail and sac; three cell types were confined to zone A, one to zone B, one to zone C and one could be found in all three zones. Two celltypes were found in the duct. (6b) Heatmap of relative gene expression of the 17 silk proteins in differentsingle cell types. Black boxes indicate the marker genes in different cell types.Fig.7. Spatial distribution of the eight major ampullate cell types (7a) H&E-stained section of the spidermajor ampullate gland, inset shows the area magnified. (7b) The same section as in (7a) with the zones(A, B and C) of the major ampullate gland indicated. The zones were identified based on the cell morphology and 11 cross-sections were obtained for this gland (9 for zone A, 1 for zone B and 1 for zone C). (7c) Deconvolution of the spatial spots using the scRNAseq data generates a pie chart for each spot, where the grayscales represent cell types identified from the scRNAseq data and shows the fraction of cells that belongs to each cell type. (7d) Relative gene expression of the silk protein genes that are markergenes for the cell types in zone A as a function of the perimeter of the gland cross-sections. Thegrayscales correspond to the three zone A cell types and the lines are the linear models of the geneexpression profiles for the respective cell types. (7e) Hematoxylin mean value as a function of theperimeter of the major ampullate gland cross sections in zone A. (7f) Average fraction of cell typesidentified from scRNAseq data in zone A (proximal, middle, and distal), zone B and zone C spots asevaluated from the deconvolution plots. (7g) Schematic figures showing the spatial distribution of theeight cell types in the major ampullate glands.Fig. 8. Origin and model of the multi-layered architecture of the major ampullate silk. (8a) Histologicalsections showing the morphology of the single layered epithelium in zone A, B and C, respectively, in theL. sclopetarius major ampullate gland (H&E staining). Arrowheads indicate basally located nuclei.Secretions from zone A, B and C forming three layers can be discerned in the lumen (indicated as I, II and III in the zone C panels). The round structures around the intracellular vesicles in the first three panelsindicate the annotated regions for image analysis in QuPath. Scale bar = 20 µm. (8b) H&E intensity plotsof the annotation objects shown in (8a). (8c) Schematic image of the major ampullate gland illustratingthe localization of the eight cell types and the layered secretions. The round structures below theschematic gland show schematic drawings of cross-sections along the gland. (8d–8g) Origin of the 17silk proteins in the major ampullate silk and their expression profiles in the major ampullate gland. Below the panels the 17 silk genes / proteins are listed. The top panel (8d) shows the level of gene expression for the 17 silk genes as determined by bulk transcriptomic data. The numbers refer to the log2 normalized expression values. Panel (8e) shows the gene expression values of the 17 silk proteins in the six cell types that are found in zone A, B and C, and the black frames indicate proteins that were identified as marker genes for these types. Panel (8f) shows the abundance of each of the 17 silk proteins in solubleextracts from major ampullate fibers incubated in 2, 4 and 8 M urea, respectively. Grayscales indicateexpression values, and the black frames indicate proteins that were identified as markers. The bottom panel (8g) shows the cumulative protein abundance in the whole silk fiber dissolved using HFIP, LiBr andurea. (8h) Model of the multi-layered architecture of the major ampullate silk. The grayscales correspondto the cell types the layers originate from.Fig.9. The results from bacterial cell lysis and IMAC purification for three constructs (MaSp4_631, MaSp4_400and MaSp2c_500) are visualized on an SDS PAGE gel. T = Total cell lysate, SN = Supernatant after centrifugation,Filt. SN = Filtered (0.2 µm) supernatant, FT = Flow-through IMAC and E = Eluate.Fig.10. Fibers spun from concentrated MaSp4_400 spidroin solution by extrusion into a coagulation bath (0.75 Macetate, pH 5) by a HPLC pump.DETAILED DESCRIPTION The present invention generally relates to recombinant spider silk proteins, and in particular such recombinant spider silk proteins comprising a REP domain derived from the repetitive region of a silk protein of Larinioides sclopetarius. The spider silk proteins of the invention are recombinant or engineered spider silk proteins, i.e., are artificial and non-naturally occurring spider silk proteins. The recombinant spider silk proteins are preferably in the form of isolated recombinant spider silk proteins. The recombinant spider silk proteinsof the invention can produce silk fibers having excellent mechanical properties in par with mechanicalproperties of silk fibers spun from NT2RepCT (WO 2018 / 002216; Andersson et al., Nat Chem Biol 11: 309-315 (2017)), which is said to have high tensile strength. However, the spider silk proteins of the invention are significantly larger in terms of number of amino acids as compared to NT2RepCT, while still being spinnable. In more detail, the spider silk proteins of the invention have significantly longer repetitiveregion (REP) domains as compared to NT2RepCT, which has a REP domain of merely 77 amino acids.In fact, REP domains of spider silk proteins of the invention can be at least 3-5 times longer, or even more, as compared to the REP domain of NT2RepCT. This means that the recombinant spider silk proteins of the invention have REP domains that are more similar to the naturally occurring spider silk proteins. The recombinant spider silk proteins of the invention comprise a repetitive region (REP) domain derived of a silk protein of Larinioides sclopetarius. L. sclopetarius, commonly called bridge-spider or gray cross- spider, is a relatively large orb-weaver spider with Holarctic distribution.L. sclopetarius creates circular orb webs unlike other orb-web spiders that construct elliptical orb webs.Additionally, their orb-webs change in shape as the spider ages. As the spider matures, the adhesive web's lower-section will continue to increase whereas the web's upper section will become proportionallysmaller. This discrepancy in web-size becomes more prominent as the spider gets larger.L. sclopetarius produces major ampullate silk fibers with impressive mechanical properties. The genomeof this spider species was sequenced, assembled, and annotated. Manual curation resulted in a spidroincatalog of 35 complete spidroin genes including 12 MaSps, 4 MiSps, 2 PySps, 4 TuSps, 4 FlSps, 3AcSps, 4 AgSps and 2 AmSps. An aspect of the invention therefore relates to a recombinant spider silk protein comprising an N-terminal (NT) domain, a repetitive region (REP) domain and a C-terminal (CT) domain. According to the invention,the REP domain comprises at least 100 amino acids derived from the repetitive region of a spider silkprotein of L. sclopetarius. Furthermore, the NT domain is derived from the NT domain of a spider silkprotein other than the spider silk protein, from which the REP domain is derived. Alternatively, or inaddition, the CT domain is derived from the CT domain of a spider silk protein other than the spider silkprotein, from which the REP domain is derived.The recombinant spider silk proteins of the present invention thereby comprise domains from at least twodifferent spider silk proteins, preferably from at least two different spider species. The REP domain isderived from the repetitive region of a silk protein of L. sclopetarius. In an embodiment, the NT domain isthen derived from the NT domain of a spider silk protein from another spider silk protein of L. sclopetarius,or preferably from a spider silk protein from another spider species, i.e., a spider species other than L.sclopetarius. The CT domain could then be derived from the CT domain of a spider silk protein from L.sclopetarius, such as from the same or different spider silk protein from L. sclopetarius as the REP domainis derived, from the same other spider species as the NT domain or indeed from yet another spiderspecies different than L. sclopetarius and the spider species, from which the NT domain is derived. Inanother embodiment, the CT domain is derived from the CT domain of a spider silk protein from another spider silk protein of L. sclopetarius, or preferably from a spider silk protein from another species, i.e., a spider species other than L. sclopetarius. The NT domain could then be derived from the NT domain of a spider silk protein from L. sclopetarius, such as from the same or different spider silk protein from L.sclopetarius as the REP domain is derived, from the same other spider species as the CT domain orindeed from yet another spider species different than L. sclopetarius and the spider species, from whichthe CT domain is derived.In an embodiment, the NT domain is derived from the NT domain a spider silk protein of a different spiderspecies than L. sclopetarius and / or the CT domain is derived from a CT domain of a spider silk proteinof a different spider species than L. sclopetarius. In a particular embodiment, the NT domain is derivedfrom the NT domain a spider silk protein of a different spider species than L. sclopetarius and the CTdomain is derived from a CT domain of a spider silk protein of a different spider species than L. sclopetarius. In this particular embodiment, the NT domain and the CT domain could be derived from the same spider silk protein of the different spider species, from different spider silk proteins of the different spider species or from different spider silk proteins from different spider species. The recombinant spider silk protein preferably comprises the REP domain arranged between the NT domain and the CT domain. Hence, the recombinant spider silk protein preferably has the general formula NT-REP-CT. As is further described herein, the recombinant spider silk protein may also comprise other amino acid sequences than the NT, REP and CT domains, including optional N-terminal and / or C-terminal tags and / or optional linkers. Hence, in an embodiment, the recombinant spider silk protein has the generalformula (X)-NT-(L1)-REP-(L2)-CT-(Y). In this embodiment, X represents an optional N-terminal tag, Yrepresents an optional C-terminal tag, L1 represents an optional first linker and L2 represents an optional second linker. As mentioned above, the recombinant spider silk proteins of the invention may contain additional amino acid sequences or domains in addition to the NT domain, the REP domain and CT domain. Such additional domains are then preferably attached to the N-terminus of the NT domain of the recombinant spider silk protein and / or to the C-terminus of the CT domain of the recombinant spider silk protein, i.e., X-NT-REP-CT, NT-REP-CT-Y or X-NT-REP-CT-Y, and / or could be provided between the NT and REP domains and / or between the REP and CT domains, i.e., NT-L1-REP-CT, NT-REP-L2-CT or NT-L1-REP- L2-CT. It is also possible to combine the N-terminal and / or C-terminal tags, X, Y, with linkers, such as X- NT-L1-REP-CT, X-NT-REP-L2-CT, X-NT-L1-REP-L2-CT, NT-L1-REP-CT-Y, NT-REP-L2-CT-Y, NT-L1- REP-L2-CT-Y, X-NT-L1-REP-CT-Y, X-NT-REP-L2-CT-Y or X-NT-L1-REP-L2-CT-Y. Illustrative, but non-limiting, examples of such additional domains X, Y are affinity tags, solubilization tags, chromatography tags, epitope tags, fluorescence tags, signal peptides or sequences, etc. Examples of domains facilitating purification include various affinity tags, such as chitin binding protein (CBP), maltose binding protein (MBP), hemagglutinin tag, Strep-tag and glutathione-S-transferase (GST), and poly(His) tags, such as His6 tag; solubilization tags, such as thioredoxin (TRX) and poly(NANP); chromatography tags, such as FLAG-tag; epitope tags, such as ALFA-tag, V5-tag, Myc-tag,HA-tag, Spot-tag, T7-tag and NE-tag; and fluorescence tags, such as GFP. An example of N-terminal tagthat can be used is defined in SEQ ID NO: 135. Illustrative examples of linkers that could be used between the NT and REP domains and / or between theREP and CT domain are various GS linkers, such as GS, SGS or GNS, and other peptide linkers. Suchlinkers may be beneficial to provide short distance between the NT and REP domains and / or between the REP and CT domain and thereby reduce the risk of any steric hindrance between the linked domains. The optional link may be very short, such as GS, SGS, GNS, or up to some tens of amino acids, preferably no more than 20 amino acids, more preferably no more than 15 amino acids. In an embodiment, the REP domain of the recombinant spider silk protein consists of at least 100 aminoacid residues derived from the repetitive region of a silk protein of L. sclopetarius. Experimental data asshown herein indicates that silk fiber could be spun from spider silk proteins comprising comparatively long REP domains, i.e., 130 amino acid residues or more, and the spun silk fibers had excellent mechanical properties. In fact, spider silk proteins with REP domains of 400 amino acid residues or even longer could be spun into silk fibers with excellent mechanical properties.In a preferred embodiment, the REP domain comprises, preferably consists of, a sequence of at least125 amino acid residues derived from the repetitive region of a silk protein of L. sclopetarius, preferablyat least 130 amino acid residues derived from the repetitive region of a silk protein of L. sclopetarius,more preferably at least 150 amino acid residues derived from the repetitive region of a silk protein of L.sclopetarius, such as at least 200 amino acid residues, at least 250 amino acid residues, at least 300amino acid residues, at least 350 amino acid residues derived from the repetitive region of a silk proteinof L. sclopetarius, and most preferably at least 400 amino acid residues derived from the repetitive regionof a silk protein of L. sclopetarius, such as at least 450 amino acid residues, at least 500 amino acidresidues, at least 550 amino acid residues or at least 600 amino acid residues derived from the repetitiveregion of a silk protein of L. sclopetarius.In an embodiment, the REP domain comprises, preferably consists of, a sequence of at least 100consecutive amino acid residues, preferably at least 125 consecutive amino acid residues and more preferably at least 130 consecutive amino acid residues derived from the repetitive region of a silk proteinof L. sclopetarius. In a particular embodiment, the REP domain comprises, preferably consists of, asequence of at least 150 consecutive amino acid residues derived from the repetitive region of a silkprotein of L. sclopetarius, such as at least 200 consecutive amino acid residues, at least 250 consecutive amino acid residues, at least 300 consecutive amino acid residues, at least 350 consecutive amino acidresidues derived from the repetitive region of a silk protein of L. sclopetarius, and most preferably at least400 consecutive amino acid residues derived from the repetitive region of a silk protein of L. sclopetarius,such as at least 450 consecutive amino acid residues, at least 500 consecutive amino acid residues, atleast 550 consecutive amino acid residues or at least 600 consecutive amino acid residues derived fromthe repetitive region of a silk protein of L. sclopetarius. The spider silk protein of the invention preferably has a molecular weight of at least 30 kDa, preferably at least 35 kDa, and more preferably at least 50 kDa. Experimental data as shown herein further show that spider silk proteins with a molecular weight of 50-60 kDa or even larger (87 kDa) could be spun into silk fibers.In an embodiment, the REP domain is derived from the repetitive region of a MaSp of L. sclopetarius. L.sclopetarius comprises twelve MaSps denoted MaSp1a - MaSp1c, MaSp2a – MaSp2f, MaSp3a –MaSp3b and MaSp4.In a particular embodiment, the REP domain is derived from a MaSp of L. sclopetarius selected from thegroup consisting of MaSp1a as defined in SEQ ID NO: 38, MaSp1b as defined in SEQ ID NO: 41, MaSp1c as defined in SEQ ID NO: 44, MaSp2a as defined in SEQ ID NO: 47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3a as defined in SEQ ID NO: 65, MaSp3bas defined in SEQ ID NO: 68 and MaSp4 as defined in SEQ ID NO: 71.In a preferred embodiment, the REP domain is derived from a MaSp of L. sclopetarius selected from thegroup consisting of a MaSp1 protein, a MaSp2 protein, a MaSp3 protein and MaSp4. For instance, theREP domain is preferably derived from a MaSp of L. sclopetarius selected from the group consisting ofMaSp1a as defined in SEQ ID NO: 38, MaSp2a as defined in SEQ ID NO: 47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3a as defined in SEQ ID NO:65 and MaSp4 as defined in SEQ ID NO: 71. In a particular preferred embodiment, the REP domain isderived from a MaSp of L. sclopetarius selected from the group consisting of MaSp2c as defined in SEQID NO: 53 and MaSp4 as defined in SEQ ID NO: 71.An example of a REP domain derived from MaSp1a of L. sclopetarius include BR_MaSp1a_400 below:BR_MaSp1a_400 (SEQ ID NO: 148) GGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGSGSQGVGQGGYGAGQGGAGAAAAAG GAGGSGQGGLGAGQGYGAGLGGQGGAGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGAAS AAAAAVGAAGGQGGYGGLGSQGAGQGGYGAGQGGATSAAAAAAGGSGGQGGYGGLGSQGAGQ SGSGSAAAAAAAGGAGGAGQGGLGAGQGYGPGLGGQRGAGQGGAASAAAAAAGGAGGQGGYG GFGSQGAGQGGYGAGQGGAASAAAAAGGAGGQGVYGGLGSQGAGQGGYGAGQGGAGSAAAAA AAVGEGGAGQGGLSAGQGYGSGLGGQGGAGQGGAASSAAAAGGSGGQGGYGGLGSQGAGQGG AASAAAAAGGAGGQGGYGGLGSQGAGQExamples of REP domains derived from MaSp2c of L. sclopetarius include BR_MaSp2_short,BR_MaSp2_long, BR_MaSp2_300 and BR_MaSp2_400 disclosed here below: BR_MaSp2_short (SEQ ID NO: 109) SAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSS GPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGG YGPGSQ BR_MaSp2_long (SEQ ID NO: 110) GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPG SSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPG GYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGP GSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPBR_MaSp2_300 (SEQ ID NO: 111)GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPG SSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPG GYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGP GSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQG GPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGBR_MaSp2_400 (SEQ ID NO: 112)GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPG SSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPG GYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGP GSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQG GPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSS AASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSG PGGYGPGSQGPSGPSGPGGYGPGSQGGPSAn example of a REP domain derived from MaSp2c of L. sclopetarius include BR_MaSp2c_500 below:BR_MaSp2c_500 (SEQ ID NO: 149) GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPG SSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPG GYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGP GSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQG GPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSS AASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSG PGGYGPGSQGPSGPSGPGGYGPGSQGGPSSGSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGP GGYGPGSQGPSGPSGPGGNGPGSQGGPSGPGGYGPGSQGPNGPGGAGSSAAVAVSSGPGGYG PGSQGGPAn example of a REP domain derived from MaSp3a of L. sclopetarius include BR_MaSp3a_400 below:BR_MaSp3a_400 (SEQ ID NO: 150) GGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAA SAAAGGSGGPRGYGGLGSQGAGQGGYGAGQGGAASAAAGGSGGPGGYGGLGSQGAGQGGYGA GQGGAASAAAASAGGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQG GYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQGSGQGGYGAGQDGASSVAAAAVSGSGGPGG YGGLGSQGAGQGRYGAGQGGADSTAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAA AAAAGGSGGPGRYGGLGSQRSGQGGYGAGQGGAASAASAAAGGSGGPRGYGGLGSQGTGQGG YGAGQSGAASAASAAAGGSGGPRGYGGLExamples of REP domains derived from MaSp4 of L. sclopetarius include BR_MaSp4_short,BR_MaSp4_long, BR_MaSp4_400 and BR_MaSp4_631 disclosed here below:BR_MaSp4_short (SEQ ID NO: 113)GPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQGPSVPGPQGPGSSVSI STSYKPDQQGPSGPSQQGPSTQVSNGPGPQAPALSTFAFSGPVPEASSGPSAQQPSFQGPAG PRPQGPGSBR_MaSp4_long (SEQ ID NO: 114)GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQ QGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYG LSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGH QAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGS BR_MaSp4_400 (SEQ ID NO: 151) GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQ QGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYG LSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGH QAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSGSVSVLSYGPGPQGPSGLSQ QGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPGSSVSIFSYG PGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTGPGP QAPALSTFAFSGPVPQGPSGPVPQGPSPQ BR_MaSp4_631 (SEQ IN NO: 168) GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQ QGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYG LSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGH QAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSNSGFQGPGSSGGALTSYGTG PQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPAL TNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQG PSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEAS SGPSTQQPSFQGPARPRPQGPGSSGSVSVLSYGPGPQGPSGLSQQGPSTQVPTGSGPQAPAL TNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPGSSVSIFSYGPGLQGPSGPVPQGPSPQG PSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGP SGPVPQGPSPQThe bold parts of BR_MaSp4_631 correspond to MaSp4_long and the underlined part of BR_MaSp4_631corresponds to MaSp4_400, which parts are interconnected by an SNS linker. Hence, in an embodiment, the REP domain comprises, preferably consists of, an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, 168. In a particular embodiment, the REP domain comprises, preferably consists of, an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, such as an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114. In another particular embodiment, the REP domaincomprises, preferably consists of, an amino acid sequence selected from the group consisting of 109 to112. In yet another particular embodiment, the REP domain comprises, preferably consists of, an amino acid sequence selected from the group consisting of SEQ ID NO: 113 to 114.In another embodiment, the REP domain is derived from the repetitive region of a MiSp of L. sclopetarius.L. sclopetarius comprises four MiSps denoted MiSpa – MiSpd.In a particular embodiment, the REP domain is derived from a MiSp of L. sclopetarius selected from thegroup consisting of MiSpa as defined in SEQ ID NO: 74, MiSpb as defined in SEQ ID NO: 77, MiSpc asdefined in SEQ ID NO: 80 and MiSpd as defined in SEQ ID NO: 83.In a further embodiment, the REP domain is derived from the repetitive region of a TuSp of L. sclopetarius.L. sclopetarius comprises four TuSps denoted TuSpa – TuSpd.In a particular embodiment, the REP domain is derived from a TuSp of L. sclopetarius selected from thegroup consisting of TuSpa as defined in SEQ ID NO: 92, TuSpb as defined in SEQ ID NO: 95, TuSpc asdefined in SEQ ID NO: 98 and TuSpd as defined in SEQ ID NO: 101.In yet another embodiment, the REP domain is derived from the repetitive region of a PySp of L.sclopetarius. L. sclopetarius comprises two PySps denoted PySpa and PySpb.In a particular embodiment, the REP domain is derived from a PySp of L. sclopetarius selected from thegroup consisting of PySpa as defined in SEQ ID NO: 86 and PySpb as defined in SEQ ID NO: 89.In another embodiment, the REP domain is derived from the repetitive region of an AcSp of L.sclopetarius. L. sclopetarius comprises three AcSps denoted AcSpa – AcSpc.In a particular embodiment, the REP domain is derived from an AcSp of L. sclopetarius selected from thegroup consisting of AcSpa as defined in SEQ ID NO: 2, AcSpb as defined in SEQ ID NO: 5, and AcSpcas defined in SEQ ID NO: 8.In a further embodiment, the REP domain is derived from the repetitive region of an AgSp of L.sclopetarius. L. sclopetarius comprises four AgSps denoted AgSpa – AgSpc and AgSp-like.In a particular embodiment, the REP domain is derived from an AgSp of L. sclopetarius selected from thegroup consisting of AgSp-like as defined in SEQ ID NO: 11, AgSp1 as defined in SEQ ID NO: 13, AgSp2aas defined in SEQ ID NO: 16 and AgSp2b as defined in SEQ ID NO: 19. In another particular embodiment,the REP domain is derived from an AgSp of L. sclopetarius selected from the group consisting of AgSp1as defined in SEQ ID NO: 13, AgSp2a as defined in SEQ ID NO: 16 and AgSp2b as defined in SEQ ID NO: 19.In yet another embodiment, the REP domain is derived from the repetitive region of a FlSp of L.sclopetarius. L. sclopetarius comprises four FlSps denoted FlSpa – FlSpc and FlSp-like.In a particular embodiment, the REP domain is derived from an FlSp of L. sclopetarius selected from thegroup consisting of FlSp-like as defined in SEQ ID NO: 26, FlSpa as defined in SEQ ID NO: 29, FlSpbas defined in SEQ ID NO: 32 and FlSpc as defined in SEQ ID NO: 35. In another particular embodiment,the REP domain is derived from an FlSp of L. sclopetarius selected from the group consisting of FlSpaas defined in SEQ ID NO: 29, FlSpb as defined in SEQ ID NO: 32 and FlSpc as defined in SEQ ID NO: 35.In an embodiment, the REP domain is derived from the repetitive region of an ampullate spider silk protein(AmSp) of L. sclopetarius. L. sclopetarius comprises two AmSps denoted AmSp-like 1 and 2.In a particular embodiment, the REP domain is derived from an AmSp of L. sclopetarius selected fromthe group consisting of AmSp-like 1 as defined in SEQ ID NO: 22 and AmSp-like 1 as defined in SEQ IDNO: 24.In an embodiment, the REP domain is derived from an amino acid sequence selected from the groupconsisting of SEQ ID NO: 2, 5, 8, 11, 13, 16, 19, 22, 24, 26, 29, 32, 35, 38, 41, 44, 47, 50, 53, 56, 59, 62, 65, 68, 71, 74, 77, 80, 83, 86, 89, 92, 95, 98 and 101.The NT domains of spider silk proteins are thought to improve the solubility of the spider silk protein andthereby enabling very high protein concentrations in the spinning dope. Furthermore, the pH dependencyof the solubility of the NT domains is an important factor for allowing rapid polymerization of the spinningdope.Some CT domains of spider silk proteins do not exhibit a pH-sensitive solubility (Hedhammar et al.,Biochemistry 47(11): 3407-3417 (2008)), but in general most CT domains having several charged aminoacid residues are in fact highly soluble and have a pH dependent solubility (Andersson et al., PLoS Biology 12(8): e1001921 (2014)).The recombinant spider silk protein of the invention could use various combinations of NT domain andCT domain together with the REP domain to form a recombinant spider silk protein that is spinnable intoa silk fiber.Illustrative, but non-limiting, examples of NT domains that could be used according to the presentinvention are listed in Table 2 in US 2019 / 0248847, the teaching of which regarding NT domains is herebyincorporated by reference. In a preferred embodiment, the NT domain of the recombinant spider silk protein is derived from the NTdomain of Euprosthenops australis MaSp1.In a particular embodiment, the NT domain comprises, preferably consists of SEQ ID NO: 115.NT domain of Euprosthenops australis MaSp1 (SEQ ID NO: 115)MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNK LQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMF AQAGMNDVSA In another embodiment, the NT domain is derived from the NT domain of a silk protein of L. sclopetarius.An illustrative, but non-limiting, example of such a NT domain is a NT domain derived from MaSp2f (SEQID NO: 61).Illustrative, but non-limiting, examples of CT domains that could be used according to the present invention are listed in Table 1 in US 2019 / 0248847, the teaching of which regarding CT domains is hereby incorporated by reference.In a preferred embodiment, the CT domain of the recombinant spider silk protein is derived from the CTdomain of Araneus ventricosus MiSp.In a particular embodiment, the CT domain comprises, preferably consists of SEQ ID NO: 116.CT domain of Araneus ventricosus MiSp (SEQ ID NO: 116)VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIY SGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG In another embodiment, the CT domain is derived from the CT domain of a silk protein of L. sclopetarius.An illustrative, but non-limiting, example of such a CT domain is a CT domain derived from MiSpc (SEQID NO: 81). In an embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, 168 and a CT domain comprising, preferably consisting of SEQ ID NO: 116. In an embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consistingof an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151and a CT domain comprising, preferably consisting of SEQ ID NO: 116.In another embodiment, the recombinant spider silk protein comprises, preferably consists of, an NTdomain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, and a CT domain comprising, preferably consisting of SEQ ID NO: 116. In a particular embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 109, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 103, 117-119.In another particular embodiment, the recombinant spider silk protein comprises, preferably consists of,an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 110, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 104, 120-122.In a further particular embodiment, the recombinant spider silk protein comprises, preferably consists of,an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 111, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 105, 123-125.In yet another particular embodiment, the recombinant spider silk protein comprises, preferably consistsof, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 112, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 106, 126-128.In another particular embodiment, the recombinant spider silk protein comprises, preferably consists of,an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 113, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 107, 129-131.In a further particular embodiment, the recombinant spider silk protein comprises, preferably consists of,an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 114, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 108, 132-134.In another particular embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 148, and a CT domain comprising, preferably consisting of SEQ ID NO: 116. Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 152, 156-158. In a further particular embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 149, and a CT domain comprising, preferably consisting of SEQ ID NO: 116. Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 153, 159-161. In yet another particular embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 150, and a CT domain comprising, preferably consisting of SEQ ID NO: 116. Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 154, 162-164. In another particular embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 151, and a CT domain comprising, preferably consisting of SEQ ID NO: 116. Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 155, 165-167. In a further particular embodiment, the recombinant spider silk protein comprises, preferably consists of, an NT domain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of SEQ ID NO: 168, and a CT domain comprising, preferably consisting of SEQ ID NO: 116. Examples of such a recombinant spider silk protein are defined in SEQ ID NO: 169-172. In an embodiment, the recombinant spider silk protein comprises, preferably consists of, an amino acid sequence selected from the group consisting of SEQ ID NO: 103-108, 117-134, 152-167, 169-172, such as selected from the group consisting of SEQ ID NO: 103-108, 117-134, 152-167. The recombinant spider silk proteins as defined in SEQ ID NO: 117, 120, 123, 126, 129, 132, 156, 159,162, 165, 170 consist of an NT domain consisting of SEQ ID NO: 115, a REP domain consisting of SEQID NO: 109-114, 148-151, 168 and a CT domain consisting of SEQ ID NO: 116.The recombinant spider silk proteins as defined in SEQ ID NO: 118, 121, 124, 127, 130, 133, 157, 160,163, 166, 171 consist of an NT domain consisting of SEQ ID NO: 115, a REP domain consisting of SEQID NO: 109-114, 148-151, 168 and a CT domain consisting of SEQ ID NO: 116 with linkers between theNT and REP domains and between the REP and CT domains. The recombinant spider silk proteins as defined in SEQ ID NO: 119, 122, 125, 128, 131, 134, 158, 161,164, 167, 172 consist of an NT domain consisting of SEQ ID NO: 115, a REP domain consisting of SEQID NO: 109-114, 148-151, 168 and a CT domain consisting of SEQ ID NO: 116 with a N-terminal tag.The recombinant spider silk proteins as defined in SEQ ID NO: 103-108, 152-155, 169 consist of an NTdomain consisting of SEQ ID NO: 115, a REP domain consisting of SEQ ID NO: 109-114,148-151, 168 and a CT domain consisting of SEQ ID NO: 116 with a N-terminal tag and linkers between the NT and REP domains and between the REP and CT domains.In another embodiment, the recombinant spider silk protein comprises, preferably consists of, an NTdomain comprising, preferably consisting of SEQ ID NO: 61, a REP domain comprising, preferably consisting of an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, 168, preferably selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, and more preferably selected from the group consisting of SEQ ID NO: 109 to 114, and a CT domain comprising, preferably consisting of SEQ ID NO: 116.In a further embodiment, the recombinant spider silk protein comprises, preferably consists of, an NTdomain comprising, preferably consisting of SEQ ID NO: 115, a REP domain comprising, preferably consisting of an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, 168, preferably selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, and more preferably selected from the group consisting of SEQ ID NO: 109 to 114, and a CT domain comprising, preferably consisting of SEQ ID NO: 81.In yet another embodiment, the recombinant spider silk protein comprises, preferably consists of, an NTdomain comprising, preferably consisting of SEQ ID NO: 61, a REP domain comprising, preferably consisting of an amino acid sequence selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, 168, preferably selected from the group consisting of SEQ ID NO: 109 to 114, 148 to 151, and more preferably selected from the group consisting of SEQ ID NO: 109 to 114, and a CT domain comprising, preferably consisting of SEQ ID NO: 81. Another aspect of the invention relates to a silk fiber made of a recombinant spider silk protein according to the invention. Hence, this aspect relates to a silk fiber, sometimes referred to as silk polymer, comprising a recombinant spider silk protein according to the invention. The silk fiber is then obtained by spinning a so-called spinning dope comprising the recombinant spider silk protein according to the invention into the silk fiber, which is further described herein. The silk fibers of the invention have very high tensile strength, also referred to as engineered (eng.)strength. In an embodiment, the silk fiber has an average tensile strength of at least 70 MPa, preferablyat least 75 MPa and more preferably at least 80 MPa.Tensile strength and other mechanical properties as referred to herein relate to average tensile strength and average values of the other mechanical properties as determined when testing a plurality of silk fibers. This means that individual silk fibers among the tested silk fibers may have a tensile strengthbelow, for instance, 70 MPa, whereas other individual silk fibers among the tested silk fibers may have atensile strength above, for instance, 70 MPa. However, the average tensile strength among the testedsilk fibers is then at least, for instance, 70 MPa.In an embodiment, the silk fiber has an average toughness modulus of at least 15 MJ / m3, preferably atleast 20 MJ / m3 and more preferably at least 40 MJ / m3.The silk fiber could have an average diameter of from one or a few µm up to several tens of µm. Forinstance, the average diameter of the silk fiber is from 1 µm up to 100 µm, preferably from 2 µm up to50 µm and more preferably from 4 up to 15 µm.In an embodiment, the silk fiber has an average Young’s modulus of at least 1.8 GPa, preferably at least1.9 GPa and more preferably at least 2.0 GPa.In an embodiment, the silk fiber also comprises a spider-silk constituting element (SpiCE) of L. sclopetarius. In a particular embodiment, the SpiCE is selected from the group consisting of SEQ ID NO: 142 to 147. The present invention also relates to a synthetic material comprising a silk fiber according to the invention. Illustrative examples of a synthetic material comprising, or made of, silk fibers of the invention include textile materials, such as filaments, yarns, ropes, and woven material. Such textile materials may benefit from the high tensile strength of the silk fiber. Other examples of synthetic materials include pliant energy absorbing materials, such as armor and bumpers. The silk fibers of the invention can also be used in medical applications, such as in sutures, compression bandages, etc. Additionally the silk fibers can be used in scaffolds and material in tissue engineering, implants and other cell scaffold-based materials. The present invention also relates to a nucleic acid molecule encoding a recombinant spider silk protein according to the invention. Nucleic acid molecule as used herein includes polynucleotide, oligonucleotide, and nucleic acid sequence, and generally means a polymer of DNA or RNA, which may be single-stranded or double- stranded, which may contain natural, non-natural or altered nucleotides, and which may contain a natural, non-natural or altered internucleotide linkage, such as a phosphoroamidate linkage or a phosphorothioate linkage, instead of the phosphodiester found between the nucleotides of an unmodified oligonucleotide. Nucleic acid molecule also includes complementary DNA (cDNA) and messenger RNA (mRNA). Illustrative, but non-limiting, examples of such nucleic acid molecules are presented in SEQ ID NO: 136-141 for the recombinant spider silk proteins in SEQ ID NO: 103-108.A further aspect of the invention relates to an expression vector comprising a nucleic acid molecule according to the invention. The expression vector comprises at least one nucleic acid molecule comprising coding sequences that can be expressed, such as transcribed and translated, in a cell, often denoted host cell, comprising the expression vector. The expression vector is in an embodiment selected among DNA molecules, RNA molecules, plasmids, episomal plasmids and virus vectors. The expression vector then comprises the nucleic acid molecule operatively coupled to a promoter to enable transcription thereof in a host cell. The promoter could be any promoter that it constitutively active or inducibly active in the host cell.The nucleic acid molecule encoding the recombinant spider silk protein is operatively controlled by thepromoter in the expression vector, i.e., is under transcriptional control of the promoter. In an embodiment,the promoter is selected from the group consisting of the human EF1^ promoter, the CMV promoter, theCAG promoter, the PGK promoter, the TRE promoter, the U6 promoter and the UAS promoter if the host cell is an eukaryotic cell, such as a human cell. Illustrative, but non-limiting, examples of a promoter that could be used if the host cell is a bacterial cell include the T5 promoter, the T7 promoter, the rhampromoter, the phoA promoter, the Sp6 promoter, the lac promoter, the AraBad promoter, the trp promoterand the Ptac promoter. If the host cell is a yeast cell the promoter could be selected from the group consisting of the CYC1 promoter, the ADH1 promoter, the TEF2 promoter, the pCYC promoter, a PGAL promoter and the GFD promoter as illustrative, but non-limiting, examples.An illustrative, but non-limiting, example of a promoter that could be used in Escherichia coli host cells isthe T7 promoter. Protein production could then be induced in the host cell by addition of isopropyl ^-D-1-thiogalactoside. Yet another aspect of the invention relates to a host cell comprising the expression vector according to the invention. The nucleic acid molecule or expression vector can then be transcribed in the host cell to produce the recombinant spider silk protein in the host cell. Various such host cells can be used according to the invention including, but not limited to, bacteria,yeast, mammalian cells, plant cells, and insect cells. It is currently preferred to produce the recombinantspider silk proteins of the invention in bacteria, such as E. coli. The recombinant spider silk protein can then be produced by the host cell, for instance, by culturing the host cell according to the invention in conditions allowing production of the recombinant spider silk protein and isolating the spider silk protein from the culture. In a particular embodiment, the spider silk protein is isolated from the cytosol of the host cells. The present invention also relates to a method for producing a silk fiber. The method comprises extruding a spinning dope comprising the recombinant spider silk protein of the present invention into an aqueous buffer having an acidic pH to induce polymerization of the recombinant spider silk protein into a silk fiber. The method also comprises isolating the silk fiber from the aqueous buffer. In an embodiment, the spinning dope comprises at least 90 mg / ml of the recombinant spider silk protein,preferably at least 125 mg / ml and more preferably at least 150 mg / ml of the recombinant spider silkprotein. In an embodiment, the aqueous buffer is an acetate buffer having a pH equal to or below 6, preferably equal to or below 5.5. In a particular embodiment, the aqueous buffer preferably also has a pH equal to or larger than 4, preferably equal to or larger than 4.5. In a preferred, embodiment, the aqueous buffer has a pH of about 5. EXAMPLEEXAMPLE 1Materials and MethodsProtein sequencesThe following minispidroins were designed with repetitive regions obtained from spidroins found in MaSpsfrom the Swedish bridge spider (BR) L. sclopetarius.His-NT-Linker1-BR_MaSp2_short-Linker2-CT (SEQ ID NO: 103) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGP GSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQG PNGPGGPGASAAVAVSSGPGGYGPGSQSGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSS AEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSA SIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp2_long-Linker2-CT (SEQ ID NO: 104) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGG PSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSA ASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGP GGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYG PGSQGPSGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASA LPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSV GQYVG His-NT-Linker1-BR_MaSp2_300-Linker2-CT (SEQ ID NO: 105) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGG PSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSA ASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGP GGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYG PGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQ GGPSGPGSQGPSGSGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAI ASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTL NVVQDSVGQYVG His-NT-Linker1-BR_MaSp2_400-Linker2-CT (SEQ ID NO: 106) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGG PSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSA ASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGP GGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYG PGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQ GGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPG SQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSSGSVTSGGYGYG TSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVS SNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp4_short-Linker2-CT (SEQ ID NO: 107) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQ GPSPQGPSVPGPQGPGSSVSISTSYKPDQQGPSGPSQQGPSTQVSNGPGPQAPALSTFAFSG PVPEASSGPSAQQPSFQGPAGPRPQGPGSSGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRL SSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLS SASIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp4_long-Linker2-CT (SEQ ID NO: 108) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQ GPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSL QGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQ QGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGS SGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVIS NIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVGCloning and protein expressionSynthetic fragments of BR_MaSp2_short; BR_MaSp2_long; BR_Masp2_300; BR__MaSp2_400;BR_MaSp4_short; and BR_MaSp4_long, were cloned in a pT7-His-NT-CT vector using EcoRI andBamHI sites located between NT and CT domains. The NT domain from MaSp1 from Euprosthenopsaustralis (SEQ ID NO: 115) was exchanged against the NT domain from MaSp2f (SEQ ID NO: 61) in thepT7-His-NT-A3I-A-CT vector using NdeI and EcoRI resitriction sites, while the CT domain from Araneusventricosus MiSp (SEQ ID NO: 116) was changed against the CT domain of MiSpc (SEQ ID NO: 81)using BamHI and HindII restriction sites. The resulting constructs were transformed in BL21(DE3)Escherichia coli. LB broth medium containing kanamycin (70 µg / ml) was inoculated with glycerol stockof E. coli containing pT7-His-NT- REP-CT vectors with respective insert. The overnight culture was usedfor a 1 / 100 inoculation LB media containing kanamycin, which was then cultured at 30°C with shaking(110 rpm) until OD600 reached 0.6, after which the temperature was lowered to 20°C and proteinexpression was induced at OD 0.8 by adding isopropylthiogalactoside (IPTG) to a final concentration of0.15 mM. The cells were cultured overnight at 20°C with shaking (110 rpm) and were then harvested bycentrifugation for 20 min at 5,000 rpm at 4°C. The pellets were resuspended in 20 mM Tris pH 8 andstored at −20°C.Protein purificationLysis was performed in a cell disrupter (T-S Series Machine, Constant Systems Limited) at 30 kPsi, afterwhich the lysate was centrifuged at 25,000 ^ g at 4°C for 30 min. The pellet was discarded, but thesupernatant was loaded onto 2 sequentially connected 20 ml HisPrep FF16 / 10 (Cytiva) columns to bindthe His-tagged proteins. The column was washed with 5 column volumes (CV) 20 mM Tris-HCl pH 8.0and 5 CV 2 mM imidazole in 20 mM Tris-HCl pH 8.0. The protein was eluted with 200 mM imidazole in20 mM Tris-HCl pH 8.0. The eluted protein was dialyzed against 20 mM Tris pH 8.0, at 4°C overnight,using a Spectra / Por dialysis membrane with a 6–8 kDa molecular-weight cutoff, while the buffer wasexchanged at least three times. SDS–PAGE (4-20%) and Coomassie Brilliant Blue staining was used todetermine the purity of the protein. Broad Range Protein Ladder (Thermo Fisher Scientific) was used asa size standard. Protein concentration was determined by recording the absorbance of a 10 ^ diluteddialyzed eluate at 280 nm.Preparation of the spinning dopePurified protein was concentrated with an Amicon Ultra-15 centrifugal filter unit (Merck-Millipore,Darmstadt, Germany) equipped with an ultracel-10 membrane (10 kDa cutoff) at 4,000 ^ g and 4°C. Theconcentration of the concentrated protein was determined by recording the absorbance at 280 nm of a333^ diluted sample, in triplicates. Protein solutions were obtained with the following concentrations:BR_MaSp2_short 315 mg / ml; BR_MaSp2_long 280 mg / ml; BR_MaSp2_300277 mg / ml;BR_MaSp2_400300 mg / ml; BR_MaSp4_short 300 mg / ml; BR_MaSp4_long 281 mg / ml. Then,respective protein solutions were transferred to a 1 mL syringe with a luer lock (BD, Franklin Lakes, NJ,USA), and stored at -20°C.Biomimetic spinning of artificial silk fibers Biomimetic spinning of the protein concentrate (spinning dope) was performed according to a method byAndersson, et al., Nature Chemical Biology 13(3): 262-264 (2017) but fine-tuned by Greco et al.Molecules 25(14): 3248 (2020) and Schmuck et al., Communications Material 3: 83 (2022). Briefly, thespinning dope was thawed at room temperature (20-25^C), and a 1 mL syringe was connected to a 27G blunt end steel needle (B. Braun, Melsungen, Germany) with an outer diameter (O.D.) of 0.40 mm. Next, polyethylene tubing (BD Intramedic, Franklin Lakes, NJ, USA) with an O.D. of 1.09 mm and an inner diameter (I.D.) of 0.38 mm was used to encase the needle. The encased needle was then inserted into polyethylene tubing with an O.D. of 1.65 mm and an I.D. of 0.76 mm. Finally, a pulled glass capillary with a tapered opening (G1 Narishige, Tokyo, Japan with an O.D. of 1.0 mm and I.D. of 0.6 mm, pulled using a Micro Electrode Puller, Stoelting Co.51217, Wood Dale, IL, USA) was inserted ~3 cm into the polyethylene tubing with an I.D. of 0.38 mm to enable a leakage free connection. Then, the syringe was placed into a neMESYS low pressure (290 N) syringe pump (Cetoni, Korbußen, Germany). Using a flow rate of 17 µl / min, the dope was extruded through the glass capillary having a tapered tip with an openingof 50 ± 10 µm into an 80 cm long coagulation bath containing a 0.75 M acetate buffer (pH 5). The fiberwas collected continuously at the end of the bath, using a rotating wheel, with a circumference of 35 cm, and a reeling speed of 59 cm s-1, at a relative humidity of < 40%. Tensile testingTensile testing was performed by first mounting a fiber onto a paper frame with 10 × 10 mm squarewindow, as previously described by Greco et al. Molecules 25(14): 3248 (2020). The fibers were attachedto the frame with a double-sided tape. The diameter of the fiber was determined by light microscopy usinga lens with 10× magnification and measuring the diameter three times on arbitrary positions for each ofthree images obtained from three distinct fiber sections. The average diameter was used to compute thecross-sectional area assuming a round cross-section. Then, the paper frame holding the fiber wasmounted in a 5943-Instron tensile tester equipped with a 5 N load cell, before the sides of the paper frame were cut. The displacement experiment speed was 6 mm / min, at a relative humidity below 40% and 22°C. The engineered stress was calculated by dividing the maximal force by the average diameter for each fiber, and the strain at break was obtained considering the total displacement divided by the gauge length. The Youngs modulus was obtained from the slope of linear elastic part of stress-strain curve, while the toughness modulus is the total area under the stress-strain curve. Values reported represent an average of at least 10 fibers tested. No outliers were removed. ResultsAll minispidroin constructs based on MaSps REP from L. sclopetarius were expressed as soluble proteinsand easily purified and concentrated to ~300 mg / ml in native conditions. In addition, all constructs werespinnable, meaning that they could be extruded through a pulled glass capillary into a 0.75 M acetatebuffer having a pH of 5 where the solid fiber formed. The fiber was continuously collected for severalminutes using a reeling speed of 59 cm / s. In addition, the fibers possess mechanical properties that arecomparable to other minispidroins, such as NT2RepCT.Table 1 – Mechanical properties of fibers made from minispidroins based on REP from L. sclopetariusEng. Young’s Toughness Molecular Strain at Diameter Construct Strength Modulus modulus weight Break (µm) (MPa) (GPa) (MJ m-3) (kDa) Br_MaSp2 short 124 ± 49 77% ± 41% 2.7 ± 0.7 72 ± 45 7.8 ± 1.2 37.5Br_MaSp2 long 90 ± 14 89% ± 7% 2.1 ± 0.3 54 ± 10 8.8 ± 0.8 46.0Br_MaSp2300 120 ± 34 60% ± 20% 2.0 ± 0.4 54 ± 23 4.0 ± 0.5 51.1Br_MaSp2400 80 ± 22 33% ± 41% 2.5 ± 0.6 20 ± 25 13.4 ± 5.9 59.3Br_MaSp4 short 117 ± 38 78% ± 23% 2.5 ± 0.5 63 ± 14 8.6 ± 2.2 39.4Br_MaSp4 long 119 ± 50 52% ± 14% 2.7 ± 1.3 46 ± 21 7.7 ± 2.1 48.7NT2RepCT 97 ± 13 112% ± 30% 2.1 ± 0.4 74 ± 15 9.1 ± 2.6Each value in Table 1 represents the average of n ≥ 10 measurements, with errors showing ± onestandard deviation. No outliers were removed.EXAMPLE 2Materials and MethodsProtein sequencesThe following minispidroins were designed with repetitive regions obtained from spidroins found in MaSpsfrom the Swedish bridge spider (BR) L. sclopetarius. His-NT-Linker1-BR_MaSp1a_400-Linker2-CT (SEQ ID NO: 152) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGSGSQ GVGQGGYGAGQGGAGAAAAAGGAGGSGQGGLGAGQGYGAGLGGQGGAGQGGAASAAAAAGGS GGQGGYGGLGSQGAGQGGAASAAAAAVGAAGGQGGYGGLGSQGAGQGGYGAGQGGATSAAAA AAGGSGGQGGYGGLGSQGAGQSGSGSAAAAAAAGGAGGAGQGGLGAGQGYGPGLGGQRGAGQ GGAASAAAAAAGGAGGQGGYGGFGSQGAGQGGYGAGQGGAASAAAAAGGAGGQGVYGGLGSQ GAGQGGYGAGQGGAGSAAAAAAAVGEGGAGQGGLSAGQGYGSGLGGQGGAGQGGAASSAAAA GGSGGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGLGSQGAGQGGSVTSGGYGYGTS AAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSN EALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp2c_500-Linker2-CT (SEQ ID NO: 153) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGG PSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSA ASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGP GGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYG PGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQ GGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPG SQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSSGSGPGGYGPGS QGPNGPGGPGSSAAVAVSSGPGGYGPGSQGPSGPSGPGGNGPGSQGGPSGPGGYGPGSQGPN GPGGAGSSAAVAVSSGPGGYGPGSQGGPSGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRLS SAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSS ASIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp3a_400-Linker2-CT (SEQ ID NO: 154) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGL GSQGAGQGGYGAGQGGAASAASAAAGGSGGPRGYGGLGSQGAGQGGYGAGQGGAASAAAGGS GGPGGYGGLGSQGAGQGGYGAGQGGAASAAAASAGGSGGRGGYGGLGSQGTGQGGAASAAAA AGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQGSGQGGYGA GQDGASSVAAAAVSGSGGPGGYGGLGSQGAGQGRYGAGQGGADSTAAAAAGGSGGQGGYGGL GSQGAGQGGYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQRSGQGGYGAGQGGAASAASAAA GGSGGPRGYGGLGSQGTGQGGYGAGQSGAASAASAAAGGSGGPRGYGGLGSVTSGGYGYGTS AAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSN EALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp4_400-Linker2-CT (SEQ ID NO: 155) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQ GPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSL QGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQ QGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGS SGSVSVLSYGPGPQGPSGLSQQGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSL QGPAGPQPQGPGSSVSIFSYGPGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGL QGNSGPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQGSVTSGGYGYGT SAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSS NEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His-NT-Linker1-BR_MaSp4_631-Linker2-CT (SEQ ID NO: 169) MGHHHHHHMSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAA QGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINE ITQLVSMFAQAGMNDVSAGNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQ GPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSL QGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQ QGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGS SNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQG PSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRIL SYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNG PGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSGSVSVLSYGPGPQGPSG LSQQGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPGSSVSIF SYGPGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTG PGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQGSVTSGGYGYGTSAAAGAGVAAGSYAGAVN RLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHV LSSASIGNVSSVGVDSTLNVVQDSVGQYVGCloning and protein expressionCloning and protein expression were performed as described above for Example 1. All minispidroinconstructs listed in Example 2 were expressed in E. coli BL-21 as soluble proteins either in a shake flaskculture as described in Example 1 or using a bioreactor (Schmuck et al., High-yield production of a super-soluble miniature spidroin for biomimetic high-performance materials, Materials Today 50: 16-23 (2021)),with an Eppendorf® BioFlo 120 equipped with a 15L single-walled glass vessel.Protein purificationAfter clarification of the bacterial lysate using centrifugation, the minispidroins were purified withchromatography as described in Example 1. When expressing the proteins using a bioreactor, the yieldsof purified proteins after IMAC were 10.6 g / L, 9.2 g / L and 5 g / L for MaSp4_400, MaSp4_631 andMaSp2c_500 respectively. In Fig. 9, the results from bacterial cell lysis and IMAC purification for thesethree constructs are visualized on an SDS PAGE gel (4-20 %). As an alternative purification strategy, theclarified lysate was treated with potassium phosphate buffer pH 8.0, to a final concentration of 330-420mM. This induced precipitation of the spidroin proteins, and the protein pellet was obtained bycentrifugation at 18000 g for 20 minutes. The pellet was dissolved in 20 mM Tris-HCl buffer adjusted topH 10. After solubilizing the protein precipitate, the solution was dialyzed to 20 mM Tris-HCl pH 8.SpinningMaSp4_400 was concentrated to 320 mg / ml as described in Example 1. An MFLX74931-30 (Masterflex)high pressure liquid chromatography (HPLC) pump was used to extrude the concentrated spidroinsolution through a PEEK capillary with an inner diameter of 75 µm and a length of 38mm, into a bathcontaining 0.75 M acetate, pH 5. In the coagulation bath, the extruded spidroin solution formed a fiber,which was collected on a rotating wheel moving at ~3 m / min, see Fig.10.ResultsAll minispidroin constructs based on MaSps REP from L. sclopetarius were expressed as soluble proteinsand easily purified to produce yields at 220-230 mg protein / L of bacterial culture in a shake flask, and upto 10 g / L using a bioreactor. The results showed that the spidroin constructs could be concentrated to>300 mg / mL and could be spun into fibers when extruding them into a coagulation bath of pH 5 with aHPLC pump.EXAMPLE 3In this Example, a complete spidroin repertoire from an orb weaving spider, Larinioides sclopetarius, ispresented herein. Spatial transcriptomics and single-cell RNA (scRNA) sequencing showed that the tailand sac of the major ampullate gland were composed of six silk-producing cell types found in the threedistinct zones (A, B and C). This, combined with image analyses of histological sections and proteomicsanalyses of sequentially dissolved fibers, indicated that the silk fiber is a three-layered structure, in whichthe core was composed of proteins belonging to the MaSp1, MaSp2 and MaSp4 families, the middle layer is dominated by MaSp3, and the outmost layer contains several uncharacterized non-spidroin proteins. Results The L. sclopetarius major ampullate fiber is composed of 17 silk proteinsIn this Example we focus on the Swedish bridge spider (L. sclopetarius) that produces major ampullatesilk fibers with impressive mechanical properties (Fig.1). First, the genome of this spider species was sequenced, assembled, and annotated (Table 2). Manual curation resulted in a spidroin catalog of 35 complete spidroin genes (Fig.2). Next, to reveal the protein composition of the major ampullate gland and silk fiber, respectively, proteomic analysis was performed using liquid chromatography tandem mass spectrometry (LC-MS / MS). In the gland samples, a total of 2,145 proteins were identified in at least one of the three replicates. Dissolving the silk fiber can be challenging and different chemical treatments can extract different proteins. Therefore, three solvents (urea, HFIP and LiBr) were used to extract proteins from the fibers and a total of 130 proteins were identified from all treatments combined. However, since the fibers are easily contaminated by other silk types during spinning and by unrelated proteins during handling, only the proteins that were also found in the major ampullate gland proteomic data, were predicted to have a signal peptide, and were found with an abundance of more than 0.1% of the total protein content of the silk fiber were considered (Fig.3a). This process rendered 17 silk proteins that were named “the 17 silk proteins” (Table 3) and corresponding genes were named “the 17 silk genes” (Table 3). Table 2 – Assembly and annotation statistics of L. sclopetarious genomeGenome assemblyAssembly size (bp) 2274070471GC % 30.5Number of Contigs 1602Longest Contig (bp) 29089857N50 (bp) 5487611N90 (bp) 1366626BUSCO complete (%) 98.4Repeat statistics Number of elements 2761064Length (bp) [% Genome] 47.50%Genome annotations Protein-coding genes 22860Transcript Isoforms 46911tRNA genes 53867BUSCO complete (%) 95.8Table 3 – Ranking of the 17 silk proteins / genes as markers in different methods usedTotal Gene / proteinProtein* RNA**Gland Silk BulkMaSp2b 4 5 208MaSp2c 5 4 7MaSp2e 28 8 26MaSp2f 8 6 8MaSp4 60 7 2966MaSp1a 1 1 6MaSp1b 3 3 12MaSp1c 106 10 56SpiCE-LMa6 707 17 97SpiCE-LMa3 135 13 30SpiCE-LMa4 294 15 378SpiCE-LMa5 482 16 198MaSp3a 2 1 25AmSp-like2 57 11 204AmSp-like1 450 14 293SpiCE-LMa2 213 12 5SpiCE-LMa1 10 9 4Differential gene expression Gene / proteinBulk*** Spatial**** Single cell*****Part Rank Zone Rank Cell type RankMaSp2b Tail 8 A 69 ZoneA_MaSp2 2MaSp2c Tail 5 A 3 ZoneA_MaSp2 6MaSp2e Tail 3 A 7 ZoneA_MaSp2 3MaSp2f Tail 1 A 4 ZoneA_MaSp2 4MaSp4 Tail 15 ZoneA_MaSp2 5ZoneA_MaSp1 2MaSp1a Tail 16 A 5ZoneA_MaSp2 47MaSp1b A 13 ZoneA_MaSp1 1MaSp1c A 14 ZoneA_MaSp1 3ZoneA_SpiCE-LMa3SpiCE-LMa6 A 102 ZoneB_MaSp3SpiCE-LMa3 ZoneA_SpiCE-LMa3 7SpiCE-LMa4 Sac 25 ZoneC_SpiCE-LMa1 68ZoneA_SpiCE-LMa3 79SpiCE-LMa5 ZoneC_SpiCE-LMa1 73ZoneB_MaSp3 1MaSp3a Sac 12 B 3ZoneABC 7AmSp-like2 Sac 28 B 4 ZoneB_MaSp3 7AmSp-like1 Sac 31 B 31 ZoneB_MaSp3 11C 16 ZoneB_MaSp3 10SpiCE-LMa2 Sac 42B 38 ZoneC_SpiCE-LMa1 11SpiCE-LMa1 Sac 46 C 7 ZoneC_SpiCE-LMa1 6 * Rank based on protein levels in the glands and the silk fibers. Rank 1 represents the most abundant protein. ** Rank based on average normalized read counts in major ampullate gland samples. Rank 1 has the most reads. *** DEG analysis on tail, sac and duct parts of the major ampullate glands. Rank is based on average log2-fold change. The higher the rank the more differential expressed the gene is. **** Marker gene identification on spots manually annotated as different zones within the spatial transcriptomic sections. Rank is based on average log2-fold change, the higher the better. ***** Marker gene identification in cell types identified with single cell data. Rank is based on average log2-fold change, the higher the better.The 17 silk proteins included nine MaSps, two Ampullate Spidroin-like proteins (AmSp-like1 and Amp-like2) and six proteins with unknown functions. The six proteins of unknown function were annotated asspider silk constituent elements of Larinioides major ampullate silk (SpiCE-LMa1–6) in order ofabundance. MaSp1 (a–c), MaSp2 (b, c, e, f), MaSp3a and MaSp4 represented 52%, 19%, 24% and 1.6%, respectively, of the protein content of the silk fiber (Fig.3a). The SpiCE-LMa proteins accounted for 3.2% while the AmSp-like proteins only constituted 0.9% (Fig. 3a). The six SpiCE-LMa proteins generally had lower molecular weight compared to the MaSps. In terms of amino acid composition, the SpiCE-LMa1 and 2 resembled the MaSp proteins and had spidroin-like repeat motifs, whereas SpiCE- LMa3–6 were rich in Cys and resembled the amino acid composition of the AmSp-like 1 and 2 (Fig.3e). The predicted secondary structure content and AlphaFold2 structural predictions of the SpiCE-LMa proteins suggested large heterogeneity within this group of proteins. The 17 silk genes are expressed in the tail and the sac Bulk RNA sequencing of whole major ampullate glands was used to verify that the RNA levels correlated with the protein abundance in the gland. Notably, the 17 silk genes ranked among the most highly expressed genes (Table 3). However, this analysis did not permit localization of the expression of specific genes to different parts of the gland. To enhance the resolution of the gene expression profiles, the major ampullate gland was sequenced after being cut into three distinct anatomical parts: tail, sac, and duct (Fig.3b). This implies that the tail samples contained transcripts from zone A while the sac samples contained transcripts from all three zones (Figs.2b, 3b). The major ampullate gene set was used toperform Partial Least Square (PLS) regression on the tail, sac and duct samples which separated thethree types of samples into distinct clusters using only two variables with high predictive relevance (Q- square = 0.857), verifying that the transcriptome in the tail, sac and duct are diverse (Fig.3c). To identify which part of the major ampullate gland the 17 silk proteins were originating from, the 17 silk genes were superimposed on the PLS plot which revealed that all the MaSp genes were expressed in the major ampullate tail, except MaSp3a which was expressed in the sac along with AmSp-like1 and 2 (Fig.3c). Two genes of unknown function, SpiCE-LMa1 and SpiCE-LMa2, were also among genes expressed in the sac. The expression profiles of the silk genes in the three parts (tail / sac / duct) were further visualized by a heatmap (Fig.3d), which clearly indicated that the genes encoding the 17 silk proteins are expressed in the tail and the sac but not in the duct. Hierarchical clustering grouped the 17 silk genes into 6 clusters based on the similarity in their expressionprofiles (Fig. 3d). Of these, cluster 1 genes showed differential expression in the tail samples and wereassigned to the tail, while all the genes in clusters 5 and 6 showed differential expression in sac samples and were assigned to the sac. Clusters 2, 3 and 4 had some genes that were differentially expressed in the tail and clustered in the same node as the cluster 1, and therefore were assigned to the tail. Based on these results, cluster 1 that contained MaSp2c, e, f genes, which were solely expressed in the tail, could be assigned to the zone A (Figs.2b, 3d) but the rest of the genes, which were also expressed in the sac, could not be assigned to one of the zones since the sac samples contained tissue from all the three zones. Expression of the silk genes is spatially resolved in the three zones To improve resolution of the silk gene expression, we next used spatial transcriptomics (10X Visium). This is an unbiased and elegant technique that allows mapping of the gene expression in tissue sections with a resolution of 50 µm. Six sections of the whole abdomen from four female individuals were used and the spots on the hematoxylin and eosin (H&E) stained sections were manually annotated to different silk glands based on histology (one section is shown in Figs.4a–4b). The sequencing data from all spots from all the sections were then isolated and visualized using Uniform Manifold Approximation and Projection (UMAP). In the UMAP, the spots assigned as silk glands clustered together and separated from the spots belonging to other tissue types. Within the silk gland cluster, the spots from different silk glands clustered together, in line with the manual annotation (Fig.4c). On the spatial sections, the spidroin expression was specific to corresponding glands and confirmed the results from the bulk-RNA expression profiles (Fig.2d), attesting to the quality of the spatial data. The spots annotated as covering major ampullate gland tissue could be further assigned as zone A, B orC based on the morphology of the epithelial cells in six of the sections (Figs.5a, 5b). This resulted in 847spots from six sections that clustered according to zone in the UMAP, indicating that the expression profiles in these zones are indeed different. Furthermore, marker genes for zone A, B and C identified from the spatial data overlapped with the bulk-RNA PLS plot of tail, sac and not the duct. The expressionof the 17 silk genes on all the spots annotated as zone A, B or C was then visualized as a heatmap (Fig.5c). Statistical analyses of the expression levels revealed that 13 of the 17 silk genes were predominantly expressed in one of the zones (Table 3). MaSp1 (a–c), MaSp2 (b,c,e,f) and SpiCE-LMa6 genes had significantly higher expression in zone A cells, MaSp3a, AmSp-like1, and AmSp-like2 in zone B cells, and SpiCE-Lma1 in zone C cells. The expression of SpiCE-LMa2 was significantly higher in both zone B and zone C. The expression of the remaining three proteins, SpiCE-LMa3, 4 and 5 was similar in all three zones (Fig.5c). The expression profiles of selected genes in one of the major ampullate glands in section 1 are shown as examples in Fig.5d. The 17 silk genes in the three zones are specifically expressed in six cell types Interestingly, the MaSp1 and MaSp2 genes were both expressed in the zone A cells, but the spatial transcriptomics analyses revealed that their expression profiles differed in different regions of zone A (Fig.5). This suggests that the epithelial zones could harbor several different cell types. To elucidate this, single-cell RNA sequencing was performed on the whole major ampullate glands isolated from 7 individuals. After QC, filtering and analysis, 9700 cells were obtained which clustered into eight groups. Based on the gene overlap with the first two components of PLS analysis of the bulk RNA gene sets, these eight clusters were found to originate from either tail, sac, or duct of the gland. In agreement with the results from using the bulk RNA data, three of the cell clusters were specific for the tail, three for the sac and two for the duct. The clusters were then compared with the spatial transcriptomic data which allowed them to be classified as zone A / B / C cells based on their gene overlap with the marker genes of the three different zones. Three clusters were identified in zone A, one in zone B, one in zone C and one cluster was found in all three zones. The clusters were assigned as cell types and named according to the zone and top marker spider silk protein gene (Fig.6a). One of the zone A cell types had MaSp1a–c as top three marker genes and was named as ZoneA_MaSp1. The second zone A cell type showed MaSp2a,b,e as the top three marker genes and was therefore annotated as ZoneA_MaSp2. This celltype included MaSp4 as the fourth marker gene. In the last zone A cell type, SpiCE-LMa3 was the topmarker gene among silk genes and was named ZoneA_SpiCE-LMa3. The top marker genes for the zone B and C cell types were MaSp3a and SpiCE-LMa1, respectively. These cell types were annotated accordingly. The ZoneB_MaSp3 cells had AmSp-like 1 and 2, two of the spidroins lacking the C-terminaldomain (Fig.2e), as top marker genes. The final cell type was mainly found in the sac but could not beassigned to a specific zone and was therefore named ZoneABC. This cell type also had MaSp3a as amarker gene. The remaining two cell types were named according to their similarity in expression profile to the bulk-RNA data from the duct as Duct_1 and Duct_2, and these two cell types did not have any of the silk proteins as marker genes (Fig.6b). A compelling observation is that all the 17 silk genes were among the top 100 marker genes in at least one of the six cell types when considering all the 22860 protein coding genes. In most cases they wereamong the top ten (Table 3 and Fig.6b). All the cell types have a distinct set of silk genes expressed asevidenced by that 12 of 17 silk genes were significantly expressed in only one of the six cell types while the remaining five genes were expressed in two cell types (Table 3). In summary, the presence of the 17 silk proteins in the fiber can be directly related to the expression of the corresponding genes in the six cell types. The six silk producing cell types can be spatially resolved along the gland To identify the spatial location of the cell types in the major ampullate gland, the scRNAseq data was combined with spatial transcriptomics data and deconvolution of the two datasets was performed. This allowed us to visualize the spatial distribution of cell types in the spots annotated as major ampullate glands on the spatial sections (Figs.7a–7c). The pattern that emerged suggested that the cell types confined to zone A may not be evenly distributed along this zone. To address this, we turned to image analysis using QuPath. Briefly, the sectioned major ampullate glands were annotated according to zones, as shown in Fig.7b, resulting in 101 cross-sections for zone A, 2 for zone B, 9 for zone C and 5 for the duct across all sections. This annotation was confirmed by digital color deconvolution in QuPath which generates values for the degree of hematoxylin and eosin staining, respectively, for each area of interest. When plotted in graphs, the H&E staining of the epithelium correlated well with the annotation of thedifferent zones. Furthermore, the perimeter of all the cross sectioned parts of the gland were determinedusing QuPath and associated with the spots on the spatial transcriptomic sections. The spatial spots corresponding to each cross-section in zone A were extracted and the mean expression of marker genesin zone A as a function of the perimeter of the cross-section was visualized (Fig.7d). This revealed higherMaSp2 expression in zone A cross-sections with smaller perimeters (proximal tail parts) while MaSp1 and SpiCE-LMa3 genes exhibited higher expression in zone A cross-sections with larger perimeters (toward the sac). Notably, a significant negative correlation was observed between hematoxylin values and the cross-section perimeter in zone A, further supporting the presence of different cell types along this zone (Fig.7e). Next, based on the perimeter values, the zone A cross-sections were split into three parts: proximal, middle and distal. The distribution of cell types in the spots corresponding to different parts of zone A was obtained by combining the scRNAseq, spatial and the QuPath data and visualized in Fig.7f. Proximal zone A regions showed higher abundance of ZoneA_MaSp2 cells, which gradually decreased in middle and distal parts. Conversely, ZoneA_MaSp1 cells were more prevalent in distal parts compared to the most proximal region. ZoneA_SpiCE-LMa3 cells were found in along the whole length zone A, but most frequently in the middle and distal parts. ZoneB_MaSp3 cells were almost exclusively found in spots annotated as zone B, affirming their identification as zone B cells. Zone C spots were dominated by ZoneC_SpiCE-LMa1 cells, with a minor fraction of other cell types. ZoneABC cells were found in all the zones with low abundance. The distribution of the different cell types along the major ampullate gland is illustrated in Fig.7g. The location of the silk producing cell types determine the composition of layers in the silk fiberThe secretions originating from zones A–C within L. sclopetarius major ampullate glands stained distinctlywith H&E and were separated in the gland lumen (Fig.8a). To confirm the zone-specific origin of these secretions, the vesicles in the cells of each zone and the secreted substances forming layers in the lumen were annotated and evaluated for staining intensity using QuPath image analysis. The obtained H&E intensity values for the vesicle content and the secreted substances within the lumen were plotted and color-coded according to their respective zones. The annotations for vesicle content and secretions from each zone clustered together, clearly indicating that the zone A secretion formed the bulk of the silk feedstock in the sac lumen, the secretion from zone B contributed a surrounding middle layer, while the secretion from zone C formed the outermost thin layer (Fig.8b). To verify that the three layers in the liquid silk feedstock persist in the major ampullate silk fiber, asolubilization protocol involving the use of different concentrations of urea (2 M, 4 M and 8 M) wasemployed. Low concentrations of urea (2 M) will not solubilize the entire silk and therefore enrich forproteins present at the surface, while the highest concentrations (8 M) will dissolve the whole fiber. LC-MS / MS analysis of the solubilized fractions provided the relative fraction for each protein in the differenturea concentrations. In 2 M urea, the relative fraction of SpiCE-LMa 1–5, and AmSp-like 1 & 2 proteinswere significantly enriched (p < 0.05) compared to the samples dissolved in 8 M urea, while the MaSp1,MaSp2b-c, MaSp2e-f, and MaSp4 were significantly enriched (p < 0.1) when using the higherconcentration of urea (Fig.8f). In supernatants from fibers that were exposed to 4M urea, we could detectseveral of the 17 silk proteins, of which MaSp3a was the most abundant. Next, we compared the protein composition in our proteomics data to the corresponding spatial gene expression in the gland. Amongthe seven proteins enriched in the 2 M urea samples, three overlapped with the marker genes for theZoneC_SpiCE-LMa1, three with ZoneB_MaSp3 cells and one with the ZoneA_SpiCE-LMa3 cells (Figs.8e, 8f). All the eight proteins enriched in the 8 M urea samples overlapped with the ZoneA_MaSp1 and ZoneA_MaSp2 cells marker genes. Interestingly, MaSp3a, the most significant marker gene ZoneB_MaSp3, showed the highest relative fraction in the 4M urea samples (Fig.8f). Advanced molecular methodologies, such as single-cell RNA sequencing and spatial transcriptomics, serve as pivotal tools in advancing our comprehension of diverse tissue biology. However, their effectiveness depends on well-annotated genomes, posing a significant challenge when investigating non-model organisms. This challenge has been particularly pronounced in investigations of spider silk production due to the nature of spidroin genes, characterized by their substantial size and repetitiveness. Addressing this limitation, we present a high-quality genome assembly of L. sclopetarius, featuring nearly complete annotations for all coding genes. Notably, our annotation reveals the presence of 35 full-length spidroins which is higher than previously reported. The spidroins exhibit variable lengths spanning 575 to 9146 amino acid residues (Fig.2). All the spidroin genes encoded proteins with signal peptides, the N-terminal domain, and a repetitive region, and had a stop codon in-frame. Manual inspection of the 3’downstream sequences revealed no sign of shifted reading frames, which means that the spidroin gene catalogue described herein is the first to encompass exclusively complete genes.Since the focus of this Example was to provide a detailed understanding of the structure and function ofthe major ampullate silk gland, as well as the fiber it produces, we next sought to determine the most abundant proteins in the fiber. This is not a trivial task, since the major ampullate fiber easily is contaminated by other silk types during spinning and collection of the fibers. To avoid these problems, we used LC-MS / MS, both on major ampullate silk fibers continuously collected from a spider mounted under a microscope and on isolated major ampullate glands. By only considering the secretory proteinsthat were found in both data sets, we could ensure that any contaminating proteins from other glandswere removed. By filtering proteins with abundance of more than 0.1%, we identified 17 proteins as being the most abundant in the major ampullate silk. Nine of the 17 came from the four distinct classes of the major ampullate spidroins, namely MaSp1-4. Together, these spidroin proteins constitute 96% of the total protein content of the fiber (Fig.8g). In addition to the MaSps, a previously unreported type of spidroin, which we named AmSp-like, was also found in the fiber. The AmSps were characterized by having an N- terminal domain that grouped with the minor and major ampullate and a repetitive region, but intriguingly,these spidroins lacked the C-terminal domain and displayed a different amino acid composition comparedto the MaSp genes (Fig.3e). Finally, six additional proteins, designated as SpiCE-LMa1 to SpiCE-LMa6 were identified as constituents of the fiber. Despite their common naming 'SpiCE-LMa', their properties are diverse in terms of molecular weight, tertiary structure predicted by Alphafold2 and amino acid composition. The amino acid composition of SpiCE-LMa1 and SpiCE-LMa2 that are expressed in zoneC was similar to that of the MaSps, yet they lack both N- and C-terminal domains. Notably, SpiCE-LMa3–6 proteins, that were expressed and secreted from zones A and B have a high cysteine content, and possibly, these could form cysteine slip knots that could contribute to fiber toughness. The over-representation of MaSps in the L. sclopetarius major ampullate silk, and the presence of a small fractionSpiCE proteins, are in line with the reported protein content in the major ampullate silk from other spider species. To determine which cell types that express the 17 silk proteins and their spatial distribution in the gland, we combined three unbiased transcriptomics techniques. First, we used bulk RNA sequencing to revealthat all the 17 silk genes are expressed in the tail and the sac and not in the duct (Fig. 3d). Twelve ofthese were significantly differentially expressed in either the tail or the sac of the gland indicating that the expression profile indeed differs along the gland. Second, single-cell RNA sequencing analysis of wholemajor ampullate glands identified eight cell types. Notably, the marker genes of six of these eight celltypes overlapped with all 17 silk genes (Table 3 and Fig.6b). By cross-referencing the marker genes ofthe scRNA cell types with the differentially expressed genes found in the bulk RNA data, we were able to assign three cell types to the tail, three to the sac, and two to the duct. The transcriptional profile of cells expressing spider silk proteins could be matched to the bulk-RNA sequencing data from the tail and sac samples, but not to the samples derived from the duct. This means that the 17 silk proteins are produced by the six cell types located in the tail and sac. In order to spatially resolve the distribution ofthe cell types, we used a third transcriptomic technique, 10X Visium. In the sections used for spatialtranscriptomics, we first manually annotated the transcriptomic spots within the major ampullate gland zone A, B, and C using the distinct H&E staining pattern and morphology of the epithelium in the respective zones (Fig 5b). Comparing the expression of genes between the zones revealed that 13 out of the 17 silk genes are differentially expressed in one of these three zones (Fig.5c). Next, by integrating the spatial transcriptomics with the single-cell data, the precise localization of the cell types was revealed. In line with the bulk RNA data, the three cell types assigned to the sac were predominately present in the zones B and C. Notably, by using the spatial transcriptomics data we could see a clear distinction between ZoneB_MaSp3 cells that were confined to zone B, while ZoneC_SpiCE-LMa1 cells were dominant in the zone C epithelium. The three cell types that were assigned to the tail were indeed found in Zone A. The tail is long and winding and by taking advantage of the observation that the cross section of the tail increases along the gland, we could generate information about the spatial location of cell types even within zone A. By determining the perimeter of each cross sectioned part of the tail, we could order themfrom proximal (small) to distal (large) (Fig. 7). We found that the ZoneA_MaSp2 cells were primarilylocalized to the proximal part of zone A, while ZoneA_MaSp1 cells were most abundant in the distal portion of zone A (closer to zone B cells). ZoneA_SpiCE-LMa3 cells were present in all parts of the zone A (Fig.7). Finally, again in line with the bulk RNA data, the two cell types assigned to the duct were not detected in the tail or sac. Taken together, our data support that the proteins that make up the major ampullate silk are produced by six cell types that have specific regional anatomical localizations in zone A, B and C, but not by the cell types confined to the duct. Next, we sought to connect the expression of the genes in the cell types along the gland to the multiplelayers observed in the silk feedstock (Fig.8a). By using digital image analysis, we showed that the H&Estaining of the layers in the dope (the liquid feedstock stored in the sac) as determined by QuPath corresponds to the staining of the intracellular vesicles in the corresponding epithelium. The inner layer of the dope matched the staining of the vesicles in the zone A, the middle layer stained as the vesicles in the zone B and the outer layer matched the staining of the vesicles in the zone C epithelial cells (Figs.8a, 8b). This indicates that the three zones indeed produce layered secretions with different proteincompositions. Moreover, since we knew the presence of different cell types in zone A, B and C, respectively, and their expression profiles (Fig.8e), we could predict the protein composition of the different layers in the fiber. Given this model (Fig.8c), the inner layer of the silk fiber contains the MaSp1, MaSp2, MaSp4, SpiCE-LMa3, and SpiCE-LMa6 proteins, the middle layer primarily contains the MaSp3, AmSp-like1–2, and SpiCE-LMa2 proteins, while the outer layer is dominated by SpiCE-LMa1 and SpiCE- LMa2 but contains no classical spidroins. In order to test the hypothesis that the layers identified in the gland lumen persist to form layers in the fiber, we ran proteomic analysis on silk fiber extracts by exposingmajor ampullate silk fibers to different concentrations of urea (2 M, 4 M, and 8 M, respectively).Supernatants from fibers incubated in 2 M urea contained proteins primarily expressed in zone B andzone C, whereas in 8 M urea, that completely dissolve the silk fibers, proteins that are expressed in zoneA were enriched (Fig.8f). When using the 4 M urea, MaSp3a, a marker gene for zone B, had the highestrelative abundance. These results allow us to present a detailed model of the composition of the three- layered major ampullate silk fiber, and to conclude that each layer has a distinct protein composition that is derived from specific cell types confined to zone A, B and C, respectively (Figs.8f-8h). Since the ZoneA_MaSp2 cells are the dominating cell type in the most proximal part of the tail, it is logical that MaSp2 proteins form the core of the fiber and that the more peripheral regions of the core are dominated by MaSp1 proteins, secreted by the ZoneA_MaSp1 cells which are located more distally inZone A. This finding contrasts to a report from Hu et al. (Hu, et al., A molecular atlas reveals the tri-sectional spinning mechanism of spider dragline silk. Nat Commun 14, 837 (2023)), which shows that thecentral core of the Trichonephila major ampullate fiber is dominated by MaSp1 proteins and MaSp2 arefound more peripherally, but is in line with work by Sponner et al. (Sponner, et al., Composition andhierarchical organisation of a spider silk. Plos One 2, e998 (2007)), who by biochemical andimmunohistochemical investigations of the Trichonephila major ampullate silk revealed the presence ofboth MaSp1 and 2 in the inner core but exclusively MaSp1 in the peripheral parts of the core. The latter study also concludes that the layer surrounding the core of the fiber (referred to as skin layer) is more tough and resistant to chemical treatment than the outmost layer (coat) and the central core. If these notions are combined with the data presented herein, a plausible conclusion is that the skin layer described by Sponner et al. is dominated by MaSp3 and corresponds to the middle layer in our model.The mechanical properties of L. sclopetarius major ampullate silk is one of the best performing silk fibersstudied. However, the presence of layered structures and three epithelial zones in the major ampullate glands of several distantly related spider species have been reported (that do not express MaSp3), which suggests that not only the specific protein composition of the fiber layers but also the layered structure in itself may be important for the fiber’s properties. For example, spider species that do not express MaSp3(e.g., Pisauridae and Agelinidae) also have three epithelial zones in the major ampullate gland, and thePisauridae spider Euprosthenops australis spins a fiber with one of the highest tensile strengths reported.In summary, the Example provides a high-resolution spatial map of the different cell types present in themajor ampullate gland of L. sclopetarius. Previously uncharacterized genes are identified that are highly and differentially expressed in the different cell types. The protein compositions of the enigmatic layered secretions in the gland are revealed and shown to correspond to the protein composition of sequentially dissolved fibers. Methods Spider SamplesL. sclopetarius adult female spiders were collected in the wild in a small habitat in Uppsala, Sweden.Taxonomic identity of the spider was verified by the Museum of Natural History, Stockholm, Sweden. The spiders were kept in big containers that allowed them to spin webs. They were fed with meal worms or Drosophila flies weekly and watered daily. Extraction and sequencing of genomic DNA High molecular weight (HMW) genomic DNA was extracted from the whole body of one individual L.sclopetarius spider using MagAttract HMW DNA Kit (Qiagen). Briefly, one adult female spider wasanesthetized using dry ice and dissected on ice. The exoskeleton was removed, and all soft tissue wascollected to extract HMW DNA. The DNA extraction was done according to the manufacturer’s protocol, except for tissue incubation in RNAse and proteinase-K at 50°C for 30 min, and elution of DNA that wasdone twice by adding an extra 100 µL of buffer AE to the beads. The purified DNA was run on a 0.5%agarose gel to assess DNA integrity. Absorbance ratios were evaluated on a Nanodrop spectrophotometer and determined to be as follows: 260 / 280: 1.82; 260 / 230: 1.91; resulting in total 36.2 µg DNA. The extracted genomic DNA was also subjected to quality check using a BioAnalyzer (Agilent)which revealed a single peak at around 11 kb. 10.3 µg DNA was used to make a 20 kb library. TheNational Genomics Infrastructure (NGI) platform at SciLifeLab, Uppsala University, performed the library preparation and sequencing using PacBio long-reads and 10X Genomics linked-reads. PacBio genomic library preparation and sequencing The QC-passed DNA samples were sent to the NGI platform at SciLifeLab, Uppsala University for library preparation and sequencing. The SMRT-bells obtained by the TPK1 kit according to manufacturer’s instructions were size-selected at 20 kb using a BluePippin instrument (SAGE) and sequenced on 60 SMRT cells of the RSII instrument, using P5-C3 chemistry. For each SMRT cells, 10 h movies were captured. A total of 798 Gbp of data with an insert size of 11 kb was produced. Library preparation and sequencing using 10X Chromium linked reads The HMW DNA was used to generate the 10X linked read libraries on 10X genomics Chromium platform (Genome Library Kit & Gel Bead Kit v2 PN-1000017, genome Chip Kit v2 PN120257) following the manufacturer’s guidelines. The 10X libraries were sequenced on Illumina NovaSeq6000 instrument(NovaSeq Control Software 1.6.0 / RTA v3.4.4) with 151 bp paired end setup using NovaSeqXp workflowin S4 flowcell. The Bcl to FastQ conversion was performed using bcl2fastq_v2.19.1.403 from the CASAVA software suite. Sanger / phred33 / Illumina 1.8+ was used as the quality scale. Bulk RNA extraction and sequencing of silk glands, head, and abdomen The spiders were anesthetized using dry ice before they were dissected on ice. After making an incision at the pedicel, the abdomen was gently pinned to a wax plate placed under a Zeiss Stemi 305 stereomicroscope, and the exoskeleton was carefully removed with micro scissors to visualize the silk glands.Phosphate buffered saline (PBS, pH 7.4) was used to wash off the excess non-silk tissue. With the help of micro tweezers, silk glands (major ampullate glands, minor ampullate glands, flagelliform glands, aggregate glands) were isolated separately by holding their ducts. The aciniform and piriform glands from these five spiders were extracted as a single sample due to difficulties in separating them owing to their small sizes. The major ampullate glands from six additional individuals were cut into three parts: tail, sac and duct and used for RNA extraction. In another preparation, after removing the exoskeleton the soft tissue from the whole abdomen was scraped out and used for extracting RNA. The RNA from head was extracted similarly after removing the legs and the thick exoskeleton. Five replicates were collected for each sample type and RNA was extracted from each replicate separately. A total of 41 RNA samples were extracted using RNeasy Plus Mini kit (QiaGen) by following manufacturer’s protocol. The integrity of the samples was estimated using Tapestation (Agilent Technologies). The transcriptome libraries from the different tissues were generated using Illumina TruSeq Stranded mRNA kit following manufacturer’s protocol and 151 bp paired end reads were sequenced using Novaseq6000 instrument. PacBio long-read Iso-Seq library construction and sequencing for major ampullate glands RNA was extracted from the major ampullate gland from one individual and homogenized in TriZol. The extracted RNA was sequenced at NGI, Uppsala University, Sweden. RNA QC was performed on the Agilent Bioanalyzer instrument, using the Eukaryote Total RNA Nano kit. The sequencing library wasprepared according to PacBio’s Procedure & Checklist – Iso-Seq™ Express Template Preparation forSequel® and Sequel II Systems, PN 101-763-800 version 02 (October 2019) using the NEBNext® Single Cell / Low Input cDNA Synthesis & Amplification Module, the Iso-Seq Express Oligo Kit, ProNex beads and the SMRTbell Express Template Prep Kit 2.0. The sample (300 ng) was first amplified to 12 cycles, followed by 3 additional cycles, according to the protocol. In the purification of amplified cDNA, the LongTranscripts workflow was applied to obtain material enriched for longer transcripts (>3 kb). The qualitycontrol of the SMRTbell libraries was performed with the Qubit dsDNA HS kit and the Agilent Bioanalyzer High Sensitivity kit. Primer annealing and polymerase binding was performed using the Sequel II binding kit 2.0. The samples were sequenced on the Sequel II instrument, using the Sequel II sequencing plate 2.0 and the Sequel® II SMRT® Cell 8M, with 24 h of movie time and 2 h of pre-extension time. Single cell preparation from major ampullate glands of L. sclopetarius In total 20 spiders were used for single cell sequencing. The first ten spiders were anesthetized in dry ice and dissected. All buffers were bubbled with carbogen during and prior to use. The major ampullate glands were taken out in ringer solution, pH 7.4. The duct was removed. The glands were washed withPBS, pH 7.4 and then incubated into pre-warmed trypsin-EDTA (Gibco, 0.5%) at 37^C for 1 min in a low-binding micro-centrifuge tube. The glands were triturated with a pipette briefly and centrifuged at 300 g for 3 min at 4^C. The pellets were resuspended in 300 µL DMEM (Gibco) containing 1% BSA (Sigma). The DMEM and the tubes were briefly bubbled with carbogen prior to resuspension. The suspension was strained through a 40 µm cell strainer and transferred to low-binding micro-centrifuge tubes. Cells were counted using Trypan blue dye. Due to problems with contaminating droplets of dope that made the separation of single cells challenging, a slightly modified protocol was used for the following ten spiders. These were processed as described above, but with the exception that the duct was not removed, and the sac was cut and kept in PBS, pH 7.4 for 10 min to allow the dope to flow out. The pieces of glands were picked up and incubated in pre warmed trypsin-EDTA (Gibco, 0.5%) at 37°C for 2 min in a low-binding micro-centrifuge tube. The suspension was triturated with fire-polished glass Pasteur pipette for 1 min, incubated in trypsin-EDTA at 37°C for 2 min, and triturated again with fire polished glass Pasteur pipette for 3 min. The suspension was centrifuged at 300 g for 3 min at 4°C and the pellet was resuspended in 200 µL DMEM containing 3% BSA. The suspension was strained through 40 µM pre-washed cell strainers into low-binding micro-centrifuge tubes. The strainer was further washed with 100 µL DMEM containing 3% BSA to reduce theloss of cells. Sample preparation for spatial transcriptomics Whole opisthosomas of spiders were flash-frozen in OCT medium on an isopentane-dry ice bath and samples were stored at –80°C until use. The samples were sectioned in a cryotome (10 µm) with knife temperature set at –23°C and sample holder temperature at –10°C. Eight sections from five individual spiders were carefully mounted on the capture areas (6.5 x 6.5 mm) of the Visium spatial slide (10X genomics) and permeabilized for 30 minutes following the 10X Visium spatial tissue optimization protocol. The libraries were constructed according to Visium spatial gene expression protocol (10X genomics). For cDNA amplification 13-16 PCR cycles were performed and for the indexing 13 PCR cycles were used. The sequencing was performed using a SP-200 flow cell on the Illumina Nova-Seq 6000. All the sections were stained with H&E for histological evaluation. Genome assembly and polishingThe Falcon and Falcon-Unzip (pb-falcon version 0.2.7) (Chin, et al., Phased diploid genome assemblywith single-molecule real-time sequencing. Nat Methods 13, 1050-1054 (2016)) de novo assemblers were used to assemble the PacBio data. The initial polishing was done using the Falcon-Unzip polishing module. The Chromium 10X linked read data was used to further polish the long-read CLR assembly. Reads were aligned to the PacBio assembly using Long Ranger (version 2.1.4) and three rounds ofpolishing were done using Pilon (Walker, et al., Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One 9, e112963 (2014)) (version 1.22) using the diploid flag. Genome size and heterozygosity estimation The raw Illumina reads from 10X genomic linked sequencing libraries were trimmed using Trimmomatic (Bolger, et al., Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114-2120(2014)), the canonical 20-mer counts were collected using Jellyfish (Marcais and Kingsford, A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764-770 (2011)).With the 20-mer histogram, GenomeScope (Vurture, et al., GenomeScope: fast reference-free genomeprofiling from short reads. Bioinformatics 33, 2202-2204 (2017)) was used to estimate the approximate genome size and heterozygosity. Mitochondrial genome assembly The mitochondrial genome was identified by mapping the assembly to an existing reference spider species, Neoscona adianta (Genbank accession: NC_029756.1) using BLAST. The identified regionsfrom the assembly were extracted using BEDTools (Quinlan and Hall, BEDTools: a flexible suite of utilitiesfor comparing genomic features. Bioinformatics 26, 841-842 (2010)). The extracted mitochondrial contigswere then annotated using MITOS web server (Bernt, et al., MITOS: improved de novo metazoanmitochondrial genome annotation. Mol Phylogenet Evol 69, 313-319 (2013)). PacBio Iso-Seq transcriptome assembly To generate full-length consensus transcript isoforms from the major ampullate gland, the raw polymerase reads were processed using SMRTlink. The subread BAM file was processed to generate the circular consensus sequence (CSS) reads. These reads were further classified into full length (FL) transcript sequences based on the criteria that they contain 5’ primer, 3’ primer and polyA tails. The FL transcript sequences were processed using IsoSeq3 platform for generating full length non-chimericreads (FLNC) which were further clustered using ICE algorithm to produce both high- and low-qualitypolished full length consensus sequences. The high-quality (HQ) sequences were used for the subsequent analysis. The high-quality transcripts were mapped to the de novo assembled L. sclopetariusgenome using minimap2 (Li, Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34,3094-3100 (2018)) (version 2.2.4) with parameters -ax splice -uf --secondary=no -C5. The alignment in the SAM format were processed into non-redundant full-length transcripts using the “collapse_isoforms_by_sam.py” script from the cDNA-Cupcake tool (https: / / github.com / Magdoll / cDNA_Cupcake). To assess the completeness of the transcriptome data,BUSCO (Simao, et al., BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210-3212 (2015)) (version 3.0.2) analysis was performed against the arthropoda_odb9 dataset. Genome annotationThe annotation of the de novo assembled L. sclopetarius genome was performed using MAKER(Cantarel, et al., MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes. Genome Res 18, 188-196 (2008)) version 3.01.02. High-confidence protein sequences (561356 proteins) were collected from the Uniprot Swiss-prot database (downloaded in November 2019) and a specific set of spidroin sequences (1051) were downloaded from NCBI (November 2019). A repeat library was created using the RepeatModeler package (version 1.0.11, https: / / www.repeatmasker.org / RepeatModeler / ). Since the spidroin sequences are highly repetitive in nature, the repeats modelled by the RepeatModeler were vetted against our specific set of spidroin data set. The repeat sequences in the assembled genome were identified using RepeatMasker (version 4.0.9, https: / / www.repeatmasker.org / ) and repeatRunner (https: / / www.yandell-lab.org / software / repeatrunner.html). The tRNAs were identified using tRNAscan (Lowe and Eddy,tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res 25, 955-964 (1997)) version1.3.1 while the conserved ncRNAs were identified using theInfernal package (Nawrocki, et al., Infernal 1.0: inference of RNA alignments. Bioinformatics 25, 1335-1337 (2009)) and the RNA family database, Rfam (Burge, et al., Rfam 11.0: 10 years of RNA families.Nucleic Acids Res 41, D226-232 (2013)) version 11. The MAKER package was executed in two runs: a) First MAKER was used to create a profile using the Uniprot Swiss-prot protein sequences, the specific set of spidroin sequences and RNA-seq data from different tissues. An in-house pipeline was used to select a set of genes from this initial evidence-basedannotation (first run) and to train Augustus (Stanke, et al., AUGUSTUS: a web server for gene finding ineukaryotes. Nucleic Acids Res 32, W309-312 (2004)) and SNAP (Korf, Gene finding in novel genomes.BMC Bioinformatics 5, 59 (2004)), b) MAKER was run a second time using the evidence from the first run and the prediction from Augustus. For the construction of gene models, the prediction from Augustus was used. Functional annotation of genes and transcripts was performed using the translated CDS features for each of the coding transcript. The protein sequences were searched against Uniprot Swiss-prot database and the specific set of spidroin sequences using BLAST to retrieve gene names and protein functions.Interproscan (5.30-69.0) (Paysan-Lafosse, et al., InterPro in 2022. Nucleic Acids Res 51, D418-D427(2023)) was used for extracting additional annotations (functional domains and sites) from various other biological databases (20 in total). Improving 3′ UTR annotation using single cell dataThe aligned BAM files from cell ranger were used for filtering BAM files using UMI-tools FilterBam (Smith,et al., UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Res 27, 491-499 (2017)) to include reads that were produced with corrected molecular barcode tag by cell ranger counts. The filtered BAM file was processed to remove PCRduplicates using UMI-tools dedup (Smith et al.2017). The Homer findPeaks was used to identify peaks(-size 50 -fragLength 100 -minDist 1). The peaks file was converted into bed file using BEDTools (Quinlanand Hall 2010). These peaks were then either annotated as 3′UTR to genes that lacked this feature orreannotated as extended 3′UTR feature to the nearest genes that were identified within 5000 bps. Identification of spidroins and manual curationFor identifying the spidroins in L. sclopetarius genome, two different approaches were used: a) thetranslated protein sequences of L. sclopetarius were scanned against the PFAM HMM profiles of N-terminal domain, C-terminal domain and Tubuliform egg casing silk strands structural domains. The HMM profile for these domains, Spidroin_N.hmm (N-terminal domain, PF16763), Spidroin_MaSp.hmm (C- terminal domain, PF11260) and RP1-2.hmm (Tubuliform egg casing silk strands domain, PF12042) were downloaded from PFAM database. To minimize the risk of false positive results, the hits with an e-value cutoff below 1e-05 were filtered out; and b) the reference spidroin sequences were downloaded from NCBI database (in November 2019). The redundant sequences were removed using CD-HIT at sequence identity cutoff of 95% and a custom database with full length spidroins (as retrieved from database), N- terminal and C-terminal domain sequences were created using BLAST package. The spidroin sequencesof L. sclopetarius were identified by homology search using BLASTp with an e-value cutoff of 1e-05against the full-length reference database. The identified sequences were confirmed for N-terminal domain and C-terminal domain. Due to the huge size and high repetitiveness of spidroins, identification of exon boundaries by assemblingtools can result in inaccuracies. Thus, the identified L. sclopetarius spidroin sequence loci and theirsurrounding regions (extending 5000–10000 bps on either end of the gene) were manually inspected ifthose were defined by the automated MAKER gene model. Web Apollo (Lee, et al., Web Apollo: a web- based genomic annotation editing platform. Genome Biol 14, R93 (2013)) genome browser was used for viewing gene models by entering the transcript identifier and identifying supporting data from bulk RNAseq and PacBio Iso-seq experiments. The extended gene sequences were searched separately against the custom N-terminal and C-terminal domain reference database using BLASTx with an e-value cutoff of 1e-05. An additional 100 bp region upstream of the identified N-terminal domain region wasscanned for signal peptide using SignalP (Almagro Armenteros, et al, SignalP 5.0 improves signal peptidepredictions using deep neural networks. Nature Biotechnology 37, 420-423 (2019)). Thus, we assume that the boundaries of spidroin genes were properly defined. The sequences for which the C-terminal domains could not be defined were kept unchanged as per the automated MAKER gene model. We also looked for multiple spidroin genes that were collapsed into single locus due to their high sequence similarity and therefore were hard to assemble as separate loci. After defining gene boundaries for every identified spidroin, the gene sequences were translated in all six translational frames and manually inspected for repetitive motifs to identify if any mis-annotations (missing exons due to incorrect readingframe) existed in the current gene model using Unipro UGENE software (Okonechnikov, et al., UniproUGENE: a unified bioinformatics toolkit. Bioinformatics 28, 1166-1167 (2012)). Based on the identified mis-annotations, either new genes were added, or the existing gene models were replaced with a corrected model. All the corrected sequences were later confirmed by mapping against the referencegenome using exonerate. GeneWise (Birney, et al., GeneWise and Genomewise. Genome Res 14, 988-995 (2004)) and Scipio (Keller, et al., Scipio: Using protein sequences to determine the preciseexon / intron structures of genes and their orthologs in closely related species. Bmc Bioinformatics 9, (2008)) were used to generate GFF file for the corrected gene models. Functional assignment to proteins with hypothetical function To predict function for proteins assigned as “hypothetical protein” (from Genome annotation), Orthofinder(Emms and Kelly, OrthoFinder: solving fundamental biases in whole genome comparisons dramaticallyimproves orthogroup inference accuracy. Genome Biol 16, 157 (2015)) (version 2.5.2) was used withdefault settings to identify gene family clusters between L. sclopetarius, Trichonephila clavipes (NCBIaccession number PRJDB10126), Trichonephila clavata (PRJDB10007), Nephila pilipes (PRJDB10128),Trichonephila inaurata madagascariensis (PRJDB10127), Argiope bruennichi (PRJNA629526), andAraneus ventricosus (PRJDB7092). The protein sequences were downloaded from NCBI. The putativetranscript isoforms were removed from the L. sclopetarius proteome dataset and the longest canonicalsequence were kept for the analysis. Proteins from each of the species were processed using Orthofinder-Diamond to assign proteins into orthogroups. In-house python scripts were used to process equivalent genes that were grouped as an orthogroup to further assign functions to hypothetical proteins based on proteins with known functions within the same orthogroup. The annotation was performed attwo levels; a) analyzing homologous cluster assigned to an orthogroup within L. sclopetarius and b)analyzing gene clusters from other spider species (T. clavipes, T. clavata, N. pilipes, T. inaurata madagascariensis, A. bruennichi, A. ventricosus) that were assigned to an orthogroup. Analysis of bulk RNA-sequencing dataThe raw sequence reads were mapped to the de novo assembled genome using STAR (Dobin, et al.,STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 (2013)) (version 2.7). The readcounts were generated using featureCounts (Liao, and Smyth, Shi, featureCounts: an efficient generalpurpose program for assigning sequence reads to genomic features. Bioinformatics 30, 923-930 (2014))from Rsubread package (Liao, et al., The R package Rsubread is easier, faster, cheaper and better foralignment and quantification of RNA sequencing reads. Nucleic Acids Research 47, (2019)) (version 2.0.0). Only genes with counts greater than 20 were kept for subsequent analysis. The read counts werenormalized using DESeq2 package (Love, et al., Moderated estimation of fold change and dispersion forRNA-seq data with DESeq2. Genome Biology 15, (2014)) (version 1.38.3) which accounts for correcting sequencing depth and library composition. All samples correlated with pairwise Pearson correlation and hierarchically clustered. PCA analysis was performed for the samples for the three parts of the major ampullate gland. The PC- space was split up into three regions: PC1<0 and PC2>0 were assigned tail, PC1>0 and PC2>0 were assigned sac, and PC2< 0 was assigned duct. A geometrical distance from origo was calculated for each gene for the loading of the first two PCs. A normalized distance was calculated by removing the mean distance and dividing by the standard deviation of the PC. Genes with a normalized distance greater than two were kept as significant which were assigned to the three different tissues based on their PC1 and PC2 loading values, i.e., in which region they were located. This set was defined as the major ampullategene set. Differential gene expression analysis was carried out between the three different regions usingDEseq2. To be assigned as a differentially expressed in one tissue it had to have a two-fold change greater than two and an adjusted p-value less than 0.001. Analysis of single cell RNA sequencing data Reads were mapped to transcripts using CellRanger (version 7.0.1, https: / / support.10xgenomics.com / single-cell-gene-expression / software / pipelines / latest / what-is-cell- ranger). Initial QC analysis removed all cells with fewer than 300 expressed genes and / or less than 500 total transcripts. Only samples with more than 500 cells were retained for further analysis. All steps wereperformed in Seurat (Satija, et al., Spatial reconstruction of single-cell gene expression data. NatureBiotechnology 33, 495-U206 (2015); Hao, et al., Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587 e3529 (2021)) (version 4.0.3, https: / / satijalab.org / seurat / ). Cells were normalized using SCT and integrated with Canonical correlation analysis distances between samples. After initial QC steps, 18539 cells were obtained from 7 samples that clustered into 23 clusters. Among these clusters, a subset was identified where the marker genes overlapped with the major ampullate tail, sac and duct gene sets. By reevaluating this specific subset and restricting the genes analyzed to the intersection of the top 2000 most variable genes from the scRNAseq data and the major ampullate gene set derived from the bulk RNA analysis, 9700 cells were obtained from 7 samples which clustered into nine groups. The smallest cluster only contained cells from one sample and was removed from further analysis. Marker genes for each cluster were identified using a subset of 400 cells per class. Analysis of spatial transcriptomic data and deconvolution Eight slides were manually annotated as silk glands based on the morphology using the annotation tool in the Loupe browser (version 6.4.1, https: / / support.10xgenomics.com / spatial-gene- expression / software / visualization / latest / what-is-loupe-browser). Reads were mapped with Space Ranger (version 1.2.0, https: / / support.10xgenomics.com / spatial-gene-expression / software / pipelines / latest / what- is-space-ranger) to the reference genome and annotation. Initial QC removed samples with less than 300genes and 500 transcripts. All steps were carried out in Seurat (Satija et al. 2015; Hao et al., 2021)(version 4.0.3). Cells were normalized using SCT and integrated with canonical correlation analysis distances between samples. The distance between the samples was visualized using UMAP. Marker genes for different classes were identified using a subset of 400 cells per class. For analysis of the major ampullate gland, five samples with good annotation of the gland were kept. Theimage files were imported in QuPath (Bankhead, et al., QuPath: Open source software for digitalpathology image analysis. Sci Rep 7, 16878 (2017)) (version 0.4.3) and the regions identified as major ampullate gland were further separated into different zones A, B and C based on H&E-staining and cell morphology using brush tool. Average Eosin, average Hematoxylin and perimeter values weredetermined for each region using QuPath default parameters.Pairwise Pearson correlation was performed for eosin, hematoxylin, and perimeter values for the zone A regions. Zone A regions were split into three classes based on their perimeter value. Regions withperimeter less than 500 pixels were assigned proximal, regions with a perimeter larger than 1000 were assigned distal and the rest were assigned to middle. Spots on the spatial transcriptomic sections were assigned to the closest region that overlapped with the zone annotation from the QuPath analysis.To identify the proportion of different cell types on each spot on the slide we used CARD (Ma and Zhou,Spatially informed cell-type deconvolution for spatial transcriptomics. Nat Biotechnol 40, 1349-1359 (2022)). Only genes from the single cell analysis with an average log2fold >2 in at least one cluster was kept to deconvolute the spatial spots. Only genes with at least 200 counts and found in at least 50 spots in the spatial transcriptomics data was kept for the deconvolution analysis. Cell type proportions were estimated for each spot in the major ampullate gland. The average proportion for the five classes zone C, zone B and the subclasses proximal, middle and distal of zone A was calculated by taking the average cell type proportions from all spots that belonged to the five classes. Sample preparation for proteomics Major ampullate glands The spiders were anesthetized and dissected as mentioned earlier. The major ampullate glands were carefully pulled out by holding the duct using micro tweezers. A cut was made in the sac of the gland allowing the dope to flow out for 15 min. The glands were washed three times with PBS and thentransferred to a low-binding 1.5 mL microtube (Axygen) containing 60 µL of 8 M urea in 20 mM Tris-HClat pH 8, vortexed and sonicated in an ultrasonic bath sonicator (VWR) for 30 min at room temperature. The samples were stored at -20°C until further use. Four biological replicates were used for the final protein sequencing. Sample aliquots were supplemented with 0.2% ProteaseMAX (Promega) in 20% ACN / 20 mM Tris-HCl, pH 8 to obtain 4M urea concentration before water bath sonication for 5 min. Proteins were reduced with 8 mM DTT incubated at 24°C for 1 h with 550 rpm and alkylated with 20 mM chloroacetamide (CAA) incubated for 1 h at RT in dark. Digestion was started with addition of 2 µg of LysC (Wako, Japan) incubated at 24°C for 2 h and completed with 2 µg sequencing grade modified trypsin (Promega) incubated at 37°C overnight (ca 16 h). Following centrifugation, the supernatants were collected, and proteolysis was stopped with 5% FA, and the samples were cleaned on a C18 Hypersep plate with 40 µL bed volume (Thermo Fisher Scientific) and dried using a vacuum concentrator (Eppendorf). Major ampullate silk fibers For collecting the silk, each spider was first anesthetized with CO2 and gently pinned down to immobilize it, without injuring the animal. By using a Zeiss Stemi 305 stereo microscope, the major ampullate silk was identified, pulled out from the anterior spinneret (Foelix, Biology of Spiders: Oxford University Press.New York 330, (1996)) with the help of a tweezer. The silk was collected by rolling it onto a frame attachedto a rotating wheel until the spider refused to spin silk. Several spiders were used to collect enough amount of silk for the experiments. After silking, the spiders were fed with fruit flies, watered, and not used again for the next two weeks. The collected silk was treated in three different ways for solubilization. About 450 µg of silk was taken for each set of samples and every treatment was done in quadruplet. Three different methods were used to prepare silk samples. In the first method, the silk was dissolved by adding 100 µL of hexafluoroisopropanol (HFIP) and brief vortexing. The samples were then sonicated in an ultrasonic bath sonicator (VWR) for 30 min at room temperature. The HFIP was evaporated on a Centrivap concentratorsystem (Labconco). The protein was resuspended in 60 µL of 8 M urea in 20 mM Tris-HCl, pH 8 andstored at -20°C until further use. In the second treatment the silk was dissolved in 60 µL of 8 M urea in20 mM Tris-HCl at pH 8, sonicated and stored as mentioned above. For the third method, the silk wasdissolved in 60 µL of 9 M LiBr in 20 mM Tris-HCl at pH 8, sonicated and stored as mentioned above.Low-binding 1.5 mL microtubes (Axygen) were used throughout the experiments. An aliquot of 30 µL samples (ca 10 µg) was taken to further preparation. From samples with the HFIP and urea methods, proteins were reduced with 3 µL of 100 mM DTT, incubated at 37°C for 3 h with 1200rpm and alkylated with 5 µL of 500 mM CAA incubated for 30 min at room temperature in dark. Half ofthe samples were supplemented with 19 µL of 50 mM Tris-HCl at pH 8.5 and digested with addition of 1 µg of LysC (Wako, Japan) incubated at 24°C for 2 h. Digestion was continued with 1 µg sequencing- grade modified trypsin (Promega) after addition of 56 µL of Tris-HCl and incubated at 37°C overnight (ca 16 h). Samples with the third method (LiBr) were prepared similarly, except for 2 µL of 500 mM DTT was used for reduction incubated at 95°C for 30 min with shaking at 12,500 rpm. Alkylation with 5 µL of 500 mM CAA (as above) was followed by digestion with LysC and trypsin as described above except for that 3 µg trypsin was used. The digestion of all samples was stopped with 6.5 µL concentrated formic acid, the samples were cleaned on a C18 Hypersep plate with 40 µL bed volume (Thermo Fisher Scientific), and dried using a vacuum concentrator (Eppendorf). Layer-wise dissolution of major ampullate silk fibers The major ampullate silk was collected by allowing each spider to fall freely from a wooden frame and the extruded silk was rolled on to the same wooden frame. The silk from several individuals was collected in pre-weighed low-binding microtubes, which were measured again to determine the weight of the collected silk. After forceful silking, the spiders were fed with fruit flies, watered, and were not used againfor the next two weeks. The collected silk was divided into 3 sets treated with: 1) 2 M urea in 50 mM Tris-HCl and 0.5M NaCl, 2) 4 M urea in 50 mM Tris-HCl and 0.5M NaCl, and 3) 8 M urea in 50 mM Tris-HCland 0.5M NaCl. Samples were sonicated at room temperature for 2 – 3 h followed by centrifugation at17,000 g. The supernatant was collected and stored at -20°C before using for proteomics analysis. Liquid Chromatography-Tandem Mass Spectrometry Data Acquisition Peptides were reconstituted in solvent A and injected on a 50 cm long EASY-Spray C18 column (Thermo Fisher Scientific) connected to an UltiMate 3000 nano-flow UPLC system (Thermo Fisher Scientific) using a 90 min long gradient: 4-26% of solvent B (98% acetonitrile, 0.1% FA) in 90 min, 26-95% in 5 min, and 95% of solvent B for 5 min at a flow rate of 300 nL / min. Mass spectra were acquired on a Q Exactive HF hybrid quadrupole orbitrap mass spectrometer (Thermo Fisher Scientific) ranging from m / z 375 to 1800 at a resolution of R=120,000 (at m / z 200) targeting 5x106ions for maximum injection time of 100 ms, followed by data-dependent higher-energy collisional dissociation (HCD) fragmentations of precursor ions with a charge state 2+ to 7+, using 45 s dynamic exclusion. The tandem mass spectra of the top 17 precursor ions were acquired with a resolution of R=30,000, targeting 2x105ions for maximum injection time of 54 ms, setting quadrupole isolation width to 1.4 Th and normalized collision energy to 28%. Analysis of proteomics data Acquired raw data files were converted to Mascot Generic File (mgf) format using an in-house developed tool, Raw2MGF (version 2.1.3), and searched with Mascot Daemon version 2.5.1 (Matrix Science Ltd., UK) against a protein database obtained from 22,856 protein entries. A maximum of two missed cleavage sites were allowed for full tryptic digestion, while setting the precursor and the fragment ion mass tolerance to 10 ppm and 0.02 Da, respectively. Carbamidomethylation of cysteine was specified as a fixed modification, while oxidation on methionine as well as deamidation of asparagine and glutamine were set as dynamic modifications. The search results were imported into Scaffold version 4.11 (Proteome Software Inc.) to calculate the contribution of each protein to the total sum of spectra(percentage of total spectra). For all proteins were there was no report of a percentage of total spectrafor a sample the percentage of total spectra value was set to 0. To identify the proteins in the silk we used MS / MS data from both the silk fibers and the glands. To determine their presence in the gland, the average percent of total spectra for each protein was calculated by taking the mean of the three gland samples. All proteins that were reported in at least one of the samples were considered to be present in the gland. The average percent of total spectra for each protein in the silk was calculated similarly by taking the mean of the nine samples, i.e., three biological replicates for the three different detergents. For a protein to be considered present in the major ampullate silk, four criteria had to be fulfilled: 1) it had to present in the gland, 2) it had to be found in at least two of the detergents, 3) it had possess a signal-peptide predicted at the 5’ end of its aa sequence, and 4) it had to have an average percent of total spectra greater than 0. After applying these criteria, 17 proteins remained. Statistical differences in the percentage of total spectra for a protein between different urea concentrations were determined using an unpaired t-test, focusing on these 17 proteins. Tissue processing for histological analysis Spiders were anesthetized with CO2, then dissected at the pedicel on an ice-cooled wax plate using 154 mM sodium chloride solution (Fresenius Kabi AG, Germany). Dissections were performed using a Leica M60 stereomicroscope with a Leica IC80 HD camera. Whole opisthosomas were fixed in 2.5% glutaraldehyde in 67 mM phosphate buffer, pH 7.2 at 4^C for 24 h, and later rinsed in 67 mM phosphate buffer. Tissues were dehydrated in graded ethanol (50, 70, 90, and 100%, for 30 minutes each), then infiltrated and embedded in water-soluble glycol methacrylate (Leica Historesin). Sections of 2 µm thickness were obtained using a Leica RM 2165 microtome with glass knives. The sections were stained with hematoxylin and eosin and mounted using Agar 100 resin. Evaluation was performed using a NikonMicrophot-FXA (Tekno Optik AB) microscope equipped with a Nikon FX-35DX camera. Images werecaptured and edited using the software Eclipse Net version 1.20.0. Tensile tests of major ampullate silk fibers Major ampullate silk fibers were reeled and mounted on carboard frames with a square window of 1 x 1 cm (gauge length 1 cm). The diameters were measured by means of light microscopy using a Nikon Eclipse Ts2R-FL inverted microscope. The diameter was measured prior to the tensile test at five locations along the fiber and then averaged. The tensile tests were performed with a 5943-Instronmachine (USA) equipped with a 5 N load cell. The used strain rate was 6 mm / min. Load displacementcurves were converted into engineering stress-strain curves assuming a circular cross-section and using the average diameter for each fiber. The embodiments described above are to be understood as a few illustrative examples of the present invention. It will be understood by those skilled in the art that various modifications, combinations and changes may be made to the embodiments without departing from the scope of the present invention.
Claims
CLAIMS1. A recombinant spider silk protein comprising an N-terminal (NT) domain, a repetitive region (REP)domain and a C-terminal (CT) domain, whereinthe REP domain comprises at least 100 amino acid residues derived from the repetitive region ofa spider silk protein of Larinioides sclopetarius; and the NT domain is derived from the NT domain of a spider silk protein other than the spider silkprotein, from which the REP domain is derived, and / or the CT domain is derived from the CT domain ofa spider silk protein other than the spider silk protein, from which the REP domain is derived.
2. The recombinant spider silk protein of claim 1, wherein the REP domain is arranged between theNT domain and the CT domain in the recombinant spider silk protein.
3. The recombinant spider silk protein of claim 2, wherein the recombinant spider silk protein has thegeneral formula (X)-NT-(L1)-REP-(L2)-CT-(Y), wherein X represents an optional N-terminal tag, Y represents an optional C-terminal tag, L1 represents an optional first linker and L2 represents an optional second linker.
4. The recombinant spider silk protein of any one of claims 1 to 3, wherein the REP domaincomprises, preferably consists of, at least 125 amino acid residues, preferably at least 130 amino acid residues, more preferably at least 150 amino acid residues, such as at least 200 amino acid residues, at least 250 amino acid residues, at least 300 amino acid residues, at least 350 amino acid residues, andmost preferably at least 400 amino acid residues derived from the repetitive domain of the spider silkprotein of L. sclopetarius, such as at least 450 amino acid residues, at least 500 amino acid residues, at least 550 amino acid residues or at least 600 amino acid residues derived from the repetitive domain of the spider silk protein of L. sclopetarius.
5. The recombinant spider silk protein of any one of claims 1 to 3, wherein the REP domaincomprises, preferably consists of, at least 100 consecutive amino acid residues, preferably at least 125 consecutive amino acids residues, more preferably at least 130 consecutive amino acid residues, suchas at least 150 consecutive amino acid residues, at least 200 consecutive amino acid residues, at least250 consecutive amino acid residues, at least 300 consecutive amino acid residues, at least 350 consecutive amino acid residues, and most preferably at least 400 consecutive amino acid residues derived from the repetitive domain of the spider silk protein of L. sclopetarius, such as at least 450 consecutive amino acid residues, at least 500 consecutive amino acid residues, at least 550 consecutiveamino acid residues or at least 600 consecutive amino acid residues derived from the repetitive domain of the spider silk protein of L. sclopetarius.
6. The recombinant spider silk protein of any one of the claims 1 to 5, wherein the REP domaincomprises at least 100 amino acid residues derived from the repetitive region of a major ampullate silkprotein (MaSp) of L. sclopetarius.
7. The recombinant spider silk protein of claim 6, wherein the REP domain is derived from a MaSpof L. sclopetarius selected from the group consisting of MaSp1a as defined in SEQ ID NO: 38, MaSp1bas defined in SEQ ID NO: 41, MaSp1c as defined in SEQ ID NO: 44, MaSp2a as defined in SEQ ID NO:47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined inSEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3aas defined in SEQ ID NO: 65, MaSp3b as defined in SEQ ID NO: 68 and MaSp4 as defined in SEQ IDNO: 71, preferably selected from the group consisting of MaSp1a as defined in SEQ ID NO: 38, MaSp2aas defined in SEQ ID NO: 47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined inSEQ ID NO: 62, MaSp3a as defined in SEQ ID NO: 65 and MaSp4 as defined in SEQ ID NO: 71, andmore preferably selected from the group consisting of MaSp2c as defined in SEQ ID NO: 53 and MaSp4as defined in SEQ ID NO: 71.
8. The recombinant spider silk protein of claim 7, wherein the REP domain comprises, preferablyconsists of, an amino acid sequence selected from the group consisting of SEQ ID NO: 109-114, 148- 151, 168, preferably selected from the group consisting of SEQ ID NO: 109-114, 148-151, and more preferably selected from the group consisting of SEQ ID NO: 109-114.
9. The recombinant spider silk protein of any one of claims 1 to 8, wherein the NT domain is derivedfrom the NT domain of a spider silk protein of a different spider species than the REP domain and / or the CT domain is derived from the CT domain of a spider silk protein of a different spider species than the REP domain.
10. The recombinant spider silk protein of any one of claims 1 to 9, wherein the NT domain is derivedfrom the NT domain of Euprosthenops australis MaSp1.
11. The recombinant spider silk protein of claim 10, wherein NT domain comprises, preferably consistsof SEQ ID NO: 115.
12. The recombinant spider silk protein of any one of claims 1 to 9, wherein the NT domain is derivedfrom the NT domain a silk protein of L. sclopetarius, preferably from the NT domain of L. sclopetariusMaSp2f.
13. The recombinant spider silk protein of claim 12, where the NT domain comprises, preferablyconsists of SEQ ID NO: 61.
14. The recombinant spider silk protein of any one of claims 1 to 13, wherein the CT domain is derivedfrom the CT domain of Araneus ventricosus MiSp.
15. The recombinant spider silk protein of claim 14, wherein the CT domain comprises, preferablyconsists of SEQ ID NO: 116.
16. The recombinant spider silk protein of any one of claims 1 to 10, wherein the CT domain is derivedfrom the CT domain of a silk protein of L. sclopetarius, preferably from the CT domain of L. sclopetariusMiSpc.
17. The recombinant spider silk protein of claim 16, where the CT domain comprises, preferablyconsists of SEQ ID NO: 81.
18. The recombinant spider silk protein of any one of claims 1 to 11, wherein the recombinant spidersilk protein comprises, preferably consists of, an amino acid sequence selected from the group consistingof SEQ ID NO: 103-108, 117-134, 152-167, 169-172, preferably selected from the group consisting ofSEQ ID NO: 103-108, 117-134, 152-167, and more preferably selected from the group consisting of SEQ ID NO: 103-108, 117-134.
19. The recombinant spider silk protein of any one of claims 1 to 18, wherein the recombinant spidersilk protein has a molecular weight of at least 30 kDa, preferably at least 35 kDa, and more preferablyselected at least 50 kDa.
20. A silk fiber made of a recombinant spider silk protein of any one of claims 1 to 19.
21. The silk fiber of claim 20, further comprising a spider-silk constituting element (SpiCE) of L.sclopetarius.
22. A synthetic material comprising a silk fiber of claim 20 or 21.
23. A nucleic acid molecule encoding a recombinant spider silk protein of any one of claims 1 to 19.
24. An expression vector comprising a nucleic acid molecule according to claim 23.
25. A host cell comprising the expression vector according to claim 24.
26. A method for producing a silk fiber comprising:extruding a spinning dope comprising a recombinant spider silk protein of any one of claims 1 to19 into an aqueous buffer having an acidic pH to induce polymerization of the recombinant spider silkprotein into a silk fiber; and isolating the silk fiber from the aqueous buffer.
Citation Information
Patent Citations
Engineered Spider Silk Proteins And Uses Thereof
US20190248847A1
Engineered spider silk proteins and uses thereof
WO2018002216A1
Silk nucleotides and proteins and methods of use
WO2020092769A2
Performance of araneus ventricosus dragline silk and silk protein gene sequences of araneus ventricosus dragline silk
CN110283243A
Methods of producing polymers of spider silk proteins
EP2243792A1