Recombinant spider silk proteins

CN122803989APending Publication Date: 2026-09-22安娜·里辛 +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202580016779.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2025-03-21
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0013]本发明的重组蜘蛛丝蛋白能够被表达、纯化并纺成丝纤维,该丝纤维具有优异的机械性能,且带有相对较长的REP结构域,其REP结构域长度是NT2RepCT的至少4至5倍。因此,本发明的重组蜘蛛丝蛋白在REP结构域长度方面更接近于天然存在的丝蛋白。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122803989A_ABST
    Figure CN122803989A_ABST
Patent Text Reader

Abstract

A recombinant spider silk protein comprising an N-terminal (NT) domain, a repetitive region (REP) domain, and a C-terminal (CT) domain. The REP domain comprises at least 100 amino acids derived from a repetitive region of a spider silk protein from Larinioides sclopetarius. The NT domain is derived from an NT domain of a spider silk protein that is different from the spider silk protein from which the REP domain is derived, and / or the CT domain is derived from a CT domain of a spider silk protein that is different from the spider silk protein from which the REP domain is derived.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to recombinant spider silk proteins, and more particularly to recombinant spider silk proteins comprising a REP domain comprising a repeating region of silk protein derived from the hard-skinned spider (Larinioides sclopetarius). Background Technology

[0002] Spiders can produce seven types of silk, each with unique mechanical properties, produced by different glands: ampulla macrocarpal, ampulla minor, flagellated, tubular, variegata, and piriformis. These silks are composed of silk proteins, named according to the gland in which they are primarily expressed: ampulla macrocarpal silk protein (MaSp), ampulla minor silk protein (MiSp), flagellated spider silk protein (FlSp), tubular spider silk protein (TuSp), variegata spider silk protein (AcSp), piriformis spider silk protein (AgSp), and piriformis spider silk protein (PySp). Spider silk proteins, also known as spider silks, have an N-terminal (NT) domain, a highly repeating region (REP), and a C-terminal (CT) domain. Their mechanical properties are believed to be determined by the REP domain. The most malleable fibers are flagellated silks, primarily composed of spider silk proteins (FlSps) carrying proline-rich REP regions, which are predicted to form spring-like structures. The strongest fiber is the macroambulatory gland filament, also known as the traction filament, which is mainly composed of spider silk protein (MaSps) carrying repeating regions consisting of repeating glycine-rich and poly-Ala repeating sequences. The tensile strength of the macroambulatory gland filament derives from the poly-Ala blocks in the MaSp, which form β-sheet crystals within the silk fiber, while the glycine-rich portions mediate the fiber's extensibility. MaSp silk is the strongest known natural fiber (approximately 150 MJ / m³).

[0003] A recombinant spider silk protein known in the art has improved solubility in water, enabling large-scale production in high yields. This protein is called NT2RepCT (WO 2018 / 002216; Andersson et al., Nature Chemical Biology 11: 309-315 (2017)). NT2RepCT contains a His6 tag, an NT domain from Eurosthenops australis MaSp1, two glycine- and polyalanine-rich tandem repeat sequences (2Rep) from Eurosthenops australis, and a CT domain from Araneus ventricosus MiSp.

[0004] WO 2020 / 092769 discloses an engineered polypeptide comprising at least two units, each containing a MaSp4 repeating unit from the Darwin's bark spider (Caerostris darwini). The MaSp4 of the Darwin's bark spider differs from other MaSp units in that it is primarily composed of the GPGPQ amino acid motif, which imparts extensibility to the spider fibers. This high extensibility has reportedly led to the high toughness of the Darwin's bark spider silk fibers.

[0005] WO 2023 / 167628 discloses a recombinant spider silk protein comprising an NT domain, a REP domain, and a CT domain. The REP domain comprises a set of domains according to the formula pA1-pG-pA2. pG represents a glycine-rich domain, and pA1 and pA2 represent alanine-rich domains. One of pA1 and pA2 is a polyalanine domain, while the other of pA1 and pA2 is a polyalanine domain in which every third or fourth alanine residue is replaced by an isoleucine or valine residue.

[0006] However, there is still a need for a recombinant spider protein with a long REP domain that can be expressed and spun into filaments. Summary of the Invention

[0007] A general objective of this invention is to provide a recombinant spider silk protein having a long REP domain and capable of producing silk fibers.

[0008] This objective, as well as other objectives, are achieved through embodiments of the present invention.

[0009] This invention is defined in the independent claims. Further embodiments of the invention are defined in the dependent claims.

[0010] One aspect of the present invention relates to a recombinant spider silk protein comprising an N-terminal (NT) domain, a repeat region (REP) domain, and a C-terminal (CT) domain. The REP domain comprises at least 100 amino acids derived from the repeat region of a spider silk protein from the hard-shelled spider *Larinioides sclopetarius*. The NT domain is derived from the NT domain of a spider silk protein not derived from the REP domain, and / or the CT domain is derived from the CT domain of a spider silk protein not derived from the REP domain.

[0011] A further aspect of the present invention relates to silk fibers made from the above-mentioned recombinant spider silk protein, synthetic materials comprising the above-mentioned silk fibers, nucleic acid molecules encoding the above-mentioned recombinant spider silk protein, expression vectors comprising the above-mentioned nucleic acid molecules, and host cells comprising the above-mentioned expression vectors.

[0012] Another aspect of the invention relates to a method for producing silk fibers. The method includes extruding a spinning solution containing the aforementioned recombinant spider silk protein into an aqueous buffer having an acidic pH to induce the recombinant spider silk protein to polymerize into silk fibers. The method further includes separating the silk fibers from the aqueous buffer.

[0013] The recombinant spider silk protein of the present invention can be expressed, purified, and spun into silk fibers with excellent mechanical properties and a relatively long REP domain, the length of which is at least 4 to 5 times that of NT2RepCT. Therefore, the recombinant spider silk protein of the present invention is closer to that of naturally occurring silk proteins in terms of REP domain length. Attached Figure Description

[0014] These embodiments, along with their further objectives and advantages, can be best understood by referring to the following description taken in conjunction with the accompanying drawings, in which: Figure 1 The stress-strain curves of the Swedish bridge spider (L. sclopetarius) and its traction silk are shown. (1a) A female L. sclopetarius spider. (1b) Stress-strain curve of the traction silk obtained from a forced-silk spider.

[0015] Figure 2The following is a catalog of spider silk proteins from *L. sclopetarius*. (2a) A schematic diagram of an orb-weaver spider, showing one group of each type of silk gland. (2b) A schematic diagram of the ampullae gland. This gland has three anatomical parts: the tail, the sac, and the duct. The tail and sac are composed of a single layer of epithelium, in which three morphologically distinct cell types were found, each localized to one of three regions (AC). (2c) A phylogenetic tree of the N-terminal domains of 35 spider silk proteins identified from *L. sclopetarius*. The numbers on the branches represent bootstrap values. Spider silk proteins shown in bold were identified through proteomic analysis of the ampullae gland and silk. (2d) A heatmap showing the expression of all spider silk proteins identified by bulk RNA sequencing in different tissues. Grayscale corresponds to the normalized counts shown in the bar (bottom left inset). (2e) A schematic diagram of spider silk protein genes. All spider silk protein genes encode proteins with a signal peptide (not shown) and an N-terminal domain. Most spider silk protein genes encode typical C-terminal domains, except for FlSp-like, which has been found to have atypical C-terminal domains, while AmSp-like1, AmSp-like2, and AgSp-like spider silk proteins completely lack C-terminal domains. Repeating motifs in each spider silk protein are represented in blocks. Table (2f) shows the number of typical MaSp repeating motifs found in each MaSp. (A)n represents a polyalanine motif, and X in GGX and GPGXX represents any amino acid residue.

[0016] Figure 3The expression of 17 silk protein genes is shown in the tail, sac, and duct of the ampulla of Vater. (3a) Relative quantification of proteins identified in the ampulla of Vater and dissolved silk fibers using LC-MS / MS proteomics, labeled according to protein class (MaSp1, MaSp2, MaSp3, MaSp4, AmSp-like, and SpiCE-LMa). (3b) Schematic diagram of the gland showing the three parts (tail, sac, and duct) separated for RNA sequencing experiments. (3c) PLS analysis of bulk RNA data distinguishes the three distinct parts (tail, sac, and duct). Diamonds represent samples, and small circles represent the superposition of the 17 silk protein genes. (3d) Heatmap showing the relative expression levels of the 17 silk protein genes in samples from the tail, sac, and duct. Gene names are shown on the right. The analysis divides the genes into six clusters based on their expression profiles, as shown in the dendrogram on the y-axis. Clusters are indicated based on whether the highest gene expression occurs in the tail or sac. The bars on the right represent protein lengths. The scale ranges from 0 (white) to 10,000 amino acids (black). (3e) The percentage (%) of different classes of amino acid residues in the 17 silk proteins. This figure shows the different classes (small nonpolar: A, G, P, S, T; hydrophobic: I, L, M, V; polar: D, E, H, K, N, Q, R; aromatic and cysteine: C, F, W, Y).

[0017] Figure 4 Spatial transcriptomics of silk glands is shown. (4a) H&E-stained sections of the spider abdomen. The inset shows a side view of the abdomen and indicates the approximate plane on which the sections were made. (4b) Spots are labeled as different silk glands based on histological morphology and spatial location. In (4a) and (4b), sp and pd represent the locations of the spinnerets and the stalk, respectively. (4c) UMAP analysis of all spots from eight sections that were manually labeled as silk glands. Each spot represents one spot in the spatial section.

[0018] Figure 5Spatial resolution of silk protein expression in regions A, B, and C is shown. (5a) A large ampulla of Vater contains several cross sections of the tail and a cross section of the sac (H&E staining). The inset shows the original image of the magnified area (white square). (5b) Based on the morphology and staining of the epithelium covering each spot, the spots corresponding to the large ampulla of Vater are labeled as regions A, B, or C. (5c) A heatmap shows the expression of 17 silk protein genes in 847 spots labeled as regions A, B, and C, respectively. Each bar on the heatmap represents a spot on the spatial section, and the black boxes indicate marker genes in different regions. (5d) Expression profiles of MaSp1a, MaSp2b, MaSp3a, SpiCE-LMa1, SpiCE-LMa2, and SpiCE-LMa3 (in order) in the three regions of the large ampulla of Vater shown in (5a) and (5b).

[0019] Figure 6 Single-cell RNA sequencing analysis of the ampulla of Vater reveals eight cell types. (6a) UMAP reveals eight cell types in the ampulla of Vater. Each point represents one cell. Cell types were annotated by correlating gene expression with bulk RNA and spatial transcriptomics data. This revealed that six cell types constitute the secretory epithelium of the tail and sac; three cell types are confined to region A, one to region B, one to region C, and one can be found in all three regions. Two cell types were found in the ducts. (6b) Heatmap of relative gene expression of 17 silk proteins in different single-cell types. Black boxes indicate marker genes in different cell types.

[0020] Figure 7 Spatial distribution of eight ampulla cell types is shown. (7a) H&E-stained section of spider ampulla, with an inset showing the magnified region. (7b) Same section as (7a), with regions (A, B, and C) of the ampulla marked. Regions were identified based on cell morphology, and 11 cross-sections were obtained from the gland (9 in region A, 1 in region B, and 1 in region C). (7c) Spatial spots were deconvolved using scRNAseq data to generate pie charts for each spot, where grayscale represents the cell type identified from the scRNAseq data and shows the proportion of cells belonging to each cell type. (7d) Relative gene expression of the silk protein gene, a marker gene for cell type A, as a function of the gland cross-sectional perimeter. Grayscale corresponds to the three cell types in region A, and lines are linear models of the gene expression profiles for each cell type. (7e) Hematoxylin average as a function of the cross-sectional perimeter of the ampulla in region A. (7f) The average proportion of cell types identified from scRNAseq data in the spots of regions A (proximal, mid, and distal), B, and C, as assessed by deconvolution plots. (7g) Schematic diagram showing the spatial distribution of eight cell types in the ampulla of Vater.

[0021] Figure 8 The origin and model of the multilayered structure of the ampulla of Vater gland filaments are shown. (8a) Histological sections showing the morphology (H&E staining) of the single-layered epithelium in regions A, B, and C of the ampulla of Vater gland of L. sclopetarius. Arrows indicate cell nuclei located at the base. Secretions from regions A, B, and C can be identified in the lumen as forming three layers (labeled I, II, and III in region C). The circular structures around intracellular vesicles in the first three images represent the annotation areas used for image analysis in QuPath. Scale bar = 20 µm. (8b) H&E intensity map of the annotated objects shown in (8a). (8c) Schematic diagram of the ampulla of Vater gland illustrating the localization and layered secretions of eight cell types. The circular structures below the gland show a schematic diagram along a cross-section of the gland. (8d-8g) Origin of 17 silk proteins in the ampulla of Vater gland filaments and their expression profiles in the ampulla of Vater gland. The 17 silk protein genes / proteins are listed below the figures. The top figure (8d) shows the expression levels of 17 silk protein genes identified using bulk transcriptome data. Numbers represent log2-normalized expression values. Figure (8e) shows the gene expression values ​​of the 17 silk proteins in six cell types found in regions A, B, and C; black boxes indicate proteins identified as marker genes for these cell types. Figure (8f) shows the abundance of each of the 17 silk proteins in soluble extracts of ampullary gland fibers incubated in 2, 4, and 8 M urea, respectively. Gray levels represent expression values, and black boxes indicate proteins identified as markers. The bottom figure (8g) shows the cumulative protein abundance in whole silk fibers dissolved using HFIP, LiBr, and urea. (8h) Model of the multilayered structure of ampullary gland filaments. Gray levels correspond to the cell types from which these layers originate.

[0022] Figure 9 Bacterial cell lysis and IMAC purification results for three constructs (MaSp4_631, MaSp4_400, and MaSp2c_500) are shown on an SDS-PAGE gel. T = total cell lysate, SN = supernatant after centrifugation, Filt. SN = supernatant after filtration (0.2 µm), FT = IMAC flow-through buffer, E = elution buffer.

[0023] Figure 10 The image shows fibers spun by extruding a concentrated MaSp4_400 spider silk protein solution into a coagulation bath (0.75 M acetate, pH 5) using an HPLC pump. Detailed Implementation

[0024] This invention generally relates to recombinant spider silk proteins, and more particularly to recombinant spider silk proteins comprising REP domains of repeating silk protein regions derived from the hard-shell spider (Larinioides sclopetarius).

[0025] The spider silk protein of the present invention is a recombinant or engineered spider silk protein, i.e., an artificial, non-naturally occurring spider silk protein. The recombinant spider silk protein is preferably in the form of isolated recombinant spider silk protein. The recombinant spider silk protein of the present invention can produce silk fibers with excellent mechanical properties, comparable to those spun from NT2RepCT (WO 2018 / 002216; Andersson et al., Nat Chem Biol 11: 309-315 (2017)), which is said to have high tensile strength. However, compared to NT2RepCT, the spider silk protein of the present invention has a significantly larger number of amino acids while still being spinnable. More specifically, compared to NT2RepCT (whose REP domain consists of only 77 amino acids), the spider silk protein of the present invention has a significantly longer repeating region (REP) domain. In fact, the REP domain of the spider silk protein of the present invention can be at least 3 to 5 times longer than that of NT2RepCT, or even longer. This means that the recombinant spider silk protein of the present invention has a REP domain that is more similar to that of naturally occurring spider silk protein.

[0026] The recombinant spider silk protein of this invention contains repeating region (REP) domains of silk protein derived from the hard-shelled spider *Larinioides sclopetarius*. *L. sclopetarius*, commonly known as the bridge spider or gray cross spider, is a relatively large orb-weaving spider distributed throughout the Holarctic region.

[0027] L. sclopetarius weaves circular orb-like webs, unlike other orb-web spiders that weave oval-shaped webs. Furthermore, the shape of their orb-like webs changes as the spider ages. As the spider matures, the lower part of the sticky web continues to increase in size, while the upper part decreases proportionally. This difference in web size becomes more pronounced as the spider grows larger.

[0028] L. sclopetarius produces large, urn-shaped glandular silk fibers with impressive mechanical properties. The genome of this spider species has been sequenced, assembled, and annotated. A catalog of spider silk proteins containing 35 complete spider silk protein genes was manually compiled, including 12 MaSp, 4 MiSp, 2 PySp, 4 TuSp, 4 FlSp, 3 AcSp, 4 AgSp, and 2 AmSp.

[0029] Therefore, one aspect of the present invention relates to a recombinant spider silk protein comprising an N-terminal (NT) domain, a repeat region (REP) domain, and a C-terminal (CT) domain. According to the invention, the REP domain comprises at least 100 amino acids derived from the repeat region of a spider silk protein from *L. sclopetarius*. Furthermore, the NT domain is derived from the NT domain of a spider silk protein other than the one derived from the REP domain. Optionally, or furthermore, the CT domain is derived from the CT domain of a spider silk protein other than the one derived from the REP domain.

[0030] Therefore, the recombinant spider silk protein of the present invention comprises domains derived from at least two different spider silk proteins, preferably from at least two different spider species. The REP domain originates from a repeating region of a silk protein from *L. sclopetarius*. In one embodiment, the NT domain originates from the NT domain of a spider silk protein from another spider silk protein of *L. sclopetarius*, or preferably from the NT domain of a spider silk protein from another spider species (i.e., a spider species other than *L. sclopetarius*). The CT domain may originate from the CT domain of a spider silk protein from *L. sclopetarius*, for example, from a spider silk protein that is the same as or different from the *L. sclopetarius* spider silk protein from which the REP domain originates; from another spider species that shares the same NT domain; or even from yet another spider species different from *L. sclopetarius* and the spider species from which the NT domain originates. In another embodiment, the CT domain is derived from a CT domain of another spider silk protein from L. sclopetarius, or preferably from a CT domain of a spider silk protein from another spider species (i.e., a spider species other than L. sclopetarius). The NT domain can be derived from an NT domain of a spider silk protein from L. sclopetarius, for example, from a spider silk protein that is the same as or different from the L. sclopetarius from which the REP domain originates; from another spider species that shares the same CT domain; or even from yet another spider species different from both L. sclopetarius and the spider species from which the CT domain originates.

[0031] In one embodiment, the NT domain originates from the NT domain of a spider silk protein from a spider species different from *L. sclopetarius*, and / or the CT domain originates from the CT domain of a spider silk protein from a spider species different from *L. sclopetarius*. In a specific embodiment, the NT domain originates from the NT domain of a spider silk protein from a spider species different from *L. sclopetarius*, and the CT domain originates from the CT domain of a spider silk protein from a spider species different from *L. sclopetarius*. In this specific embodiment, the NT domain and the CT domain may originate from the same spider silk protein from different spider species, from different spider silk proteins from different spider species, or from different spider silk proteins from different spider species.

[0032] The recombinant spider silk protein preferably includes a REP domain disposed between the NT domain and the CT domain. Therefore, the recombinant spider silk protein preferably has the general formula NT-REP-CT.

[0033] As further described herein, the recombinant spider silk protein may contain other amino acid sequences in addition to the NT, REP, and CT domains, including optional N-terminal and / or C-terminal tags and / or optional linkers. Thus, in one embodiment, the recombinant spider silk protein has the general formula (X)-NT-(L1)-REP-(L2)-CT-(Y). In this embodiment, X represents an optional N-terminal tag, Y represents an optional C-terminal tag, L1 represents an optional first linker, and L2 represents an optional second linker.

[0034] As described above, the recombinant spider silk protein of the present invention may contain additional amino acid sequences or domains in addition to the NT domain, REP domain, and CT domain. These additional domains are preferably attached to the N-terminus of the NT domain and / or the C-terminus of the CT domain of the recombinant spider silk protein, i.e., X-NT-REP-CT, NT-REP-CT-Y, or X-NT-REP-CT-Y, and / or may be located between the NT and REP domains and / or between the REP and CT domains, i.e., NT-L1-REP-CT, NT-REP-L2-CT, or NT-L1-REP-L2-CT. The N-terminal and / or C-terminal labels X and Y can also be combined with the connector, for example, X-NT-L1-REP-CT, X-NT-REP-L2-CT, X-NT-L1-REP-L2-CT, NT-L1-REP-CT-Y, NT-REP-L2-CT-Y, NT-L1-REP-L2-CT-Y, X-NT-L1-REP-CT-Y, X-NT-REP-L2-CT-Y or X-NT-L1-REP-L2-CT-Y.

[0035] Schematic but non-limiting examples of these additional domains X and Y include affinity tags, solubilization tags, chromatographic tags, epitope tags, fluorescent tags, signal peptides, or sequences.

[0036] Examples of domains that facilitate purification include various affinity tags such as chitin-binding protein (CBP), maltose-binding protein (MBP), hemagglutinin tags, Strep-tags, and glutathione S-transferase (GST); and polyhistidine (His) tags, such as the His6 tag; solubilizing tags such as thioredoxin (TRX) and poly(NANP); chromatographic tags such as the FLAG tag; epitope tags such as the ALFA tag, V5 tag, Myc tag, HA tag, Spot tag, T7 tag, and NE tag; and fluorescent tags such as GFP. One example of an N-terminal tag that can be used is defined in SEQ ID NO: 135.

[0037] Schematic examples of connectors that can be used between the NT and REP domains and / or between the REP and CT domains include various GS connectors, such as GS, SGS, or GNS, as well as other peptide connectors. Such connectors may be advantageous in providing shorter distances between the NT and REP domains and / or between the REP and CT domains, thereby reducing the risk of any steric hindrance between the connected domains. The optional connectors can be very short, such as GS, SGS, or GNS, or up to tens of amino acids, preferably no more than 20 amino acids, and more preferably no more than 15 amino acids.

[0038] In one embodiment, the REP domain of the recombinant spider silk protein consists of at least 100 amino acid residues from a repeating region of silk protein derived from L. sclopetarius. Experimental data presented herein demonstrate that silk fibers can be spun from spider silk proteins containing relatively long REP domains (i.e., 130 or more amino acid residues), and the spun silk fibers exhibit excellent mechanical properties. Indeed, spider silk proteins with REP domains containing 400 amino acid residues or even longer can also be spun into silk fibers with excellent mechanical properties.

[0039] In a preferred embodiment, the REP domain comprises, preferably, a sequence of at least 125 amino acid residues derived from a repeating region of silk protein from L. sclopetarius, more preferably, at least 130 amino acid residues derived from a repeating region of silk protein from L. sclopetarius, more preferably, at least 150 amino acid residues derived from a repeating region of silk protein from L. sclopetarius, for example, at least 200, at least 250, at least 300, or at least 350 amino acid residues derived from a repeating region of silk protein from L. sclopetarius, and most preferably, at least 400 amino acid residues derived from a repeating region of silk protein from L. sclopetarius, for example, at least 450, at least 500, at least 550, or at least 600 amino acid residues derived from a repeating region of silk protein from L. sclopetarius.

[0040] In one embodiment, the REP domain comprises, preferably, a sequence of at least 100 consecutive amino acid residues derived from a repeating region of silk fibroin from L. sclopetarius, more preferably at least 125 consecutive amino acid residues, and more preferably at least 130 consecutive amino acid residues. In a specific embodiment, the REP domain comprises, preferably, a sequence of at least 150 consecutive amino acid residues derived from a repeating region of silk fibroin from L. sclopetarius, for example, at least 200, at least 250, at least 300, or at least 350 consecutive amino acid residues derived from a repeating region of silk fibroin from L. sclopetarius, and most preferably, at least 400 consecutive amino acid residues derived from a repeating region of silk fibroin from L. sclopetarius, for example, at least 450, at least 500, at least 550, or at least 600 consecutive amino acid residues derived from a repeating region of silk fibroin from L. sclopetarius.

[0041] The spider silk protein of the present invention preferably has a molecular weight of at least 30 kDa, more preferably at least 35 kDa, and more preferably at least 50 kDa. Experimental data presented herein further demonstrate that spider silk proteins with a molecular weight of 50-60 kDa or even greater (87 kDa) can also be spun into silk fibers.

[0042] In one implementation, the REP domain is derived from a repeating region of the MaSp of L. sclopetarius. L. sclopetarius contains twelve MaSps, named MaSp1a-MaSp1c, MaSp2a-MaSp2f, MaSp3a-MaSp3b, and MaSp4, respectively.

[0043] In one specific implementation, the REP domain is derived from the MaSp of L. sclopetarius selected from the group consisting of: MaSp1a as defined in SEQ ID NO: 38, MaSp1b as defined in SEQ ID NO: 41, MaSp1c as defined in SEQ ID NO: 44, MaSp2a as defined in SEQ ID NO: 47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3a as defined in SEQ ID NO: 65, MaSp3b as defined in SEQ ID NO: 68, and MaSp4 as defined in SEQ ID NO: 71.

[0044] In a preferred embodiment, the REP domain is derived from the MaSp of L. sclopetarius selected from the group consisting of MaSp1, MaSp2, MaSp3, and MaSp4. For example, the REP domain is preferably derived from the MaSp of L. sclopetarius selected from the group consisting of MaSp1a as defined in SEQ ID NO: 38, MaSp2a as defined in SEQ ID NO: 47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3a as defined in SEQ ID NO: 65, and MaSp4 as defined in SEQ ID NO: 71. In a particularly preferred embodiment, the REP domain is derived from the MaSp of L. sclopetarius selected from the group consisting of MaSp2c as defined in SEQ ID NO: 53 and MaSp4 as defined in SEQ ID NO: 71.

[0045] An example of the REP structure field of MaSp1a from L. sclopetarius includes BR_MaSp1a_400 as shown below: BR_MaSp1a_400 (SEQ ID NO: 148) GGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGSGSQGVGQGGYGAGQGGAGAAAAAGGAGGSGQGGLGAGQGYGAGLGGQGGAGQGGAASAAAA AGGSGGQGGYGGLGSQGAGQGGAASAAAAAVGAAGGQGGYGGLGSQGAGQGGYGAGQGGATSAAAAAAGGSGGQGGYGGLGSQGAGQSGSGSAAAAAAAG GAGGAGQGGLGAGQGYGPGLGGQRGAGQGGAASAAAAAAGGAGGQGGYGGFGSQGAGQGGYGAGQGGAASAAAAAGGAGGQGVYGGLGSQGAGQGGYGAG QGGAGSAAAAAAAVGEGGAGQGGLSAGQGYGSGLGGQGGAGQGGAASSAAAAGGSGGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGLGSQGAGQ Examples of the REP structure fields of MaSp2c derived from L. sclopetarius include BR_MaSp2_short, BR_MaSp2_long, BR_MaSp2_300, and BR_MaSp2_400, as disclosed below: BR_MaSp2_short (SEQ ID NO: 109) SAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQ BR_MaSp2_long (SEQ ID NO: 110) GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGP BR_MaSp2_300 (SEQ ID NO:111) GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSG BR_MaSp2_400 (SEQ ID NO:112) GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPS An example of the REP domain of MaSp2c derived from L. sclopetarius includes BR_MaSp2c_500 as shown below: BR_MaSp2c_500 (SEQ ID NO: 149) GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSSGSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGPSGPSGPGGNGPGSQGGPSGPGGYGPGSQGPNGPGGAGSSAAVAVSSGPGGYGPGSQGGP An example of the REP domain of MaSp3a derived from L. sclopetarius includes BR_MaSp3a_400 as shown below: BR_MaSp3a_400 (SEQ ID NO: 150) GGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAASAAAGGSGGPRGYGGLGSQGAGQGGYGAGQGGAASAAAGGSGGPGGYGGLGSQGAGQGGYGAGQGGAASAAAASAGGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQGSGQGGYGAGQDGASSVAAAAVSGSGGPGGYGGLGSQGAGQGRYGAGQGGADSTAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQRSGQGGYGAGQGGAASAASAAAGGSGGPRGYGGLGSQGTGQGGYGAGQSGAASAASAAAGGSGGPRGYGGL Examples of the REP domain of MaSp4 from L. sclopetarius include BR_MaSp4_short, BR_MaSp4_long, BR_MaSp4_400 and BR_MaSp4_631 disclosed below: BR_MaSp4_short (SEQ ID NO: 113) GPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPDQQGPSGPSQQGPSTQVSNGPGPQAPALSTFAFSGPVPEASSGPSAQQPSFQGPAGPRPQGPGS BR_MaSp4_long (SEQ ID NO: 114) GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGS BR_MaSp4_400 (SEQ ID NO:151) GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSGSVSVLSYGPGPQGPSGLSQQGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPGSSVSIFSYGPGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQ BR_MaSp4_631 (SEQ IN NO:168) GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSS GSVSVLSYGPGPQGPSGLSQQGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPG SSVSIFSYGPGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTGPGPQAPALS TFAFSGPVPQGPSGPVPQGPSPQ The bold part in BR_MaSp4_631 corresponds to MaSp4_long, and the underlined part in BR_MaSp4_631 corresponds to MaSp4_400. These two parts are connected to each other through an SNS header.

[0046] Therefore, in one embodiment, the REP domain comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, SEQ ID NO: 168. In one specific embodiment, the REP domain comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, for example, an amino acid sequence selected from the group consisting of: SEQ ID NO: 109 to 114. In another specific embodiment, the REP domain comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 109 to 112. In yet another specific embodiment, the REP domain comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 113 to 114.

[0047] In another embodiment, the REP domain is derived from a repeating region of the MiSp of L. sclopetarius. L. sclopetarius contains four MiSps, named MiSpa-MiSpd.

[0048] In one specific implementation, the REP domain is derived from the MiSp of L. sclopetarius selected from the group consisting of: MiSpa as defined in SEQ ID NO: 74, MiSpb as defined in SEQ ID NO: 77, MiSpc as defined in SEQ ID NO: 80, and MiSpd as defined in SEQ ID NO: 83.

[0049] In another embodiment, the REP domain is derived from a repeating region of the TuSp of L. sclopetarius. L. sclopetarius contains four types of TuSp, named TuSpa-TuSpd.

[0050] In one specific implementation, the REP domain is derived from TuSp of L. sclopetarius selected from the group consisting of: TuSpa as defined in SEQ ID NO: 92, TuSpb as defined in SEQ ID NO: 95, TuSpc as defined in SEQ ID NO: 98, and TuSpd as defined in SEQ ID NO: 101.

[0051] In yet another implementation, the REP domain is derived from a repeating region of the PySp of L. sclopetarius. L. sclopetarius contains two PySps, named PySpa and PySpb, respectively.

[0052] In one specific implementation, the REP domain is derived from PySp of L. sclopetarius selected from the group consisting of PySpa as defined in SEQ ID NO: 86 and PySpb as defined in SEQ ID NO: 89.

[0053] In another embodiment, the REP domain is derived from a repeating region of the AcSp of L. sclopetarius. L. sclopetarius contains three AcSps, named AcSpa-AcSpc.

[0054] In one specific implementation, the REP domain is derived from the AcSp of L. sclopetarius selected from the group consisting of: AcSpa as defined in SEQ ID NO: 2, AcSpb as defined in SEQ ID NO: 5, and AcSpc as defined in SEQ ID NO: 8.

[0055] In yet another implementation, the REP domain is derived from a repeating region of AgSp from L. sclopetarius. L. sclopetarius contains four types of AgSp, named AgSpa-AgSpc and AgSp-like.

[0056] In one specific embodiment, the REP domain is derived from AgSp of L. sclopetarius selected from the group consisting of AgSp-like as defined in SEQ ID NO: 11, AgSp1 as defined in SEQ ID NO: 13, AgSp2a as defined in SEQ ID NO: 16, and AgSp2b as defined in SEQ ID NO: 19. In another specific embodiment, the REP domain is derived from AgSp of L. sclopetarius selected from the group consisting of AgSp1 as defined in SEQ ID NO: 13, AgSp2a as defined in SEQ ID NO: 16, and AgSp2b as defined in SEQ ID NO: 19.

[0057] In yet another implementation, the REP domain is derived from a repeating region of the FlSp of L. sclopetarius. L. sclopetarius contains four types of FlSp, named FlSpa-FlSpc and FlSp-like, respectively.

[0058] In one specific embodiment, the REP domain is derived from FlSp of L. sclopetarius selected from the group consisting of: FlSp-like as defined in SEQ ID NO: 26, FlSpa as defined in SEQ ID NO: 29, FlSpb as defined in SEQ ID NO: 32, and FlSpc as defined in SEQ ID NO: 35. In another specific embodiment, the REP domain is derived from FlSp of L. sclopetarius selected from the group consisting of: FlSpa as defined in SEQ ID NO: 29, FlSpb as defined in SEQ ID NO: 32, and FlSpc as defined in SEQ ID NO: 35.

[0059] In one embodiment, the REP domain is derived from a repeating region of the ampulla sericulture silk protein (AmSp) of L. sclopetarius. L. sclopetarius contains two AmSp proteins, designated AmSp-like 1 and AmSp-like 2.

[0060] In one specific implementation, the REP domain is derived from the AmSp of L. sclopetarius selected from the group consisting of AmSp-like 1 as defined in SEQ ID NO: 22 and AmSp-like 1 as defined in SEQ ID NO: 24.

[0061] In one embodiment, the REP domain is derived from an amino acid sequence selected from the group consisting of: SEQ ID NO: 2, 5, 8, 11, 13, 16, 19, 22, 24, 26, 29, 32, 35, 38, 41, 44, 47, 50, 53, 56, 59, 62, 65, 68, 71, 74, 77, 80, 83, 86, 89, 92, 95, 98 and 101.

[0062] The NT domain of spider silk protein is believed to improve its solubility, thereby achieving extremely high protein concentrations in spinning solutions. Furthermore, the pH dependence of NT domain solubility is a crucial factor in achieving rapid polymerization of the spinning solution.

[0063] Some spider silk proteins’ CT domains do not exhibit pH-sensitive solubility (Hedhammar et al., Biochemistry 47(11): 3407-3417 (2008)), but in general, most CT domains with multiple charged amino acid residues are actually highly soluble and pH-dependent (Andersson et al., PLoSBiology 12(8): e1001921 (2014)).

[0064] The recombinant spider silk protein of the present invention can be used in various combinations of NT and CT domains with REP domains to form recombinant spider silk protein that can be spun into silk fibers.

[0065] Illustrative but non-limiting examples of NT domains that can be used in this invention are listed in Table 2 of US 2019 / 0248847, the teachings of which regarding NT domains are incorporated herein by reference.

[0066] In a preferred embodiment, the NT domain of the recombinant spider silk protein is derived from the NT domain of Eurosthenops australis MaSp1.

[0067] In one specific implementation, the NT structure domain comprises, preferably, SEQ ID NO: 115.

[0068] The NT domain of the Australian spider MaSp1 (SEQ ID NO: 115) MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA In another embodiment, the NT domain is derived from the NT domain of silk fibroin from L. sclopetarius. An illustrative but non-limiting example of such an NT domain is the NT domain derived from MaSp2f (SEQ ID NO: 61).

[0069] Illustrative but non-limiting examples of CT domains that can be used in this invention are listed in Table 1 of US 2019 / 0248847, the teachings of which on CT domains are incorporated herein by reference.

[0070] In a preferred embodiment, the CT domain of the recombinant spider silk protein is derived from the CT domain of Araneus ventricosus MiSp.

[0071] In one specific implementation, the CT domain comprises, preferably, SEQ ID NO: 116.

[0072] The CT domain of *MiSp* (SEQ ID NO: 116) VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG In another embodiment, the CT domain is derived from the CT domain of silk protein from L. sclopetarius. An illustrative but non-limiting example of such a CT domain is the CT domain derived from MiSpc (SEQ ID NO: 81).

[0073] In one embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, amino acid sequences selected from the group consisting of: SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, SEQ ID NO: 168; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0074] In one embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, amino acid sequences selected from the group consisting of SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0075] In another embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, amino acid sequences selected from the group consisting of SEQ ID NO: 109 to 114; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0076] In one specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain, preferably composed of SEQ ID NO: 115; a REP domain, preferably composed of SEQ ID NO: 109; and a CT domain, preferably composed of SEQ ID NO: 116.

[0077] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 103 and SEQ ID NO: 117-119.

[0078] In another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain, preferably composed of SEQ ID NO: 115; a REP domain, preferably composed of SEQ ID NO: 110; and a CT domain, preferably composed of SEQ ID NO: 116.

[0079] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 104 and SEQ ID NO: 120-122.

[0080] In yet another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 111; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0081] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 105 and SEQ ID NO: 123-125.

[0082] In yet another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 112; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0083] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 106, SEQ ID NO: 126-128.

[0084] In another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 113; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0085] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 107 and SEQ ID NO: 129-131.

[0086] In yet another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 114; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0087] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 108 and SEQ ID NO: 132-134.

[0088] In another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 148; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0089] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 152, SEQ ID NO: 156-158.

[0090] In yet another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 149; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0091] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 153 and SEQ ID NO: 159-161.

[0092] In yet another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 150; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0093] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 154, SEQ ID NO: 162-164.

[0094] In another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 151; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0095] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 155, SEQ ID NO: 165-167.

[0096] In yet another specific embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, SEQ ID NO: 168; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0097] Examples of such recombinant spider silk proteins are defined by SEQ ID NO: 169-172.

[0098] In one embodiment, the recombinant spider silk protein comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 103-108, 117-134, 152-167, 169-172, for example, an amino acid sequence selected from the group consisting of: SEQ ID NO: 103-108, 117-134, 152-167.

[0099] The recombinant spider silk protein defined as SEQ ID NO: 117, 120, 123, 126, 129, 132, 156, 159, 162, 165, 170 consists of the following: the NT domain composed of SEQ ID NO: 115, the REP domain composed of SEQ ID NO: 109-114, 148-151, 168, and the CT domain composed of SEQ ID NO: 116.

[0100] The recombinant spider silk protein defined as SEQ ID NO: 118, 121, 124, 127, 130, 133, 157, 160, 163, 166, 171 consists of the following: an NT domain consisting of SEQ ID NO: 115, a REP domain consisting of SEQ ID NO: 109-114, 148-151, 168, and a CT domain consisting of SEQ ID NO: 116, and has a linker between the NT and REP domains and between the REP and CT domains.

[0101] The recombinant spider silk protein defined as SEQ ID NO: 119, 122, 125, 128, 131, 134, 158, 161, 164, 167, 172 consists of the following: an NT domain composed of SEQ ID NO: 115, a REP domain composed of SEQ ID NO: 109-114, SEQ ID NO: 148-151, SEQ ID NO: 168, and a CT domain composed of SEQ ID NO: 116, and has an N-terminal tag.

[0102] The recombinant spider silk protein as defined by SEQ ID NO: 103-108, 152-155, 169 consists of the following: an NT domain consisting of SEQ ID NO: 115, a REP domain consisting of SEQ ID NO: 109-114, SEQ ID NO: 148-151, SEQ ID NO: 168, and a CT domain consisting of SEQ ID NO: 116, and has an N-terminal tag and a linker between the NT and REP domains and between the REP and CT domains.

[0103] In another embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 61; a REP domain comprising, preferably, amino acid sequences selected from the group consisting of SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, SEQ ID NO: 168, preferably selected from SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, more preferably selected from SEQ ID NO: 109 to 114; and a CT domain comprising, preferably, SEQ ID NO: 116.

[0104] In yet another embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 115; a REP domain comprising, preferably, amino acid sequences selected from the group consisting of SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, SEQ ID NO: 168, preferably selected from SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, more preferably selected from SEQ ID NO: 109 to 114; and a CT domain comprising, preferably, SEQ ID NO: 81.

[0105] In yet another embodiment, the recombinant spider silk protein comprises, preferably, the following: an NT domain comprising, preferably, SEQ ID NO: 61; a REP domain comprising, preferably, amino acid sequences selected from the group consisting of SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, SEQ ID NO: 168, preferably selected from SEQ ID NO: 109 to 114, SEQ ID NO: 148 to 151, more preferably selected from SEQ ID NO: 109 to 114; and a CT domain comprising, preferably, SEQ ID NO: 81.

[0106] Another aspect of the invention relates to silk fibers made from recombinant spider silk protein according to the invention. Therefore, this aspect relates to a silk fiber (sometimes also called a silk polymer) comprising the recombinant spider silk protein according to the invention. The silk fiber is obtained by spinning a so-called spinning solution comprising the recombinant spider silk protein according to the invention into a silk fiber, as will be further described below.

[0107] The filaments of the present invention have very high tensile strength, also known as engineered strength. In one embodiment, the average tensile strength of the filaments is at least 70 MPa, preferably at least 75 MPa, and more preferably at least 80 MPa.

[0108] The tensile strength and other mechanical properties described herein refer to the average of the average tensile strength and other mechanical properties determined when testing multiple filaments. This means that, among the filaments tested, the tensile strength of individual filaments may be lower than, for example, 70 MPa, while the tensile strength of other individual filaments tested may be higher than, for example, 70 MPa. However, the average tensile strength of the tested filaments is at least, for example, 70 MPa.

[0109] In one embodiment, the average toughness modulus of the filament fiber is at least 15 MJ / m³, preferably at least 20 MJ / m³, and more preferably at least 40 MJ / m³.

[0110] The average diameter of the filaments can range from one or a few micrometers to tens of micrometers. For example, the average diameter of the filaments is 1 µm to 100 µm, preferably 2 µm to 50 µm, and more preferably 4 µm to 15 µm.

[0111] In one embodiment, the average Young's modulus of the filament fiber is at least 1.8 GPa, preferably at least 1.9 GPa, and more preferably at least 2.0 GPa.

[0112] In one embodiment, the silk fiber further comprises spider silk constituent elements (SpiCEs) of L. sclopetarius. In one specific embodiment, the SpiCEs are selected from the group consisting of SEQ ID NO: 142 to 147.

[0113] The present invention also relates to synthetic materials comprising silk fibers according to the present invention.

[0114] illustrative examples of synthetic materials comprising or made from the silk fibers of the present invention include textile materials such as filaments, yarns, ropes, and braided materials. Such textile materials can benefit from the high tensile strength of the silk fibers. Other examples of synthetic materials include flexible energy-absorbing materials such as armor and bumpers. The silk fibers of the present invention can also be used in medical applications such as sutures, pressure bandages, etc. Furthermore, the silk fibers can be used in scaffolds and materials in tissue engineering, implants, and other cell-based scaffold materials.

[0115] The present invention also relates to nucleic acid molecules encoding recombinant spider silk proteins according to the present invention.

[0116] The term "nucleic acid molecule" as used in this article includes polynucleotides, oligonucleotides, and nucleic acid sequences, typically referring to polymers of DNA or RNA. These can be single-stranded or double-stranded, and may contain natural, non-natural, or modified nucleotides, as well as natural, non-natural, or modified nucleotide linkages, such as aminophosphate linkages or thiophosphate linkages, to replace the phosphodiester linkages between nucleotides in unmodified oligonucleotides. Nucleic acid molecules also include complementary DNA (cDNA) and messenger RNA (mRNA).

[0117] An illustrative but non-limiting example of such nucleic acid molecules is SEQ ID NO: 136-141 for recombinant spider silk protein in SEQ ID NO: 103-108.

[0118] Another aspect of the invention relates to an expression vector comprising a nucleic acid molecule according to the invention.

[0119] The expression vector comprises at least one nucleic acid molecule containing a coding sequence that can be expressed (e.g., transcribed and translated) in a cell containing the expression vector (typically referred to as a host cell). In one embodiment, the expression vector is selected from DNA molecules, RNA molecules, plasmids, episome plasmids, and viral vectors.

[0120] The expression vector also contains a nucleic acid molecule operatively linked to a promoter to enable transcription in a host cell. The promoter can be any promoter that is constitutively or inducibly active in the host cell.

[0121] The nucleic acid molecule encoding recombinant spider silk protein is operatively controlled by a promoter in the expression vector, i.e., under the transcriptional control of the promoter. In one embodiment, if the host cell is a eukaryotic cell, such as a human cell, the promoter is selected from the group consisting of: human EF1α promoter, CMV promoter, CAG promoter, PGK promoter, TRE promoter, U6 promoter, and UAS promoter. If the host cell is a bacterial cell, illustrative but non-limiting examples of promoters that may be used include the T5 promoter, T7 promoter, rham promoter, phoA promoter, and Sp6 promoter. lac Startup, AraBad startup trp Promoters and Ptac promoters. If the host cell is a yeast cell, the promoters may be selected from the following group, as illustrative but not limiting examples: CYC1 promoter, ADH1 promoter, TEF2 promoter, pCYC promoter, PGAL promoter, and GFD promoter.

[0122] An illustrative but non-limiting example of a promoter that can be used in Escherichia coli host cells is the T7 promoter. Protein production can then be induced in host cells by adding isopropyl-β-D-1-thiogalactoside.

[0123] Another aspect of the invention relates to a host cell comprising an expression vector according to the invention.

[0124] The nucleic acid molecule or expression vector can then be transcribed in a host cell to produce recombinant spider silk protein in the host cell.

[0125] According to the present invention, a variety of such host cells can be used, including but not limited to bacteria, yeast, mammalian cells, plant cells, and insect cells. Currently, the recombinant spider silk protein of the present invention is preferably produced in bacteria (such as Escherichia coli).

[0126] The recombinant spider silk protein can then be produced by the host cells, for example, by culturing host cells according to the invention under conditions that allow for the production of recombinant spider silk protein, and isolating the spider silk protein from the culture. In one specific embodiment, the spider silk protein is isolated from the cytoplasm of the host cells.

[0127] The present invention also relates to a method for producing silk fibers. The method includes extruding a spinning solution containing recombinant spider silk protein according to the invention into an aqueous buffer having an acidic pH to induce the recombinant spider silk protein to polymerize into silk fibers. The method further includes separating the silk fibers from the aqueous buffer.

[0128] In one embodiment, the spinning solution contains at least 90 mg / ml of the recombinant spider silk protein, preferably at least 125 mg / ml, more preferably at least 150 mg / ml of the recombinant spider silk protein.

[0129] In one embodiment, the aqueous buffer solution is an acetate buffer solution with a pH equal to or lower than 6, preferably equal to or lower than 5.5. In a specific embodiment, the aqueous buffer solution preferably also has a pH equal to or greater than 4, preferably equal to or greater than 4.5. In a preferred embodiment, the pH of the aqueous buffer solution is approximately 5.

[0130] Example Example 1 Materials and Methods protein sequence The following minispidroins were designed, with repeating regions derived from spider silk proteins found in MaSp of Swedish bridge spider (BR) L. sclopetarius.

[0131] His -NT- Connector 1 -BR_MaSp2_short- Connector 2 -CT(SEQ ID NO: 103) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNS SAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQ SGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT - Connector 1 -BR_MaSp2_long- Connector 2-CT(SEQ ID NO:104) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNS GGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGP SGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp2_300- Connector 2 -CT(SEQ ID NO:105) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSG SGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp2_400- Connector 2 -CT(SEQ ID NO:106) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPS SGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp4_short- Connector 2 -CT(SEQ ID NO:107) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNS GPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPDQQGPSGPSQQGPSTQVSNGPGPQAPALSTFAFSGPVPEASSGPSAQQPSFQGPAGPRPQGPGS SGSVTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp4_long- Connector 2 -CT(SEQ ID NO: 108) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNS GFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGS SGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG Cloning and Protein Expression Synthetic fragments of BR_MaSp2_short, BR_MaSp2_long, BR_MaSp2_300, BR_MaSp2_400, BR_MaSp4_short, and BR_MaSp4_long were cloned into the pT7-His-NT-CT vector using EcoRI and BamHI sites located between the NT and CT domains. Using NdeI and EcoRI restriction sites, the NT domain (SEQ ID NO: 115) of MaSp1 from Eurosthenopsaustralis was replaced with the NT domain from MaSp2f (SEQ ID NO: 61) in the pT7-His-NT-A3I-A-CT vector; simultaneously, using BamHI and HindIII restriction sites, the CT domain (SEQ ID NO: 116) of MiSp from Araneus ventricosus was replaced with the CT domain (SEQ ID NO: 81) of MiSpc. The obtained construct was transformed into BL21(DE3) *Escherichia coli*. *E. coli* glycerol strains containing the corresponding insert fragment pT7-His-NT-REP-CT vector were inoculated into LB broth containing kanamycin (70 µg / ml). The overnight culture was inoculated at a 1 / 100 ratio into LB broth containing kanamycin and then cultured at 30°C with shaking (110 rpm) until OD... 600 Reaching 0.6, the temperature was then lowered to 20°C, and at OD... 600 Protein expression was induced by adding isopropyl-β-D-thiogalactopyranoside (IPTG) to a final concentration of 0.15 mM when the concentration reached 0.8. Cells were cultured overnight at 20°C with shaking (110 rpm), and then harvested by centrifugation at 4°C and 5,000 rpm for 20 minutes. The pellet was resuspended in 20 mM Tris pH 8 buffer and... Store at 20℃.

[0132] Protein purification Lysis was performed using a TS Series Machine (Constant Systems Limited) at 30 kPsi, followed by centrifugation of the lysate at 25,000 × g for 30 min at 4 °C. The precipitate was discarded, and the supernatant was loaded onto two tandemly linked 20 mL HisPrep FF16 / 10 (Cytiva) columns to bind His-tagged proteins. The columns were washed with 5 column volumes (CV) of 20 mM Tris-HCl pH 8.0 buffer and 5 CV of 20 mM Tris-HCl pH 8.0 buffer containing 2 mM imidazole. Proteins were eluted with 20 mM Tris-HCl pH 8.0 buffer containing 200 mM imidazole. The eluted protein was dialyzed overnight at 4 °C against 20 mM Tris pH 8.0 buffer using a Spectra / Por dialysis membrane with a molecular weight cutoff of 6–8 kDa, changing the buffer at least three times during the process. Protein purity was determined by SDS-PAGE (4–20%) and Coomassie Brilliant Blue staining. The Broad Range Protein Ladder (Thermo Fisher Scientific) was used as the size standard. Protein concentration was determined by recording the absorbance of 10-fold diluted dialysis eluent at 280 nm.

[0133] Preparation of spinning solution The purified protein was concentrated at 4,000 × g and 4 °C using an Amicon Ultra-15 centrifugal filtration unit (Merck-Millipore, Darmstadt, Germany) equipped with an ultrafiltration membrane-10 (10 kDa molecular weight cutoff). The concentration of the concentrated protein was determined by recording the absorbance (in triplicate) of the 333-fold diluted sample at 280 nm. The obtained protein solution concentrations were as follows: BR_MaSp2_short 315 mg / ml; BR_MaSp2_long 280 mg / ml; BR_MaSp2_300 277 mg / ml; BR_MaSp2_400 300 mg / ml; BR_MaSp4_short 300 mg / ml; BR_MaSp4_long 281 mg / ml. Subsequently, the corresponding protein solutions were transferred to 1 mL syringes with Luer locks (BD, Franklin Lake, New Jersey, USA) and stored at -20 °C.

[0134] Biomimetic spinning of rayon fibers Based on the optimized schemes of Andersson et al., Nature Chemical Biology 13(3): 262-264 (2017), and referring to the optimization schemes of Greco et al., Molecules 25(14): 3248 (2020) and Schmuck et al., CommunicationsMaterial 3: 83 (2022), biomimetic spinning was performed on the protein concentrate (spinning solution). In short, the spinning solution was thawed at room temperature (20-25℃), and a 1 mL syringe was connected to a 27G flat-head steel needle (B. Braun, Melsungen, Germany) with an outer diameter (OD) of 0.40 mm. Next, the needle was wrapped with a polyethylene tube (BD Intramedic, Franklin Lake, NJ, USA) with an outer diameter of 1.09 mm and an inner diameter (ID) of 0.38 mm. Then, the wrapped needle was inserted into a polyethylene tube with an outer diameter of 1.65 mm and an inner diameter of 0.76 mm. Finally, a drawn glass capillary with a tapered opening (G1 Narishige, Tokyo, Japan, 1.0 mm outer diameter, 0.6 mm inner diameter, drawn using a microelectrode drawing instrument, Stoelting Co. 51217, Wooddale, Illinois, USA) was inserted approximately 3 cm into a polyethylene tube with an inner diameter of 0.38 mm to achieve a leak-free connection. Then, a syringe was placed in a neMESYS low-pressure (290 N) syringe pump (Cetoni, Kolbsen, Germany). The spinning solution was extruded through a glass capillary with a tapered tip and an opening of 50 ± 10 µm into an 80 cm long coagulation bath containing 0.75 M acetate buffer (pH 5) at a flow rate of 17 µl / min. At the end of the bath, fibers were continuously collected using a 35 cm circumference rotating wheel at a winding speed of 59 cm / s under conditions of relative humidity <40%.

[0135] Tensile test Tensile tests were performed according to the method previously described in Greco et al., Molecules 25(14): 3248 (2020). First, the fibers were mounted on a paper frame with a 10×10 mm square window. The fibers were secured to the frame using double-sided tape. The fiber diameter was determined using an optical microscope equipped with a 10× magnification lens, and the diameter was measured three times at arbitrary locations in each of three images obtained from three different fiber segments. The cross-sectional area was calculated using the average diameter, assuming a circular cross-section. The paper frame with the fibers fixed was then mounted on a 5943-Instron tensile testing apparatus equipped with a 5 N load cell, and the paper frame was subsequently cut on both sides. Displacement tests were conducted at a rate of 6 mm / min under conditions of relative humidity below 40% and 22°C. Engineering stress was calculated by dividing the maximum force by the average diameter of each fiber, and fracture strain was obtained by removing the total displacement over the gauge length. Young's modulus was obtained from the slope of the linear elastic portion of the stress-strain curve, while the ductile modulus was the total area under the stress-strain curve. The reported value is the average of at least 10 tested fibers. Outliers were not removed.

[0136] result All mini spider silk protein constructs based on *L. sclopetarius* MaSp REP were expressed as soluble proteins and were readily purified and concentrated to approximately 300 mg / ml under native conditions. Furthermore, all constructs were spinnable, meaning they could be extruded through drawn glass capillaries into 0.75 M acetate buffer at pH 5, where they formed solid fibers. Fibers could be continuously collected for several minutes using a winding speed of 59 cm / s. Moreover, the mechanical properties of these fibers were comparable to other mini spider silk proteins, such as NT2RepCT.

[0137] Table 1 - Mechanical properties of fibers made from mini spider silk protein from L. sclopetarius REP

[0138] Each value in Table 1 represents the average of n ≥ 10 measurements, with error expressed as ± one standard deviation. Outliers have not been removed.

[0139] Example 2 Materials and Methods protein sequence The following mini spider silk proteins were designed, with repeating regions derived from spider silk proteins found in the MaSp of Swedish bridge spider (BR) L. sclopetarius.

[0140] His -NT- Connector 1 -BR_MaSp1a_400- Connector 2 -CT(SEQ ID NO:152) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNS GGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGSGSQGVGQGGYGAGQGGAGAAAAAGGAGGSGQGGLGAGQGYGAGLGGQGGAGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGAASAAAAAVGAAGGQGGYGGLGSQGAGQGGYGAGQGGATSAAAAAAGGSGGQGGYGGLGSQGAGQSGSGSAAAAAAAGGAGGAGQGGLGAGQGYGPGLGGQRGAGQGGAASAAAAAAGGAGGQGGYGGFGSQGAGQGGYGAGQGGAASAAAAAGGAGGQGVYGGLGSQGAGQGGYGAGQGGAGSAAAAAAAVGEGGAGQGGLSAGQGYGSGLGGQGGAGQGGAASSAAAAGGSGGQGGYGGLGSQGAGQGGAASAAAAAGGAGGQGGYGGLGSQGAGQ GGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp2c_500- Connector 2 -CT (SEQ ID NO:153) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNSGGPGASAAVAVSSGPGGYGPGSPGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGGPSGPGSQGPSGPGGPGSSSAASGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSGPGGYGPGSQGPNGPGGPGASAAVAVSSGPGGYGPGSQGPSGPSGPGGYGPGSQGGPSSGSGPGGYGPGSQGPNGPGGPGSSAAVAVSSGPGGYGPGSQGPSGPSGPGGNGPGSQGGPSGPGGYGPGSQGPNGPGGAGSSAAVAVSSGPGGYGPGSQGGP SGS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp3a_400- Connector 2 -CT (SEQ ID NO:154) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNSGGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAASAAAGGSGGPRGYGGLGSQGAGQGGYGAGQGGAASAAAGGSGGPGGYGGLGSQGAGQGGYGAGQGGAASAAAASAGGSGGRGGYGGLGSQGTGQGGAASAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQGSGQGGYGAGQDGASSVAAAAVSGSGGPGGYGGLGSQGAGQGRYGAGQGGADSTAAAAAGGSGGQGGYGGLGSQGAGQGGYGAGQGGAASAAAAAAGGSGGPGRYGGLGSQRSGQGGYGAGQGGAASAASAAAGGSGGPRGYGGLGSQGTGQGGYGAGQSGAASAASAAAGGSGGPRGYGGL GSV TSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp4_400- Connector 2 -CT (SEQ ID NO:155) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSGSVSVLSYGPGPQGPSGLSQQGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPGSSVSIFSYGPGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQ GS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG His -NT- Connector 1 -BR_MaSp4_631- Connector 2 -CT (SEQ ID NO:169) MGHHHHHH MSHTTPWTNPGLAENFMNSFMQGLSSMPGFTASQLDDMSTIAQSMVQSIQSLAAQGRTSPNKLQALNMAFASSMAEIAASEEGGGSLSTKTSSIASAMSNAFLQTTGVVNQPFINEITQLVSMFAQAGMNDVSA GNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAFSPGGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSI STSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSNSGFQGPGSSGGALTSYGTGPQGPSRPGSQGPSPQGPNGPRPQGPGSSVTVLTSYGPGPQGPSQQGPSTQVQTGTGPQDPALTNYAF SGPGSQGPSGPSSQQQSLQGQAGPQPQGPGSSVRILSYGLSQQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYKPGQQGPSGPSQQGPSTQVSNGPGHQAPALSTFAFSGPVPEASSGPSTQQPSFQGPARPRPQGPGSSGSVSVLSYGPGPQGP SGLSQQGPSTQVPTGSGPQAPALTNYAFSGPGPQGPSGPSPQQPSLQGPAGPQPQGPGSSVSIFSYGPGLQGPSGPVPQGPSPQGPSVPGPQGPGSSVSISTSYGPGLQGNSGPSQQEPSTQGPTGPGPQAPALSTFAFSGPVPQGPSGPVPQGPSPQ GS VTSGGYGYGTSAAAGAGVAAGSYAGAVNRLSSAEAASRVSSNIAAIASGGASALPSVISNIYSGVVASGVSSNEALIQALLELLSALVHVLSSASIGNVSSVGVDSTLNVVQDSVGQYVG Cloning and Protein Expression Cloning and protein expression were performed as described in Example 1 above. All the mini spider silk protein constructs listed in Example 2 were expressed as soluble proteins in E. coli BL-21, either by shake-flask culture as described in Example 1 or by using a bioreactor (Schmuck et al., Materials Today 50: 16-23 (2021)) equipped with an Eppendorf® BioFlo 120 with a 15L single-layer glass container.

[0141] Protein purification After clarifying the bacterial lysate by centrifugation, the mini spider silk protein was purified by chromatography according to the method described in Example 1. When expressing the protein using a bioreactor, the yields of purified protein after IMAC were: 10.6 g / L for MaSp4_400, 9.2 g / L for MaSp4_631, and 5 g / L for MaSp2c_500. Figure 9 Visualizations of bacterial cell lysis and IMAC purification results for these three constructs on SDS-PAGE gels (4–20%) are shown. As an alternative purification strategy, the clarified lysate was treated with potassium phosphate buffer at pH 8.0 to a final concentration of 330–420 mM. This induced precipitation of spider silk proteins, which was obtained by centrifugation at 18,000 g for 20 min. The precipitate was dissolved in 20 mM Tris-HCl buffer adjusted to pH 10. After dissolution, the solution was dialyzed to 20 mM Tris-HCl at pH 8.

[0142] spinning MaSp4_400 was concentrated to 320 mg / ml as described in Example 1. Using an MFLX74931-30 (Masterflex) high-performance liquid chromatography (HPLC) pump, the concentrated spider silk protein solution was extruded through a 75 µm inner diameter, 38 mm length PEEK capillary into a bath containing 0.75 M acetate at pH 5. In the coagulation bath, the extruded spider silk protein solution formed fibers and was collected on a rotating wheel at a speed of approximately 3 m / min. See [link to example]. Figure 10 .

[0143] result All mini spider silk protein constructs based on MaSp REP from L. sclopetarius were expressed as soluble proteins and were easily purified, yielding 220–230 mg protein / L bacterial culture in shake flasks and up to 10 g / L using a bioreactor. The results showed that these spider silk protein constructs could be concentrated to >300 mg / mL and spun into fibers when extruded into a coagulation bath at pH 5 using an HPLC pump.

[0144] Example 3 In this embodiment, a complete spider silk protein library of the orb-weaver spider *Larinioides sclopetarius* is presented. Spatial transcriptomics and single-cell RNA (scRNA) sequencing revealed that the tail and sac of the ampulla of Vater gland are composed of six silk-producing cell types distributed in three distinct regions (A, B, and C). This, combined with image analysis of histological sections and proteomic analysis of sequentially dissolved fibers, indicates that the silk fibers have a three-layered structure, with the core composed of proteins belonging to the MaSp1, MaSp2, and MaSp4 families, the middle layer dominated by MaSp3, and the outermost layer containing several uncharacterized non-spider silk proteins.

[0145] result The large ampullae of L. sclopetarius is composed of 17 types of silk proteins. In this embodiment, we focus on the Swedish bridge spider (L. sclopetarius), whose large, urn-shaped glandular silk fibers possess excellent mechanical properties. Figure 1 First, the genome of this spider species was sequenced, assembled, and annotated (Table 2). A spider silk protein catalog containing 35 complete spider silk protein genes was manually compiled. Figure 2 Next, to reveal the protein composition of the ampulla of Vater and the silk fibers, proteomics analysis was performed using liquid chromatography-tandem mass spectrometry (LC-MS / MS). In the gland samples, a total of 2,145 proteins were identified in at least one replicate out of three replicates. Dissolving the silk fibers can be challenging, as different chemical treatments can extract different proteins. Therefore, three solvents (urea, HFIP, and LiBr) were used to extract proteins from the fibers, and a total of 130 proteins were identified from all combinations of treatments. However, because the fibers are easily contaminated by other silk types during spinning and by irrelevant proteins during treatment, only proteins that were also found in the ampulla of Vater proteomics data, predicted to have signal peptides, and whose abundance was greater than 0.1% of the total protein content of the silk fibers were considered. Figure 3 a). This process screened out 17 silk proteins, which were named "17 silk proteins" (Table 3), and the corresponding genes were named "17 silk protein genes" (Table 3).

[0146] Table 2 - Statistics on the assembly and annotation of the L. sclopetarius genome.

[0147] Table 3-17 Ranking of Silk Proteins / Genes as Biomarkers in Different Application Methods

[0148] *Ranked based on protein levels in glands and silk fibers. Rank 1 represents the highest protein content.

[0149] **Ranking based on the average normalized read count in the ampulla of Vater samples. Rank 1 has the most reads.**

[0150] ***DEG analysis of the tail, cyst, and ductal portions of the ampulla of Vater. Ranking is based on mean log2-fold change. Higher rankings indicate more significant differences in gene expression.

[0151] ****Identification of marker genes for manually annotated spots in spatial transcriptomics slices, representing different regions. Ranking is based on mean log2 fold change, with higher values ​​being better.

[0152] *****Identification of marker genes in cell types identified from single-cell data. Ranking is based on average log2-fold change, higher is better.

[0153] These 17 silk proteins include nine MaSp proteins, two ampullary gland spider silk protein-like proteins (AmSp-like1 and AmSp-like2), and six proteins of unknown function. These six proteins of unknown function are labeled, in order of abundance, as spider silk building blocks of Larinioides ampullary gland silk (SpiCE-LMa1-6). MaSp1 (ac), MaSp2 (b, c, e, f), MaSp3a, and MaSp4 account for 52%, 19%, 24%, and 1.6% of the silk fiber protein content, respectively. Figure 3 a). SpiCE-LMa protein accounted for 3.2%, while AmSp-like protein accounted for only 0.9% ( Figure 3 a). Compared to MaSp, the six SpiCE-LMa proteins generally have lower molecular weights. In terms of amino acid composition, SpiCE-LMa1 and 2 are similar to MaSp proteins, possessing spider-silk-like repetitive motifs, while SpiCE-LMa3-6 are Cys-rich and have amino acid compositions similar to AmSp-like 1 and 2. Figure 3 e). The predicted secondary structure content of SpiCE-LMa protein and the AlphaFold2 structure prediction indicate that there is significant heterogeneity within this type of protein.

[0154] Seventeen filiform genes are expressed in the tail and sac. Batch RNA sequencing of the entire ampulla of Vater was used to validate the correlation between RNA levels and protein abundance in the gland. Notably, the 17 silk protein loci were among the highest-ranking genes (Table 3). However, this analysis did not allow for the localization of specific gene expression to different parts of the gland. To improve the resolution of gene expression profiling, the ampulla of Vater was dissected into three distinct anatomical regions: the tail, the sac, and the duct, before sequencing. Figure 3 b). This means that the tail sample contains transcripts from region A, while the sac sample contains transcripts from all three regions (b). Figure 2 b, 3b). Partial least squares (PLS) regression analysis was performed on tail, sac, and duct samples using the ampulla of Vater gene set. Using only two variables with high predictive relevance (Q squared = 0.857), the three sample types were classified into different clusters, confirming transcriptomic diversity in the tail, sac, and duct. Figure 3 c). To determine the origin of the 17 silk proteins from the ampulla of Vater, the 17 silk genes were superimposed on a PLS map. The results showed that all MaSp genes were expressed in the tail of the ampulla of Vater, except for MaSp3a, which was expressed in the sac along with AmSp-like1 and 2. Figure 3 c). Two genes with unknown functions, SpiCE-LMa1 and SpiCE-LMa2, also belong to the gene group expressed in the sac. The expression profiles of filament genes in the three sites (tail / sac / duct) were further visualized using heatmaps. Figure 3 d) clearly shows that the genes encoding 17 silk proteins are expressed in the tail and sac, but not in the duct.

[0155] Hierarchical clustering divides the 17 filament genes into 6 clusters based on the similarity of their expression profiles. Figure 3 d). Among them, genes in cluster 1 showed differential expression in the tail samples, therefore belonging to the tail region; while all genes in clusters 5 and 6 showed differential expression in the sac samples, therefore belonging to the sac region. Some genes in clusters 2, 3, and 4 showed differential expression in the tail and clustered with cluster 1 at the same node, therefore also belonging to the tail region. Based on these results, cluster 1, containing MaSp2c, e, and f genes, is expressed only in the tail and can be classified as region A (…). Figure 2 (b, 3d), but the remaining genes expressed in the cyst cannot be attributed to one of the regions because the cyst sample contains tissue from all three regions.

[0156] The expression of the silk gene is spatially resolved in three regions. To improve the resolution of filament gene expression, we next used spatial transcriptomics (10X Visium). This is an unbiased and sophisticated technique that enables gene expression mapping at a resolution of 50 µm in tissue sections. Six whole abdominal sections from four female individuals were used, and spots on hematoxylin and eosin (H&E) stained sections were manually annotated as different filament glands based on histology (one section is shown in...). Figure 4 (a-4b). Subsequently, unified manifold approximation and projection (UMAP) was used to separate and visualize sequencing data from all spots across all sections. In UMAP, spots annotated as filamentous glands clustered together, separated from spots belonging to other tissue types. Within filamentous gland clusters, spots from different filamentous glands clustered together, consistent with manual annotation ( Figure 4 c). On spatial slices, the expression of spider silk proteins was specific to the corresponding glands, confirming the results of the bulk RNA expression profiling. Figure 2 d) proves the quality of the spatial data.

[0157] The spots annotated as covering the ampulla of Vater gland tissue can be further classified into areas A, B, or C based on the morphology of the epithelial cells in the six sections. Figure 5 (a, 5b). This yielded 847 spots from six slices, clustered by region in UMAP, indicating that the expression profiles in these regions were indeed distinct. Furthermore, marker genes identified from the spatial data in regions A, B, and C overlapped with bulk RNA PLS maps of the tail and sac (but not the duct). The expression of 17 filogenes on all spots annotated as regions A, B, or C was then visualized as a heatmap. Figure 5 c). Statistical analysis of expression levels showed that 13 of the 17 filogenes were mainly expressed in one of the regions (Table 3). MaSp1 (ac), MaSp2 (b, c, e, f), and SpiCE-LMa6 genes were significantly expressed in region A cells; MaSp3a, AmSp-like1, and AmSp-like2 were significantly expressed in region B cells; and SpiCE-LMa1 was significantly expressed in region C cells. SpiCE-LMa2 expression was significantly higher in both regions B and C. The expression of the other three proteins, SpiCE-LMa3, 4, and 5, was similar in all three regions. Figure 5 c). Figure 5 d shows the expression profile of selected genes in a large ampulla of Vater as an example in slice 1.

[0158] Seventeen silk genes in three regions were specifically expressed in six cell types. Interestingly, both MaSp1 and MaSp2 genes are expressed in cells in region A, but spatial transcriptomics analysis shows that their expression profiles differ in different regions of region A. Figure 5 This suggests that the epithelial region may house several different cell types. To clarify this, single-cell RNA sequencing was performed on the entire ampulla of Vater isolated from seven individuals. After quality control (QC), filtering, and analysis, 9700 cells were obtained, which clustered into eight groups. Based on gene overlap of the first two components with the bulk RNA gene set PLS analysis, these eight clusters were found to originate from the tail, sac, or duct of the gland. Consistent with results using bulk RNA data, three cell clusters were tail-specific, three sac-specific, and two duct-specific. These clusters were then compared with spatial transcriptomics data, classifying them into region A / B / C cells based on their gene overlap with marker genes in three different regions. Three clusters were identified in region A, one in region B, one in region C, and one cluster was found in all three regions. These clusters were assigned to cell types and named according to the region and the top marker spider silk protein gene. Figure 6 a) One cell type in region A has MaSp1a-c as the first three marker genes, named region A_MaSp1. The second cell type in region A has MaSp2a, b, and e as the first three marker genes, therefore annotated as region A_MaSp2. This cell type includes MaSp4 as the fourth marker gene. In the last cell type in region A, SpiCE-LMa3 is the top marker gene in the silk genome, named region A_SpiCE-LMa3. The top marker genes for cell types in regions B and C are MaSp3a and SpiCE-LMa1, respectively. These cell types are annotated accordingly. Region B_MaSp3 cells have AmSp-like 1 and 2 (two spider silk proteins lacking C-terminal domains, Figure 2 e) is the top marker gene. The last cell type is mainly located in the sac but cannot be assigned to a specific region, so it is named region ABC. This cell type also has MaSp3a as a marker gene. The remaining two cell types are named Duct_1 and Duct_2 based on the similarity of their expression profiles to bulk RNA data from the ducts. These two cell types do not have any silk proteins as marker genes. Figure 6 b).

[0159] A striking observation is that, when considering all 22,860 protein-coding genes, all 17 filogenes are among the top 100 marker genes in at least one of the six cell types. In most cases, they are in the top ten (Table 3 and...). Figure 6(b) All cell types possess a unique set of expressed silk genes, evidenced by the fact that 12 of the 17 silk protein genes are significantly expressed in only one of the six cell types, while the remaining five are expressed in two cell types (Table 3). In summary, the presence of the 17 silk proteins in the fiber can be directly attributed to the expression of the corresponding genes in the six cell types.

[0160] Six types of mitogenic cells can be spatially distinguished along the gland. To determine the spatial location of cell types within the ampulla of Vater, scRNA-seq data were combined with spatial transcriptomics data, and the two datasets were deconvolved. This enabled us to visualize the spatial distribution of cell types within spots annotated as ampulla of Vater on spatial slices. Figure 7 The patterns presented suggest that cell types confined to region A may not be uniformly distributed along that region. To address this issue, we turned to QuPath for image analysis. In short, as Figure 7 As shown in b, the ampulla of Vater glands in the sections were annotated according to region, yielding 101 cross-sections in region A, 2 cross-sections in region B, 9 cross-sections in region C, and 5 cross-sections of the ducts across all sections. This annotation was confirmed by digital color deconvolution in QuPath, which generated hematoxylin and eosin staining values ​​for each target region. When plotted, the H&E staining of the epithelium showed good correlation with the annotations of different regions. Furthermore, the perimeter of all cross-sectional portions of the gland was determined using QuPath and correlated with spots on the spatial transcriptome sections. Spatial spots corresponding to each cross-section in region A were extracted, and the average expression of marker genes in region A as a function of the cross-sectional perimeter was visualized. Figure 7 d). This indicates that MaSp2 expression is higher in the smaller perimeter cross-section of region A (proximal tail portion), while MaSp1 and SpiCE-LMa3 gene expression is higher in the larger perimeter cross-section of region A (towards the sac). Notably, a significant negative correlation was observed between hematoxylin content and cross-sectional perimeter in region A, further supporting the presence of different cell types along this region. Figure 7 e). Next, based on the perimeter value, the cross-section of region A was divided into three parts: proximal, mid-terminal, and distal. By combining scRNAseq, spatial, and QuPath data, the distribution of cell types in the corresponding spots of different parts of region A was obtained and visualized. Figure 7In region A, the proximal region shows a higher abundance of region A_MaSp2 cells, gradually decreasing in the middle and distal portions. Conversely, region A_MaSp1 cells are more prevalent in the distal portion compared to the proximal region. Region A_SpiCE-LMa3 cells are found along the entire length of region A, but are most common in the middle and distal portions. Region B_MaSp3 cells are found almost entirely in the spots annotated as region B, confirming their identification as region B cells. Spots in region C are predominantly composed of region C_SpiCE-LMa1 cells, with a small number of other cell types. Regions ABC cells are found in all regions, but at a lower abundance. The distribution of different cell types along the ampulla of Vater is illustrated in the diagram. Figure 7 g.

[0161] The location of the filament-producing cell type determines the composition of each layer in the silk fiber. Secretions from the AC region of the greater ampulla of Vater gland of L. sclopetarius showed significant differences after H&E staining and were separated within the glandular ducts. Figure 8 a). To confirm the region-specific origin of these secretions, vesicles in each region's cells and the secreted material forming the cambium in the lumen were annotated, and staining intensity was assessed using QuPath image analysis. The H&E intensity values ​​of the obtained vesicle contents and luminal secretions were plotted by region and color-coded according to their respective regions. The annotations of vesicle contents and secretions in each region were clustered together, clearly indicating that secretions in region A form the bulk of the silk-like material in the lumen, secretions in region B contribute to the surrounding intermediate layer, and secretions in region C form the outermost thin layer (…). Figure 8 b).

[0162] To verify whether the three-layered structure in the liquid filament feedstock persists in the macropogonal glandular filaments, a solubilization scheme involving different concentrations of urea (2 M, 4 M, and 8 M) was employed. Low concentrations of urea (2 M) did not dissolve the entire filament, thus enriching proteins present on the surface, while the highest concentration (8 M) dissolved the entire fiber. LC-MS / MS analysis of the soluble fractions provided the relative proportions of each protein at different urea concentrations. In 2 M urea, the relative proportions of SpiCE-LMa 1-5 and AmSp-like 1 and 2 proteins were significantly enriched compared to samples dissolved in 8 M urea (p < 0.05), while MaSp1, MaSp2b-c, MaSp2e-f, and MaSp4 were significantly enriched when using higher urea concentrations (p < 0.1). Figure 8f). In the fiber supernatant exposed to 4 M urea, several of the 17 silk proteins were detected, with MaSp3a being the most abundant. Next, we compared the protein composition from the proteomic data with the corresponding spatial gene expression in the glands. Of the seven proteins enriched in the 2 M urea sample, three overlapped with the marker gene of region C_SpiCE-LMa1, three overlapped with region B_MaSp3 cells, and one overlapped with region A_SpiCE-LMa3 cells (f). Figure 8 e, 8f). All eight proteins enriched in the 8 M urea sample overlapped with marker genes in regions A_MaSp1 and A_MaSp2 cells. Interestingly, MaSp3a (the most important marker gene in region B_MaSp3) showed the highest relative proportion in the 4 M urea sample. Figure 8 f).

[0163] Advanced molecular methodologies, such as single-cell RNA sequencing and spatial transcriptomics, are key tools for advancing our understanding of the biology of various tissues. However, their effectiveness depends on well-annotated genomes, which poses a significant challenge when studying non-model organisms. This challenge is particularly pronounced in studies of spider silk production, given the inherently large size and repetitive nature of spider silk protein genes. To address this limitation, we present a high-quality genome assembly of *L. sclopetarius*, characterized by near-complete annotation of almost all coding genes. Notably, our annotation reveals the presence of 35 full-length spider silk proteins, a higher number than previously reported. These spider silk proteins vary in length, spanning from 575 to 9146 amino acid residues. Figure 2 All spider silk protein genes encode proteins with signal peptides, N-terminal domains, and repeating regions, and contain in-frame stop codons. Manual examination of the 3' downstream sequences revealed no signs of shifted reading frames, meaning that the spider silk protein gene catalog described herein is the first catalog to contain only complete genes.

[0164] Since the focus of this embodiment is to provide a detailed understanding of the structure and function of the ampullae-gyne silk gland and the fibers it produces, we next attempt to identify the most abundant proteins in the fibers. This is not a simple task, as ampullae-gyne fibers are easily contaminated by other silk types during spinning and collection. To avoid these problems, we used liquid chromatography-tandem mass spectrometry (LC-MS / MS) on both continuously collected ampullae-gyne silk fibers from spiders fixed under a microscope and on isolated ampullae-gyne glands. By considering only secretory proteins found in both datasets, we can ensure the removal of any contaminating proteins from other glands. By filtering proteins with an abundance greater than 0.1%, we identified 17 proteins as the most abundant proteins in ampullae-gyne silk. Nine of these 17 are from four different classes of ampullae-gyne spider silk proteins, namely MaSp1-4. These spider silk proteins together account for 96% of the total protein content of the fibers (…). Figure 8 g). In addition to MaSp, a previously unreported type of spider silk protein was also found in the fiber, which we named AmSp-like. AmSp is characterized by an N-terminal domain grouped with the small and large ampullae glands and a repeating region, but interestingly, these spider silk proteins lack a C-terminal domain and show a different amino acid composition compared to the MaSp gene. Figure 3 e). Finally, six additional proteins were identified as components of the fiber, named SpiCE-LMa1 to SpiCE-LMa6. Although they are collectively named "SpiCE-LMa," they exhibit diverse properties in terms of molecular weight, tertiary structure predicted by AlphaFold2, and amino acid composition. SpiCE-LMa1 and SpiCE-LMa2, expressed in region C, have amino acid compositions similar to MaSp, but they simultaneously lack both N-terminal and C-terminal domains. Notably, the SpiCE-LMa3-6 proteins expressed and secreted from regions A and B have high cysteine ​​content, suggesting that these proteins may form cysteine ​​knots, thus contributing to fiber toughness. The overrepresentation of MaSp and the presence of low levels of SpiCE proteins in the ampulla of Vater filaments of L. sclopetarius are consistent with reported protein content in the ampulla of Vater filaments of other spider species.

[0165] To determine the cell types expressing 17 filament proteins and their spatial distribution within glands, we combined three unbiased transcriptomics techniques. First, we used batch RNA sequencing to reveal that all 17 filament genes were expressed in the tail and sac regions, but not in the ducts. Figure 3d). Twelve of these genes showed significant differential expression in the tail or sac portion of the gland, indicating that the expression profile indeed differs along the glandular contour. Secondly, single-cell RNA sequencing analysis of the entire ampulla of Vater identified eight cell types. Notably, six of these eight cell types showed overlap between their marker genes and all 17 filogenes (Table 3 and...). Figure 6 b). By cross-aligning scRNA cell type marker genes with differentially expressed genes identified in bulk RNA data, we were able to classify three cell types as tail, three as sac, and two as duct. The transcriptomic profiles of cells expressing spider silk proteins matched bulk RNA sequencing data from tail and sac samples, but not from duct samples. This means that 17 silk proteins are produced by six cell types located in the tail and sac. To spatially resolve the distribution of cell types, we used a third transcriptomics technique, 10X Visium. In the slices used for spatial transcriptomics, we first manually annotated transcriptomic spots in regions A, B, and C of the ampulla of Vater using unique H&E staining patterns and epithelial morphology in each region. Figure 5 b). Comparison of gene expression across regions showed that 13 out of 17 silk protein genes were differentially expressed in one of the three regions. Figure 5 c). Next, by integrating spatial transcriptomics with single-cell data, the precise localization of cell types was revealed. Consistent with the bulk RNA data, the three cell types belonging to the sac were predominantly located in regions B and C. Notably, using spatial transcriptomics data, we could clearly see a distinct difference between the B_MaSp3 cells confined to region B and the C_SpiCE-LMa1 cells that dominate the epithelium in region C. The three cell types belonging to the tail were indeed found in region A. The tail is long and curved, and by observing the increase in the cross-section of the tail along the gland, we could even generate information about the spatial location of cell types within region A. By measuring the circumference of each cross-sectional portion of the tail, we could sort them from proximal (small) to distal (large). Figure 7 We found that region A_MaSp2 cells were mainly located in the proximal part of region A, while region A_MaSp1 cells were most abundant in the distal part of region A (near region B cells). Region A_SpiCE-LMa3 cells were present in all parts of region A. Figure 7 Finally, consistent with the bulk RNA data, neither of the two cell types belonging to the duct was detected in the tail or sac. In summary, our data support that the proteins constituting the ampulla of Vater gland filaments are produced by six cell types with specific regional anatomical localizations in regions A, B, and C, rather than by cell types confined to the duct.

[0166] Next, we attempted to link gene expression in cell types along the glands with the multilayered structure observed in the silk feedstock. Figure 8 a). Using digital image analysis, we demonstrated that the H&E staining of the middle layer of the spinning solution (liquid raw material stored in vesicles), as determined by QuPath, corresponds to the staining of intracellular vesicles in the corresponding epithelial cells. The inner layer of the spinning solution matches the staining of vesicles in region A, the middle layer matches the staining of vesicles in region B, and the outer layer matches the staining of vesicles in epithelial cells in region C. Figure 8 a, 8b). This indicates that the three regions indeed produce stratified secretions with different protein compositions. Furthermore, we know the presence of different cell types and their expression profiles in regions A, B, and C ( Figure 8 e) allows us to predict the protein composition of different layers within the fiber. Based on this model ( Figure 8 c) The inner layer of the silk fibers contains MaSp1, MaSp2, MaSp4, SpiCE-LMa3, and SpiCE-LMa6 proteins; the middle layer mainly contains MaSp3, AmSp-like1-2, and SpiCE-LMa2 proteins; and the outer layer is predominantly composed of SpiCE-LMa1 and SpiCE-LMa2, but lacks classic spider silk proteins. To verify the hypothesis that the layer identified in the glandular lumen continuously forms a layer within the fiber, we performed proteomic analysis on the silk fiber extracts by exposing the ampulla of Vater gland silk fibers to different concentrations of urea (2 M, 4 M, and 8 M, respectively). The fiber supernatant incubated in 2 M urea contained proteins mainly expressed in regions B and C, while in 8 M urea that completely dissolved the silk fibers, proteins expressed in region A were enriched (…). Figure 8 f). When using 4 M urea, MaSp3a (a marker gene in region B) had the highest relative abundance. These results enabled us to propose a detailed model of the three-layered filamentary structure of the ampulla of Vater gland and conclude that each layer has a unique protein composition derived from specific cell types confined to regions A, B, and C, respectively. Figure 8 f-8h).

[0167] Since region A_MaSp2 cells are the dominant cell type in the most proximal part of the tail, it is logical that MaSp2 proteins form the core of the filament, while the more peripheral region of the core is dominated by MaSp1 proteins, which are secreted by region A_MaSp1 cells located more distal to region A. This finding contrasts with the report by Hu et al. (Hu, et al., Amolecular atlas reveals the tri-sectional spinning mechanism of spiderdragline silk. Nat Commun 14, 837 (2023)), which showed that the central core of the ampullae of Trichonephila fibers is dominated by the MaSp1 protein, while MaSp2 is more peripheral. This is consistent with the work of Sponner et al. (Sponner, et al., Composition and hierarchical organisation of a spidersilk. Plos One 2, e998 (2007)), who, through biochemical and immunohistochemical studies of the ampullae of Trichonephila fibers, revealed the simultaneous presence of MaSp1 and MaSp2 in the inner core, but with only MaSp1 in the peripheral portion of the core. The latter study also concluded that the layer surrounding the fiber core (called the cortex) is more resilient and chemically resistant than the outermost layer (coating) and the central core. Combining the above viewpoints with the data in this paper, a reasonable conclusion is that the cortex described by Sponner et al. is predominantly MaSp3, corresponding to the middle layer in our model. The mechanical properties of the ampullae gland silk from *L. sclopetarius* are among the best studied. However, in the ampullae glands of several distantly related spider species (these species do not express MaSp3), there have been reports of layered structures and three epithelial regions, suggesting that not only the specific protein composition of the fiber layer, but also the layered structure itself may have significant implications for fiber performance. For example, the ampullae glands of spider species that do not express MaSp3 (such as *Euprosthenops* and *Fungitrachidae*) also have three epithelial regions, and the fibers spun by the *Euprosthenops australis* have one of the highest reported tensile strengths.

[0168] In summary, this embodiment provides a high-resolution spatial map of different cell types in the ampulla of Vater gland of L. sclopetarius. Previously uncharacterized genes highly differentially expressed in different cell types were identified. The protein composition of the mysteriously layered secretions within the gland was revealed, showing a correspondence with the protein composition of sequentially dissolved fibers.

[0169] method Spider samples An adult female *L. sclopetarius* spider was collected in the wild from a small habitat in Uppsala, Sweden. The spider's taxonomic identity was verified by the Natural History Museum of Stockholm, Sweden. The spider was kept in a large container that allowed it to spin its web. It was fed weekly with mealworms or fruit flies and given water daily.

[0170] Genomic DNA extraction and sequencing High molecular weight (HMW) genomic DNA was extracted from the entire body of a *L. sclopetarius* spider using the MagAttract HMW DNA Kit (Qiagen). Briefly, an adult female spider was anesthetized with dry ice and dissected on ice. The exoskeleton was removed, and all soft tissue was collected for HMW DNA extraction. DNA extraction was performed according to the manufacturer's protocol, except that the tissue was incubated with RNase and proteinase-K at 50°C for 30 min, and eluted twice with an additional 100 µL of buffer AE to the magnetic beads. The purified DNA was electrophoresed on a 0.5% agarose gel to assess DNA integrity. The absorbance ratios were assessed on a Nanodrop spectrophotometer, with the following results: 260 / 280: 1.82; 260 / 230: 1.91; a total of 36.2 µg of DNA was obtained. The extracted genomic DNA was also quality-checked using a BioAnalyzer (Agilent), which showed a single peak at approximately 11 kb. A 20 kb library was prepared using 10.3 µg of DNA. Library preparation and sequencing were performed using the National Genomics Infrastructure (NGI) platform of Uppsala University's SciLifeLab, employing PacBio long reads and 10X Genomics linkage reads.

[0171] PacBio genome library preparation and sequencing DNA samples that passed quality control were sent to the NGI platform of SciLifeLab at Uppsala University for library preparation and sequencing. Using the TPK1 kit obtained according to the manufacturer's instructions, SMRT-bell samples were screened at 20 kb size using a BluePippin instrument (SAGE), and sequencing was performed on 60 SMRT cells using the P5-C3 chemical reagent on an RSII instrument. For each SMRT cell, 10 hours of sequencing movies were captured. A total of 798 Gbp of data was generated, with an insert size of 11 kb.

[0172] Library preparation and sequencing using 10X Chromium linked reads 10X linked read libraries were generated using HMW DNA on the 10X Genomics Chromium platform according to the manufacturer's instructions (Genome Library Kit & Gel Bead Kit v2 PN-1000017, genome Chip Kit v2PN120257). The 10X libraries were sequenced in an S4 flow cell using a NovaSeqXp workflow with 151 bp paired-end settings on an Illumina NovaSeq 6000 instrument (NovaSeq Control Software 1.6.0 / RTA v3.4.4). Bcl-to-FastQ conversion was performed using bcl2fastq_v2.19.1.403 from the CASAVA software suite. Sanger / phred33 / Illumina 1.8+ was used as the quality scale.

[0173] Batch RNA extraction and sequencing from silk glands, head, and abdomen Spiders were anesthetized with dry ice before dissection on ice. After incision at the stalk, the abdomen was gently fixed onto a wax plate placed under a ZeissStemi 305 stereomicroscope, and the exoskeleton was carefully removed with micro-scissors to expose the silk glands. Excess non-silk tissue was washed away with phosphate-buffered saline (PBS, pH 7.4). Using micro-tweezers, the silk glands (major ampulla, minor ampulla, flagellated, and conglomerate glands) were separated by grasping the ducts. The varicella and piriform glands from these five spiders were difficult to separate due to their small size and were extracted as a single sample. The major ampulla from six other individuals was cut into three parts: the tail, the sac, and the duct, for RNA extraction. In another preparation, after removing the exoskeleton, soft tissue from the entire abdomen was scraped for RNA extraction. RNA from the head was extracted in a similar manner after removing the legs and thick exoskeleton. Five replicates were collected for each sample type, and RNA was extracted separately from each replicate.

[0174] Forty-one RNA samples were extracted using the RNeasy Plus Mini kit (Qiagen) according to the manufacturer's instructions. Sample integrity was assessed using Tapestry (Agilent Technologies). Transcriptome libraries were generated from different tissues using the Illumina TruSeq Stranded mRNA kit according to the manufacturer's instructions, and 151 bp paired-end reads were sequenced using a Novaseq 6000 instrument.

[0175] Construction and sequencing of PacBio long-read Iso-Seq libraries for the ampulla of Vater RNA was extracted from the ampulla of Vater of an individual and homogenized in TriZol. The extracted RNA was sequenced at NGI at Uppsala University, Sweden. RNA quality control (QC) was performed on an Agilent Bioanalyzer instrument using the Eukaryote Total RNA Nano kit. Sequencing libraries were prepared using the NEBNext® Single Cell / Low Input cDNA Synthesis & Amplification Module, Iso-Seq Express Oligo Kit, ProNex magnetic beads, and SMRTbell ExpressTemplate Prep Kit 2.0, following PacBio's Procedure & Checklist - Iso-Seq™ ExpressTemplate Preparation for Sequel® and Sequel II Systems, PN 101-763-800 version 02 (October 2019). Samples (300 ng) were first amplified for 12 cycles, followed by 3 additional cycles. For the purification of amplified cDNA, the Long Transcripts workflow was applied to obtain material enriched with longer transcripts (>3 kb). SMRTbell libraries were quality controlled using the Qubit dsDNA HS kit and the Agilent Bioanalyzer High Sensitivity kit. Primer annealing and polymerase binding were performed using the Sequel II binding kit 2.0. Samples were sequenced on a Sequel II instrument using a Sequel II sequencing plate 2.0 and a Sequel® II SMRT® Cell 8M for a movie time of 24 hours and a pre-extension time of 2 hours.

[0176] Single-cell preparation from the ampulla of Vater gland of L. sclopetarius Twenty spiders were used for single-cell sequencing. The first ten spiders were anesthetized and dissected on dry ice. All buffers were bubbled with carbon gas before and during use. The ampulla of Vater glands were removed from Ringer's solution at pH 7.4. The ducts were removed. The glands were washed with PBS (pH 7.4) and then incubated in pre-warmed trypsin-EDTA (Gibco, 0.5%) at 37°C for 1 min in low-binding microcentrifuge tubes. The glands were briefly triturated with a pipette and centrifuged at 300 g for 3 min at 4°C. The pellet was resuspended in 300 µL DMEM (Gibco) containing 1% BSA (Sigma). Before resuspending, the DMEM and tubes were briefly bubbled with carbon gas. The suspension was filtered through a 40 µm cell filter and transferred to low-binding microcentrifuge tubes. Cells were counted using trypan blue dye.

[0177] Because contaminating spinneret droplets made single-cell isolation difficult, a slightly modified protocol was used for the subsequent 10 spiders. These spiders were treated as described above, but instead of removing the ducts, the cysts were cut open and placed in PBS (pH 7.4) for 10 minutes to allow the spinneret to flow out. Glandular fragments were collected and incubated for 2 minutes at 37°C in pre-warmed trypsin-EDTA (Gibco, 0.5%) in low-binding microcentrifuge tubes. The suspension was pipetted for 1 minute using a flame-polished glass Pasteurized pipette, incubated for 2 minutes at 37°C in trypsin-EDTA, and then pipetted for another 3 minutes using a flame-polished glass Pasteurized pipette. The suspension was centrifuged at 300 g for 3 minutes at 4°C, and the pellet was resuspended in 200 µL DMEM containing 3% BSA. The suspension was filtered through a 40 µM pre-washed cell filter into low-binding microcentrifuge tubes. The filter was further washed with 100 µL DMEM containing 3% BSA to minimize cell loss.

[0178] Sample preparation for spatial transcriptomics The entire metatarsal region of the spider was rapidly frozen in OCT embedding medium in an isopentane-dry ice bath and stored at -80°C until use. Samples were sectioned to a thickness of 10 µm using a cryostat with the blade temperature set to -23°C and the sample holder temperature set to -10°C. Eight sections from five individual spiders were carefully mounted on the capture region (6.5 × 6.5 mm) of a Visium space slide (10X Genomics) and permeated for 30 min according to the 10X Visium space tissue optimization protocol. Libraries were constructed according to the Visium space gene expression protocol (10X Genomics). cDNA amplification was performed for 13–16 PCR cycles, and indexing for 13 PCR cycles. Sequencing was performed on an Illumina Nova-Seq 6000 using an SP-200 flow cell. All sections were stained with H&E for histological evaluation.

[0179] Genome assembly and correction PacBio data were assembled using the de novo assembly procedure Falcon and Falcon-Unzip (pb-falcon version 0.2.7) (Chin, et al., Phaseddiploid genome assembly with single-molecule real-time sequencing. NatMethods 13, 1050-1054 (2016)). Initial correction was performed using the Falcon-Unzip correction module. Long read CLR assembly was further corrected using Chromium 10X linked read data. Reads were aligned to PacBio assemblies using Long Ranger (version 2.1.4), and three rounds of correction were performed using diploid markers using Pilon (Walker, et al., Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One 9, e112963 (2014)) (version 1.22).

[0180] Genome size and heterozygosity estimation Raw Illumina reads from 10X genome linkage sequencing libraries were trimmed using Trimmomatic (Bolger, et al., Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114-2120 (2014)). Canonical 20-mer counts were collected using Jellyfish (Marcais and Kingsford, Afast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764-770 (2011)). Approximate genome size and heterozygosity were estimated using GenomeScope (Vurture, et al., GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics 33, 2202-2204 (2017)) based on the 20-mer histogram.

[0181] Mitochondrial genome assembly The mitochondrial genome was identified by aligning the assembled results with the existing reference spider species *Neosconaadianta* (Genbank accession number: NC_029756.1) using BLAST. Identified regions were extracted from the assembled results using BEDTools (Quinlan and Hall, BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842 (2010)). The extracted mitochondrial contigs were then annotated using the MITOS web server (Bernt, et al., MITOS: improved de novo metazoanmitochondrial genome annotation. Mol Phylogenet Evol 69, 313-319 (2013)).

[0182] PacBio Iso-Seq transcriptome assembly To generate full-length concordant transcript isoforms from the ampulla of Vater, raw polymerase reads were processed using SMRTlink. Subread BAM files were processed to generate circular concordant (CCS) reads. These reads were further classified into full-length (FL) transcript sequences based on the presence of a 5' primer, a 3' primer, and a polyA tail. FL transcript sequences were processed using the IsoSeq3 platform to generate full-length non-chimeric reads (FLNC), and further clustered using the ICE algorithm to produce high-quality and low-quality corrected full-length concordant sequences. High-quality (HQ) sequences were used for subsequent analysis. High-quality transcripts were aligned to the de novo assembled L. sclopetarius genome using minimap2 (Li, Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094-3100 (2018)) (version 2.2.4) with the parameter -ax splice-uf--secondary=no-C5. The alignment results in SAM format were processed into non-redundant full-length transcripts using the "collapse_isoforms_by_sam.py" script from the cDNA-Cupcake tool (https: / / github.com / Magdoll / cDNA_Cupcake). To assess the completeness of the transcriptome data, the arthropoda_odb9 dataset was analyzed using BUSCO (Simao, et al., BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210-3212 (2015)) (version 3.0.2).

[0183] Genome annotation The de novo assembled L. sclopetarius genome was annotated using MAKER (Cantarel, et al., MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes. Genome Res 18, 188-196 (2008)) version 3.01.02. High-confidence protein sequences (561,356 proteins) were collected from the Uniprot Swiss-prot database (downloaded November 2019), and a specific set of spider silk protein sequences (1,051 sequences) was downloaded from NCBI (November 2019).

[0184] A library of repetitive sequences was created using the RepeatModeler package (version 1.0.11, https: / / www.repeatmasker.org / RepeatModeler / ). Since spider silk protein sequences are inherently highly repetitive, the repetitive sequences modeled by RepeatModeler were reviewed against our specific spider silk protein dataset. Repetitive sequences in the assembled genome were identified using RepeatMasker (version 4.0.9, https: / / www.repeatmasker.org / ) and RepeatRunner (https: / / www.yandell-lab.org / software / repeatrunner.html). tRNAs were identified using tRNAscan (Lowe and Eddy, tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res 25, 955-964 (1997)) version 1.3.1, while conserved ncRNAs were identified using the Infernal package (Nawrocki, et al., Infernal 1.0: inference of RNA alignments. Bioinformatics 25, 1335-1337 (2009)) and the RNA family gene bank, Rfam (Burge, et al., Rfam 11.0: 10 years of RNA families. NucleicAcids Res 41, D226-232 (2013)) version 11.

[0185] The MAKER package is executed in two rounds: a) First, a configuration file is created using Uniprot Swiss-Prot protein sequences, a specific set of spider silk protein sequences, and RNA-seq data from different tissues. An internal workflow is used to select a set of genes from the initial evidence-based annotation (first round) for training Augustus (Stanke, et al., AUGUSTUS: a web server for gene finding in eukaryotes. Nucleic Acids Res 32, W309-312 (2004)) and SNAP (Korf, Gene finding in novel genomes. BMC Bioinformatics 5, 59 (2004)); b) A second round of MAKER is run using the evidence from the first round and the predictions from Augustus. The predictions from Augustus are used to build the gene model.

[0186] Functional annotation of genes and transcripts was performed using the translational CDS features for each coding transcript. Protein sequences were retrieved using BLAST by alignment against the Uniprot Swiss-Prot database and specific spider silk protein sequence sets to obtain gene names and protein functions. Additional annotations (functional domains and sites) were extracted from 20 other biological databases using Interproscan (5.30–69.0) (Paysan-Lafosse, et al., InterPro in 2022. Nucleic Acids Res 51, D418–D427 (2023)).

[0187] Improving 3' UTR Annotation Using Single-Cell Data The alignment BAM files from Cell Ranger were filtered using UMI-tools FilterBam (Smith, et al., UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Res 27, 491-499 (2017)) to include reads with molecular barcode tags corrected for Cell Ranger counts. The filtered BAM files were then processed using UMI-tools dedup (Smith et al., 2017) to remove PCR duplicates. Peaks were identified using Homer findPeaks (-size 50-fragLength 100-minDist 1). The peak files were converted to BED files using BEDTools (Quinlan and Hall, 2010). These peaks were then annotated with the 3'UTR of genes lacking this feature, or re-annotated with the extended 3'UTR feature of the most recently identified gene within 5000 bp.

[0188] Identification and artificial processing of spider silk proteins To identify spider silk proteins in the L. sclopetarius genome, two different methods were used: a) The translated protein sequences of L. sclopetarius were scanned against the PFAMHMM maps of the N-terminal domain, C-terminal domain, and tubular gland eggshell silk chain domain. The HMM maps of these domains, namely Spidroin_N.hmm (N-terminal domain, PF16763), Spidroin_MaSp.hmm (C-terminal domain, PF11260), and RP1-2.hmm (tubular gland eggshell silk chain domain, PF12042), were downloaded from the PFAM database. To minimize the risk of false positives, matches with an e-value cutoff below 1e-05 were filtered out; b) Reference spider silk protein sequences were downloaded from the NCBI database (November 2019). Redundant sequences were removed using CD-HIT with a 95% sequence identity cutoff, and a custom database containing full-length spider silk protein sequences (retrieved from a database), N-terminal, and C-terminal domain sequences was created using the BLAST package. Spider silk protein sequences of *L. sclopetarius* were identified by performing a homology search on the full-length reference database using BLASTp with an e-value cutoff of 1e-05. The identified sequences were then confirmed for their N-terminal and C-terminal domains.

[0189] Due to the large size and high reproducibility of spider silk proteins, identifying exon boundaries using assembly tools can be inaccurate. Therefore, manual examination was performed on the L. sclopetarius spider silk protein sequence sites and their surrounding regions (extending 5000-10000 bp at either end of the gene) defined by the automated MAKER gene model. The gene model was viewed using the Web Apollo (Lee, et al., Web Apollo: a web-based genomic annotation editing platform. GenomeBiol 14, R93 (2013)) genome browser, by entering transcript identifiers and identifying supporting data from batch RNAseq and PacBio Iso-seq experiments. The extended gene sequence was searched against custom N-terminal and C-terminal domain reference databases using BLASTx with an e-value cutoff of 1e-05. Signal peptides were scanned using SignalP (Almagro Armenteros, et al., SignalP 5.0 improves signal peptide predictions using deep neural networks. Nature Biotechnology 37, 420-423 (2019)) in an additional 100 bp region upstream of the identified N-terminal domain. Therefore, we assumed that the boundaries of spider silk protein genes were correctly defined. For sequences where the C-terminal domain could not be defined, the automated MAKER gene model remained unchanged. We also identified multiple spider silk protein genes that were merged into a single site due to high sequence similarity, making it difficult to assemble them into independent sites. After defining gene boundaries for each identified spider silk protein, the gene sequences were translated in all six reading frames, and repetitive motifs were manually checked using Unipro UGENE software (Okonechnikov, et al., Unipro UGENE: a unified bioinformatics toolkit. Bioinformatics 28, 1166-1167 (2012)) to identify any mis-annotations (deleted exons due to incorrect reading frames) in the current gene model. Based on the identified erroneous annotations, either new genes are added or existing gene models are replaced with corrected models. All corrected sequences are then confirmed by alignment to a reference genome using exonerate.GeneWise (Birney, et al., GeneWise and Genomewise. Genome Res 14, 988-995 (2004)) and ScipioKeller, et al., Scipio: Using protein sequences to determine the precise exon / intron structures of genes and their orthologs in closely related species. Bmc Bioinformatics 9, (2008)) were used to generate GFF files for the corrected gene model.

[0190] Functional assignment of proteins with hypothesized functions To predict the function of proteins annotated as "presumed proteins" (from genome annotation), gene family clusters among *L. sclopetarius*, *Trichonephila clavipes* (NCBI accession number PRJDB10126), *Trichonephila clavata* (PRJDB10007), *Nephila pilipes* (PRJDB10128), *Trichonephila inauratamadagascariensis* (PRJDB10127), *Argiope bruennichi* (PRJNA629526), ​​and *Araneus ventricosus* (PRJDB7092) were identified using default settings using Orthofinder (Emms and Kelly, OrthoFinder: solving fundamental biases in wholegenome comparisons dramatically improves orthogroup inference accuracy. Genome Biol 16, 157 (2015)) (version 2.5.2) at default settings. Protein sequences were downloaded from NCBI. Putative transcript isoforms were removed from the *L. sclopetarius* proteome dataset, retaining only the longest canonical sequences for analysis. Proteins for each species were processed using Orthofinder-Diamond to classify them into orthologs. Equivalent genes classified into the same ortholog were processed using an internal Python script to further assign functions to hypothetical proteins based on their known functions within the same ortholog. Annotation was performed at two levels: a) analyzing homologous clusters within *L. sclopetarius* belonging to orthologs; b) analyzing gene clusters from other spider species (*T. clavipes*, *T. clavata*, *N. pilipes*, *T. inaurata madagascariensis*, *A. bruennichi*, *A. ventricosus*) belonging to orthologs.

[0191] Batch RNA sequencing data analysis Raw sequence reads were aligned to the de novo genome using STAR (Dobin, et al., STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 (2013)) (version 2.7). Read counts were generated using featureCounts (Liao, and Smyth, Shi, featureCounts: an efficient general purpose program for assigning sequencereads to genomic features. Bioinformatics 30, 923-930 (2014)) from the Rsubread package (Liao et al., The R package Rsubread is easier, faster, cheaper and better for alignment and quantification of RNA sequencing reads. NucleicAcids Research 47, (2019)) (version 2.0.0). Only genes with counts greater than 20 were retained for subsequent analysis. Read counts were normalized using the DESeq2 package (Love, et al., Moderatedestimation of fold change and dispersion for RNA-seq data with DESeq2. GenomeBiology 15, (2014)) (version 1.38.3), which corrects for sequencing depth and library composition. All samples underwent pairwise Pearson correlation and hierarchical clustering.

[0192] PCA analysis was performed on samples from the three parts of the ampulla of Vater. The PC space was divided into three regions: PC1<0 and PC2>0 were assigned to the tail, PC1>0 and PC2>0 to the cyst, and PC2<0 to the duct. For the loadings of the first two PCs, the geometric distance from each gene to the origin was calculated. The normalized distance was calculated by removing the mean distance and dividing by the standard deviation of the PCs. Genes with a normalized distance greater than 2 were retained as significant genes and assigned to the three different tissues based on their PC1 and PC2 loadings (i.e., their respective regions). This set was defined as the ampulla of Vater genome. Differential gene expression analysis was performed between the three different regions using DESeq2. To be considered differentially expressed in a tissue, a fold change greater than 2 and an adjusted p-value less than 0.001 were required.

[0193] Single-cell RNA sequencing data analysis Reads were aligned to transcripts using CellRanger (version 7.0.1, https: / / support.10xgenomics.com / single-cell-gene-expression / software / pipelines / latest / what-is-cell-ranger). Initial QC analysis removed all cells with fewer than 300 expressed genes and / or fewer than 500 total transcripts. Only samples with more than 500 cells were retained for further analysis. All steps were performed in Seurat (Satija, et al., Spatialreconstruction of single-cell gene expression data. Nature Biotechnology 33,495-U206 (2015); Hao, et al., Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587 e3529 (2021)) (version 4.0.3, https: / / satijalab.org / seurat / ). Cells were normalized using SCT and integrated using canonical correlation analysis distance between samples.

[0194] Following the initial QC steps, 18,539 cells were obtained from 7 samples and clustered into 23 clusters. Within these clusters, a subset was identified whose marker genes overlapped with the ampulla of Vater, sac, and duct gene sets. By re-evaluating this specific subset and limiting the analyzed genes to the intersection of the top 2000 most variable genes from scRNAseq data and the ampulla of Vater gene set from bulk RNA analysis, 9,700 cells were obtained from the 7 samples and clustered into nine groups. The smallest cluster, containing only cells from a single sample, was removed for further analysis. Marker genes for each cluster were identified using a subset of 400 cells per category.

[0195] Spatial Transcriptomics Data Analysis and Deconvolution Eight slides were manually annotated morphologically as filages using the annotation tools in the Loupe browser (version 6.4.1, https: / / support.10xgenomics.com / spatial-gene-expression / software / visualization / latest / what-is-loupe-browser). Reads were aligned to reference genomes and annotations using Space Ranger (version 1.2.0, https: / / support.10xgenomics.com / spatial-gene-expression / software / pipelines / latest / what-is-space-ranger). Initial QC removed samples with fewer than 300 genes and fewer than 500 transcripts. All steps were performed in Seurat (Satija et al. 2015; Hao et al., 2021) (version 4.0.3). Cells were normalized using SCT and integrated using canonical correlation analysis distances between samples. Distances between samples were visualized using UMAP. Marker genes for different categories were identified using a subset of 400 cells per category.

[0196] For the analysis of ampulla of Vater glands, five well-annotated samples of glands were retained. Image files were imported into QuPath (Bankhead, et al., QuPath: Open source software for digital pathology imageanalysis. Sci Rep 7, 16878 (2017)) (version 0.4.3), and the brush tool was used to further divide the regions identified as ampulla of Vater glands into three distinct regions (A, B, and C) based on H&E staining and cell morphology. The mean eosin value, mean hematoxylin value, and perimeter value for each region were determined using QuPath's default parameters.

[0197] Pairwise Pearson correlation analysis was performed on the eosin, hematoxylin, and perimeter values ​​of region A. Based on the perimeter value, region A was divided into three categories: regions with a perimeter less than 500 pixels were classified as proximal, regions with a perimeter greater than 1000 pixels as distal, and the rest as mid-region. Spots on the spatial transcriptomics slices were assigned to the nearest regions overlapping with the region annotations in the QuPath analysis.

[0198] To determine the proportion of different cell types in each spot on the slide, we used CARD (Ma and Zhou, Spatially informed cell-type deconvolution for spatial transcriptomics. Nat Biotechnol 40, 1349-1359 (2022)). Only genes with an average log2 fold change >2 within at least one cluster in the single-cell analysis were retained for spatial spot deconvolution. Only genes counted at least 200 and appearing in at least 50 spots in the spatial transcriptomics data were retained for deconvolution analysis. The proportion of cell types in each spot within the ampulla of Vater was estimated. The average proportion of the five categories was calculated by averaging the proportions of cell types in all spots belonging to the five categories (regions C, B, and the proximal, mid, and distal subclasses of region A).

[0199] Proteomics Sample Preparation ampulla The spider was anesthetized and dissected as described above. The large ampulla of Vater gland was carefully pulled out using micro-forceps to grasp the duct. An incision was made in the glandular sac to allow the spinning solution to drain for 15 minutes. The gland was washed three times with PBS and then transferred to a low-binding 1.5 mL microtube (Axygen) containing 60 µL of 8 M urea (dissolved in 20 mM Tris-HCl, pH 8), vortexed, and sonicated at room temperature for 30 minutes in a water-water-wash sonicator (VWR). The sample was stored at -20°C until further use. Final protein sequencing was performed using four biological replicates.

[0200] Before 5 minutes of sonication in a water bath, samples were aliquoted and replenished with a 20% ACN / 20 mM Tris-HCl (pH 8) solution containing 0.2% ProteaseMAX (Promega) to obtain a 4M urea concentration. Protein reduction was performed by incubation with 8 mM DTT at 24°C and 550 rpm for 1 hour, followed by alkylation with 20 mM chloroacetamide (CAA) at room temperature in the dark for 1 hour. Digestion was initiated by adding 2 µg LysC (Wako, Japan) and incubating at 24°C for 2 hours, followed by adding 2 µg sequencing-grade modified trypsin (Promega) and incubating overnight (approximately 16 hours) at 37°C to complete digestion. After centrifugation, the supernatant was collected, and protein hydrolysis was terminated with 5% FA. Samples were washed on C18 Hypersep plates (Thermo Fisher Scientific) with a 40 µL bed volume and dried using a vacuum concentrator (Eppendorf).

[0201] Large ampulla gland fibers To collect the silk, each spider was first anesthetized with CO2 and gently restrained to prevent movement without harming it. Using a Zeiss Stemi 305 stereomicroscope and with the aid of tweezers, the large urn-shaped glandular silk was identified and pulled from the front spinneret (Foelix, Biology of Spiders: Oxford University Press. New York 330, (1996)). The silk was collected by winding it around a frame attached to a rotating reel until the spider refused to continue spinning. Sufficient silk was collected from multiple spiders for the experiment. After silk collection, the spiders were fed fruit flies and given water, but not used for the next two weeks.

[0202] The collected filaments were solubilized using three different methods. Approximately 450 µg of filament was used for each sample, with four replicates for each treatment. Filament samples were prepared using three different methods. In the first method, 100 µL of hexafluoroisopropanol (HFIP) was added and briefly vortexed to dissolve the filaments. The samples were then sonicated in a water-water-wash sonicator (VWR) at room temperature for 30 minutes. The HFIP was evaporated using a Centrivap concentration system (Labconco). The protein was resuspended in 60 µL of 8 M urea (dissolved in 20 mM Tris-HCl, pH 8) and stored at -20 °C until further use. In the second treatment, the filaments were dissolved in 60 µL of 8 M urea (dissolved in 20 mM Tris-HCl, pH 8), sonicated, and stored as described above. In the third method, the filaments were dissolved in 60 µL of 9 M LiBr (dissolved in 20 mM Tris-HCl, pH 8), sonicated, and stored as described above. The entire experiment used low-binding 1.5 mL microtubes (Axygen).

[0203] Aliquots (approximately 10 µg) of 30 µL sample were aliquoted for further preparation. For samples obtained using the HFIP and urea methods, protein reduction was performed by incubation at 37°C and 1200 rpm for 3 hours with 3 µL of 100 mM DTT, followed by alkylation by incubation at room temperature and in the dark for 30 minutes with 5 µL of 500 mM MCAA. Half of the samples were then digested with 19 µL of 50 mM Tris-HCl (pH 8.5) and 1 µg of LysC (Wako, Japan) at 24°C for 2 hours. After adding 56 µL of Tris-HCl, digestion was continued with 1 µg of sequencing-grade modified trypsin (Promega) and incubated overnight at 37°C (approximately 16 hours). Samples using the third method (LiBr) were prepared similarly, except that reduction was performed by incubation at 95°C and 12,500 rpm for 30 minutes with 2 µL of 500 mM DTT. Alkylation was performed using 5 µL of 500 mM CAA (as above), followed by digestion with LysC and trypsin as described above, except that the amount of trypsin used was 3 µg. Digestion of all samples was terminated with 6.5 µL of concentrated formic acid. Samples were purified on C18 Hypersep plates (Thermo Fisher Scientific) with a bed volume of 40 µL and dried using a vacuum concentrator (Eppendorf).

[0204] Layer-by-layer dissolution of large ampulla glandular fibers Large urn gland silk was collected by allowing each spider to fall freely from a wooden frame, and the extruded silk was wound onto the same frame. Silk from multiple individuals was collected into pre-weighed low-binding microtubes and weighed again to determine the weight of the collected silk. After forced silk extraction, the spiders were fed fruit flies and given water, and then not used for two weeks. The collected silk was divided into three groups and treated with the following solutions: 1) 50 mM Tris-HCl and 0.5 M NaCl containing 2 M urea, 2) 50 mM Tris-HCl and 0.5 M NaCl containing 4 M urea, and 3) 50 mM Tris-HCl and 0.5 M NaCl containing 8 M urea. The samples were sonicated at room temperature for 2–3 hours, followed by centrifugation at 17,000 g. The supernatant was collected and stored at -20°C for proteomics analysis.

[0205] Liquid chromatography-tandem mass spectrometry data acquisition The peptide was reconstituted in solvent A and loaded onto a 50 cm long EASY-Spray C18 column (Thermo Fisher Scientific) connected to a UltiMate 3000 nanofluidic UPLC system. Gradient elution was performed for 90 minutes: solvent B (98% acetonitrile, 0.1% FA) was increased from 4% to 26% over 90 minutes, from 26% to 95% over 5 minutes, and held at 95% solvent B for 5 minutes at a flow rate of 300 nL / min. Mass spectrometry data were acquired on a Q Exactive HF hybrid quadrupole orbital trap mass spectrometer (Thermo Fisher Scientific) with a mass range of m / z 375 to 1800, a resolution R = 120,000 (at m / z 200), and a target size of 5 × 10⁻⁶. 6 Ions were implanted for a maximum time of 100 ms, followed by data-dependent high-energy collisional dissociation (HCD) fragmentation of precursor ions in charge states 2+ to 7+. + The dynamic exclusion time was 45 seconds. Tandem mass spectrometry of the first 17 precursor ions was acquired at a resolution R=30,000, with a target size of 2×10⁻⁶. 5 Ions, maximum implantation time 54 ms, quadrupole isolation width set to 1.4 Th, normalized collision energy set to 28%.

[0206] Proteomics data analysis The acquired raw data files were converted to Mascot Universal File (mgf) format using the internally developed tool Raw2MGF (version 2.1.3), and a protein database of 22,856 protein entries was searched using Mascot Daemon version 2.5.1 (Matrix Science Ltd., UK). For complete trypsin digestion, a maximum of two missed cleavage sites were allowed, with mass tolerances of 10 ppm for precursor ions and 0.02 Da for fragment ions. Cysteine ​​carbamoyl methylation was set as a fixed modification, while methionine oxidation and asparagine and glutamine deamidation were set as dynamic modifications. The search results were imported into Scaffold version 4.11 (Proteome Software Inc.) to calculate the contribution of each protein to the total spectrum (percentage of the total spectrum). For samples where the percentage of the total spectrum was not reported, the percentage value for that protein was set to 0.

[0207] To identify proteins in the filaments, we used MS / MS data from the filament fibers and glands. To determine their presence in the glands, the average percentage of each protein in the total spectrum was calculated by averaging three gland samples. All proteins reported in at least one sample were considered present in the glands. Similarly, the average percentage of each protein in the total spectrum was calculated by averaging nine samples (i.e., three biological replicates for three different detergents). For a protein to be considered present in the ampulla of Vater gland filaments, four criteria must be met: 1) it must be present in the glands; 2) it must be found in at least two detergents; 3) its amino acid sequence must have a predicted signal peptide at the 5' end; and 4) its average percentage in the total spectrum must be greater than 0. After applying these criteria, 17 proteins were retained. Unpaired t-tests were used to determine the statistical differences in the percentage of proteins in the total spectrum among different urea concentrations, focusing on these 17 proteins.

[0208] Tissue processing for histological analysis Spiders were anesthetized with CO2 and then dissected at the peduncle using 154 mM sodium chloride solution (Fresenius Kabi AG, Germany) on ice-cooled paraffin plates. Dissection was performed using a Leica M60 stereomicroscope equipped with a Leica IC80 HD camera. The entire metastomium was fixed at 4°C for 24 hours in 67 mM phosphate buffer (pH 7.2) containing 2.5% glutaraldehyde, followed by rinsing in 67 mM phosphate buffer. The tissue was dehydrated in a gradient of ethanol (50%, 70%, 90%, and 100%, 30 min each), then infiltrated and embedded in water-soluble ethylene glycol methacrylate (Leica Historesin). Sections were obtained to a thickness of 2 µm using a Leica RM 2165 microtome equipped with a glass scalpel. Sections were stained with hematoxylin and eosin and mounted with Agar100 resin. Evaluation was performed using a Nikon Microphot-FXA (Tekno Optik AB) microscope equipped with a Nikon FX-35DX camera. Image capture and editing were performed using Eclipse Net version 1.20.0 software.

[0209] Tensile test of large-ampullary gland fiber Large urn-shaped glandular fibers were wound and mounted on a cardboard frame with a 1×1 cm square window (gauge length 1 cm). The diameter was measured using an optical microscope via a Nikon Eclipse Ts2R-FL inverted microscope. The diameter was measured at five locations along the fiber leading edge of the tensile test and the average value was taken. Tensile tests were performed using a 5943-Instron instrument (USA) equipped with a 5 N load cell. The strain rate used was 6 mm / min. Assuming a circular cross-section and using the average diameter of each fiber, the load-displacement curve was converted into an engineering stress-strain curve.

[0210] The above embodiments should be understood as several illustrative examples of the present invention. Those skilled in the art will understand that various modifications, combinations, and changes can be made to the embodiments without departing from the scope of the present invention.

Claims

1. A recombinant spider silk protein comprising an N-terminal (NT) domain, a repeat region (REP) domain, and a C-terminal (CT) domain, wherein... The REP domain contains at least 100 repeating amino acid residues derived from the silk protein of the hard-shelled spider; and The NT domain is derived from the NT domain of spider silk protein, which is different from the spider silk protein from which the REP domain originates, and / or the CT domain is derived from the CT domain of spider silk protein, which is different from the spider silk protein from which the REP domain originates.

2. The recombinant spider silk protein of claim 1, wherein in the recombinant spider silk protein, the REP domain is disposed between the NT domain and the CT domain.

3. The recombinant spider silk protein of claim 2, wherein the recombinant spider silk protein has the general formula (X)-NT-(L1)-REP-(L2)-CT-(Y), wherein X represents an optional N-terminal tag, Y represents an optional C-terminal tag, L1 represents an optional first connector, and L2 represents an optional second connector.

4. The recombinant spider silk protein according to any one of claims 1 to 3, wherein the REP domain comprises, preferably at least 125 amino acid residues, more preferably at least 130 amino acid residues, more preferably at least 150 amino acid residues, for example, at least 200 amino acid residues, at least 250 amino acid residues, at least 300 amino acid residues, at least 350 amino acid residues, and most preferably at least 400 amino acid residues, for example, at least 450 amino acid residues, at least 500 amino acid residues, at least 550 amino acid residues, or at least 600 amino acid residues, of the repeating domain of the spider silk protein derived from L. sclopetarius.

5. The recombinant spider silk protein according to any one of claims 1 to 3, wherein the REP domain comprises, preferably at least 100 consecutive amino acid residues, more preferably at least 125 consecutive amino acid residues, more preferably at least 130 consecutive amino acid residues, for example, at least 150 consecutive amino acid residues, at least 200 consecutive amino acid residues, at least 250 consecutive amino acid residues, at least 300 consecutive amino acid residues, at least 350 consecutive amino acid residues, and most preferably at least 400 consecutive amino acid residues, for example, at least 450 consecutive amino acid residues, at least 500 consecutive amino acid residues, at least 550 consecutive amino acid residues, or at least 600 consecutive amino acid residues, of the repeating domain of the spider silk protein derived from L. sclopetarius.

6. The recombinant spider silk protein according to any one of claims 1 to 5, wherein the REP domain comprises at least 100 amino acid residues from the repeating region of the large ampullae gland silk protein (MaSp) of L. sclopetarius.

7. The recombinant spider silk protein of claim 6, wherein the REP domain is derived from the MaSp of L. sclopetarius selected from the group consisting of: MaSp1a as defined in SEQ ID NO: 38, MaSp1b as defined in SEQ ID NO: 41, MaSp1c as defined in SEQ ID NO: 44, MaSp2a as defined in SEQ ID NO: 47, MaSp2b as defined in SEQ ID NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3a as defined in SEQ ID NO: 65, MaSp3b as defined in SEQ ID NO: 68, and MaSp4 as defined in SEQ ID NO: 71, preferably from the group consisting of L. sclopetarius: MaSp1a as defined in SEQ ID NO: 38, MaSp2a as defined in SEQ ID NO: 47, and MaSp4 as defined in SEQ ID NO:

68. MaSp2b as defined in NO: 50, MaSp2c as defined in SEQ ID NO: 53, MaSp2d as defined in SEQ ID NO: 56, MaSp2e as defined in SEQ ID NO: 59, MaSp2f as defined in SEQ ID NO: 62, MaSp3a as defined in SEQ ID NO: 65, and MaSp4 as defined in SEQ ID NO: 71, more preferably MaSp from L. sclopetarius of the group consisting of MaSp2c as defined in SEQ ID NO: 53 and MaSp4 as defined in SEQ ID NO:

71.

8. The recombinant spider silk protein of claim 7, wherein the REP domain comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 109-114, SEQ ID NO: 148-151, SEQ ID NO: 168, more preferably selected from the group consisting of: SEQ ID NO: 109-114, SEQ ID NO: 148-151, and more preferably selected from the group consisting of: SEQ ID NO: 109-114.

9. The recombinant spider silk protein according to any one of claims 1 to 8, wherein the NT domain is derived from the NT domain of a spider silk protein from a different spider species than the REP domain, and / or the CT domain is derived from the CT domain of a spider silk protein from a different spider species than the REP domain.

10. The recombinant spider silk protein according to any one of claims 1 to 9, wherein the NT domain is derived from the NT domain of the Australian spider MaSp1.

11. The recombinant spider silk protein of claim 10, wherein the NT domain comprises, preferably, SEQ ID NO:

115.

12. The recombinant spider silk protein according to any one of claims 1 to 9, wherein the NT domain is derived from the NT domain of silk protein of L. sclopetarius, preferably from the NT domain of L. sclopetarius MaSp2f.

13. The recombinant spider silk protein of claim 12, wherein the NT domain comprises, preferably, SEQ ID NO:

61.

14. The recombinant spider silk protein according to any one of claims 1 to 13, wherein the CT domain is derived from the CT domain of the spider *Misp. lumbricoides*.

15. The recombinant spider silk protein of claim 14, wherein the CT domain comprises, preferably, SEQ ID NO:

116.

16. The recombinant spider silk protein according to any one of claims 1 to 10, wherein the CT domain is derived from the CT domain of silk protein of L. sclopetarius, preferably from the CT domain of L. sclopetarius MiSpc.

17. The recombinant spider silk protein of claim 16, wherein the CT domain comprises, preferably, SEQ ID NO:

81.

18. The recombinant spider silk protein according to any one of claims 1 to 11, wherein the recombinant spider silk protein comprises, preferably, an amino acid sequence selected from the group consisting of: SEQ ID NO: 103-108, SEQ ID NO: 117-134, SEQ ID NO: 152-167, SEQ ID NO: 169-172, more preferably selected from the group consisting of: SEQ ID NO: 103-108, SEQ ID NO: 117-134, SEQ ID NO: 152-167, and more preferably selected from the group consisting of: SEQ ID NO: 103-108, SEQ ID NO: 117-134.

19. The recombinant spider silk protein according to any one of claims 1 to 18, wherein the molecular weight of the recombinant spider silk protein is at least 30 kDa, preferably at least 35 kDa, and more preferably at least 50 kDa.

20. A silk fiber made from recombinant spider silk protein according to any one of claims 1 to 19.

21. The silk fiber of claim 20, further comprising a spider silk constituent element (SpiCE) of L. sclopetarius.

22. A synthetic material comprising the silk fibers as described in claim 20 or 21.

23. A nucleic acid molecule encoding the recombinant spider silk protein according to any one of claims 1 to 19.

24. An expression vector comprising the nucleic acid molecule according to claim 23.

25. A host cell comprising the expression vector according to claim 24.

26. A method for producing silk fibers, comprising: The spinning solution containing the recombinant spider silk protein according to any one of claims 1 to 19 is extruded into an aqueous buffer solution with an acidic pH to induce the recombinant spider silk protein to polymerize into silk fibers; as well as The silk fibers are separated from the aqueous buffer solution.

Citation Information

Patent Citations

  • Engineered Spider Silk Proteins And Uses Thereof

    US20190248847A1

  • Engineered spider silk proteins and uses thereof

    WO2018002216A1

  • Silk nucleotides and proteins and methods of use

    WO2020092769A2

  • Recombinant spider silk proteins

    WO2023167628A1