Expanded artificial scaffold protein, artificial fibrosome, preparation method and application

By constructing an extended artificial scaffold protein through a covalent self-assembly system, the problem of low cohesin content in heterologous hosts was solved, achieving efficient enzyme loading and cellulose degradation, and making it suitable for efficient cellulase assembly in various expression systems.

CN121652289APending Publication Date: 2026-03-13ZHEJIANG FORESTRY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-13

Smart Images

  • Figure CN121652289A_ABST
    Figure CN121652289A_ABST
Patent Text Reader

Abstract

The invention discloses an expanded artificial scaffold protein, an artificial fibrosome and a preparation method and application thereof, and belongs to the technical field of enzyme engineering. According to the invention, the scaffold proteins containing different numbers of adhesion modules are designed, and the two scaffold proteins are modularly assembled by using a covalent linkage system, so that the oversized artificial scaffold protein of which the number of adhesion modules can reach 19 is successfully constructed, and the scale of the oversized artificial scaffold protein exceeds that of the known natural scaffold protein. The scaffold protein can efficiently carry various catalytic subunits through specific interaction of an adhesion module and a docking module to form an artificial fibrosome. Experiments show that the artificial cellulosome can efficiently degrade cellulose substrates such as crystalline cellulose and the like especially when the scaffold protein contains an extremely high amount of cohesin. The invention provides a novel efficient enzyme complex tool for efficiently degrading lignocellulose biomass, and has a wide application prospect in the fields of biofuel, bio-based chemical production and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enzyme engineering technology, and in particular to an extended artificial scaffold protein, an artificial cellulosome, its preparation method, and its application. Background Technology

[0002] Cellulose is the most abundant renewable biomass resource on Earth, and its efficient degradation and transformation are crucial for sustainable development. In nature, anaerobic bacteria (such as *Ruminiclostridium cellulolyticum* and *Clostridium thermocellum*) efficiently degrade cellulose by producing a supramolecular multi-enzyme complex called a "cellulose fibrous body." The cellulose fibrous body consists of a non-catalytic scaffoldin and multiple catalytic subunits. The scaffoldin contains multiple cohesin modules and a carbohydrate-binding module (CBM), while each catalytic subunit contains a docking module (dockerin). Through high-affinity cohesin-dockerin interactions, multiple enzymes are precisely assembled on the scaffoldin and anchored to the cellulose substrate via the CBM, resulting in a powerful "proximity effect" and synergistic effect that greatly improves degradation efficiency.

[0003] Due to the low yield and complex composition of natural cellulose bodies, researchers are dedicated to constructing "artificial cellulose bodies" or "designed cellulose bodies" in order to achieve high yields and superior performance in heterologous hosts. However, a major technical bottleneck in this field is that natural scaffold proteins have large molecular weights and contain highly repetitive cohesin sequences, making their expression in heterologous hosts difficult and resulting in low yields. Existing techniques can typically construct scaffold proteins with a limited number of cohesins (generally ≤8), restricting further improvements in their enzyme loading and catalytic activity. Although some studies have attempted to increase the number of cohesins through strategies such as constructing linker scaffold proteins or splicing intimatic peptides, these methods are often cumbersome, inefficient, or lack versatility.

[0004] Therefore, there is an urgent need in the field for a universal, efficient and flexible method to construct large artificial scaffold proteins with a greater number of cohesins and customizable structures, and then assemble artificial fibrosomes with performance exceeding that of natural systems. Summary of the Invention

[0005] The purpose of this invention is to provide an extended artificial scaffold protein, artificial fibrous bodies, their preparation methods, and applications to solve the problems existing in the prior art. Based on the SpyTag / SpyCatcher and SnoopTag / SnoopCatcher covalent self-assembly system, a super-large artificial scaffold protein with up to 19 cohesins was successfully constructed. This scaffold protein can efficiently carry multiple catalytic subunits (such as cellulase, hemicellulase, etc.) through specific cohesin-dockerin interactions, forming artificial fibrous bodies with well-defined structures and high enzyme loading. These artificial fibrous bodies exhibit efficient degradation of lignocellulose biomass.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] This invention provides an extended artificial scaffold protein, which is obtained by self-assembling a scaffold protein containing at least 5 adhesion modules and one carbohydrate-binding module through a covalent self-assembly system. The covalent self-assembly system is one or two of SpyTag, SpyCatcher, SnoopTag, and SnoopCatcher, and the adhesion modules are derived from Clostridium fibrinolyticum.

[0008] Preferably, the number of adhesive modules is 5, 7, 11 or 19.

[0009] The present invention also provides a nucleic acid molecule encoding the extended artificial scaffold protein described above.

[0010] The present invention also provides a biomaterial comprising the aforementioned nucleic acid molecules, the biomaterial comprising a recombinant vector and recombinant bacteria. Furthermore, the host cell of the recombinant bacteria is preferably a prokaryotic cell or a eukaryotic cell, more preferably *Escherichia coli* or yeast.

[0011] The present invention also provides an artificial fibrous body, comprising the extended artificial scaffold protein and a catalytic subunit that binds to the extended artificial scaffold protein through an adhesion module-docking module interaction, wherein the catalytic subunit is derived from Clostridium fibrinolyticum.

[0012] Preferably, the catalytic subunit includes one or more of glycoside hydrolases, carbohydrate esterases, and polysaccharide lyases. More specifically, the glycoside hydrolases include enzymes from the GH5, GH8, GH9, GH44, and GH48 families.

[0013] Preferably, the catalytic subunit comprises a core subunit and an auxiliary subunit, wherein the molar ratio of the core subunit to the auxiliary subunit is 4:1.

[0014] Preferably, the core subunits include Cel9E, Cel44O, and Cel48F, and the auxiliary subunits include Cel5A, Cel8C, and Cel9G.

[0015] The present invention also provides a method for preparing the aforementioned artificial fibrous bodies, comprising the following steps:

[0016] The adhesion module and carbohydrate-binding module of the scaffold protein, as well as the catalytic subunit, are expressed respectively;

[0017] The artificial fibrous body is obtained by mixing and self-assembling the adhesion module, the carbohydrate binding module, and the catalytic subunit using one or two of SpyTag, SpyCatcher, SnoopTag, and SnoopCatcher.

[0018] The present invention also provides the use of the extended artificial scaffold protein or the artificial fibrous body in any of the following:

[0019] (1) Application in the degradation of cellulose or polysaccharides;

[0020] (2) Application in the degradation of biomaterials containing cellulose or polysaccharides.

[0021] The present invention discloses the following technical effects:

[0022] (1) Modularity and flexibility: By utilizing the two pairs of orthogonal covalent linkage systems of Spy / Snoop, scaffold proteins of different sizes and structures can be designed and assembled flexibly like building blocks, and the amount of cohesin can be precisely controlled.

[0023] (2) Ultra-high enzyme loading: A scaffold protein containing 19 cohesins was successfully constructed, which is the highest reported in the world to date. It can carry enzyme molecules far exceeding those in the natural system, greatly increasing the local enzyme concentration.

[0024] (3) Highly efficient synergistic catalysis: Experiments have shown that the artificial fibrosomes assembled from this scaffold protein have a significantly higher degradation efficiency for crystalline cellulose than the free enzyme mixture, especially the fibrosomes containing 19 cohesins, which have an activity up to 1.94 times that of the free enzyme.

[0025] (4) Technical universality: This method does not depend on a specific host and can be implemented in multiple expression systems such as E. coli, laying the foundation for large-scale production. Attached Figure Description

[0026] Figure 1The results of the interaction analysis between cohesin and dockerin from R. cellulolyticum are as follows: (A) Phylogenetic analysis of 63 Dockerin from R. cellulolyticum, with seven major groups highlighted in different colors; (B) Phylogenetic analysis of 8 Cohesin from the scaffold protein CipC; (C) The interaction between Cohesin and GST-Dockerin (from R. cellulolyticum enzymes, as marked in red in Figure A) was analyzed by ELISA; (D) The interaction between Cohesin and GST-Dockerin was confirmed by isothermal thermodynamic titration (ITC).

[0027] Figure 2 The design, assembly, and optimization results of the extended scaffold protein (19Coh-CBM) based on the Spy / Snoop system are shown in the diagram. (A) A schematic diagram of the covalent linkages of the extended scaffold protein 19-Coh-CBM is shown, which are mediated by the interaction between SpyTag / SpyCatcher and SnoopTag / SnoopCatcher, and are achieved through the three sub-assemblies shown in the figure. (BC) Reactions were carried out through the interaction of SpyTag / SpyCatcher (B) and SnoopTag / SnoopCatcher (C), thereby forming isopeptide bonds between the sub-assemblies in a reaction ratio of 1:2 to 6:1. (D) The assembly product of 19Coh-CBM was analyzed by SDS-PAGE using these three sub-assemblies. (E) 19Coh-CBM was further purified by size exclusion chromatography (SEC) and analyzed by SDS-PAGE. The scaffold protein assembly referred to as 11Coh-CBM was used as a control marker.

[0028] Figure 3 To assemble scaffold proteins with different amounts of cohesin using the SpyTag / SpyCatcher system; (A) Schematic diagram of scaffold protein assembly mediated by SpyTag / SpyCatcher interaction; (BC) Analysis of the assembly products of various scaffold proteins by SDS-PAGE (B) and by SEC (C); (D) Evaluation of the purity of the scaffold protein constructs after SEC results by SDS-PAGE.

[0029] Figure 4The following are the results of expression, purification, and synergistic analysis of six key catalytic subunits: (A) Schematic diagram of the structures of the six selected key subunits; (B) SDS-PAGE analysis of purified GH44O, GH5A, GH9E, GH9G, GH8C, and GH48F; (CD) Measurement of reducing sugar yield of each catalytic subunit on crystalline cellulose (C) and carboxymethyl cellulose (D) substrates; (E) Comparison of the activity of any one enzyme (missing enzymes are marked on the x-axis) with an equimolar mixture of the five enzymes on crystalline cellulose, and the activity of a mixture of all six enzymes. The properties were compared; based on their contribution rate, Cel9E, Cel44O, and Cel48F were defined as core subunits of cellulose, while Cel9G, Cel8C, and Cel5A were defined as auxiliary subunits; (F) The activity of mixtures of different proportions of core and auxiliary groups on crystalline cellulose was measured; the individual activity of each group was used as a control; each reaction was performed in triplicate; data are presented as mean ± standard deviation, where (* indicates P < 0.05, ** indicates P < 0.01, *** indicates P < 0.001, **** indicates P < 0.0001, ANOVA);

[0030] Figure 5 The assembly and characterization results of the artificial fibrous bodies are as follows: (A) The assembly of microfibrils composed of Cel5A and 3Coh-CBM in different ratios was analyzed by non-denaturing polyacrylamide electrophoresis; (BC) The assembly products were purified by SEC analysis (B), and the products of the two SEC peaks were identified by SDS-PAGE (C); (D) The particle size of the assembly products was further analyzed by DLS; (E) The assembly of the scaffold protein 11Coh-CBM with the six catalytic subunits was analyzed by non-denaturing polyacrylamide electrophoresis; (F) The interaction between 11Coh-CBM and crystalline cellulose was analyzed by pull-down assay; (G) Transmission electron microscopy (TEM) images of the free enzyme and the artificial fibrous bodies (3Coh-CBM complex and 11Coh-CBM complex).

[0031] Figure 6 This is a schematic diagram of various scaffold protein and enzyme combinations used in this invention; scaffold proteins (lacking or containing CBM domains) with different numbers of adhesion proteins are constructed using different subunit combinations (as shown); these scaffold proteins are assembled with six catalytic subunits (Cel5A, Cel8C, Cel9G, Cel48F, Cel9E, and Cel44O) to form the desired artificial fibrous bodies;

[0032] Figure 7To compare the degradation activities of different artificial cellulose bodies on crystalline cellulose, the amount of reducing sugars produced by various artificial cellulose bodies on the substrate crystalline cellulose was compared. The artificial cellulose bodies contained an equimolar mixture (A) with a total catalytic subunit concentration of 0.5 μM, or a core subunit to accessory subunit ratio of 4:1, and contained (B) or (C) free β-glucosidase (0.01 μM). (mg); Free enzyme was used as a control; The degree of cellulose degradation co-activity was defined as the ratio of sugar produced by fibrosome scaffold proteins to sugar produced by free enzyme; (D) The amount of reducing sugar produced by fibrosomes assembled from various CBM-deficient scaffold proteins containing a total concentration of 0.5 μmol catalytic subunits (core enzyme to coenzyme ratio of 4:1); Each reaction was performed in triplicate; Data are presented as mean ± standard deviation, where (* indicates P < 0.05, ** indicates P < 0.01, *** indicates P < 0.001, **** indicates P < 0.0001, ANOVA); (E) Comparison of transmission electron microscopy images showing the substrate morphology remaining after digestion by free enzyme, 3Coh-CBM, or 11Coh-CBM. Detailed Implementation

[0033] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0034] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.

[0035] The primer sequences involved in the following examples are shown in Tables 1-3 below.

[0036] Table 1

[0037]

[0038] Table 2

[0039]

[0040] Table 3

[0041]

[0042] The following examples involve protein sequences:

[0043] SpT-CBM-1Coh (433 aa, SEQ ID NO.1; single underscore sequence is SpT; double underscore sequence is CBM; dashed line sequence is Coh):

[0044] M VPTIVMVDAYKRYKGSGESGGTGVVS VQFNNGSSPASSNSIYARFKVTNTSGSPINLADLKLRYYYT QDADKPLTFWCDHAGYMSGSNYIDATSKVTGSFKAVSPAVTNADHYLEVALNSDAGSLPAGGSIEIQTRFARNDWS NFDQSNDWSYTAAGSYMDWQKISAFWGGTLAYGSTPDGGNPPPQDPTINPTSISAKAGSFADTKITLTPNGTFNG ISELQSSQYTKGTNEVTLLASYLNTLPENTTKTLTFDFGVGTKNPKLTI TVLPKDIPGDS LKVTVGTANGKPGDTV TVPVTFADVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVF ANIKFKLKSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TIDPGTQPTKEHHHHHH*.

[0045] SpT-3Coh-CBM-SpT (754 aa, SEQ ID NO.2; single underscore sequence is SpT; double underscore sequence is CBM; dashed line sequence is Coh):

[0046] MGWSHPQFEK VPTIVMVDAYKRYK GSGESGSG LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNF YLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTF KDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSASGVVS VQFNNGSSPASSNSIYARFKVTNTSGSPINLADLK LRYYYTQDADKPLTFWCDHAGYMSGSNYIDATSKVTGSFKAVSPAVTNADHYLEVALNSDAGSLPAGGSIEIQTRF ARNDWSNFDQSNDWSYTAAGSYMDWQKISAFVGGTLAYGSTPDGGNPPPPQDPTINPTSISAKAGSFADTKITLTPN GNTFNGISELQSSQYTKGTNEVTLLASYLNTLPENTTKTLTFDGVGTKNPKLTI TVLPKDIPGDGS LKVTVGTAN GKPGDTVTVPVTFADVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDEL ITADGVFANIKFKLKSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSKL LKVTVGTANGKP GDTVTVPVTFADVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITA DGVFANIKFKLKSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV GSGESGSG VPTIVMVDAYKRYK LEHHHHHH*.

[0047] SpT-Coh (176 aa, SEQ ID NO.3; single underline sequence is SpT; dashed line sequence is Coh)

[0048] MGWSHPQFEK VPTIVMVDAYKRYK GSGESGSG LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNF YLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTF KDGGAFGDGTMSKIASVTKTNGSV LEHHHHHH*.

[0049] SpT-2Coh (325 aa, SEQ ID NO.4; single underline sequence is SpT; dashed line sequence is Coh):

[0050] MGWSHPQFEK VPTIVMVDAYKRYK GSGESGSG LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNF YLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTF KDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSAS LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNFYLG YDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTFKDG GAFGDGTMSKIASVTKTNGSV LEHHHHHH*.

[0051] SpT-4Coh (623 aa, SEQ ID NO.5; single underline sequence is SpT; dashed line sequence is Coh)

[0052] MGWSHPQFEK VPTIVMVDAYKRYK GSGESGSG LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNF YLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTF KDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSAS LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNFYLG YDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTFKDG GAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSGS LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNFYLGYDA SLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTFKDGGAF GDGTMSKIASVTKTNGSV TVLPKDIPGDSKL LKVTVGTANGKPGDTVTVPVTFADVAKMKNVGTCNFYLGYDASLL EVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAKTTTPVTFKDGGAFGDG TMSKIASVTKTNGSV LEHHHHHH*.

[0053] SpC-4Coh-SnT (736 aa, SEQ ID NO.6; the bold underlined sequence is SpC; the dashed sequence is Coh; the bold dashed sequence is SnT):

[0054] MGGS VTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEGRELAGATMELRDSSGKTISTWISDGHVK DFYLYPGKYTFVETAAPDGYEVATAITFTVNEQGQVTVNGEATKGDAHT GSGESGSG LKVTVGTANGKPGDTTVP VTFADVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANI KFKLKSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSAS LKVTVGTANGKPGDTVTVPVTF ADVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFK LKSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSGS LKVTVGTANGKPGDTVTVPVTFADV AKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKS VTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSKL LKVTVGTANGKPGDTVTVPVTFADVAKM KNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTA KTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV GSGESGSG KLGYYIEFYKVEK LEHHHHHH*.

[0055] SnC-4Coh (707 aa, SEQ ID NO.7; single underline sequence is SnC; dashed line sequence is Coh):

[0056] MGGS KPLRGAVFSLQKQHPDYPDIYGAIDQNGTYQNVRTGEDGKLTFKNLSDGKYRLFENSEPAGYKP VQNKPIVAFQIVNGEVRDVTSIVPQDIPATYEFTNGKHYITNEPIPPKGSGESGSG LKVTVGTANGKPGDTVTVPV TFADVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIK FKLKSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSAS LKVTVGTANGKPGDTVTVPVTFA DVAKMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKL KSVTAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSGS LKVTVGTANGKPGDTVTVPVTFADVA KMKNVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSV TAKTTTPVTFKDGGAFGDGTMSKIASVTKTNGSV TVLPKDIPGDSKL LKVTVGTANGKPGDTVTVPVTFADVAKMK NVGTCNFYLGYDASLLEVVSVDAGPIVKNAAVNFSSSASNGTISFLFLDNTITDELITADGVFANIKFKLKSVTAK TTTPVTFKDGGAFGDGTMSKIASVTKTNGSV *

[0057] Cel48F (695 aa, SEQ ID NO.8; single underscore sequence is GH48; double underscore sequence is Doc):

[0058] MG YQDRFESMYSKIKDPANGYFSEQGIPYHSIETLMVEAPDYGHVTTSEAMSYYMWLEAMHGRFSGDF TGFDKSWSVTEQYLIPTEKDQPNTSMSRYDANKPATYAPEFQDPSKYPSPLDTSQPVGRDPINSQLTSAYGTSMLY GMHWILDVDNWYGFGARADGTSKPSYINTFQRGEQESTWETIPQPCWDEHKFGGQYGFLDLFTKDTGTPAKQFKYT NAPDADARAVQATYWADEQWAKEQGKSVSTSVGKATKMGDYLRYSFFDKYFRKIGQPSQAGTGYDAAHYLLSWYYAW GGGIDSTWSWIIGSSHNHHFGYQNPFAAWVLSTDANFKPKSSNGASDWAKSLDRQLEFYQWLQSAEGAIAGGATNSW NGRYEAVPSGTSTFYGMGYVENPVYADPGSNTWFGMQVWSMQRVAELYYKTGDARAKKLLDKWAKWINGEIKFNAD GTFQIPSTIDWEGQPDTWNPTQGYTGNANLHVKVVNYGTDLGCASSLANTLTYYAAKSGDETSRQNAQKLLDAMWN NYSDSKGISTVEQRGDYHRFLDQEVFVPAGWTGKMPNGDVIKSGVKFIDIRSKYKQDPEWQTMVAALQAGQVPTQR LHRFWAQSEFAVANGVYAILF PDQGPEKL LGDVNGDETVDAIDLAILKKYLLNSSTTINTANADAMNSDNAIDAIDY ALLKKALL GTILEHHHHHH*.

[0059] Cel9G (723 aa; single underscore sequence is GH9; wavy line sequence is CBM; double underscore sequence is Doc):

[0060] MGSSHHHHHHSSGLVPRGSHMASAGTYNYG EALQKSIMFYEFQRSGDLPADKRDNWRDDSGMKDGSDV GVDLTGGWYDAGDHVKFNLPMSYTSAMLAWSLYEDKDAYDKSGQTKYIMDGIKWANDYFIKCNPTPGVYYYQVGDG GKDHSWWGPAEVMQMERPSFKVDASKPGSAVCASTAASLASAAVVFKSSDPTYAEKCISHAKNLFDDMADKAKSDAG YTAASGYYSSSSFYDDLSWAAVWLYLATNDSTYLDKAESYVPNWGKEQQTDIIAYKWGQCWDDVHYGAELLLAKLT NKQLYKDSIEMNLDFWTTGVNGTRVSYTPKGLAWLFQWGSLRHATTQAFLAGVYAEWEGCTPSKVSVYKDFLKSQI DYALGSTGRSFVVGYGVNPPQHPHRTAHGSWTDQMTSPTYHRHTIYGALVGGPDNADGYTDEINNYVNNEIACDY NAGFTGALAKMYKHS GGDPIPNFKAIEKITNDEVII KAGLNSTGPNYTEIKAVVYNQTGWPARVTDKISFKYFMDL SEIVAAGIDPLSLVTSSNYSEGKNTKVSGVLPWDVSNNVYYVNVD LTGENIYPGGQSACRREVQFRIAAPQGTTYWNPKNDFSYDGLPTTTSTVNTVTNIPVYDNGVKVFGNEPAGGSENPDPEILGSGESGSG LGDVNGDETVDAIDLAILK KYLLNSSTTINTANADMNSDNAIDAIDYALLKKALL GTILEHHHHHH*.

[0061] Cel9E (914 aa; wavy line sequence CBM; single underscore sequence is GH9; double underscore sequence is Doc):

[0062] MGSSHHHHHHSSGLVPRGSHMASMTGGQQMGRGSEFELRRQACGRVGA GDLIRNHTFDNVRVGLPWHVV ESYPAKASFEITSDGKYKITAQKIGEAGKGERWDIQFRHRGLALQQGHTYTVKFTVTASRACKIYPKIGDQGDPYD EYWNMNQQWNFLELQANTPKTVTQTFTQTKGDKKNVEFAFHLAPDKTTSEAQNPASFQPITYTFDEIYIQDPQFAGYTEDPPEPTNVVRLNQVGFYPNADKIATVATSSTTPINWQLVNSTGAAVLTGKSTVKGADRASGDNVHIIDFSSYTTPGTDYKIVTDVSVTKAGDNESMKFNIGDDLFT QMKYDSMKYFYHNRSAIPIQMPYCDQSQWARPAGHTTDILAPD PTKDYKANYTLDVTGGWYDAGDHGKYVVNGGIATWTVMNAYERALHMGGDTSVAPFKDGSLNIPESGNGYPDILDE ARYNMKTLLNMQVPAGNELAGMAHHKAHDERWTALAVRPDQDTMKRWLQPPSTAATLNLAAIAAQSSRLWKQFDSA FATKCLTAAETAWDAAVAHPEIYATMEQGAGGGAYGDNYVLDDFYWAACELYATTGSDKYLNYIKSSKHYLEMPTE LTGGENTGITGAFDWGCTAGMGTITLALVPTKLPAADVATAKANIQAAADKFISISKAQGYGVPLEEKVISSPFDA SVVKGFQWGSNSFVINEAIVMSYAYEFSDVNGTKNNKYINGALTAMDYLLGRNPNIQSYITGYGDNPLENPHHRFW AYQADNTFPKPPPGCLSGGPNSGLQDPWVKGSGWQPGERPAEKCFMDNIESWSTNEITINWNAPLVWI SAYLDEKGPEIGGSVTPPTNGSGESGSG LGDVNGDETVDAIDLAILKKYLLNSSTTINTANADMNSDNAIDAIDYALLKKALL GTILEHHHHHH*.

[0063] Cel5A (483 aa; single underscore sequence is GH5; double underscore sequence is Doc):

[0064] MGSSHHHHHHSSGLVPRGSHMASDASLIPNLQIPQKNIPNNDGMNFVKGLLRGWNLGNTFDAFNGTNITNELD YETSWSGIKTTKQMIDAIKQKGFNTVRIPVSWHPHVSGSDYKISDVWMNRVQEVVNYCIDNKMYVILNTHH DVDKVKGYFPSSQYMASSKKYITSVWAQIAARFANYDEHLIFEGMNEPRLVGHANEWWPELTNSDVVDSINCINQL NQDFVNTVRATGGKNASRYLMCPGYVASPDGATNDYFRMPNDISGNNNKIIVSVHAYCPWNFAGLAMADGGTNAWN INDSKDQSEVTWFMDNIYNKYTSRGIPVIIGECGAVDKNNLKTRVEYMSYYVAQAKARGILCILWDNNNFS GTGELFGFFDRRSCQFKFPEIIDGMVKYAFEAKTDPDPVIVYGSGESGSG LGDVNGDETVDAIDLAILKKYLLNSSTTINT ANADMNSDNAIDAIDYALLKKALL GTILEHHHHHH*.

[0065] Cel8C (465 aa; single underscore sequence is GH8; double underscore sequence is Doc):

[0066] MGSSHHHHHHSSGLVPRGSHMASIPFPYDAKYPNGAYSCLADSQ SIGNNLVRSEWEQWKSAHITSNGA RGYKRVQRDASTNYDTVSEGLGYGLLLSVYFGEQQLFDDLYRYVKVFLNSNGLMSWRIDSSGNIMGKDSIGAATDA DEDIAVSLVFAHKKWGTSGGFNYQTEAKNYINNIYNKMVEPGTYVIKAGDTWGGSNVTNPSYFAPAWYRIFADFTG NSGWINVANKCYEIADKARNSNTGLVPDWCTANGTPASGQGFDFYYDAIRYQWRAAIDYSWYGTAKAKTHCDAISN FFKNIGYANIKDGYTISGSQISSNHTATFVSCAAAAAMTGTDTTYAKNIYNECVKVKDSGNYTYFGNTLRMMVLLY TTG NFPNLYTYNSQPKPDLGSGESGSG LGDVNGDETVDAIDLAILKKYLLNSSTTINTANADMNSDNAIDAIDYAL LKKALL GTILEHHHHHH*.

[0067] Cel44O (588 aa; single underscore sequence is GH44; double underscore sequence is Doc):

[0068] MGINVSIDTTAERAAISPYIYGGNWEFNNAKLTAKRFGGNRTTGYNWENNY SNAGSDWQQSSDTYMLT SNKIPEDKWSEPGVVITDFHDKNLAAGEPYSLVTLQAAGYVSADANGTVAEDEVAPSERWKEVKFKKDAPLSLTPD TTDNYVYMDELVNLLVNKYGSASTATGIKGYAIDNEPALWSGTHPRMHPNNATCAEVIDKNINLAKTVKGVDPSAE TFGLVAYGFAAYNDFQSATDWKDLKGNYTWFLDYYLDSMKKASTEAGTRLIDALDLHWYPEA KGGGQRICFGEDPTNILCNKARLQAARTLWDPTYKEDSWIAQWCSFGLPLIPKVQESIDKYNPGTKLAFTEYSYGADNHITGGIAEADVLGVFGKYGVYLATVWGGGSYTAAGVNIYTNY DGNGSKYGDTKVKAETSDVENSSVYASVDSKDDSKLHVILINKNYDSPMTVNFGINSDKQYTSGRVWSFDRSSANITEKDAIDAISGNKLTYTIPALTVCHIVLDSSAQTTLGSGESGSG LG DVNGDETVDAIDLAILKKYLLNSSTTINTANADMNSDNAIDAIDYALLKKALL GTILEHHHHHH*.

[0069] Example 1: Screening of cohesin-dockerin combinatorial interactions with strong interactions

[0070] 1. Expression of cohesin and dockerin proteins

[0071] Three cohesin genes (Coh-1, Coh-7, and Coh-8) and nine dockerin genes from different clusters were cloned from *R. cellulolyticum*. After constructing recombinant plasmids by introducing them into vectors, these plasmids were transformed into *E. coli* for protein expression, followed by purification to obtain the corresponding proteins. The specific steps are as follows:

[0072] (1) Plasmid construction

[0073] DNA sequences encoding nine dockerin domains (Doc-Cel48F, Doc-Xyn10A, Doc-LipA, Doc-Cel9G, Doc-Abf62A, Doc-Cel9E, Doc-CE4A, Doc-Xyn141A, and Doc-Xyn11A) and three cohesin domains (Coh-1, Coh-7, and Coh-8) were cloned from *Clostridium fibrinolyticum* using the corresponding primers (Table 1). The PCR products of Dockerin and cohesin were ligated to the pGEX-6P-1 and pET28a vectors, respectively, via BamHI / XhoI and BamHI / NcoI restriction sites.

[0074] DNA sequences encoding the cellulase catalytic domains (Cel5A, Cel8C, Cel9G, Cel48F, Cel9E, and Cel44O) were cloned from *Clostridium fibronectin* using the corresponding primers (Table 3) and fused with the Doc-Cel48F sequence via SOE-PCR. The fusion fragments were ligated to the pET28a vector using the following restriction enzyme sites: NocⅠ / XhoⅠ for Cel48F and Cel44O, NheⅠ / XhoⅠ for Cel8C, Cel9G, and Cel5A, and NotⅠ / XhoⅠ for Cel9E. The resulting cellulases and their subunit assemblies have amino acid sequences shown in Tables 4-9.

[0075] Table 4 Docker module sequence

[0076]

[0077] Table 5

[0078]

[0079] Table 6

[0080]

[0081] Table 7

[0082]

[0083] Table 8

[0084]

[0085] Table 9

[0086]

[0087] Note: *Genes are from the cip-cel operon; **Genes are from the xyl-doc gene cluster; genes encoding the dockerin module were selected to interact with adhesion proteins, and these genes are underlined.

[0088] (2) Protein expression and purification

[0089] Recombinant plasmids encoding various dockerins and cohesins were expressed in *E. coli* BL21(DE3) star. Transformed cells were cultured in LB medium at 37°C. When the optical density reached 0.6, 500 mM isopropyl-β-D-thiogalactoside was added for induction. Cells were then cultured at 16°C for another 16 hours. Cells were collected by centrifugation at 10,000 × g and lysed using a high-pressure cell disruptor at 4°C and 800 bar for 2 minutes.

[0090] The recombinant protein was first purified according to the manufacturer's instructions using a nickel-nitrotriacetic acid column or a glutathione-S-transferase column. The purity of the recombinant protein was determined by 12% polyacrylamide gel SDS-PAGE. Protein concentration was determined using the BCA method. The protein was stored in 50% glycerol at -20 °C.

[0091] 2. Screening of strongly interacting cohesin-dockerin pairs

[0092] The specificity of Cohesin with dockerin was determined by affinity-based ELISA: dockerin (GSTDoc), fused with glutathione transferase, was interacted with 1 μg / mL of the required 15 nM CBM-Coh on an ELISA plate at concentrations ranging from 1 ng / mL to 1000 ng / mL.

[0093] The interaction between cohesin and dockerin was measured using a MicroCal iTC200 isothermal titration calorimeter. All samples were dialyzed in Tris buffer (50 mM Tris-HCl, 100 mM NaCl, 10 mM CaCl2, pH 7.4). 400 μL of cohesin sample (20 μM) was added to the sample cell, and 70 μL of dockerin sample (200 μM) was added to the syringe. The titration procedure was as follows: 0.2 μL of dockerin protein was added in the first injection, followed by 1.3 μL in each subsequent injection. The binding parameters were determined by fitting the experimental binding isotherms using a single-site model.

[0094] like Figure 1 As shown, sequence analysis revealed significant differences in Dockerin sequences among the different cellulases. They were clustered into seven distinct groups (I-VII), but no significant differences were observed among the adhesion proteins observed in native fibrosome scaffold proteins. Figure 1 (A and B in the middle). Figure 1 The results from the ITC assay showed that the interaction forces between different docking proteins and adhesion proteins varied significantly. ITC measurements revealed that the Ka value of the interaction between Doc-Cel48F and Coh-1 was 1.10E8 ± 2.51E8 M. -1 ( Figure 1 Therefore, Doc-Cel48F and Coh-1 were chosen as components for constructing artificial fibrous bodies to ensure that all catalytic subunits could be effectively anchored.

[0095] Example 2: Modular design and assembly of extended scaffold proteins

[0096] 1. Scaffold protein design and assembly

[0097] (1) Design three core building blocks:

[0098] Subunit A (SpC-4Coh-SnT): The N-terminus is SpyCatcher, which connects 4 Coh-1s, and the C-terminus is SnoopTag.

[0099] Subunit B (SnC-4Coh): The N-terminus is SnoopCatcher, which connects to 4 Coh-1s.

[0100] Subunit C (SpT-3Coh-CBM-SpT): Both the N-terminus and C-terminus are SpyTag, with 3 Coh-1 and 1 CBM in the middle.

[0101] (2) Assembly of 19Coh-CBM scaffold protein:

[0102] First, subunit A and subunit B are mixed in the optimal molar ratio (e.g., 3:1) and the intermediate SpC-8Coh (containing 8 cohesins) is formed through the SnoopTag / SnoopCatcher reaction.

[0103] Then, excess SpC-8Coh is mixed with subunit C, and the two SpC-8Coh molecules are attached to the two ends of subunit C through the SpyTag / SpyCatcher reaction, finally forming the target product 19Coh-CBM (4+4+3+4+4=19 cohesins).

[0104] The target product was purified by SEC as follows: The protein sample after reaction was filtered through a 0.22 μm filter and loaded into an NGC Scout10 + BioFrac (1 mL loading loop). SEC was performed using Superdex Hiload 200 (pg). Filtered Exchange Buffer (Tris 20 mM, NaCl 300 mM, CaCl2 10 mM, pH=8) was used as the mobile phase. The flow rate was 1 min / mL, and the temperature was 16 ℃. Peaks with UV (280 nM mAU) > 4 were collected. The collected peaks were then subjected to SDS-PAGE, and fractions with the desired molecular weight were concentrated.

[0105] 2. Results and Analysis

[0106] Figure 2Figure A shows a schematic diagram of the covalent linkages of the extended scaffold protein 19-Coh-CBM. These linkages are mediated by interactions between SpyTag / SpyCatcher and SnoopTag / SnoopCatcher, and are achieved through the assembly of the three basic scaffold proteins shown in the figure. Because SpT-3Coh-CBM-SpT has SpyTags at both ends, the reaction produces both single- and double-ended coupling products. However, the yield of the double-ended coupling product gradually increases with the increase of the SpC-4Coh-SnT ratio. When the molar ratio of SpC-4Coh-SnT to SpT-3Coh-CBM-SpT reaches 6:1, 95% of the reaction product is the double-ended coupling product. Figure 2 (Middle B). The SnoopTag / SnoopCatcher-mediated reaction of SnC-4Coh with SpC-4Coh-SnT generates a new protein product of approximately 150 kDa. When the molar ratio of SnC-4Coh to SpC-4Coh-SnT reaches 3:1, 90% of SpC-4Coh-SnT reacts with SnC-4Coh. Figure 2 (C). Under optimal conditions, 19Coh-CBM was successfully synthesized via a two-step reaction mediated by SnoopTag / SnoopCatcher and SpyTag / SpyCatcher interactions. However, in addition to 19Coh-CBM, excess reactants SpC-8Coh and SnC-4Coh in the reaction mixture also produced two impurities (C). Figure 2 (D). To obtain pure 19Coh-CBM, the reaction product was further purified by SEC. SDS-PAGE results showed that the molecular weight of the purified 19Coh-CBM was higher than that of 11Coh-CBM used as a control label, and the estimated purity was 90%. Figure 2 (E).

[0107] Similarly, by designing different subunit combinations and utilizing a one-step SpyTag / SpyCatcher reaction, a series of scaffold proteins, including 5Coh-CBM, 7Coh-CBM, and 11Coh-CBM, were successfully assembled. The results are shown in [Figure number missing]. Figure 3 The primers used to amplify different subunits are shown in Tables 2-3.

[0108] Example 3: Preparation and optimization of catalytic subunits

[0109] 1. Experimental Methods

[0110] Six key cellulases (Cel5A, Cel8C, Cel9G, Cel48F, Cel9E, and Cel44O) were selected, and their natural dockerin was replaced with the high-affinity Doc-Cel48F. These enzymes were successfully expressed and purified in *E. coli*. The activities of each enzyme were determined by enzyme activity assays, and the core and accessory subunits were identified through deletion experiments.

[0111] (1) Enzyme activity assay

[0112] Enzyme activity assays were performed in acetate buffer (50 mM acetate, 24 mM CaCl2, pH 6). Cellulase activity was measured at 40°C for 24 hours with 5% crystalline cellulose or for 1 hour with 1% sodium carboxymethyl cellulose. Cellulase mixtures and artificial cellulosomes were measured at 40°C for 48 hours with 1% crystalline cellulose. The concentration of released sugars was estimated using the dinitrobenzoic acid (DNS) method with glucose as the standard (refer to Ghose, 1987). Absorbance was measured at 490 nm.

[0113] (2) Missing experiment

[0114] One of the six key cellulases was missing, and the other five enzymes were mixed in equal molar ratios. The mixture was then compared with the control group—a mixture of all six enzymes—on crystalline cellulose to determine the core subunits and auxiliary subunits based on their contribution rates.

[0115] 2. Results and Analysis

[0116] The results are as follows Figure 4 As shown, enzyme activity and deletion experiments revealed that Cel9E, Cel44O, and Cel48F are the core subunits for degrading crystalline cellulose, while Cel5A, Cel8C, and Cel9G are the auxiliary subunits. Further optimization experiments determined that a 4:1 ratio of core to auxiliary subunits resulted in the optimal synergistic degradation effect of the enzyme mixture.

[0117] Example 4: Assembly and characterization of artificial fibrous bodies

[0118] 1. Experimental Methods

[0119] Assembly validation: The purified scaffold protein (such as 3Coh-CBM or 11Coh-CBM) was mixed with an excess of the catalytic subunit (see [link to documentation]). ​ The assay was performed using non-denaturing PAGE, SEC, dynamic light scattering (DLS), and cellulose binding pull-down assays.

[0120] (1) Non-denaturing polyacrylamide gel electrophoresis

[0121] The assembly between the catalytic subunit containing dockerin and the scaffold protein containing cohesin was determined by changes in lane bands during non-denaturing polyacrylamide gel electrophoresis. Both proteins were incubated in buffer (20 mM Tris, 24 mM CaCl2, 2 mM EDTA, pH 7.4) at 37°C for 2 hours. 20 μL of the assembly product was added to 5 μL of loading buffer (without SDS and β-mercaptoethanol) and then loaded onto a non-denaturing gel for analysis.

[0122] (2) Dynamic light scattering

[0123] The particle size distribution of assembled artificial cellulosomes was measured by dynamic light scattering. Protein samples were diluted to a final concentration of 0.5 mg / mL with buffer (20 mM Tris-HCl, 10 mM CaCl2, pH 7.4) and measured using a NanoBrook ZetaPlus instrument at 25 °C.

[0124] (3) Cellulose binding pull-down experiment

[0125] For the cellulose binding pull-down assay, the total reaction volume was 200 μL, containing 200 pmol of each cellulase, with or without 150 pmol of scaffold protein 11Coh-CBM. The mixture was placed in incubation buffer (50 mM acetate, 12 mM CaCl2, 2 mM EDTA, pH 5), supplemented with 10% cellobiose to prevent direct interaction between cellulase and cellulose. The tubes were incubated at 37°C for 3 hours. Then, 30 mg of crystalline cellulose was gently added, and the mixture was incubated at 4°C for 1 hour.

[0126] The supernatant was collected by centrifugation at 14,000×g for 5 minutes, and 5× SDS-PAGE loading buffer was added to a total volume of 60 μL. The cellulose precipitate was washed twice with washing buffer (500 mM NaCl, 20 mM Tris, 10 mM CaCl2, 0.1% Tween-20, pH 8) and resuspended in incubation buffer containing 5× SDS-PAGE loading buffer, with a total volume of 60 μL. The sample was boiled for 10 minutes before SDS-PAGE analysis.

[0127] (4) Transmission electron microscopy imaging

[0128] Cellulose granule complexes were prepared as described above, adsorbed onto a copper grid, and then negatively stained with 2% uranium acetate. Electron micrographs were acquired using a Tecnai T120C microscope equipped with a Ceta camera at an accelerating voltage of 120 kV. The image acquisition magnification was 67,000x, and the defocus was set in the range of -1.4 to -1 μm. For imaging of residual crystalline cellulose, the sample was prepared as a suspension, dropped onto a copper grid, and stained with 2% phosphotungstic acid. Electron micrographs were acquired using a Hitachi transmission electron microscope system at an accelerating voltage of 80 kV.

[0129] (5) Pull-down experiment

[0130] A mixture of 11Coh-CBM and six catalytic subunits at equimolar concentrations was incubated for 3 hours in the presence of crystalline cellulose. The composition of the supernatant (unbound fraction) and crystalline cellulose (bound fraction) was detected by SDS-PAGE. A mixture of free catalytic subunits was used as a control in the absence of scaffold protein.

[0131] 2. Results and Analysis

[0132] The formation of complexes between each catalytic subunit and the mini-scaffold protein 3Coh-CBM at different molar ratios of cohesins and dockerins was analyzed by non-denaturing polyacrylamide gel electrophoresis (PAGE). The results showed that a single complex master band was observed at different molar ratios, and its migration rate varied. When the molar ratio of the catalytic subunit to 3Coh-CBM exceeded 3:1, a band of the free catalytic subunit Cel5A began to appear, indicating that each cohesin module of 3Coh-CBM could bind to the catalytic subunit. ​ (A)

[0133] Following non-denaturing PAGE analysis, the assembly of excess Cel5A with the 3Coh-CBM mini-scaffold protein was characterized by SEC and dynamic light scattering (DLS). SEC results (Figure 5, B) showed a second peak (P1) preceding the 3Coh-CBM retention volume, in addition to the Cel5A peak (P2). SDS-PAGE identification confirmed that peak P1 contained both Cel5A and 3Coh-CBM in a 3:1 molar ratio, indicating that the assembled complex formed the correct molar ratio (Figure 5, C). When excess Cel5A assembled with 3Coh-CBM, the resulting complex also exhibited a narrower particle size distribution, ranging from 30 to 40 nm, which is larger than that of Cel5A alone (Figure 5, D).

[0134] When the amount of dockerin exceeded that of cohesin, all catalytic subunits were detectable in free form in the gel (Figure 5, E). Therefore, the interaction between 11Coh-CBM cohesin and dockerin (Doc-Cel48F) of different catalytic subunits appeared to be random and independent of the catalytic domain. To investigate the interaction between the 11Coh-CBM-mediated assembled complex and crystalline cellulose, affinity precipitation experiments were performed. Complexes formed by 11Coh-CBM with six molar concentrations of cellulases (Cel5A, Cel8C, Cel9G, Cel48F, Cel9E, and Cel44O), as well as enzyme-only samples, were interacted with crystalline cellulose. The samples were then centrifuged to separate the bound and unbound components, and the unbound components were analyzed by SDS-PAGE. The results (Figure 5, F) indicated that the enzymes were present in the unbound components in the absence of the scaffold protein, while only a small amount of Cel9E was present in the bound components, due to the weak interaction between its CBM and crystalline cellulose.

[0135] Transmission electron microscopy revealed that the assembled artificial fibrous bodies were irregularly clustered, with their size increasing with the amount of cohesin (e.g., approximately 15 nm for the 3Coh-CBM complex and approximately 40 nm for the 11Coh-CBM complex). (See...) ​ G.

[0136] The above results confirm the successful formation of a complex with a larger molecular weight, more uniform particle size, and higher thermal stability.

[0137] Example 5: Evaluation of the degradation activity of artificial fibrous bodies

[0138] Artificial fibrosomes, assembled from scaffold proteins carrying the same total enzyme amount (0.5 μM) but different numbers of cohesin (1, 3, 5, 7, 11, 19), were used to degrade 1% of crystalline cellulose.

[0139] like ​ As shown, the results revealed that the activity of the enzyme exhibited a "decreasing then increasing" trend with the amount of cohesin: 1-Coh-CBM and 19-Coh-CBM showed the highest activities (1.94 and 1.84 times that of the free enzyme, respectively), while intermediate amounts (such as 3-Coh-CBM and 11-Coh-CBM) showed relatively low activities. This indicates that when the enzyme is moderately anchored (1-Coh) or forms a super-large complex (19-Coh), it is more conducive to its dynamic turnover on the substrate, thereby achieving efficient degradation. Using an optimized ratio (core:helper = 4:1) of enzyme mixture or adding β-glucosidase can further enhance the activity of all assemblies.

[0140] Conclusion: This invention successfully establishes a general platform technology based on the Spy / Snoop covalent self-assembly system for constructing ultra-large, high-performance artificial fiber bodies, which has great application potential in the field of biomass energy.

[0141] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. An extended artificial scaffold protein, characterized in that, The extended artificial scaffold protein is obtained by self-assembling a scaffold protein containing at least 5 adhesion modules and one carbohydrate-binding module through a covalent self-assembly system. The covalent self-assembly system is one or two of SpyTag, SpyCatcher, SnoopTag, and SnoopCatcher, and the adhesion modules are derived from Clostridium fibrinolyticum.

2. The extended artificial scaffold protein as described in claim 1, characterized in that, The number of adhesive modules is 5, 7, 11, or 19.

3. A nucleic acid molecule encoding the extended artificial scaffold protein of claim 1 or 2.

4. A biomaterial comprising the nucleic acid molecule of claim 3, characterized in that, The biomaterials include recombinant vectors and recombinant bacteria.

5. An artificial fibrous corpuscle, characterized in that, It includes the extended artificial scaffold protein as described in claim 1 or 2 and a catalytic subunit that binds to the extended artificial scaffold protein through an adhesion module-docking module interaction, the catalytic subunit being derived from Clostridium fibrinolyticum.

6. The artificial fiber corpuscle as described in claim 5, characterized in that, The catalytic subunit includes one or more of glycoside hydrolases, carbohydrate esterases, and polysaccharide lyases.

7. The artificial fibrous body as described in claim 6, characterized in that, The catalytic subunit comprises a core subunit and an auxiliary subunit, wherein the molar ratio of the core subunit to the auxiliary subunit is 4:

1.

8. The artificial fibrous body as described in claim 7, characterized in that, The core subunits include Cel9E, Cel44O, and Cel48F, and the auxiliary subunits include Cel5A, Cel8C, and Cel9G.

9. A method for preparing artificial fibrous bodies as described in any one of claims 5-8, characterized in that, Includes the following steps: The adhesion module and carbohydrate-binding module of the scaffold protein, as well as the catalytic subunit, are expressed respectively; The artificial fibrous body is obtained by mixing and self-assembling the adhesion module, the carbohydrate binding module, and the catalytic subunit using one or two of SpyTag, SpyCatcher, SnoopTag, and SnoopCatcher.

10. The use of the extended artificial scaffold protein as described in any one of claims 1-2 or the artificial fibrous body as described in any one of claims 5-8 in any of the following: (1) Application in the degradation of cellulose or polysaccharides; (2) Application in the degradation of biomaterials containing cellulose or polysaccharides.