Composition and method for producing human milk oligosaccharides

The recombinant β-hexosyltransferase enzyme addresses the limitations of traditional GOS production by efficiently converting lactose to LacNAc-rich GOS, enhancing prebiotic effectiveness and scalability for food applications.

JP2026053481APending Publication Date: 2026-03-25NORTH CAROLINA STATE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current methods for producing galactooligosaccharides (GOS) using β-galactosidase face challenges such as high lactose concentration requirements, competitive inhibition by glucose and galactose, and inefficient production of LacNAc-rich GOS, which affects their prebiotic effectiveness.

Method used

Utilization of a recombinant β-hexosyltransferase (rBHT) enzyme, derived from Hamamotoa (Sporobolomyces) singularis, with specific modifications to enhance secretion and activity, allowing efficient conversion of lactose to LacNAc-rich GOS, overcoming limitations of traditional β-galactosidase methods.

Benefits of technology

The rBHT enzyme achieves high yields of LacNAc-rich GOS, improving the prebiotic properties and facilitating large-scale production with enhanced stability and secretion efficiency, suitable for applications in food products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053481000001_ABST
    Figure 2026053481000001_ABST
Patent Text Reader

Abstract

The present invention provides a polypeptide that catalyzes the hydrolysis of lactose β-(1-4) glycoside linkages for converting lactose and N-acetylglucosamine (GlcNAc) into a galactooligosaccharide (GOS) composition rich in N-acetyllactosamine (LacNAc). [Solution] A functional recombinant β-hexosyl-transferase (rBHT) polypeptide is provided, which has at least 90% sequence identity with a specific sequence and includes an N-terminal shortening of at least one amino acid with respect to the specific sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Detailed description of the invention

[0001] [Technical Field] Cross-reference of related applications This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 026,776 filed on 19 May 2020 and U.S. Provisional Patent Application No. 63 / 030,054 filed on 26 May 2020, both of which are incorporated herein by reference in their entirety for any purpose.

[0002] Integration by referencing electronically submitted documents The computer-readable nucleotide / amino acid sequence listing, submitted concurrently with this specification and identified as follows: 17,886-byte ASCII (text) filename "38389-601_SEQUENCE_LISTING_ST25", created on 18 May 2021.

[0003] Technical field This disclosure provides compositions and methods related to the production of human milk oligosaccharides (HMOs). In particular, this disclosure provides compositions and methods for converting lactose and N-acetylglucosamine (GlcNAc) into N-acetyllactosamine (LacNAc)-rich galactooligosaccharide (GOS) compositions using a novel β-hexosyltransferase (BHT) enzyme.

[0004] [Background technology] The complex interactions of diet, a healthy gut microbiome, and health are driving the development of strategies to promote the selective growth of beneficial microorganisms in the human gastrointestinal tract. Probiotics are microorganisms that have a positive impact on human health due to their potent anti-pathogenic and anti-inflammatory properties.

[0005] Furthermore, years of probiotic research have shown that selective prebiotics, when present in the diet, can promote selective modification of the gut microbiota and its associated biochemical activity. Prebiotics added to the diets of infants or adults are involved in the prevention of diseases such as allergies, lactose intolerance, and food sensitivities. Prebiotics are indigestible oligosaccharides (NDOs) with a dual capacity. Firstly, they reduce the efficiency of gut colonization by harmful bacteria, and secondly, they act as selective substrates to promote growth, thereby increasing the number of specific probiotic bacteria. In addition, a growing body of research shows that probiotics work best when combined with prebiotics.

[0006] Galactooligosaccharides (GOS) are considered one of the preferred prebiotic options. In the gastrointestinal tract, GOS are enzyme-resistant and pass through the small intestine undigested. In the large intestine, GOS ferments, promoting the growth of intestinal bifidobacteria, Lactobacilli species such as Lactobacillus acidophilus and L. casei, thus acting as a prebiotic. GOS are non-digestible oligosaccharides due to the conformation of their anomeric C atom (C1 or C2), and their glycosidic bonds allow them to avoid hydrolysis by digestive enzymes in the stomach or small intestine. Free oligosaccharides are found in the milk of all placental mammals, providing a natural example of prebiotic intake during infancy. The composition of human milk oligosaccharides (HMOs) is so complex that it is rare to find alternative sources containing oligosaccharides with similar compositions. Improved colon health in breastfed infants is attributed to the presence of GOS in breast milk. In fact, infant formula supplemented with GOS replicated the bifidogenic effect of human milk in terms of the metabolic activity and bacterial count of the colon microbiota. Among non-milk oligosaccharides, GOS has attracted particular attention because its structure is similar to the core molecule of HMOs. However, the concentration and composition of GOS vary depending on the method and enzymes used for its production, which can affect its prebiotic effect and the growth of colon probiotic strains. Traditionally, GOS has been produced using β-galactosidase from mesophilic or thermophilic microorganisms. β-galactosidase requires a high initial concentration of lactose to drive the reaction away from lactose hydrolysis and towards GOS synthesis. Because lactose is highly soluble at high temperatures, thermally stable β-galactosidases exhibiting high initial rates and extended half-lives have been used to reach a favorable equilibrium in the transgalactosylation reaction. However, competitive inhibition by glucose and / or galactose is another obstacle that can persist and be overcome by incorporating cells into the reaction.

[0007] The basidiomycete yeast Hamamotoa (Sporobolomyces) singularis (formerly Bullera singularis) cannot grow using galactose, but it does grow on lactose due to the activity of β-hexosyltransferase (BHT, EC3.2.1.21). Multiple studies have shown that BHT possesses transgalactosylation even at low lactose concentrations and with very limited lactose hydrolysis. Furthermore, this enzyme is not thought to be inhibited by lactose concentrations above 20%, and has the potential to convert lactose to GOS at a theoretical maximum of 75%. Unlike β-galactosidase, BHT from Hamamotoa (Sporobolomyces) singularis simultaneously performs glycosylhydrolase and β-hexosyltransferase activity, converting lactose to GOS without extracellular accumulation of galactose. The transgalactosylation reaction requires two molecules of lactose. The first molecule is hydrolyzed, and the second molecule acts as a galactose acceptor, resulting in the trisaccharide galactosyllactose (β-D-Gal(1-4)-β-D-Gal(1-4)-β-D-Glc) and residual glucose. Galactosyl-lactose also acts as an acceptor for the new galactose, producing the tetrasaccharide galactosylgalactosyllactose (β-D-Gal(1-4)-β-D-Gal(1-4)-β-D-Gal(1-4)-β-D-Glc), and similarly producing pentasaccharides and subsequent sugars. The trisaccharides, tetrasaccharides, and pentasaccharides accumulated in H. singularis are collectively referred to as GOS.

[0008] For practical benefits, recombinant secreted BHT may offer several advantages over native enzymes, including improved large-scale production and purification. Currently, the purification of active enzymes from Hamamotoa (Sporobolomyces) singularis requires cell lysis and subsequent multiple chromatographic steps. Previous attempts to express recombinant β-hexosyl-transferase in E. coli BL21 resulted in high levels of production, but the enzyme was inactive and insoluble.

[0009] [Summary of the Invention] Embodiments of the present disclosure include a functional recombinant β-hexosyl-transferase (rBHT) polypeptide having at least 90% sequence identity with SEQ ID NO: 1 and comprising at least one N-terminal truncated form of SEQ ID NO: 1.

[0010] In some embodiments, the polypeptide has at least 95% sequence identity with SEQ ID NO: 1. In some embodiments, the polypeptide further includes at least one additional amino acid substitution. In some embodiments, the polypeptide includes an N-terminal shortening that is about 1 to about 81 amino acids long. In some embodiments, the N-terminal shortening is about 1 to about 56 amino acids long. In some embodiments, the polypeptide has at least 90% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9.

[0011] In some embodiments, the polypeptide further comprises a signal sequence. In some embodiments, the signal sequence is non-natural. In some embodiments, the signal sequence comprises an amino acid sequence derived from a yeast protein. In some embodiments, the signal sequence comprises an amino acid sequence derived from a protein derived from one of Komagataella (Pichia) pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Hansenula (Ogataea) polymorpha, or Kluyveromyces lactis. In some embodiments, the signal sequence comprises a polypeptide having at least 90% sequence identity with at least one of the following: α-conjugation factor signal sequence (MFα) (SEQ ID NO: 29), invertase (IV) signal sequence (SEQ ID NO: 30), glucoamylase (GA) signal sequence (SEQ ID NO: 31), or inulinase (IN) signal sequence (SEQ ID NO: 32) derived from Saccharomyces cerevisiae. In some embodiments, the polypeptide has at least 90% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least one asparagine residue at positions 289, 297, 431, and / or 569 relative to SEQ ID NO: 1.

[0012] In some embodiments, the polypeptide is soluble or membrane-bound. In some embodiments, about 1% to about 50% of the polypeptide is soluble. In some embodiments, the polypeptide catalyzes the hydrolysis of lactose β-(1-4) glycoside linkages. In some embodiments, the catalysis of the hydrolysis of lactose β-(1-4) glycoside linkages by the polypeptide produces a composition containing LacNAc-rich GOS.

[0013] Embodiments of this disclosure also include nucleic acid molecules encoding any of the polypeptides described above. Embodiments of this disclosure also include vectors containing any one of these nucleic acid molecules.

[0014] Embodiments of this disclosure also include methods for producing a GOS composition from lactose in a host cell using any of the polypeptides described above. In some embodiments, the GOS composition includes LacNAc-rich GOS and / or GlcNAc-free GOS.

[0015] In some embodiments of this method, the host cell is one or more of the following: yeast cells, fungal cells, mammalian cells, insect cells, plant cells, or algal cells. In some embodiments, the host cell includes any cell of the genus Komagataella.

[0016] In some embodiments of this method, the host cells include one or more cells derived from Komagataella (Pichia) pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Hansenula (Ogataea) polymorpha, or Kluyveromyces lactis, Aspergillus spp., and Trichoderma reesei. In some embodiments, this method produces a LacNAc-rich GOS yield with a total GOS concentration of at least 10% of the initial lactose concentration and at least 50% of the initial lactose concentration.

[0017] Embodiments of the present disclosure also include compositions comprising any of the above polypeptides, and / or compositions comprising one or more LacNAc-rich GOS using any of the above polypeptides.

[0018] In some embodiments, the composition is a food. In some embodiments, the food comprises one or more of the following: infant formula, yogurt, dairy products, milk-based beverages, fruit beverages, hydration beverages, energy beverages, fruit preparations, and meal replacement beverages.

[0019] Other aspects and embodiments of this disclosure will become apparent in light of the following detailed description.

[0020] [Brief explanation of the drawing] [Figure 1] This figure shows the predicted structural post-translational modifications and disordered vs. ordered secondary motifs of β-hexosyltransferase from H. singularis. The glycosylation, phosphorylation, and secondary structure of the BHT protein were predicted using various algorithms. Described are the structural elements, conserved regions, and functional domains of BHT using the PSIPRED and GlobplotGlobular prediction tools. Disordered regions were predicted using the Phyre2, IUPRED2A, DISOPRED3, GlobplotDisorder, and PONDR algorithms. The phosphorylation servers DisPhos1.3 and NetPhosYeast1.0 indicate phosphorylation sites. GlycoEP shows N-glycosylation (red line) and O-glycosylation (black line), but the C-mannosylation site was not predicted. The numbers below each prediction line indicate the BHT amino acid residue number.

[0021] [Figure 2] Figure 2A: Comparison of enzymatic activity of rBHT variants is based on the amount of secreted soluble protein produced by recombinant K. pastoris strains carrying the truncated variant of rBht-HIS under AOX1 promoter control (final culture (OD) 600nm (Normalized over) (A) Graphs illustrating the generated chimeric genes, including combinations of the leader domain and the ORF of the rBht variant. Specific tags, mutations, and deletions are shown. Figure 2B: Comparison of enzymatic activity of rBHT variants is shown by the amount of secreted soluble protein (final culture (OD)) produced by recombinant K. pastoris strains carrying the truncated variant of rBht-HIS under AOX1 promoter control. 600nm (Normalized over ) The protein concentrations (B) of soluble secreted proteins secreted by the following recombinant strains were compared: Column 1, GS115::rBht (1-594) -HIS; column 2, GS115::MFα-rBht (1-594) -HIS; column 3, GS115::MFα-rBht (23-594)-HIS; Column 4, GS115::MFα-rBht (23-594) (N289Q)-HIS; Column 5, GS115::MFα-rBht (23-594) (N297Q)-HIS; Column 6, GS115::MFα-rBht (23-594) (N431Q)-HIS; Column 7, GS115::MFα-rBht (23-594) (N569Q)-HIS; Column 8, GS115::MFα-rBht (32-594) -HIS; Column 9, GS115::MFα-rBht (54-594) -HIS; Column 10, GS115::MFα-rBht (57-594) -HIS; Column 11, GS115::MFα-rBht (82-594) -HIS; Column 12, GS115::MFα-rBht (95-594) -HIS; Column 13, GS115::MFα-rBht (103-594) -HIS; Column 14, GS115::MFα-rBht (111-594) -HIS; Column 15, GS115::IV-rBht (54-594) -HIS; Column 16, GS115::GA-rBht (54-594) -HIS; Column 17, GS115::IN-rBht (54-594) -HIS; Column 18, GS115::MFα (Δ57-70) -rBht (23-594) -HIS; Column 19, GS115::MFα (Δ57-70) -rBht (57-594) -HIS; Column 20, GS115(His + ) control. Figure 2C: The comparison of the enzymatic activities of rBHT variants is the amount of secreted soluble protein (normalized with respect to the final culture (OD 600nm )) produced by recombinant K. pastoris strains harboring truncated variants of rBht-HIS under the control of the AOX1 promoter. The enzymatic activities (C) of the soluble secreted proteins secreted by the following recombinant strains were compared: Column 1, GS115::rBht (1-594) -HIS; Column 2, GS115::MFα-rBht (1-594) -HIS; Column 3, GS115::MFα-rBht (23-594) -HIS; Column 4, GS115::MFα-rBht (23-594)(N289Q)-HIS; row 5, GS115::MFα-rBht (23-594) (N297Q)-HIS; row 6, GS115::MFα-rBht (23-594) (N431Q)-HIS; Row 7, GS115::MFα-rBht (23-594) (N569Q)-HIS; row 8, GS115::MFα-rBht (32-594) -HIS; column 9, GS115::MFα-rBht (54-594) -HIS; column 10, GS115::MFα-rBht (57-594) -HIS; column 11, GS115::MFα-rBht (82-594) -HIS; column 12, GS115::MFα-rBht (95-594) -HIS; column 13, GS115::MFα-rBht (103-594) -HIS; column 14, GS115::MFα-rBht (111-594) -HIS;Column 15, GS115::IV-rBht (54-594) -HIS; column 16, GS115::GA-rBht (54-594) -HIS; column 17, GS115::IN-rBht (54-594) -HIS; column 18, GS115::MFα (Δ57-70) -rBht (23-594) -HIS; column 19, GS115::MFα (Δ57-70) -rBht (57-594) -HIS; Column 20, GS115(His + ) contrast.

[0022] [Figure 3] Figure 3A: This figure shows Coomassi-stained SDS-PAGE (10%) separation and Western blot. The figure shows cell-free extracts (soluble secreted proteins) of proteins expressed by different recombinants of K. pastoris GS115. (A) Separated proteins produced by SDS-PAGE exposed to anti-HIS antiserum; Lane 1, GS115::MFα-rBht-HIS; Lane 2, GS115::MFα-rBht (23-594) -HIS; Lane 3, GS115::MFα-rBht (32-594) -HIS; Lane 4, GS115::αMF-rBht (54-594) -HIS; Lane 5, GS115::αMF-rBht (57-594)-HIS; Lane 6, GS115::αMF-rBht (82-594) -HIS; Lane 7, GS115::αMF-rBht (95-594) -HIS; Lane 8, GS115::MFα-rBht (103-594) -HIS; Lane 9, GS115::MFα-rBht (111-594) -HIS; Lane 10, GS115 control containing an empty pPIC9 vector. Equal amounts were loaded into each lane to aid comparison. The total protein (ng) loaded into each well is shown above in (A). "---" indicates that the concentration could not be measured. M indicates a lane containing a molecular weight protein marker, and (kDa) is shown on the left side of the panel. Figure 3B: Figure showing Coomassie-stained SDS-PAGE (10%) separation and Western blot. The figure shows cell-free extracts (soluble secreted proteins) of proteins expressed by different recombinants of K. pastoris GS115. (B) Separated proteins produced by Western blot were exposed to anti-HIS antiserum; Lane 1, GS115::MFα-rBht-HIS; Lane 2, GS115::MFα-rBht (23-594) -HIS; Lane 3, GS115::MFα-rBht (32-594) -HIS; Lane 4, GS115::αMF-rBht (54-594) -HIS; Lane 5, GS115::αMF-rBht (57-594) -HIS; Lane 6, GS115::αMF-rBht (82-594) -HIS; Lane 7, GS115::αMF-rBht (95-594) -HIS; Lane 8, GS115::MFα-rBht (103-594) -HIS; Lane 9, GS115::MFα-rBht (111-594) -HIS; Lane 10, GS115 control containing an empty pPIC9 vector. Equal amounts were loaded into each lane to aid comparison. The total protein (ng) loaded into each well is shown above (B). "---" indicates that the concentration could not be measured. M indicates a lane containing a molecular weight protein marker, and (kDa) is shown on the left side of the panel.

[0023] [Figure 4] This figure shows the enzyme kinetic parameters of rBHT variants tested at 20°C, 30°C, 42°C, and 55°C. kcat / km versus temperature. The enzyme assay was performed using 0.3 μg of rBHT. (23-594) -HIS, rBHT (32-594) -HIS, rBHT (54-594) -HIS and rBHT (57-594) -The experiment was conducted in the presence of HIS within the ONP-Glu substrate concentration range (0.08 to 10.4 mM) described in "Methods". Km and kcat were calculated from the initial rate of ONP-Glu cleavage using the Hill formula. The values ​​are the mean ± standard deviation (SD) of three independent measurements.

[0024] [Figure 5] Figure 5A: This figure shows an example of N-acetyllactosamine (LacNAc) production at a lactose / N-acetylglucosamine ratio of 1:2. The recombinant BHT (rBHT) polypeptide of this disclosure can catalyze the repeated addition of galactose (Gal from lactose) to N-acetylglucosamine (GlcNAc). Figure 5B: This figure shows an example of N-acetyllactosamine (LacNAc) production at a lactose / N-acetylglucosamine ratio of 1:2. It shows the enzymatic reaction catalyzed by rBHT. An example of a time course study of galactosyl-lactose (Gal-lactose), galactosyl-N-acetalactosamine (Gal-LacNAc), and N-acetyllactosamine (LacNAc) synthesis is shown. -1 The assay was performed using lactose. The assay contained approximately 20 g / L lactose and approximately 10 g / L N-acetylglucosamine (GlcNAc) in 5 mM sodium phosphate buffer (pH 5.0), which was incubated at 30°C. Samples were periodically removed and analyzed by HPLC, and detected by ELSD and PDA.

[0025] [Figure 6A] 6m4e(HsBglA (23-594)This figure shows multiple secondary structure alignments of -HIS) with structurally homologous GH1 proteins. (A) The proteins found to be most structurally homologous from the PDB database include 2E3ZA (BGL1A), 3AHYB (TrBgl2), 5BWFA (ThBgl), 4MDOA (HiBG), and 5JBOA (ThBgl2) (Table 4). The primary sequence alignment is shown below. rBHT (23-594) The secondary structural elements of -HIS] and their names are shown above the alignment. β strands are indicated by black arrows, α helix structures by coils, exact α turns (TTT letters), β turns (TT letters), and η are 3 10 Refers to a helix random coil. The secondary structural element numbers of the (α / β)-Tim barrel structure are indicated above the structural alignment as (α1-α8) and (β1-β8). HsBglA (23-594) -In the analysis of the HIS unstructured region, the amino acid numbers of HsBglA from the amino terminus to the carboxyl terminus include the deletion signal sequence (residues 1-22) indicated by dashed arrows and the unstructured region missing from the crystal structure (residues 23-53) indicated by dotted lines. The amino acids were aligned using ClustalO based on % sequence similarity. Identical residues are enclosed in white on a black background, conservative changes in a gray box, and insertions are highlighted with a purple background. Catalytic acid / base nucleophilic residues are indicated with an asterisk. HsBglA (23-594) - Glycosylation sites in HIS are indicated by triangles. Figure 6A shows the predicted phosphorylation sites and potential O-glycosylation sites at the N-terminus shown in Figure 1, indicated by squares and circles, respectively. The consensus sequence is shown below the aligned sequence. The image was generated using the ENDscript2.0 web server (http: / / endscript.ibcp.fr / ESPript / ENDscript / )(5) and HsBglA was generated using data obtained from the Dali protein structure comparison server (http: / / ekhidna2.biocenter.helsinki.fi / dali / ). (23-594)-Derived from a comparison of the 3D crystal structure and the crystal structure in the Protein Databank, based on HIS (PDB ID: M6E4) (Holm, 2019).

[0026] [Figure 6B] 6m4e(HsBglA (23-594) This figure shows multiple secondary structural alignments of HsBglA (HsBglA) with structurally homologous proteins. (23-594) -HIS (PDBID:M6E4) The four extension loops A, B, C, and D are colored blue, green, yellow, and red, respectively, and are shown as arrows of the same color that form the substrate binding pocket entrances, shown above the secondary structure of (A). Generated using PyMOL (https: / / pymol.org / 2 / ).

[0027] [Figure 6C] 6m4e(HsBglA (23-594) This figure shows multiple secondary structural alignments of HsBglA (HsBglA) with structurally homologous proteins. (23-594) The degree of conservation of -HIS(PDBID:M6E4) is represented by a color gradient from red to blue. Deep red indicates more conserved residues, while deeper blue indicates more variable residues. Generated using PyMOL (https: / / pymol.org / 2 / ).

[0028] [Figure 7] Figure 7A: This figure shows SAXS data for BHT at 1 mg / ml (red) and 4 mg / ml (blue). The SAXS data is shown as a logarithmic plot (left). I(Q) is an arbitrary unit. Figure 7B: The P(r) curve calculated from the SAXS data is normalized to a maximum height of 1.0.

[0029] [Modes for carrying out the invention] This disclosure provides compositions and methods related to the production of human milk oligosaccharides (HMOs). In particular, this disclosure provides compositions and methods for converting lactose and N-acetylglucosamine (GlcNAc) into N-acetyllactosamine (LacNAc)-rich galactooligosaccharide (GOS) compositions using a novel β-hexosyltransferase (BHT) enzyme.

[0030] Hamamotoa (Sporobolomyces) singularis encodes an industrially important inducible membrane-bound β-hexosyltransferase (BHT), which, when heterologously expressed by Komagataella (Pichia) pastoris, is partially secreted and soluble. BHT secretion is determined by a 22-amino acid signal sequence, part of a novel amino-terminal region (1-110 amino acids), and is predicted to be glycosylated at four arginine positions of catalytic glycosylhydrolase (GH1) within the carboxyl-terminal domain. To evaluate the role of each N-glycosylation site in the generation of a biologically active soluble enzyme, the activities of N-glycosylating recombinant enzyme variants (e.g., N289Q, N297Q, N431Q, and N569Q) produced by Komagataella (Pichia) pastoris were comparatively analyzed. Functional analysis of four deglycosylated soluble variants revealed a measurable decrease in the activity of total recombination (rBHT) (a 58–97% reduction), indicating that glycosylation at all four sites is crucial for the generation of the active enzyme. Furthermore, in silico structural predictions revealed the presence of a disordered segment within a novel amino-terminal region (1–110 amino acids) preceding the catalytic C-terminal GH1 domain. Deletion analysis was performed targeting the segment surrounding the putative disordered region to generate eight shortened N-terminal domain enzyme variants. The effect of enzyme shortening on the ratio of membrane-bound variants to secreted soluble enzyme variants was evaluated. Fusion of the MFα signaling sequence of the active soluble shortened variants with the modified MFα type generated by Komagataella (Pichia) pastoris was compared for secretion titer, stability, and enzyme kinetics. Remarkably, the deletion of up to 56 amino acids in the N-terminus produced a fully functional secreted soluble enzyme variant, and approximately 65% ​​of the total secreted active enzyme was membrane-bound under the experimental conditions described herein.

[0031] Hamamotoa (Sporobolomyces) singularis (H. Singularis) expresses extracellular membrane-bound glycosylated β-hexosyltransferase (BHT) under inducible conditions. BHT catalyzes the hydrolysis of cellobiose β-(1-4) glycoside linkages and possesses attractive enzymatic transgalactosylation ability in the presence of lactose, enabling the synthesis of galactooligosaccharides (GOS), which are considered prebiotics and widely used as functional food additives. Therefore, there is growing interest in the important role of this novel enzyme in catalyzing transgalactosylation reactions.

[0032] In recent years, heterologous expression of biologically inactive rBHT by Escherichia coli (E. coli) has suggested that post-translational modifications such as glycosylation are necessary to obtain the active enzyme. However, it remains unclear whether all potential glycosylation sites within the carboxyl-terminal domain and / or N-terminal region motifs are involved in the generation of biologically active rBHT. The novel N-terminal region has no known sequence homologs, and its characteristics are still unclear. The carbohydrate portion of glycoproteins is generally thought to promote protein folding, oligomerization, protection from proteolysis, secretion, intracellular transport, cell surface expression, and enzymatic activity.

[0033] Komagataella (Pichia) pastoris (K. pastoris) is commonly used as a eukaryotic host for recombinant protein production due to its post-translational modification and secretory capabilities. As will be recognized by those skilled in the art based on this disclosure, Komagataella (Pichia) pastoris (K. pastoris) is also known as Kamagataella phaffi. As will be further described herein, the various compositions and methods of this disclosure are applicable to any host cell, including but not limited to yeast cells, fungal cells, mammalian cells, insect cells, plant cells, or algal cells. In some embodiments, the host cell includes any cell of the genus Komagataella.

[0034] In K. pastoris, N-glycans form a high-mannose heterologous oligosaccharide starting with the addition of the core unit Glc3Man9GlcNAc2 at asparagine in the recognition sequence Asn-X-Ser / Thr (Glc = glucose; GlcNAc = N-acetylglucosamine; Man = mannose). Heterologous expression of rBHT by K. pastoris resulted in glycosylated extracellular cell wall or membrane-bound enzymes. Surprisingly, the natural protein leader directed a fraction of the enzyme to be secreted into culture broth as an active soluble enzyme. Previous studies have demonstrated that K. pastoris can secrete soluble, biologically active rBHT into culture broth, opening up the possibility of a simple downstream recovery process protocol. Therefore, we conducted experiments to recover, purify, and evaluate the activity and stability of the soluble active enzyme and compare it to membrane-bound rBHT.

[0035] The predicted protein contains 594 amino acids, including an amino-terminal region of 1–110 amino acids, no known sequence homologs, followed by a carboxyl-terminal glycosyl hydrolase family 1 (GH1) catalytic domain. The N-terminus also contains a 22-amino acid secretion signal peptide, which restricts secretion when fused to the α-conjugation factor (MFα) signal sequence derived from Saccharomyces cerevisiae upstream of the entire open reading frame. Experiments showed that this restriction could be partially removed by replacing the native BHT signal sequence (1–22aa) with the MFα signal sequence. As a result, the activity of the biologically active soluble enzyme in the culture broth unexpectedly increased 53-fold, and the K. pastoris membrane-associated form of this enzyme also increased. These results demonstrate that the BHT signal sequence influences membrane-bound localization versus secretion of the soluble enzyme into the culture medium. Previous results did not address the role of the N-terminal region outside the initial 22-amino acid signal peptide, but as further described herein, we have established a system that can evaluate this issue using deletion mutagenesis within a novel 1-110 N-terminal domain.

[0036] The secretion of soluble proteins by K. pastoris is highly protein-dependent and remains a common bottleneck in the production process, as is well recognized in this art. One reason for this limitation is thought to be improper folding, which can be improved by overexpressing folding helper proteins. Alternative approaches to improving secretion include redesigned strains and mutagenesis. Furthermore, several studies have shown that altering glycosylation and cell transport-related genes increases the secretion of soluble recombinant proteins.

[0037] In this disclosure, experiments (using site-directed mutagenesis and progressive deletion analysis) were conducted to address whether the secretion of soluble active rBHT is regulated by post-translational N-glycosylation modifications embedded within the C-terminal GH1 domain and / or limited by function contained within a novel 110 N-terminal region (amino acids 23-110). Overall analysis of rBHT expression for each modified or shortened enzyme variant was complemented by analysis of enzyme activity, measured as the ratio of soluble enzyme to membrane-associated enzyme. Finally, the results of this disclosure further demonstrate the uniqueness of the N-terminus by presenting comparative sequence and structural analysis with homologous GH1 proteins, whose coordinates are available in the Protein Databank (PDB), using recently derived crystal structures of the BHT enzymes.

[0038] Based on the industrial applications and importance of BHT, there is a strong desire to improve the secretion efficiency of soluble active enzymes. In recent years, structural information for 90% of BHT enzymes has become available, and these findings were compared with other GH1 family members. From the obtained data, in silico structural predictions of the enzyme were confirmed, showing two different structural domains: a novel 110 N-terminal domain including the signal sequence and putative disordered region, and a conserved carboxyl GH1 domain. From these data, various glycosylation and phosphorylation sites were also predicted. Therefore, three general categories of protein structural modification were performed: 1) site-directed mutagenesis of four glycosylation sites; 2) shortening of 110 N-terminal regions; and 3) substitution and modification of secretion signals. In the first group of modifications, site-directed mutagenesis targeted glycosylation sites, and their importance to enzyme activity was confirmed. In the second group of modifications, it was shown that the removal of up to 56 N-terminal amino acids did not affect enzyme activity, and that these residues do not play a significant role in the secretion of soluble active rBHT. In the third group of modifications, it was shown that altering the MFα signal sequence increased the ratio of secreted soluble protein to membrane-associated protein (0.67) (Table 1).

[0039] Investigating the correlation between rBHT N-glycosylation and corresponding enzymatic properties is a crucial step in evaluating enzyme stability, activity, and even production. Post-translational modifications such as N-glycosylation are involved in protein folding in the ER and play a vital role in the secretion of heterologous proteins. However, not all predicted N-glycosylation sequences in polypeptides are glycosylated in vivo. While several algorithms are available to predict N- and O-glycosylation sites, the effects of enhancing or removing putative sites on expression and secretion can only be confirmed in vivo. In silico analysis has recently suggested that the BHTGH1 domain, confirmed by three-dimensional structure, contains four N-glycosylation sites (HsBglA, PDB ID: M6E4). Importantly, single-site substitution of asparagine with glutamine showed a strong correlation with the expression of the active enzyme. Surprisingly, however, the ratio of secreted soluble enzyme to cell membrane association activity was found to be related to BHT (23-594)(N569Q) -HIS increased from 0.40 to 0.66. In particular, these substitutions increased the rBHT of the parent strain. (23-594) Compared to HIS, the secreted soluble protein decreased dramatically from 58% to 97%, and cell membrane-bound active protein decreased from 75% to 95%. This wide range of activity, expressed as the percentage of fully active enzymes, indicates that the absence of even one N-glycosylation site is sufficient to reduce the titer of the active enzyme, and that a fully functional enzyme is only obtained when all four tested sites are glycosylated.

[0040] The experiments also investigated whether the secretion of soluble active protein is affected by the presence of disordered N-terminal segments, and whether their removal has functional importance to the catalytic activity of the shortened secreted soluble rBHT variant. Little is known about the novel 110 N-terminal region of BHT, but so far it is a fragment that lacks homology to other known proteins. Based on the predicted disordered segment of the novel 110 N-terminal domain, deletion chimeras were generated by fusing the MFα signaling sequence. Heterologous expression of the N-terminal shortening containing amino acids 1-56 resulted in equivalent enzyme kinetic parameters for secreted solubility, stability, and bioactive enzymes, but the secretion process was inactivated by further deletion of the N-terminus of the disordered segment (Figure 2; Figure 3; Table 1). Therefore, the activity and stability of BHT are independent of the 56 amino acids at the N-terminus, but the effect on secretion, as explained regarding the N-glycosylation site, can only be confirmed in vivo. For example, at amino acid 56, the carboxyl-terminal boundary of the disordered region predicted by IUPRED2A can be removed, but the disordered region predicted downstream was necessary to obtain the active enzyme.

[0041] Essentially disordered proteins (IDPs) exist by exchanging conformations rather than adapting to a clearly defined structure. Disordered regions can be distinguished from ordered regions based on their amino acid sequence, and in most cases, disordered proteins are not evolutionarily conserved, rather their disordered structure is maintained. IDPs are involved in multiple cellular functions, including transcription, translation, regulation, and signal transduction, and are abundant in phosphorylation sites. Often, IDPs are involved in the binding of DNA or RNA and to other proteins, and can assist in the assembly of multiprotein complexes. Furthermore, IDPs are infrequent in enzymes, and significant deviations occur as output within the GH1 domain depending on the server, although there are no disordered regions within the GH1 domain when using the more stringent server DISOPRED3.

[0042] Furthermore, as further described herein, structural modifications were made by substituting the secretory signal, considering that a truncated active polypeptide of BHT had been previously detected at residues 17 or 22 in protein cell extracts from the cell membrane of H. singularis. This finding suggested that this fragment was cleaved to form mature BHT. Using K. pastoris, the results demonstrated that BHT amino acids 1–22 act as a functional native signal sequence. Substitution of this MFα signal sequence was demonstrated to enable the secretion of a soluble active rBHT variant, although approximately 71% of the secreted enzyme still associated with the membrane (Table 1), which is consistent with previous results.

[0043] It should be noted that the persistent partial localization of rBHT by the cell membrane after removal of the N-terminal disorder region suggests that the binding site to the cell membrane may be located in either the unbiased cleavage of MFα, or possibly in a novel N-terminal region or within 57-110 amino acids of the BHT GH1 domain. Most secreted proteins in eukaryotes contain an N-terminal signal sequence that directs the protein to an intracellular or extracellular location. The ability of peptide sequences with minimal sequence homology to function as signal peptides makes it possible to replace the original signal sequence with a signal peptide sequence found in yeast. By comparing four signal sequences, it was revealed that the secreted BHT peptide continues to associate with the cell membrane.

[0044] The cleavage of signal peptides has been shown to be crucial for the assembly and secretion of functional prolipoproteins across the E. coli membrane. One study showed that untreated consensus MFα-α-interferon accumulated in the periplasmic space and cell wall, and its secretion into the culture medium and intracellular accumulation could be mitigated by the Glu-Ala dipeptide between MFα and α-interferon. Furthermore, deletion of amino acids 57-70 in the pro region of MFα was shown to increase the secretion of horseradish peroxidase and lipase by at least 50%. Therefore, based on these results, cleavage of signal peptides by signal peptidases may be necessary for the final assembly and secretion of soluble rBHT. Variant GS115::MFα (Δ57-70) -rBht (23-594) -HIS and GS115::MFα (Δ57-70) -rBht (57-594) -The same modification to MFα seen in HIS is GS115::MFα-rBht (23-594) -Compared to HIS, secretion increased by 58%, resulting in a 40% increase in the proportion of soluble secreted to the associated membrane (Table 1).

[0045] The crystal structure of BHT is generally similar to that of GH1 family proteins, but the N-terminus (residues 1-110) has no known homologs, and residues 23-54 were not defined within the structure. This region has previously been proposed to be unstructured and structurally dynamic. As further described herein, deletion analysis was performed on the unstructured N-terminal domain based on in silico results. In light of these results, the features within the first 56 residues of the N-terminus are likely to play a limited role in cell association activity, but are not required for enzyme folding, secretion, or activity.

[0046] To rationally redesign the enzyme, it is essential to determine the possible regulatory mechanisms of the unstructured N-terminal region of BHT. According to the results of this disclosure, the homologous structure contains a conserved C-terminal catalytic domain but lacks the highly essentially disordered N-terminal domain observed in BHT in in silico analysis (Figure 1). Consistent with in silico predictions, recently published BHT was degraded by X-ray crystallography (HsBglA, PDB:6M4E). (23-594) In the three-dimensional structure of -HIS, residues 23–54 within the N-terminus do not possess a detectable electron density. This is consistent with the unstructured residues in this region predicted in silico (Figure 1). The overall structure of the C-terminal catalytic domain is similar to the classical GH1 structure, which is also confirmed by the crystal structure. However, certain elements (Figure 6) have been found in addition to the intrinsic amino acids within the catalytic nucleophile, which may provide a handle to different catalytic properties of BHT for future studies.

[0047] All of the above data further document the role of the N-terminal disorder region beyond 56 residues in maintaining active rBHT, suggesting that the basis for the partial selective sequestration of cell wall-bound rBHT lies in the inefficient processing of the signaling secretory sequence. Overall, the results of this disclosure using K. pastoris improved the secretory titer of soluble rBHT by removing 56 endogenous N-terminal amino acids while fusing to truncated MFα.

[0048] The section headings used in this section and throughout the disclosure herein are for structural purposes only and are not intended to limit the scope of the information.

[0049] 1.Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. In case of any conflict, including definitions, this document shall prevail. Preferred methods and materials are described below, but similar or equivalent methods and materials described herein may also be used in the practice or testing of the disclosure. The phrase "in one embodiment" may, but not necessarily, refer to the same embodiment as used herein. Furthermore, the phrase "in another embodiment" may, but not necessarily refer to a different embodiment as used herein. Thus, the various embodiments of the disclosure, as described below, can be readily combined without departing from the scope or spirit of the embodiments provided herein. All publications, patent applications, patents, and other references referenced herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and are not intended to be limiting.

[0050] The terms “comprise,” “include,” “having,” “has,” “can,” “contain,” and their variations are intended, when used herein, to be open transitional phrases, terms, or words that do not preclude the possibility of further actions or structures. The singular forms “a,” “an,” and “the” include multiple referents unless the context indicates otherwise. This disclosure also contemplates other embodiments or elements present herein that “comprising,” “consisting of,” and “consisting essentially of,” whether expressly described herein or otherwise.

[0051] In the enumeration of numerical ranges as described herein, each numerical value having a similar degree of precision between them is explicitly intended. For example, in the range 6–9, the numerical values ​​7 and 8 are intended in addition to 6 and 9, and in the range 6.0–7.0, the numerical values ​​6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly intended.

[0052] As used herein, "correlated with" means to compare or contrast with.

[0053] As used herein, the term “nucleic acid molecule” refers to any nucleic acid-containing molecule, including but not limited to DNA or RNA. This term includes, but is not limited to, 4-acetylcytosine, 8-hydroxy-N6-methyladenosine, aziridinylcytosine, pseudoisocytosine, 5-(carboxyhydroxylmethyl)uracil, 5-fluorouracil, 5-bromouracil, 5-carboxymethylaminomethyl-2-thiouracil, 5-carboxymethylaminomethyluracil, dihydrouracil, inosine, N6-isopentenyladenine, 1-methyladenine, 1-methylpseuduracil, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-methyladenine, 7-methylguanine, 5-methylam The sequences include any known base analogues of DNA and RNA, such as nomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylkeosin, 5'-methoxycarbonylmethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, oxybutoxosin, pseudouracil, keuosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, N-uracil-5-oxyacetate methyl ester, uracil-5-oxyacetic acid, pseudouracil, keuosin, 2-thiocytosine, and 2,6-diaminopurine.

[0054] The term “gene” refers to a nucleic acid (e.g., DNA) sequence containing a coding sequence for the production of polypeptides, precursors, or RNA (e.g., rRNA, tRNA, sRNA, microRNA, lincRNA). Polypeptides may be coded by the full-length coding sequence or by any portion of the coding sequence, as long as the desired activity or functional properties (e.g., enzymatic activity, ligand binding, signaling, immunogenicity, etc.) of the full-length or fragment are preserved. The term also encompasses the coding region of a structural gene and sequences located adjacent to the coding region at both the 5' and 3' ends, at a distance of approximately 1 kb or more at either end, such that the gene corresponds to the length of full-length mRNA. Sequences located on the 5' side of the coding region and present on mRNA are referred to as the 5' untranslated sequence. Sequences located on the 3' side or downstream of the coding region and present on mRNA are referred to as the 3' untranslated sequence. The term “gene” encompasses both the cDNA and genomic forms of genes. The genome type or clone of a gene contains coding regions separated by non-coding sequences, which are called "introns," "intervening regions," or "intervening sequences." Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA), and they may contain regulatory elements such as enhancers. Introns are removed from the nucleus or primary transcript, or "removed by splicing," and therefore are not present in messenger RNA (mRNA) transcripts. mRNA functions during translation to determine the sequence or order of amino acids in nascent polypeptides.

[0055] As used herein, the term “heterogene” refers to a gene that does not exist in its natural environment. For example, heterogenes include genes introduced from one species to another. Heterogenes also include genes specific to an organism that have been modified in some way (e.g., mutation, addition in multiple copies, linking to a non-natural regulatory sequence). Heterogenes are distinguished from endogenous genes in that heterogene sequences are typically linked to DNA sequences that do not naturally associate with gene sequences within a chromosome, or to parts of a chromosome not found in nature (e.g., genes expressed at loci where genes do not normally express themselves).

[0056] As used herein, the term “oligonucleotide” refers to a short, single-stranded polynucleotide chain. Oligonucleotides are typically less than approximately 300 residues in length (e.g., 15–100), but as used herein, the term is intended to encompass longer polynucleotide chains as well. Oligonucleotides are often named according to their length. For example, a 24-residue oligonucleotide is called a “24-mer.” Oligonucleotides can form secondary and tertiary structures by self-hybridization or hybridization to other polynucleotides. Such structures include, but are not limited to, double-stranded, hairpin, cruciate, bent, and triple-stranded structures.

[0057] As used herein, "peptide" and "polypeptide" refer to polymer compounds of two or more amino acids linked via a main chain by a peptide amide bond (-C(O)NH-), unless otherwise specified. The term "peptide" typically refers to a short amino acid polymer (e.g., a chain with fewer than 25 amino acids), while the term "polypeptide" typically refers to a longer amino acid polymer (e.g., a chain with more than 25 amino acids).

[0058] As used herein, the term “fragment” refers to a peptide or polypeptide resulting from the excision or “fragmentation” of a larger entity (e.g., a protein, polypeptide, enzyme, etc.), or a peptide or polypeptide prepared to have the same sequence as such. Therefore, a fragment is a partial sequence of the whole entity (e.g., a protein, polypeptide, enzyme, etc.) from which it is made and / or designed. A peptide or polypeptide that is not a partial sequence of an existing whole protein is not a fragment (e.g., not a fragment of an existing protein).

[0059] As used herein, the term “sequence identity” refers to the degree to which two polymer sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have the same contiguous composition of monomeric subunits. The term “sequence similarity” refers to the degree to which two polymer sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have similar polymer sequences. For example, similar amino acids are those that share the same biophysical characteristics and can be grouped into families; for example, acidic (e.g., aspartic acid, glutamic acid), basic (e.g., lysine, arginine, histidine), nonpolar (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), and uncharged polar (e.g., glycine, asparagine, glutamine, cysteine, serine, threonine, tyrosine). "Percent sequence identity" (or "percent sequence similarity") is calculated by: (1) comparing two optimally aligned sequences within a comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, the specified window); (2) determining the number of positions containing identical (or similar) monomers (e.g., the same amino acid present in both sequences, similar amino acids present in both sequences) to obtain the number of matching positions; (3) dividing the number of matching positions by the total number of positions within the comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, the specified window); and (4) multiplying the result by 100 to obtain the sequence identity percentage or sequence similarity percentage. For example, if peptides A and B are both 20 amino acid lengths and have all identical amino acids except for one position, then peptides A and B have 95% sequence identity. If the amino acids at the non-identical positions share the same biophysical characteristics (e.g., both are acidic), then peptides A and B will have 100% sequence similarity. As another example, if peptide C is 20 amino acids long and peptide D is 15 amino acids long, and 14 of the 15 amino acids in peptide D are identical to some of peptide C, then peptides C and D have 70% sequence identity, but peptide D has 93.3% sequence identity relative to the optimal comparison window of peptide C.For the purposes of calculating “percent sequence identity” (or “percent sequence similarity” in this specification), any gap in the aligned sequences is treated as a mismatch at that location.

[0060] In some embodiments, substitutions may be conservative amino acid substitutions. Examples of conservative amino acid substitutions that are unlikely to affect biological activity include: serine to alanine, isoleucine to valine, glutamic acid to aspartic acid, serine to threonine, glycine to alanine, threonine to alanine, asparagine to serine, valine to alanine, glycine to serine, phenylalanine to tyrosine, proline to alanine, arginine to lysine, asparagine to aspartic acid, isoleucine to leucine, valine to leucine, glutamic acid to alanine, glycine to aspartic acid, and the reverse. See, for example, Neurath et al., The Proteins, Academic Press, New York (1979). The relevant parts of that work are incorporated herein by reference. Furthermore, the substitution of one amino acid within a group with another amino acid within the same group is a conservative substitution, and the groups are as follows: (1) alanine, valine, leucine, isoleucine, methionine, norleucine, and phenylalanine; (2) histidine, arginine, lysine, glutamine, and asparagine; (3) aspartic acid and glutamic acid; (4) serine, threonine, alanine, tyrosine, phenylalanine, tryptophan, and cysteine; (5) glycine, proline, and alanine.

[0061] The terms "homology" and "homological" refer to the degree of identity. Partial homology or complete homology may exist. A partially homologous sequence is a sequence that is less than 100% identical to another sequence.

[0062] As used herein, the terms “complementary” or “complementarity” are used in reference to polynucleotides (e.g., sequences of nucleotides such as oligonucleotides or target nucleic acids) that are related by base pairing rules. For example, the sequence “5'-AGT-3'” is complementary to the sequence “3'-TCA-5'”. Complementarity may be “partial,” in which case only some of the nucleic acid bases match according to base pairing rules. Alternatively, “complete” or “whole” complementarity may exist between nucleic acids. The degree of complementarity between nucleic acid chains has a significant impact on the efficiency and strength of hybridization between nucleic acid chains. Complementarity is particularly important in amplification reactions and detection methods that depend on binding between nucleic acids. Both terms may be used in reference to individual nucleotides, particularly in the context of polynucleotides. For example, a particular nucleotide in an oligonucleotide may be noted for its complementarity or lack thereof to a nucleotide in another nucleic acid chain, in contrast to or in comparison to the complementarity between the rest of the oligonucleotide and the nucleic acid chain.

[0063] In some contexts, the term “complementarity” and related terms (e.g., “complement,” “complementary,” “complementary”) refer to nucleotides of a nucleic acid sequence that can bond to another nucleic acid sequence via hydrogen bonding, for example, nucleotides capable of base pairing by Watson-Crick base pairing or other base pairing. Nucleotides that can form complementary base pairs, for example, are cytosine and guanine, thymine and adenine, adenine and uracil, and guanine and uracil. Percent complementarity does not need to be calculated over the entire length of the nucleic acid sequence. Percent complementarity may be limited to a specific region of the nucleic acid sequence that is a base pair, for example, starting with the nucleotide of the first base pair and ending with the nucleotide of the last base pair. As used herein, a complement of a nucleic acid sequence refers to an oligonucleotide in “antiparallel association” when aligned with a nucleic acid sequence such that the 5' end of one sequence pairs with the 3' end of the other sequence. Certain bases not commonly found in natural nucleic acids may be included in the nucleic acids of this disclosure, for example, inosine and 7-deazaguanine. Complementarity does not need to be perfect, and a stable double helix may contain mismatched base pairs or mismatched bases. Those skilled in nucleic acid technology can empirically determine the stability of a double helix by considering several variables, such as the length of the oligonucleotide, the base composition and sequence of the oligonucleotide, the ionic strength, and the occurrence rate of mismatched base pairs.

[0064] Therefore, in some embodiments, “complementarity” means that in a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides, the first nucleic acid sequence is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical to the second nucleic acid sequence, or that the two sequences hybridize under stringent hybridization conditions. “Fully complementary” means that each nucleic acid base of the first nucleic acid can be paired with each other at the corresponding position of the second nucleic acid. For example, in a particular embodiment, each nucleic acid base of an oligonucleotide complementary to a given nucleic acid has a nucleic acid base sequence identical to that of the nucleic acid complement in a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleic acid bases.

[0065] As used herein, “double-stranded nucleic acid” may be a portion of a nucleic acid, a region of a long nucleic acid, or the entire nucleic acid. “Double-stranded nucleic acid” may include, for example, double-stranded DNA, double-stranded RNA, double-stranded DNA / RNA hybrids, etc. Single-stranded nucleic acids having a secondary structure (e.g., base-paired secondary structure) and / or higher-order structures include “double-stranded nucleic acid”. For example, a triple structure is considered “double-stranded”. In some embodiments, any base-paired nucleic acid is a “double-stranded nucleic acid”.

[0066] The term “isolated,” when used in reference to nucleic acids, such as “isolated oligonucleotide” or “isolated polynucleotide,” refers to a nucleic acid sequence that is identified and isolated from at least one component or contaminant that is normally associated with its natural source. Isolated nucleic acids thus exist in a form or context different from that which they are found naturally. In contrast, nucleic acids that are not isolated as nucleic acids, such as DNA and RNA, are found in the state in which they are found naturally. For example, a given DNA sequence (e.g., a gene) is found on a host cell chromosome adjacent to neighboring genes, and an RNA sequence, such as a specific mRNA sequence encoding a particular protein, is found in a cell as a mixture with many other mRNAs encoding many proteins. However, an isolated nucleic acid encoding a given protein includes, for example, nucleic acids in a cell that normally expresses a given protein, where the nucleic acid is located in a different chromosomal location than that of a natural cell, or is otherwise adjacent to nucleic acid sequences that are different from those found naturally. Isolated nucleic acids, oligonucleotides, or polynucleotides can exist in single-stranded or double-stranded form. When isolated nucleic acids, oligonucleotides, or polynucleotides are used to express proteins, these oligonucleotides or polynucleotides may contain a minimum sense strand or coding strand (i.e., the oligonucleotide or polynucleotide may be single-stranded), but may also contain both a sense strand and an antisense strand (i.e., the oligonucleotide or polynucleotide may be double-stranded).

[0067] As used herein, the terms “purified” or “for purification” also refer to the removal of components (e.g., contaminants) from a sample. For example, antibodies are purified by removing contaminating non-immunoglobulin proteins. They are also purified by removing immunoglobulins that do not bind to the target molecule. Removing non-immunoglobulin proteins and / or immunoglobulins that do not bind to the target molecule will increase the percentage of target-reactive immunoglobulins in the sample. In another example, recombinant polypeptides are expressed in bacterial host cells, and these polypeptides are purified by removing host cell proteins, thereby increasing the percentage of recombinant polypeptides in the sample.

[0068] Preferred methods and materials will be described below, but similar or equivalent methods and materials described herein may also be used in the implementation or testing of this disclosure. All publications, patent applications, patents, and other references referenced herein are incorporated in their entirety by reference. The materials, methods, and examples disclosed herein are illustrative and not intended to be limiting.

[0069] 2. Recombinant β-hexosyl-transferase (rBHT) polypeptide Embodiments of this disclosure provide compositions and methods related to the production of human milk oligosaccharides (HMOs). In particular, this disclosure provides compositions and methods for converting lactose and N-acetylglucosamine (GlcNAc) into N-acetyllactosamine (LacNAc)-rich galactooligosaccharide (GOS) compositions using a novel β-hexosyltransferase (BHT) enzyme.

[0070] As will be recognized by those skilled in the art based on this disclosure, recombinant rBHT protein, or rBHT protein, comprises the full-length rBHT protein and any fragments and / or variants thereof, including proteins encoded by naturally occurring allele variants of the rBHT gene, as well as recombinant-produced rBHT proteins which may include some sequence changes compared to naturally occurring rBHT proteins. Recombinant proteins can be proteins resulting from a genetic engineering process that generally involves the use of nucleic acid molecules and corresponding recombinant nucleic acid molecules encoding peptides that are inserted into engineered host cells to express the corresponding peptides. That is, host cells are transfected, transformed, or transduced with recombinant polynucleotide molecules, thereby modifying the cells to express a desired polypeptide (e.g., rBHT).

[0071] According to these embodiments, the disclosure includes a functional recombinant β-hexosyl-transferase (rBHT) polypeptide that has at least 90% sequence identity with SEQ ID NO: 1 and comprises at least one N-terminal shortening of an amino acid relative to SEQ ID NO: 1. In some embodiments, the polypeptide has at least 95% sequence identity with SEQ ID NO: 1. In some embodiments, the polypeptide further comprises at least one additional amino acid substitution.

[0072] In some embodiments, the polypeptide includes an N-terminal shortening of about 1 to about 81 amino acids in length. In some embodiments, the N-terminal shortening is about 1 to about 56 amino acids in length. In some embodiments, the polypeptide is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, Includes N-terminal shortenings with amino acid lengths of 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 83, 74, 75, 76, 77, 78, 79, 80, or 81.

[0073] In some embodiments, the polypeptide contains at least 90% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 91% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 92% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 93% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 94% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 95% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 96% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 97% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide contains at least 98% sequence identity with any of SEQ ID NOs: 3, 5, 7, and 9. In some embodiments, the polypeptide has at least 99% sequence identity with any of sequence numbers 3, 5, 7, and 9.

[0074] As will be recognized by those skilled in the art based on this disclosure, soluble secretory proteins and proteins expressed on the cell surface may contain an N-terminal signal sequence, which is generally a hydrophobic sequence that mediates the insertion of the protein across the endoplasmic reticulum (ER) membrane in eukaryotic cells. Type I transmembrane proteins also contain a signal sequence. The signal sequences used herein may include an amino-terminal hydrophobic sequence that is generally enzymatically removed after the insertion of part or all of the protein into the lumen of the ER membrane. Thus, the signal sequence may be present as part of the precursor form of a secretory or transmembrane protein, but is generally absent in the mature form of the protein. When a protein is said to contain a signal sequence, it should be understood that the precursor form of the protein is likely to contain a signal sequence, while the mature form of the protein is likely not. The signal sequence may include a residue immediately upstream of the cleavage site (position 1), which is important for this enzymatic cleavage, and another residue at position 3. (See, for example, Nielsen et al. 1997 Protein Eng 10(1) 1-6; von Heijne 1983 Eur J Biochem 133(1) 7-21; von Heijne 1985 J Mol Biol 184 99-105, which describe signal sequences and methods for their identification.) In some embodiments, the rBHT polypeptides of this disclosure may be soluble or membrane-bound. In some embodiments, about 1% to about 50% of the polypeptide is soluble. In some embodiments, about 1% to about 45% of the polypeptide is soluble. In some embodiments, about 1% to about 40% of the polypeptide is soluble. In some embodiments, about 1% to about 35% of the polypeptide is soluble. In some embodiments, about 1% to about 30% of the polypeptide is soluble. In some embodiments, about 1% to about 25% of the polypeptide is soluble. In some embodiments, about 1% to about 20% of the polypeptide is soluble. In some embodiments, about 1% to about 15% of the polypeptide is soluble. In some embodiments, about 1% to about 10% of the polypeptide is soluble.

[0075] In accordance with embodiments of this disclosure, any signal peptide(s) or signal sequence(s), including signal sequences derived from peptides(s) or polypeptides(s) derived from prokaryotes, eukaryotes, fungi, mammals, insects, yeasts, or plants, can be included in the rBHT polypeptide of this disclosure. In some embodiments, but not limited to, signal sequences(s) can be included in the rBHT polypeptide of this disclosure, include those described in Ahmad, M., et. Al., (2014) "Protein expression in Komagataella (Pichia) pastoris: recent achievements and perspectives for heterologous protein production" Applied Microbiology and Biotechnology 98(12):5301-5317.

[0076] In some embodiments, the rBHT polypeptide of the Disclosure comprises a signal sequence that is non-native or exogenous with respect to a host cell engineered to express the rBHT polypeptide. In some embodiments, the rBHT polypeptide of the Disclosure comprises a signal sequence that is native or endogenous with respect to a host cell engineered to express the rBHT polypeptide. In any case, the signal sequence may be in its native form / sequence, or may be truncated and / or may contain at least one amino acid substitution with respect to its native form / sequence.

[0077] In some embodiments, the signal sequence includes an amino acid sequence derived from a yeast protein. In some embodiments, the signal sequence includes an amino acid sequence derived from a protein derived from one of Komagataella (Pichia) pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Hansenula (Ogataea) polymorpha, or Kluyveromyces lactis. In some embodiments, the signal sequence includes a polypeptide having at least 90% sequence identity with at least one of the following: α-conjugation factor signal sequence (MFα) (SEQ ID NO: 29), invertase (IV) signal sequence (SEQ ID NO: 30), glucoamylase (GA) signal sequence (SEQ ID NO: 31), or inulinase (IN) signal sequence (SEQ ID NO: 32) derived from Saccharomyces cerevisiae. In some embodiments, the polypeptide has at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% sequence identity with any of SEQ ID NOs: 29, 30, 31, or 32. In some embodiments, the polypeptide has at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% sequence identity with any of SEQ ID NOs: 29. In some embodiments, the polypeptide has at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% sequence identity with any of SEQ ID NOs: 30. In some embodiments, the polypeptide has at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% sequence identity with any of the sequence numbers 31. In some embodiments, the polypeptide has at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% sequence identity with any of the sequence numbers 32.

[0078] As further described herein, the rBHT polypeptides of this disclosure include a signal sequence (or a functional fragment thereof) from any of SEQ ID NOs: 29, 30, 31, or 32. According to this embodiment, the rBHT polypeptide can contain at least 90% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide contains at least 91% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 92% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 93% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 94% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 95% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 96% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 97% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70.In some embodiments, the polypeptide has at least 98% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. In some embodiments, the polypeptide has at least 99% sequence identity with any of SEQ ID NOs: 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70.

[0079] The rBHT polypeptides of this disclosure may be glycosylated to varying degrees or not glycosylated at all. For example, the rBHT polypeptides of this disclosure may contain one or more N-linked or O-linked glycosylation sites, in addition to those already found in proteins or polypeptides including any of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70. Those skilled in the art will recognize, based on this disclosure, that asparagine residues that are part of the sequence Asn Xxx Ser / Thr (where Xxx is any amino acid other than proline) can function as N-glycan attachment sites. Furthermore, there are serine and threonine residues that can function as O-linked glycosylation sites. Glycosylation can extend the in vivo half-life or modify the biological activity. Variants of rBHT proteins also include proteins containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 more N-linked and / or O-linked glycosylation sites than present in the corresponding wild-type protein or polypeptide, provided that the resulting protein or polypeptide maintains its function as a glycosyl hydrolase and β-hexosyltransferase. Variant rBHT polypeptides also include those having 1, 2, 3, 4, or 5 fewer N-linked and / or O-linked glycosylation sites than present in the corresponding wild-type protein or polypeptide, provided that the resulting protein or polypeptide maintains its function as a glycosyl hydrolase and β-hexosyltransferase. In some embodiments, the rBHT polypeptides of this disclosure contain at least one asparagine residue at positions 289, 297, 431, and 569 of SEQ ID NO: 1. In some embodiments, the rBHT polypeptide of the present disclosure contains at least two asparagine residues at positions 289, 297, 431, and 569 relative to SEQ ID NO: 1. In some embodiments, the rBHT polypeptide of the present disclosure contains at least three asparagine residues at positions 289, 297, 431, and 569 relative to SEQ ID NO: 1.In some embodiments, the rBHT polypeptide of the present disclosure contains asparagine residues at positions 289, 297, 431, and 569 relative to SEQ ID NO: 1.

[0080] Embodiments of the present disclosure include secreted soluble variants of the rBHT polypeptide described herein, as well as variants containing a transmembrane domain that can be expressed on the cell surface. Such proteins can be isolated as part of a purified protein preparation in which the rBHT polypeptide constitutes at least 80% or at least 90% of the proteins present in the preparation. The rBHT polypeptides of the present disclosure include proteins and polypeptides comprising the amino acid sequences shown in SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70, as well as fragments, derivatives, and variants thereof, such as fusion proteins.

[0081] The rBHT polypeptides of this disclosure may be fusion proteins comprising at least one rBHT polypeptide and may include an amino acid sequence and / or fragment that is any of the variants and / or fragments of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70 (as listed above), and at least one other moiety. The other moiety may also be a non-protein moiety, e.g., a polyethylene glycol (PEG) moiety, or a cytotoxic, cell proliferation inhibitory, luminescent, and / or radioactive moiety. PEG adhesion has been shown to extend the in vivo half-life of at least some proteins. Furthermore, the cytotoxic, cell proliferation inhibitory, luminescent, and / or radioactive moieties have been fused to antibodies for diagnostic or therapeutic purposes. Various polypeptides other than rBHT polypeptides (or fragments thereof) can be fused to rBHT polypeptides for a variety of purposes, such as extending the in vivo half-life of a protein, facilitating the identification, isolation, and / or purification of a protein, increasing the activity of a protein, and promoting the oligomerization of a protein.

[0082] Many polypeptides can facilitate the identification and / or purification of recombinant fusion proteins in which they are part. Examples include polyarginine, polyhistidine, or HAT® (Clontech), a natural sequence of non-adjacent histidine residues with high affinity for immobilized metal ions. rBHT proteins containing these polypeptides can be purified by affinity chromatography using, for example, TALON® resin (Clontech) containing immobilized nickel or immobilized cobalt ton. See, for example, Knol et al. 1996J Biol Chem 27(26) 15358-15366. Polyarginine-containing polypeptides can be effectively purified by ion-exchange chromatography. Other useful polypeptides include, for example, antigen-identifying peptides described in U.S. Patent No. 5,011,912 and Hopp et al. 1988 Bio / Technology 6 1204. One such peptide is the FLAG® peptide. This provides a highly antigenic epitope that is reversibly bound by a specific monoclonal antibody, enabling rapid assay and easy purification of the expressed recombinant fusion protein. The mouse hybridoma, named 4E11, produces a monoclonal antibody that binds to the FLAG peptide in the presence of a specific divalent metal cation, as described in U.S. Patent No. 5,011,912. The 4E11 hybridoma cell line is deposited in the American Type Culture Collection under depositary number HB9259. The monoclonal antibody that binds to the FLAG peptide can be used as an affinity reagent to recover polypeptide purification reagents containing the FLAG peptide.Other suitable protein tags and affinity reagents include those described in the GST-Bind® system (Novagen), which utilizes the affinity of glutathione-S-transferase fusion protein to immobilized glutathione; those described in the T7-TAG® affinity purification kit, which utilizes the affinity of the amino-terminal 11 amino acids of the T7 gene 10 protein to monoclonal antibodies; or those described in the STREP-TAG® system (Novagen), which utilizes the affinity of a modified form of streptavidin to protein tags. Some of the protein tags described above, like others, are described in Sassenfeld 1990 TIBTECH8:88-93, Brewer et al., Purification and Analysis of Recombinant Proteins, pp.239-266, Seetharam and Sharma (eds.), Marcel Dekker, Inc. (1991), and Brewer and Sassenfeld, Protein Purification Applications, pp.91-111, Harris and Angal (eds.), Press, Inc., Oxford England (1990). The portions of these references describing the protein tags are incorporated herein by reference. Furthermore, fusions of two or more of the tags described herein, such as the fusion of the FLAG tag and the polyhistidine tag, can be fused to the rBHT polypeptide of this disclosure.

[0083] In some embodiments, the rBHT polypeptides of this disclosure also include affinity tags that can be used as part of the means for producing the polypeptide. In addition to the 6X-HIS tags further described herein, various purification methods may be used, such as antigen tags (e.g., FLAG (Sigma-Aldrich, Hopp et al. 1988 Nat Biotech 6:1204-1210), hemagglutanin (HA) (Wilson et al., 1984 Cell 37:767), intein fusion expression system (New England Biolabs, USA) Chong et al. 1997 Gene 192(2), 271-281, or maltose-binding protein (MBP)), glutathione S-transferase (GST) / glutathione, polyHis / Ni or Co (Gentz ​​et al., 1989 PNAS USA 86:821-824). Fusion proteins containing a GST tag at the N-terminus of the protein are also described in U.S. Patent No. 5,654,176 (Smith). Magnetic separation techniques such as Strepavidin-DynaBeads® (Life Technologies, USA) can also be used. Alternatively, an optically cleavable linker may be used (e.g., Patent No. 7,595,198 (Olejnik & Rothchild)). Many other systems are known in the art and are suitable for use in embodiments of this disclosure.

[0084] 3. Nucleic acid constructs Embodiments of this disclosure also include nucleic acid molecules encoding any of the rBHT polypeptides described herein. Embodiments of this disclosure also include vectors containing any one of these nucleic acid molecules. In some embodiments, isolated nucleic acids, such as DNA and RNA molecules, include polypeptides encoding the rBHT polypeptides described herein and comprising the amino acid sequence of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70, and fragments and / or variants thereof. In some embodiments, these nucleic acids are useful for producing recombinant proteins having glycosyl hydrolase and β-hexosyl-transferase activity. Such nucleic acids may be modified genomic DNA or cDNA. In some cases, the nucleic acid may include an uninterrupted open reading frame encoding the rBHT protein. The nucleic acid molecules of this disclosure include DNA and RNA in both single-stranded and double-stranded forms, as well as their corresponding complementary sequences. In the case of nucleic acids isolated from naturally occurring sources, isolated nucleic acids are those isolated from adjacent gene sequences present in the genome of the organism from which the nucleic acid was isolated. In the case of chemically synthesized nucleic acids, such as oligonucleotides, or enzymatically synthesized nucleic acids from a template, such as polymerase chain reaction (PCR) products or cDNA, the nucleic acids resulting from such processes are understood to be isolated nucleic acids. An isolated nucleic acid molecule refers to a nucleic acid molecule in the form of a distinct fragment or as a component of a larger nucleic acid construct.

[0085] This disclosure includes nucleic acids or fragments thereof containing the sequences of SEQ ID NOs: 2, 4, 6, 7, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, and 69, or nucleic acids that hybridize to nucleic acids containing the nucleotide sequences of SEQ ID NOs: 2, 4, 6, 7, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, and 69 under moderately stringent conditions, and optionally under more highly stringent conditions, and full-length rBHT The cDNA contains a nucleotide sequence (SEQ ID NO: 1), and the nucleic acid encodes a protein that can act as a glycosyl hydrolase and a β-hexosyltransferase. Hybridization techniques are well known in the art and are described in Sambrook, J., E.F. Fritsch, and T. Maniatis (Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, chapters 9 and 11, 1989) and Current Protocols in Molecular Biology (F.M. Ausubel et al., eds., John Wiley & Sons, Inc., sections 2.10 and 6.3-6.4, 1995).

[0086] 4. Production method Embodiments of the present disclosure include methods for producing a composition containing GOS from lactose in a host cell ("GOS composition") using any of the rBHT polypeptides described herein. As further described herein, the rBHT polypeptides of the present disclosure are functional in that they exhibit the ability to produce any GOS composition (plural) from lactose (but not limited to GOS with or without GlcNAc, and GOS compositions rich in LacNAc) by catalyzing the hydrolysis of β-(1-4) glycoside linkages. As will be recognized by those skilled in the art based on this disclosure, GOS generally refers to galactose-containing polysaccharides having two or more sugar units; for example, Gal-Gal or [Gal]n-Glc (1 ≤ n ≤ 8), e.g., β-D-Gal(1 → 4)-β-D-Gal(1 → 4)-β-D-Glc, β-D-Gal(1 → 4)-β-D-Gal(1 → 4)-β-D-Gal(1 → 4)-β-D-Gal(1 → 4)-β-D-Glc.

[0087] In some embodiments, the GOS produced using the rBHT polypeptide of this disclosure comprises one or more N-acetyllactosamine (LacNAc) units. In one embodiment, the GOS can be produced by incubating a host cell expressing the rBHT polypeptide in a medium containing a disaccharide substrate, such as lactose. In one embodiment, the GOS is produced from lactose concurrently with a glucose elimination system. The glucose elimination system may be a generally recognized safe (GRAS) organism. In some embodiments, the host cell is one or more of yeast cells, fungal cells, mammalian cells, insect cells, plant cells, or algal cells. In some embodiments, the host cell includes one or more cells derived from Komagataella (Pichia) pastoris (also known as Kamagataaella phaffi), Saccharomyces cerevisiae, Yarrowia lipolytica, Hansenula (Ogataea) polymorpha, or Kluyveromyces lactis, Aspergillus spp., and Trichoderma reesei. In some embodiments, the host cells include any cells of the genus Komagataella. In some embodiments, GOS contains N-acetyllactosamine (LacNAc). In some embodiments, the method yields a LacNAc-rich GOS at least 10% of the initial lactose concentration and a total GOS concentration of at least 50% of the initial lactose concentration. In some embodiments, the method yields a LacNAc-rich GOS at least 10% of the initial lactose concentration and a total GOS concentration of at least 60% of the initial lactose concentration. In some embodiments, the method yields a LacNAc-rich GOS at least 10% of the initial lactose concentration and a total GOS concentration of at least 70% of the initial lactose concentration. In some embodiments, the method yields a LacNAc-rich GOS at least 10% of the initial lactose concentration and a total GOS concentration of at least 75% of the initial lactose concentration.For example, using an initial lactose-to-GlcNAc ratio of 1:8, the method provided herein, using approximately 200 g of lactose and an rBHT polypeptide (e.g., a total cell membrane binding enzyme) having approximately 25 g of GlcNAc, produces approximately 25 g of LacNAc and approximately 100 g of GOS. The initial lactose-to-GlcNAc ratio is in the range of approximately 1:20 to approximately 20:1.

[0088] In some embodiments, the rBHT polypeptide of this disclosure is useful for producing LacNAc and related compositions. Prebiotic LacNAc is considered one of the most important building blocks for producing higher-order human milk oligosaccharides (HMOs). However, the biocatalytic activity of LacNAc is advantageous because viable industrial production routes by chemical synthesis suffer from low yields. As will be further described herein, the main differences between the biological synthesis of LacNAc using the enzyme BHT and other biosynthetic routes are low cost and high purity. Embodiments of this disclosure demonstrate that LacNAc production by the rBHT polypeptide described herein is more suitable for industrial scale compared to other processes. As shown in the example in Figure 5, LacNAc is produced by mixing lactose and GlcNAc with the rBHT polypeptide. The results of this disclosure demonstrate a yield of at least about 25 g / L of LacNAc from about 25 g / L of GlcNAc and about 200 g / L of lactose in a single synthesis step, with tin initially present in the reaction mixture when the reaction ratio of lactose to GlcNAc was about 1:8.

[0089] In some embodiments, the rBHT polypeptides of this disclosure can be used to produce GOS compositions that do not contain N-acetylglucosamine (GlcNAc). Embodiments of this disclosure include materials and methods for producing GOS compositions that do not contain GlcNAc, which include reacting an rBHT polypeptide having the amino acid sequence provided herein with lactose under preferred conditions to produce GOS. Similar compositions and methods are described in related U.S. Patents 10,513,695 and 9,783,789, both of which are incorporated herein by reference.

[0090] The rBHT polypeptides of this disclosure can be prepared using various means known in the art. For example, the nucleic acid molecule encoding the rBHT polypeptide described herein can be introduced into a vector that can be introduced into a host cell. Vectors containing nucleic acids encoding the rBHT polypeptide and host cells are included in embodiments of this disclosure. Host cells containing nucleic acids encoding the rBHT polypeptide can be cultured under conditions that enable the expression of the rBHT polypeptide. The expressed rBHT polypeptide can then be obtained from the culture medium in which the cells are cultured, or from the cells themselves, and purified by any of the many suitable means known in the art. Furthermore, genetic engineering methods for producing rBHT polypeptides include the expression of polynucleotide molecules in cell-free expression systems, cell hosts, tissues, and animal models by known methods.

[0091] The vector may include a selection marker and an origin of replication for replication in a host. The vector may further include a suitable transcriptional or translational regulatory sequence, such as one derived from a mammalian, microorganism, virus, or insect gene, operably ligated to the nucleic acid encoding the rBHT polypeptide. Examples of such regulatory sequences include transcription promoters, operators, or enhancers, mRNA ribosome binding sites, and appropriate sequences that control transcription and translation. A nucleotide sequence is operably ligated if the regulatory sequence is functionally relevant to the DNA encoding the target protein. Thus, if a promoter nucleotide sequence directs the transcription of the rBHT protein-coding sequence, the promoter nucleotide sequence is operably ligated to the rBHT polypeptide sequence. If the rBHT polypeptide is a fusion protein, a nucleic acid sequence encoding part of the fusion protein, such as a signal sequence, may be part of the vector, and the nucleic acid encoding the rBHT polypeptide can be inserted into such a vector, thereby encoding a protein containing the added signal sequence and rBHT polypeptide.

[0092] Suitable host cells for rBHT polypeptide expression include prokaryotic cells, yeast cells, plant cells, insect cells, and higher eukaryotic cells. The regulatory sequences within the vector are selected to be operable within the host cells. Suitable prokaryotic host cells include bacteria of the genera Escherichia, Bacillus, and Salmonella, as well as members of the genera Pseudomonas, Streptomyces, and Staphylococcus. For expression in prokaryotic cells, e.g., E. coli, the polynucleotide molecule encoding the rBHT polypeptide includes an N-terminal methionine residue to facilitate recombinant polypeptide expression. The N-terminal methionine can optionally be cleaved from the expressed polypeptide. Suitable yeast host cells, but not limited to these, include cells of genera such as Saccharomyces, Pichia (Komagataella), and Kluyveromyces. In some embodiments, the host cells include any cells derived from the genus Pichia (Komagataella). Preferred yeast hosts are S. cerevisiae and P. pastoris (also known as Kamagataaella phaffi). Suitable systems for expression in insect host cells are described, for example, in the overview by Luckow and Summers (1988 BioTechnology 6 47-55), the relevant portions of which are incorporated herein by reference. Suitable mammalian host cells include COS-7 strain monkey kidney cells (Gluzman et al., 1981 Cell 23 175-182), baby hamster kidney (BHK) cells, Chinese hamster ovary (CHO) cells (Puck et al., 1958 PNAS USA 60 1275-1281), CV-1 (Fischer et al., 1970 Int J Cancer 5 21-27), human kidney-derived 293 cells (American Type Culture Collection (ATCC) catalog number CRL-10852), and human cervical cancer cells (HELA) (ATCC CCL2). Relevant portions of the references mentioned in this paragraph are incorporated herein by reference.

[0093] Expression vectors for use in cell hosts generally contain one or more phenotypic selection marker genes. Such genes encode, for example, proteins that confer antibiotic resistance or proteins that supply nutritional requirements. A wide variety of such vectors are readily available from commercial sources. Examples include pGEM vectors (Promega), pSPORT vectors, and pPROEX vectors (InVitrogen, Life Technologies, Carlsbad, Calif.), Bluescript vectors (Stratagene), and pQE vectors (Qiagen). Yeast vectors often contain a replication initiation sequence, an autonomous replication sequence (ARS), a promoter region, a sequence for polyadenylation, a sequence for transcription termination, and a selection marker gene derived from a yeast plasmid. Vectors that can replicate in both yeast and E. coli (called shuttle vectors) are also available. In addition to the above characteristics of yeast vectors, shuttle vectors also contain sequences for replication and selection in E. coli. Direct secretion of target polypeptides expressed in the yeast host can be achieved by including a nucleotide sequence encoding the yeast α-factor reader sequence at the 5' end of the rBHT-coding nucleotide sequence. Brake 1989 Biotechnology 13 269-280.

[0094] Examples of expression vectors suitable for use in mammalian host cells include pcD A3.1 / Hygro (Invitrogen), pDC409 (McMahan et al. 1991 EMBO J10: 2821-2832), and pSVL (Pharmacia Biotech). Expression vectors for use in mammalian host cells may contain transcriptional and translational regulatory sequences derived from the viral genome.

[0095] Commonly used promoter and enhancer sequences for expressing rBHT RNA include, but are not limited to, those derived from human cytomegalovirus (CMV), adenovirus 2, polyomavirus, and Simianvirus 40 (SV40). Methods for constructing mammalian expression vectors are described, for example, in Okayama and Berg (1982 Mol Cell Biol 2:161-170), Cosman et al. (1986 Mol Immunol 23:935-941), Cosman et al. (1984 Nature 312:768-771), EP-A-0367566, and WO91 / 18982. Relevant portions of these references are incorporated herein by reference. Furthermore, as will be recognized by those skilled in the art based on this disclosure, the reaction mixture may be used as the final product by any spray-drying, lyophilization, or other concentration method. If whole cells are used instead of pure enzyme, cell separation techniques may be required.

[0096] 5. Composition Embodiments of this disclosure include compositions comprising any of the polypeptides described herein and / or one or more GOS (e.g., GOS containing or not containing GlcNAc, and GOS rich in LacNAc) produced using any of the polypeptides described herein. In some embodiments, the composition is a food. In some embodiments, the food is, but is not limited to, infant formula, yogurt, dairy products, milk-based beverages, fruit beverages, hydration beverages, energy beverages, fruit preparations, and meal replacement beverages.

[0097] As will be recognized by those skilled in the art based on this disclosure, GOS compositions, such as GOS containing or not containing GlcNAc and GOS compositions rich in LacNAc, are widely used as prebiotic supplements in foods and beverages worldwide. These highly valuable non-digestible sugars can mimic human milk oligosaccharides (HMOs) by having a positive effect on the growth and metabolism of gastrointestinal (GI) bacteria (probiotics). Adding prebiotics to the diet has shown substantial improvements in the host's overall health by reducing gastrointestinal discomfort, managing the immune system, and decreasing pathogenic and opportunistic bacteria and viruses. Embodiments of this disclosure demonstrate novel materials and methods for developing prebiotics to produce LacNAc from pure lactose and GlcNAc, and to significantly increase the concentration of secreted soluble rBHT.

[0098] In some embodiments, the disclosure includes the production of a food or dietary supplement containing a LacNAc-rich GOS composition using rBHT protein or cells expressing rBHT. The food may be a dairy food such as yogurt, cheese, or fermented dairy products. rBHT or cells expressing rBHT may be partially added to the food or dietary supplement. rBHT can be dried using spray drying, which is a rapid and gentle method for obtaining the smallest amount of temperature-sensitive substance in powder form. Dried rBHT can also be encapsulated using the ability of a spray dryer to coat particles, immobilize solid materials in a matrix, and produce microcapsules (www.buchi.com / Mini_Spray_Dryer_B-290.179.0 DOT html). Other drug delivery applications using functional GRAS encapsulating agents and techniques may be used. Dried rBHT tablets and powder forms can be analyzed for the activity rate of rBHT after rehydration with lactose-containing buffers and dairy products.

[0099] Any rBHT polypeptide described herein may be delivered in the form of a composition, i.e., with one or more additional components, such as a physiologically acceptable carrier, excipient, or diluent. For example, in addition to the soluble rBHT polypeptide described herein, a composition may include buffers, antioxidants such as ascorbic acid, low molecular weight polypeptides (such as those having fewer than 10 amino acids), proteins, amino acids, carbohydrates, e.g., glucose, sucrose, or dextrin, chelating agents, e.g., EDTA, glutathione, and / or other stabilizers, excipients, and / or preservatives. The composition may be formulated as a liquid or lyophilized powder. Further examples of components that may be used in pharmaceutical formulations can be found in Remington's Pharmaceutical Sciences, 16 th This is shown in Ed., Mack Publishing Company, Easton, Pa., (1980), and the relevant portions thereof are incorporated herein by reference.

[0100] Compositions containing the therapeutic molecules described above can be administered by any suitable means, including but not limited to parenteral, topical, oral, nasal, vaginal, enteral, or pulmonary (by inhalation) administration. When administered by injection, the composition(s) may be administered intra-articular, intravenous, intra-arterial, intramuscular, intraperitoneal, or subcutaneously by bolus injection or serial infusion. Topical administration, i.e., at the site of disease, is intended, and transdermal delivery and sustained release from implants, skin patches, or suppositories are intended. Inhalation delivery includes, for example, nasal or oral inhalation, the use of a nebulizer, and inhalation in aerosol form. Administration by suppositories inserted into a body cavity can be achieved, for example, by inserting the composition in solid form into a selected body cavity and allowing it to dissolve. Other alternatives include eye drops, oral preparations such as pills, lozenges, syrups, and chewing gum, and topical preparations such as lotions, gels, sprays, and ointments. In most cases, therapeutic molecules, which are polypeptides, can be administered topically, by injection, or by inhalation.

[0101] The therapeutic molecules described above may be administered in any dose, frequency, and duration that may be effective in treating the condition being treated. The dose depends on the molecular properties of the therapeutic molecule and the nature of the disorder being treated. Treatment may be continued for as long as necessary to achieve the desired outcome. The periodicity of treatment may be constant or variable throughout the treatment period. For example, treatment may be initially given weekly and then every other week. Treatments having periods of several days, weeks, months, or years are included in embodiments of this disclosure. Treatment may be interrupted and then resumed.

[0102] Maintenance doses can be administered after initial treatment. The dosage is calculated as milligrams per kilogram of body weight (mg / kg) or milligrams per square meter of skin surface (mg / m²), regardless of height or weight. 2 ), or can be measured as a fixed dose. These are standard dose units in the art. Human skin surface area is calculated from height and weight using a standard formula. For example, therapeutic rBHT protein can be administered in doses of approximately 0.05 mg / kg to approximately 10 mg / kg, or approximately 0.1 mg / kg to approximately 1.0 mg / kg. Alternatively, doses of approximately 1 mg to approximately 500 mg can be administered. Or doses of approximately 5 mg, 10 mg, 15 mg, 20 mg, 25 mg, 30 mg, 35 mg, 40 mg, 45 mg, 50 mg, 55 mg, 60 mg, 100 mg, 200 mg, or 300 mg can be administered.

[0103] 6. Materials and Methods Strains and media. The strain GS115 (Invitrogen Life Technologies, Thermo Fisher Scientific) and the growth and maintenance of the media have been described previously. E. coli XL1-Blue was used as the cloning host (Agilent Technologies, Thermo Fisher Scientific). The plasmid pPIC9 (Invitrogen Life Technologies, Thermo Fisher Scientific) was used to construct an expression vector containing the codon-optimized Bht (rBht variant) (GenBank accession number JF29828).

[0104] Plasmid construction, expression, and purification of rBHT shortened variants. All molecular biology protocols were performed as previously described. Briefly, the plasmid constructed for the expression of the rBHT variant in K. pastoris encoding the shortened variant was generated by PCR amplification of the codon-optimized rBht open reading frame in pPIC9-MFα-rBht (1-594) -HIS using primers purchased from Integrated DNA Technologies (IDT Coralville, IA, USA) (listed in Table 5). The bacterial strains and K. pastoris strains used in this study are shown in Table 4. Bacteria were grown at 37 °C in Luria-Bertani (LB) medium containing the antibiotic ampicillin (100 μg / ml) (Thermo Fisher Scientific).

[0105] Mutagenesis and cloning. The plasmid encoding the shortened rBHT variant was used as a template with HotStar® Taq (Qiagen, Hilden, Germany) and pJB110 (pPIC9-MFα-rBht (1-594)-HIS) were generated by PCR amplification using primers purchased from Integrated DNA Technologies (IDT, Coralville, IA, USA). Restriction sites were included in the primers as appropriate to facilitate cloning (listed in Table 5). Briefly, primer pairs for sequences encoding truncated rBHT variants encoding amino acids 32 - 594 (primers: JBB21 / JBB5), 54 - 594 (primers: JBB22 / JBB5), 57 - 594 (primers: JBB23 / JBB5), 82 - 594 (primers: JBB24 / JBB5), 95 - 594 (primers: JBB25 / JBB5), and 103 - 594 (primers: JBB26 / JBB5). The amplicons were digested with XhoI-NotI and cloned into pPIC9 (Invitrogen Life Technologies, Thermo Fisher Scientific) to generate pJB123 (pPIC9-MFα-rBht (32-594) -HIS), pJB124 (pPIC9-MFα-rBht (54-594) -HIS), pJB125 (pPIC9-MFα-rBht (57-594) -HIS), pJB126 (pPIC9-MFα-rBht (82-594) -HIS), pJB127 (pPIC9-MFα-rBht (95-594) -HIS), and pJB128 (pPIC9-MFα-rBht (103-594) -HIS).

[0106] Plasmids encoding pJB134 (pPIC9-IV-rBht (54-594) -HIS), pJB135 (pPIC9-GA-rBht (54-594) -HIS), and pJB136 (pPIC9-IN-rBht (54-594) -HIS) were generated using pJB124 (pPIC9-MFα-rBht (54-594)The amplicon was generated using -HIS. The amplicon was digested with XhoI-NotI and cloned into pPIC9 (Invitrogen Life Technologies, Thermo Fisher Scientific).

[0107] Site-directed mutagenesis was performed using the QuickChange Site-Directed Mutagenesis Kit (Agilent Technologies Santa Clara, CA, USA) according to the manufacturer's instructions, using complementary oligonucleotides designed to incorporate the desired base change at the putative N-glycosylation site, (pJB112, pPIC9-MFα-rBht (23-594) Using -HIS) as a template, constructs containing single amino acid exchanges from asparagine to glutamine were generated using oligonucleotide primers with nucleotide substitutions (Table 5) (N289Q (primer: JBB27 / JBB28), N297Q (primer: JBB29 / JBB30), N431Q (primer: JBB31 / JBB32), and N569Q (primer: JBB33 / JBB34)).

[0108] Site-directed mutagenesis was also used to remove amino acids 57-70 from MFα using the primer set JBB35 / JBB36 (Table 5), resulting in pJB133(pPIC9-MFα(Δ57-70)-rBht (23-594) -HIS) and pJB137(pPIC9-MFα (Δ57-70) -rBht (57-594) -HIS) (Table 4) was generated. DNA fragments from restriction enzyme digests were purified from agarose gels using the QIAquick gel extraction kit (Qiagen, Hilden, Germany). All mutations were confirmed by restriction digestion to detect restriction sites in primers and by Sanger sequencing performed by the NC State University Genomic Sciences Laboratory (Raleigh, NC, USA) using primers JBB3, JBB4, 5'AOX1, 3'AOX1, and α factor (Table 1).

[0109] Transformation and expression of K. pastoris. K. pastoris was transformed with a linearized plasmid according to the Invitrogen Pichia Expression Kit manual (Invitrogen, USA). Plasmid integration and muting of histidine-positive colonies. + The phenotype was confirmed by sequencing of PCR products generated using primers 5'AOX1 and 3'AOX1 (Invitrogen Pichia expression kit). As previously mentioned, single-copy integration was confirmed. Expression and purification were previously described. Briefly, the filtered culture medium was purified using an AKTApurifier and a HISTrap® HP Nickel column (GE Healthcare, Life Sciences). The purified protein was quantified by the Bradford protein assay (Thermo Fisher Scientific).

[0110] SDS-PAGE and Western immunoblotting analysis. Proteins were analyzed by SDS-PAGE using a 10% degraded gel and visualized with Coomassie and silver staining (Bio-Rad, Hercules, CA). Immunoblotting was performed by probing with a 1:10,000 dilution of anti-HIS antibody (GenScript, Piscataway, NJ) followed by a 1:10,000 dilution of alkaline phosphatase-conjugated goat anti-mouse antibody (GenScript, Piscataway, NJ). Detection was performed with 1-Step® NBT / BCIP substrate solution according to the manufacturer's instructions (Thermo Fisher Scientific).

[0111] Enzyme assay. ONP-Glu activity was measured using the previously described method (see, for example, Dagher, SF, and Bruno-Barcena, JM (2016) A novel N-terminal region of the membrane β-hexosyltransferase: its role in secretion of soluble protein by Pichia pastoris. Microbiology 162, 23-34).

[0112] Sequence analysis. Alignment was generated using the ClustalX algorithm (http: / / www.clustal.org / ) and the Jalview algorithm. Using NCBI blastp, the sequences of the top five homologous proteins were selected (https: / / blast.ncbi.nlm.nih.gov / ): Glycoside hydrolase family 1 protein, Glycoside hydrolase family 1 protein [Sphaerobolus stellatus SS14], accession number BAD95570.1, Glycoside hydrolase [Violaceomyces palustris], accession number KIJ57308.1, Glycoside hydrolase [Violaceomyces palustris], accession number PWN48553.1, Virtual protein PFL1_06098 [Anthracocystis flocculosa PF-1], accession number XP_007881827.1, Glycoside hydrolase [Testicularia cyperi], accession number PWZ03736.1 and Glycoside hydrolase family 1 protein [Gymnopus luxurians [FD-317 M1] Accession number KIK57390.1.

[0113] Secondary structure prediction. Consensus prediction of BHT's secondary structure was performed using the PSIPRED server (protein structure prediction) and NPS@server (network protein sequence analysis). Signal sequences were predicted using the SignalP 5.0 algorithm. Protein dysfunction was predicted using a consensus of six methods: Dispred3, Phyre2, IUPred2A, PONDR-VSL2, and GlobPlot (prediction of protein dysfunction and globularity), and PHYRE2. Domain boundaries were predicted using the DomPred server and Pfam version 32.0.

[0114] N-glycosylation prediction. The prediction of N- and O-glycosylation sites of BHT was performed using the GlycoEP server (see, for example, Chauhan, JS, Rao, A., and Raghava, GPS (2013) In silico Platform for Prediction of N-, O- and C-Glycosites in Eukaryotic Protein Sequences. PLOS ONE 8, e67008).

[0115] Prediction of phosphorylation sites. BHT phosphorylation sites were predicted using DEPP (Disorder enhanced phosphorylation predictor), also known as DisPhos1.3 (http: / / www.dabi.temple.edu / disphos / ), and NetPhosYeast1.0 (http: / / www.cbs.dtu.dk / services / NetPhosYeast / ).

[0116] Structural modeling program. Structural diagrams and superpositions were created using PyMOL (http: / / www.schrodinger.com / pymol / ). Dimers exist in crystal-asymmetric units. However, monomers were considered for structural analysis. BHT (23-594)Structural comparisons of HIS with other known structures were performed using Dali (http: / / ekhidna2.biocenter.helsinki.fi / dali / ). The Dali Server for the PDB90 database was used for protein structure alignment. Alignments were visualized using the ESPript / ENDscript program (http: / / espript.ibcp.fr / ESPript / ESPript / ). Protein sequences were obtained from the UniProt database (https: / / www.uniprot.org / ) and aligned using the Clustal Omega tool.

[0117] Size exclusion chromatography. To determine molecular weight, NTA-purified samples were subjected to size exclusion chromatography (Superdex200 10 / 300GL, GE Healthcare) and equilibrated with SEC buffer (100mM Tris pH 7.5, 200mM sodium chloride). The protein samples equilibrated with SEC buffer were used as the column. BHT (23-594) -The mass of HIS was calculated based on the standards of the High Molecular Weight Gel Calibration Kit (CytivaLife Sciences®).

[0118] Small-angle X-ray scattering: Data collection and analysis. rBHT (23-594) -HIS samples, 1 mg / ml and 4 mg / ml sodium phosphate buffer, pH 5 were measured using a Rigaku Bio-SAXS2000. This instrument uses CuK α For radiation (λ=1.54 Å), the range is 0.01~0.67 Å. -1 Collimation was performed to provide a sufficient Q range. Measurements were performed at ambient temperature. Samples were measured for a total of 40 minutes with 5-minute scans. Data were corrected for transmittance and sample background. Reduction, averaging, and buffer subtraction were performed using Rigaku SAXSLab 3.1.0b14 (Figure 7A).

[0119] 7. Examples As will be obvious to those skilled in the art, other suitable modifications and adaptations of the methods disclosed herein are readily applicable and recognizable and may be made using suitable equivalents without departing from the scope of this disclosure or the aspects and embodiments disclosed herein. Having described this disclosure in detail, the following examples will provide a clearer understanding of the disclosure, and these are intended only to illustrate some aspects and embodiments of the disclosure and should not be considered to limit the scope of the disclosure. All journal references, U.S. patents, and publication disclosures referenced herein are incorporated herein by reference in their entirety.

[0120] This disclosure has several embodiments, as illustrated by the following non-limiting embodiments.

[0121] Example 1 In silico analysis of BHT. After translation, proteins can be modified by various post-translational modifications (PTMs). These modifications may include, for example, glycosylation and phosphorylation. PTMs alter the conformation of proteins, thereby affecting their stability, activity, intracellular distribution, and secretion. BHT (23-594) Although the crystal structure of -HIS(6M4E) has been elucidated in recent years, the initial portion of the novel N-terminal region (residues 23-54) has not been modeled, and no known structure exists for it. To achieve accurate predictions of the BHT N-terminal structure and PTM, comprehensive in silico predictions were performed using various comparison methods (Figure 1).

[0122] Example 2 Site-directed mutagenesis of predicted N-glycosylation sites within the BHT GH1 conserved domain. Glycosylation is one of the central post-translational modifications of proteins, primarily occurring by linking glycans to the nitrogen atom of an asparagine residue (N-linking) or to the hydroxyl oxygen of a serine, threonine, or tyrosine residue (O-linking), but also occurring through C-mannosylation, phosphorylated serine glycosylation, and glycation (GPI anchor formation). N-glycosylation has been shown to affect enzyme activity, stability, and cell surface expression, as previously outlined. Therefore, extensive searches and alignment analyses conducted to identify BHT homologs predicted 25 potential N-linked glycosylation sites. Four of these are located within the GH1 domain and are predicted to have highly conserved glycosylation consensus sites (Asn-X-Ser / Thr), which are at the locations of N289LTY, N297STS, N431QSD, and N569QSD (Figure 1; GlycoEP analysis), suggesting a high probability of functionally related glycosylation. In recent years, rBht (23-594) The crystal structure of -HIS (HsBglA, PDB:6M4E) was confirmed. Of these, both N431QSD were predicted to be N-glycosylated (Figure 1; GlycoEP), and phosphorylated at serine 433 within N431QSD (NetPhosYeast1.0). Therefore, four N-glycosylation sites of BHT were analyzed to help narrow down the putative region responsible for membrane-associated rBHT and to determine if these sites have functional importance. Asparagine residues (N289, N297, N431, and N569) are responsible for rBHT (23-594) Site-directed mutagenesis using -HIS as a template independently induced mutations in glutamine residues, and glycosylation was disabled as described in Materials and Methods.

[0123] The result was the non-mutant variant GS115::MFα-rBht (23-594) - Compared to the activity of HIS, there are three variants; GS115::MFα-rBht (23-594)(N431Q) -HIS, GS115::MFα-rBht (23-594)(N289Q)-HIS and GS115::MFα-rBht (23-594)(N297Q) -Significant decrease in soluble enzyme activity secreted from HIS was observed (90%, 95%, and 97%). GS115::MFα-rBht (23-594)(N569Q) -HIS variant is GS115::MFα-rBht (23-594) - Compared to HIS, it showed a non-strict decrease in activity (58%) (Figure 2; Table 1). Parent strain GS115::MFα-rBht (23-594) -When compared to HIS, GS115::MFα-rBht (23-594)(N289Q) -HIS, GS115::MFα-rBht (23-594)(N297Q) -HIS, GS115::MFα-rBht (23-594)(N431Q) -HIS, and GS115::MFα-rBht (23-594)(N569Q) -HIS cell membrane association activity also decreased by 81%, 95%, 84%, and 75%, respectively (Figure 2 and Table 1). In particular, membrane binding-related activity decreased significantly but was not completely eliminated, and GS115::MFα-rBht (23-594)(N569Q) Regarding HIS, the ratio of secretion to cell membrane association activity increased from 0.40 to 0.66 (Table 1). This suggests that glycosylation affects catalytic activity but does not completely determine cell membrane localization.

[0124] [Table 1]

[0125] Example 3 Expression and secretion of a truncated N-terminal rBHT variant by K. pastoris. The novel BHTN-terminal 110 region shows no homology to known proteins. Therefore, in silico structural prediction was performed in this region, primarily showing the majority of the disordered fragment using five available prediction tools (Figure 1). Secondary structure and globular domain were predicted using the known PSIPRED and Globplot methods. By comparison, the disorder datasets derived from Phyre2, IUPRED2A, DISOPRED3, GlobplotDisorder, and PONDR (Figure 1) show possible disorder boundaries between residues 18–42, 43–57, 87–96, and 96–110. Combining different disorder predictors enhances the reliability of the predicted region to use different definitions of disorder.

[0126] The primary function of the disordered region is thought to be its ability to fold upon membrane contact and specific ligand binding. The approach of this disclosure utilized this information to progressively and selectively delete predicted disordered fragments and determine whether they have an effect on limiting the secretion of soluble active rBHT. Schematic diagrams of complete rBHT and eight rBHT shortened variants of the enzyme generated and tested in this disclosure are shown in Figure 2A. These rBHT variants are related to rBHT (1-594) - These are created by gradually removing the N-terminal amino acid blocking group from the HIS parent sequence, and these rBHT variants include 1-22 rBHT, as shown in Figure 2A. (23-594) -HIS, 1-31 rBHT (32-594) -HIS, 1~53 rBHT (54-594) -HIS, 1-56 rBHT (57-594) -HIS, 1~81 rBHT (82-594) -HIS, 1~94 rBHT (95-594) -HIS, 1~102 rBHT (103-594) -HIS and 1-110 rBHT (111-594) -HIS is one example. As mentioned above, K. pastoris secretory membrane-associated enzymes and soluble active enzymes were evaluated for each shortened variant after methanol induction.

[0127] To investigate the presence of secreted soluble rBHT truncated protein variants, the culture medium broth was first examined by Coomassi staining SDS-PAGE (Figure 3A), followed by Western blot analysis (Figure 3B). (23-594) -HIS, rBHT (32-594) -HIS, rBHT (54-594) -HIS and rBHT (57-594) -HIS was clearly detectable by Coomassi staining (Figure 3A) and Western blotting (Figure 3B). rBHT (82-594) -HIS, rBHT (95-594) -HIS, rBHT (103-594) -HIS, rBHT (111-594) - The HIS protein band was not detectable by Western blotting (Figure 2B), nor by silver staining (data not shown but available upon request). This suggests that 57 downstream residues are important for processing the secreted protein. Consistent with previous results, rBHT (1-594) -HIS variants were barely visible in Western blot (Figure 3B) (8). No protein bands were detected when broth medium of induced GS115 transformed with a pPIC9 empty vector was used as a negative control.

[0128] The most noteworthy finding was rBHT detected by SDS-PAGE. (32-594) -HIS and rBHT (54-594)-There is a mobility shift of approximately 30 kDa with HIS, which is likely due to the deletion of the predicted phosphorylation site and surrounding acidic residues (Y37 (LTSNYETPS), T39 (SNYETPSPT), S41 (YETPSPTAI), T43 (TPSPTAIPL), T50 (PLEPTPTAT), T52 (EPTPTATGT)) (Figure 1; DisPhos3.1), which is known to delay proteins in SDS-PAGE. The algorithm DisPhos1.3 (DEPP) uses impairment information to help improve and identify phosphorylation and non-phosphorylation sites (http: / / www.pondr.com / pondr-tut2.html). Furthermore, the accuracy of DEPP reaches 76.0+ / -0.3%, 81.3+ / -0.3%, and 83.3+ / -0.3% for serine, threonine, and tyrosine, respectively. The observation that the amino acid characteristics in regions adjacent to phosphorylation sites are essentially similar to those in disordered regions suggests that disorder within and around potential phosphorylation sites may be a prerequisite for phosphorylation. Furthermore, transmembrane disordered proteins are rich in phosphorylated residues and interact with more partners than their structured counterparts.

[0129] Following the results above, the cell concentration (OD) when assayed at 42°C using ONP-Glu as the substrate was obtained. 600nm The concentration and activity of soluble protein normalized to ) were compared (Figure 2). The truncated protein variant rBHT (23-594) -HIS, rBHT (32-594) -HIS and rBHT (54-594) No significant difference in enzyme activity was detected between -HIS and the shortened variant rBHT. (57-594) -HIS is rBHT (23-594) - Compared to HIS, it showed a 38% increase in enzyme activity in the culture medium.

[0130] Further testing focused on the ability to drive secretion from associated membranes into a soluble form. The secretory enzyme activity associated with the membrane was rBHT. (23-594) -HIS, rBHT (32-594) -HIS and rBHT(54-594) -HIS and rBHT (57-594) - It is constant in HIS, and variant rBHT (23-594) -HIS, rBHT (32-594) -HIS and rBHT (54-594) -HIS showed no significant difference in the ratio of secretory soluble enzyme activity to membrane-associated enzyme activity. However, the activity that showed membrane association remained relatively constant, while rBHT (57-594) - The ratio of secretory enzyme activity to membrane-associated enzyme activity of the HIS variant is the same as that of the variant rBHT. (23-594) -HIS, rBHT (32-594) -HIS and rBHT (54-594) - Compared to HIS, it increased by 25-38% (Table 1). Bioactive rBHT (82-594) -HIS, rBHT (95-594) -HIS and rBHT (103-594) To further evaluate whether the HIS variant is produced and secreted, even in small amounts, corresponding cell lines were induced, the culture broth was concentrated 100-fold, and then affinity chromatography was performed using nickel resin. However, the protein could not be eluted, and / or activity was detected from soluble or cell-associated rBHT from these deletion variants (data not shown but available upon request). The activity assay results indicated that amino acid residues 1–56 are not required for the expression and secretion of the active enzyme. This finding is consistent with SDS-PAGE and Western blot data (Figure 3).

[0131] Example 4 Evaluation of alternative signal sequences. Testing alternative signal sequences other than the widely used MFα was considered difficult given the ever-increasing number of options. Therefore, rBht (54-594)Chimeras were generated by merging the variant into the following open reading frames (ORFs): glucoamylase (GA), invertase (IV), and inulinase (IN) signal sequences. Under experimental conditions, the results showed lower levels of soluble and membrane-associated active proteins compared to the MFα signal sequences routinely used throughout this disclosure (Table 1). Therefore, we decided to focus our investigation on MFα. Deletion of amino acids 57-70 in the MFα pro region increased reporter protein secretion by at least 50%. GS115::MFα (Δ57-70) -rBht (23-594) -HIS and dGS115::MFα (Δ57-70) -rBht (57-594) - To express the HIS variant, amino acids 57-70 are removed from MFα, resulting in GS115::MFα-rBht (23-594) -HIS and GS115::MFα-rBht (57-594) Compared to expression from HIS, we obtained increases in soluble enzyme secretion of 58% and 31%, respectively (Table 1).

[0132] These experiments suggested that maintaining BHT amino acids 57–110 from the BHT N-terminal domain is necessary for enzyme activity, secretion, and stability. These findings also highlight the unbalanced secretion of soluble bicellularly associated rBHT, and that the balance shifts to the active soluble secreted form when either 56 amino-terminal amino acids are deleted or the MFα signaling sequence is modified (Table 1).

[0133] Kinetic parameters of secreted soluble rBHT variants. After purification to homogeneity using carboxy-6x histidine epitope and nickel affinity chromatography, the active soluble secreted rBHT variant was functionally characterized by a standard kinetic assay. SDS-PAGE separation and subsequent detection with anti-HIS monoclonal antibody under reducing conditions demonstrated that the isolated protein was essentially homogeneous (Figure 3A-3B). rBHT (23-594) -HIS, rBHT (32-594) -HIS, rBHT(54-594) -HIS, rBHT (57-594) We investigated the kinetic parameters characteristic of each active secreted soluble variant, including -HIS. To obtain a complete kinetic profile, a key parameter to evaluate is the effect of temperature on enzyme activity. Therefore, using ONP-Glu as the substrate, we investigated rBHT (23-594) - The assay was performed at the optimal temperature of 42°C (8) for HIS, as well as at lower temperatures (20 and 30°C) and higher temperatures (55°C). The resulting kcat / km values ​​for each of the four shortened enzyme variants indicate the optimal temperature. Surprisingly, all enzyme shortened variants retained similar affinity for the substrate ONP-Glu (Km) and turnover activity (kcat), indicating that shortening does not affect the catalytic integrity of the enzyme (Figure 4).

[0134] Example 5 Production of N-acetyllactosamine (LacNAc). As shown in Figure 5A, the rBHT polypeptide of this disclosure can catalyze the repeated addition of galactose (Gal from lactose) to N-acetylglucosamine (GlcNAc). Figure 5B includes representative results demonstrating the enzymatic reaction catalyzed by rBHT. Time course studies of galactosyl-lactose and N-acetylglucosamine (LacNAc) synthesis were performed using total cell membrane-bound proteins (1U rBHT.g -1 The assay was performed using lactose. The assay contained approximately 20 g / L lactose and approximately 10 g / L N-acetylglucosamine (GlcNAc) in 5 mM sodium phosphate buffer (pH 5.0), which was incubated at 30°C. Samples were periodically removed and analyzed by HPLC, and detected by ELSD and PDA.

[0135] The data provided herein offer an efficient solution for generating LacNAc on a cost-competitive, industrial scale. The ability of the rBHT polypeptide of this disclosure to synthesize LacNAc using lactose as a donor and N-acetylglucosamine as an acceptor has been demonstrated (Figure 5). These data provide evidence that this enzyme is an essential and novel tool for achieving the synthesis of LacNAc (Galβ1-4GlcNAc) at concentrations exceeding gram concentrations, which is considered a human milk oligosaccharide (HMO)-like sugar. These catalytic reactions are highly regioselective, forming β-galactosyl linkages at position 4 of GlcNAc and position 1 of D-galactose, and synthesizing various complex carbohydrates directly from soluble GlcNAc. The resulting product contained Galβ(1,4)GlcNAc(LacNAc, Figure 5A, Panel B) disaccharide and Galβ(1,4)Galβ(1,4)GlcNAc(galactosyl-LacNAc, Figure 5A, Panel C) trisaccharide, and was produced by two sequential transgalactosylations (Figures 5A-5B).

[0136] Example 6 Sequence and structural BHT homologs. β-glucosidase GH1 family members are classified as a single domain with an (α / β)8TIM barrel topology in CaZy classification (http: / / www.cazy.org / GH1_characterized.html). However, BHT folds into two domains (Figure 1). The main domain is an (α / β)8TIM barrel, starting at residue 116 and extending to residue 547 (HsBglA, PDB:6M4E) (17). This domain has eight parallel β strands forming a central barrel connected by eight external α helices common to GH1 family members (http: / / www.cazypedia.org / index.php / Glycoside_Hydrolase_Family_1). To identify similar structures, a heuristic PDB search was performed using the Dali server, and rBHT (23-594)-The structure of HIS(HsBglA, PDB:6M4E) was used as a query to perform a search across all deposits in the Protein Data Bank. The structure of the Dali server PDB90 database shows that β-glucosidase BGL1A from the basidiomycete Phanerochaete chrysosporium (PDB:2B3Z-A) is rBHT (23-594) -It was found to be the closest structural match to the C domain of HIS, with the highest Z score of 49.3 (Table 2). In this case, 450 of the 460 amino acids Cα are rBHT (23-594) -HIS structure and 34% and sequence identity could be superimposed. rBHT (23-594) -To directly compare structural similarities and differences with HIS, the top five structures with a Z score exceeding 46.5 and an rmsd of less than 1.8 Å were selected. Interestingly, the top five structurally identical structures are also fungal β-glucosidases, and the alignment for all five is rBHT. (23-594) -Specific to the C-terminal domain of HIS (Table 2).

[0137] In addition to Phanerochaete chrysosporium BGL1A (PDB:2B3Z-A), this list also includes β-glucosidase from Trichoderma reesei (PDB:3AHY-B), β-1,4-glucosidase from Trichoderma harzianum (PDB:5BWF-A), β-glucosidase from Humicola insolens (PDB:4MDO-A), and β-glucosidase from Trichoderma harzianum (PDB:5JBO-A). These are the top five structures and rBHT. (23-594) -The primary sequence alignment with HIS shows that the core GH1 structure is shared, while its N-terminus is distinct and unique (Figure 6A). Sequence identity remained largely unchanged from 31% to 34%, but the nucleophile and common acid / base residues in the enzyme sequence were well aligned. These structures and rBHT (23-594) - The superposition of the HIS(HsBglA,PDB:6M4E) structure is rBHT (23-594)The C-terminal core of -HIS outlines the catalytic pocket located at the center of the typical barrel of the GH1 protein (Figure 6A-6B). Except for the variability seen in loops A-D, it is almost identical to other GH1 structures. The most notable features are the long insertion in loop C (residues 423-433, NGIANCIRNQS) and the short insertions in loops A (Y147), B (Q282, N283, L290), C (S460 and A461), and D (L510, Y511, Q512) and T244, G245, G327, T328, G374, K489, and the insertion of P573 (Figure 6A). The inserted residues are located on the surface of the structure (Figure 6B). Interestingly, in all five structural homologs, the non-(α / β)8TIM barrel β8 in loop C and β9 in loop D are much longer than those found in HsBglA (PDB:6M4E), and the third β sheet is substituted with α17 in loop C (Figure 6A). It is also noteworthy that all four established N-glycosylation sites (N289, N297, N431, and N569) are devoid of N-glycosylation residues (Figure 6A).

[0138] The CaZy database indicates that GenBank contains over 40,000 GH1 proteins, with over 270 PDB structures available. GH1 protein sequences from the RCSB PDB database (rcsb.org) were extracted using SANSparallel (http: / / ekhidna2.biocenter.helsinki.fi / cgi-bin / sans / sans.cgi) and integrated with the Clustal Omega program. The amino acid sequences of 60 GH1 genes structurally homologous to the C-terminal domain of BHT and with a Z-score greater than 40 showed 60% sequence identity with each other, 27-33% identity with BHT, and no N-terminus match. Furthermore, Blast analysis of the N-terminal domain amino acid sequence did not confirm any match. From this sequence and structure comparison data, we concluded that the N-terminal structure of BHT observed in the 6M4E structure is novel for the GH1 protein and currently does not have a structural homolog close to it in the PDB database.

[0139] [Table 2]

[0140] Example 7 The effect of N-terminal deletion on rBHT dimerization. rBHT (23-594) -HIS exists as a dimer in solution, measured by a size exclusion chromatography (SEC) column packed with cefacryl S-200, and rBHT (23-594) -HIS is the calculated M corresponding to the dimeric state. w It was shown to elute as a single peak with a 150 kDa and a retention time of 12.5 minutes (data is not shown but can be made available upon request). The dimer structure in solution was further verified by micro-X-ray scattering (SAXS) (Figure 7A). Guinier and P(r) analyses were performed using PRIMUS and GNOM, respectively. To optimize the P(r) calculation, D max The values ​​were manually selected using GNOM (Figure 7B). These D max The values ​​are approximate, approximately ±2-3 Å. Molecular weight was calculated using the method described by Rambo and Tainer. The data is shown in Table 3. At a molecular weight of approximately 169 kDa obtained from SAXS, it was confirmed that BHT forms dimers in solution (Table 3). The R of the dimer in solution... g and D max These are 39 Å and 124 Å, respectively. The deposited X-ray crystal structures (6M4E, 6M4F, and 6M55) also suggest that BHT forms dimers. The R of the 6M4F crystal dimer (molecule A and molecule C) calculated using Crysol g and D max These values ​​are 34 Å and 110 Å, respectively. These values ​​are in good agreement with SAXS experimental data. The disordered N-terminus leads to a more expanded dimer in solution, rBHT (23-594) -HIS was concluded to be likely to function as a dimer. Similarly, rBHT (32-594) -HIS, rBHT (54-594) -HIS and rBHT (57-594)-SEC analysis of HIS also shows dimerization (data not shown, but can be made available upon request), which suggests that the unstructured region spanning residues 23-56 is not involved in dimerization.

[0141] [Table 3]

[0142] 8. Array Sequences related to various embodiments of this disclosure are provided in the following table.

[0143] [Table 4-1]

[0144] [Table 4-2]

[0145] [Table 4-3]

[0146] [Table 4-4]

[0147] [Table 5-1]

[0148] [Table 5-2]

[0149] [Table 5-3]

[0150] Table 5-4

[0151] Table 5-5

[0152] Table 5-6

[0153] Table 5-7

[0154] Table 5-8

[0155] Table 5-9

[0156] Table 5-10

[0157] Table 5-11

[0158] Table 5-12

[0159] Table 5-13

[0160] Table 5-14

[0161] Table 5-15

[0162] Table 5-16

[0163] Table 5-17

[0164] Table 5-18

[0165] Table 5-19

[0166] Table 5-20

[0167] Table 5-21

[0168] Table 5-22

[0169] Table 5-23

[0170] Table 5-24

[0171] Table 5-25

[0172] Table 5-26

[0173] Table 5-27

[0174] Table 5-28

[0175] Table 5-29

[0176] Table 5-30

[0177] Table 5-31

[0178] Table 5-32

[0179] Table 5-33

[0180] Table 5-34

[0181]

Table 5-35

[0182]

Table 5-36

[0183]

Table 5-37

[0184]

Table 5-38

[0185]

Table 5-39

[0186]

Table 5-40

[0187]

Table 5-41

[0188] The stocks and plasmids related to the embodiments of the present disclosure are provided in the following table.

[0189]

Table 6-1

[0190]

Table 6-2

[0191] It should be understood that the above detailed description and accompanying examples are illustrative only and should not be construed as limiting the scope of this disclosure as defined solely by the attached claims and their equivalents.

[0192] All publications and patents described in the above specification are incorporated herein by reference as if they were expressly described herein. Various changes and modifications to the disclosed embodiments will be obvious to those skilled in the art and can be made without departing from the spirit and scope thereof.

[0193] One aspect of the present invention can also be configured as follows. [1] A functional recombinant β-hexosyl-transferase (rBHT) polypeptide having at least 90% sequence identity with SEQ ID NO: 1 and comprising at least one N-terminal shortening of one amino acid with respect to SEQ ID NO: 1. [2] The polypeptide according to [1], comprising at least 95% sequence identity with sequence number 1. [3] The polypeptide according to [1] or [2], further comprising at least one additional amino acid substitution. [4] The polypeptide according to any one of [1] to [3], wherein the N-terminal shortening is approximately 1 to approximately 81 amino acids in length. [5] The polypeptide according to any one of [1] to [4], wherein the N-terminal shortening is approximately 1 to approximately 56 amino acids in length. [6] A polypeptide according to any one of [1] to [5], comprising at least 90% sequence identity with any of sequence numbers 3, 5, 7, and 9. [7] A polypeptide according to any one of [1] to [6], further comprising a signal sequence. [8] The polypeptide according to [7], wherein the signal sequence is non-natural. [9] The polypeptide according to [7] or [8], wherein the signal sequence comprises an amino acid sequence derived from a yeast protein.

[10] The polypeptide according to any one of [7] to [9], wherein the signal sequence comprises an amino acid sequence derived from a protein derived from one of Komagataella (Pichia) pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Hansenula (Ogataea) polymorpha, or Kluyveromyces lactis.

[11] The polypeptide according to any one of [7] to

[10] , wherein the signal sequence comprises a polypeptide having at least 90% sequence identity with at least one of the following: α-conjugation factor signal sequence (MFα) (SEQ ID NO: 29), invertase (IV) signal sequence (SEQ ID NO: 30), glucoamylase (GA) signal sequence (SEQ ID NO: 31), or inulinase (IN) signal sequence (SEQ ID NO: 32) derived from Saccharomyces cerevisiae.

[12] A polypeptide according to any one of [7] to

[11] , comprising at least 90% sequence identity with any of sequence numbers 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, or 70.

[13] A polypeptide according to any one of [1] to

[12] , comprising at least one asparagine residue at positions 289, 297, 431, and / or 569 of SEQ ID NO: 1.

[14] A polypeptide according to any one of [1] to

[13] , which is soluble or membrane-bound.

[15] The polypeptide according to

[14] , wherein about 1% to about 50% of the polypeptide is soluble.

[16] A polypeptide according to any one of [1] to

[15] that catalyzes the hydrolysis of lactose β-(1-4) glycoside linkages.

[17] The polypeptide according to

[16] , wherein the catalyst for hydrolysis of lactose β-(1-4) glycoside linkage by the polypeptide produces a composition containing LacNAc-rich GOS. A nucleic acid molecule encoding any of the polypeptides described in any of

[18] , [1], to

[17] . A vector containing the nucleic acid described in

[19] and

[18] . A method for producing a GOS composition from lactose in host cells using any of the polypeptides described in

[20] [1] to

[17] .

[21] The method according to

[20] , wherein the GOS composition comprises LacNAc-rich GOS and / or GlcNAc-free GOS.

[22] The method according to

[21] , wherein the host cell is one or more of yeast cells, fungal cells, mammalian cells, insect cells, plant cells, or algal cells.

[23] The method according to

[22] , wherein the host cells include one or more cells derived from Komagataella (Pichia) pastoris, Saccharomyces cerevisiae, Yarrowia lipolytica, Hansenula (Ogataea) polymorpha, or Kluyveromyces lactis, Aspergillus spp., and Trichoderma reesei.

[24] The method according to any one of

[21] to

[23] , wherein the yield of LacNAc-rich GOS is at least 10% of the initial lactose concentration, and the total GOS concentration is at least 50% of the initial lactose concentration. A composition comprising one or more LacNAc-rich GOS produced using any of the polypeptides described in

[25] [1] to

[17] and / or any of the polypeptides described in [1] to

[17] .

[26] The composition described in

[25] , which is a food product.

[27] The composition according to

[26] , wherein the food comprises one or more of the following: infant formula, yogurt, dairy products, milk-based beverages, fruit beverages, hydration beverages, energy beverages, fruit preparations, and meal replacement beverages. [Brief explanation of the drawing]

[0194] [Figure 1]This figure shows the predicted structural post-translational modifications and disordered vs. ordered secondary motifs of β-hexosyltransferase from H. singularis. The glycosylation, phosphorylation, and secondary structure of the BHT protein were predicted using various algorithms. Described are the structural elements, conserved regions, and functional domains of BHT using the PSIPRED and GlobplotGlobular prediction tools. Disordered regions were predicted using the Phyre2, IUPRED2A, DISOPRED3, GlobplotDisorder, and PONDR algorithms. The phosphorylation servers DisPhos1.3 and NetPhosYeast1.0 indicate phosphorylation sites. GlycoEP shows N-glycosylation (red line) and O-glycosylation (black line), but the C-mannosylation site was not predicted. The numbers below each prediction line indicate the BHT amino acid residue numbers. [Figure 2]Figure 2A: Comparison of enzymatic activity of rBHT variants is shown as the amount of secreted soluble protein produced by recombinant K. pastoris strains carrying the truncated rBht-HIS variant under AOX1 promoter control (normalized to the final culture (OD600nm)). (A) Illustrated by a graph of the generated chimeric genes, including combinations of the leader domain and the ORF of the rBht variant. Specific tags, mutations, and deletions are indicated. Figure 2B: Comparison of enzymatic activity of rBHT variants is shown as the amount of secreted soluble protein produced by recombinant K. pastoris strains carrying the truncated rBht-HIS variant under AOX1 promoter control (normalized to the final culture (OD600nm)).The protein concentrations (B) of soluble secreted proteins secreted by the following recombinant strains were compared: Column 1, GS115::rBht(1-594)-HIS; Column 2, GS115::MFα-rBht(1-594)-HIS; Column 3, GS115::MFα-rBht(23-594)-HIS; Column 4, GS115::MFα-rBht(23-594)(N289Q)-HIS; Column 5, GS11 5::MFα-rBht(23-594)(N297Q)-HIS; Row 6, GS115::MFα-rBht(23-594)(N431Q)-HIS; Row 7, GS115::MFα-rBht (23-594)(N569Q)-HIS; Row 8, GS115::MFα-rBht(32-594)-HIS; Row 9, GS115::MFα-rBht(54-594)-HIS; Row 10, G S115::MFα-rBht(57-594)-HIS;Column 11, GS115::MFα-rBht(82-594)-HIS;Column 12, GS115::MFα-rBht(95-594)- HIS;Column 13, GS115::MFα-rBht(103-594)-HIS;Column 14, GS115::MFα-rBht(111-594)-HIS;Column 15, GS115::IV-rBh t(54-594)-HIS; Column 16, GS115::GA-rBht(54-594)-HIS; Column 17, GS115::IN-rBht(54-594)-HIS; Column 18, GS115::MFα(Δ57-70)-rBht(23-594)-HIS; Column 19, GS115::MFα(Δ57-70)-rBht(57-594)-HIS; Column 20, GS115(His+) control. Figure 2C: Comparison of enzymatic activity of rBHT variants is shown as the amount of secreted soluble protein produced by recombinant K. pastoris strains carrying the truncated variant of rBht-HIS under AOX1 promoter control (normalized for the final culture (OD600nm)).The enzymatic activity (C) of soluble secreted proteins secreted by the following recombinant strains was compared: Column 1, GS115::rBht(1-594)-HIS; Column 2, GS115::MFα-rBht(1-594)-HIS; Column 3, GS115::MFα-rBht(23-594)-HIS; Column 4, GS115::MFα-rBht(23-594)(N289Q)-HIS; Column 5, GS115: :MFα-rBht(23-594)(N297Q)-HIS;Row 6, GS115::MFα-rBht(23-594)(N431Q)-HIS;Row 7, GS115::MFα-rBht( 23-594)(N569Q)-HIS; Row 8, GS115::MFα-rBht(32-594)-HIS; Row 9, GS115::MFα-rBht(54-594)-HIS; Row 10, GS 115::MFα-rBht(57-594)-HIS;Column 11, GS115::MFα-rBht(82-594)-HIS;Column 12, GS115::MFα-rBht(95-594)- HIS;Column 13, GS115::MFα-rBht(103-594)-HIS;Column 14, GS115::MFα-rBht(111-594)-HIS;Column 15, GS115::IV-rBh t(54-594)-HIS; Column 16, GS115::GA-rBht(54-594)-HIS; Column 17, GS115::IN-rBht(54-594)-HIS; Column 18, GS115::MFα(Δ57-70)-rBht(23-594)-HIS; Column 19, GS115::MFα(Δ57-70)-rBht(57-594)-HIS; Column 20, GS115(His+) control. [Figure 3]Figure 3A: This figure shows Coomassi-stained SDS-PAGE (10%) separation and Western blot. The figure shows cell-free extracts (soluble secreted proteins) of proteins expressed by different recombinants of K. pastoris GS115. (A) Separated proteins produced by SDS-PAGE exposed to anti-HIS antiserum; Lane 1, GS115::MFα-rBht-HIS; Lane 2, GS115::MFα-rBht(23-594)-HIS; Lane 3, GS115::MFα-rBht(32-594)-HIS; Lane 4, GS115::αMF-rBht(54-594)-HIS; Lane 5, GS115::αM F-rBht(57-594)-HIS; lane 6, GS115::αMF-rBht(82-594)-HIS; lane 7, GS115::αMF-rBht(95-594)-HIS; lane 8, GS115::MFα-rBht(103-594)-HIS; lane 9, GS115::MFα-rBht(111-594)-HIS; lane 10, GS115 control containing an empty pPIC9 vector. Equal amounts were loaded into each lane to aid comparison. The total protein (ng) loaded into each well is shown above in (A). "---" indicates that the concentration could not be measured. M indicates a lane containing a molecular weight protein marker, and (kDa) is shown on the left side of the panel. Figure 3B: Figure showing Coomassie-stained SDS-PAGE (10%) separation and Western blot. The figure shows cell-free protein extracts (soluble secreted proteins) expressed by different recombinant strains of K. pastoris GS115.(B) Separated proteins generated by Western blotting were exposed to anti-HIS antiserum; Lane 1, GS115::MFα-rBht-HIS; Lane 2, GS115::MFα-rBht(23-594)-HIS; Lane 3, GS115::MFα-rBht(32-594)-HIS; Lane 4, GS115::αMF-rBht(54-594)-HIS; Lane 5, GS115::αM F-rBht(57-594)-HIS; lane 6, GS115::αMF-rBht(82-594)-HIS; lane 7, GS115::αMF-rBht(95-594)-HIS; lane 8, GS115::MFα-rBht(103-594)-HIS; lane 9, GS115::MFα-rBht(111-594)-HIS; lane 10, GS115 control containing an empty pPIC9 vector. Equal amounts were loaded into each lane to aid comparison. The total protein (ng) loaded into each well is shown above in (B). "---" indicates that the concentration could not be measured. M indicates a lane containing a molecular weight protein marker, and (kDa) is shown on the left side of the panel. [Figure 4] This figure shows the enzyme kinetic parameters of rBHT variants tested at 20°C, 30°C, 42°C, and 55°C. kcat / km versus temperature. Enzyme assays were performed in the presence of 0.3 μg of rBHT(23-594)-HIS, rBHT(32-594)-HIS, rBHT(54-594)-HIS, and rBHT(57-594)-HIS within the ONP-Glu substrate concentration range (0.08–10.4 mM) described in "Methods". Km and kcat were calculated from the initial rate of ONP-Glu cleavage using the Hill formula. The values ​​are the mean ± standard deviation (SD) of three independent measurements. [Figure 5]Figure 5A: This figure shows an example of N-acetyllactosamine (LacNAc) production at a lactose / N-acetylglucosamine ratio of 1:2. The recombinant BHT (rBHT) polypeptide of this disclosure can catalyze the repeated addition of galactose (Gal from lactose) to N-acetylglucosamine (GlcNAc). Figure 5B: This figure shows an example of N-acetyllactosamine (LacNAc) production at a lactose / N-acetylglucosamine ratio of 1:2. It shows the enzymatic reaction catalyzed by rBHT. Examples of time-course studies of galactosyl-lactose (Gal-lactose), galactosyl-N-acetalactosamine (Gal-LacNAc), and N-acetyllactosamine (LacNAc) synthesis were performed using whole cell membrane-bound protein (1U rBHT.g-1 lactose). The assay involved a 5 mM sodium phosphate buffer (pH 5.0) containing approximately 20 g / L lactose and approximately 10 g / L N-acetylglucosamine (GlcNAc), incubated at 30°C. Samples were periodically removed and analyzed by HPLC, and detected by ELSD and PDA. [Figure 6A]This figure shows multiple secondary structure alignments of 6m4e(HsBglA(23-594)-HIS) with structurally homologous GH1 proteins. (A) The proteins found to be most structurally homologous from the PDB database include 2E3ZA(BGL1A), 3AHYB(TrBgl2), 5BWFA(ThBgl), 4MDOA(HiBG), and 5JBOA(ThBgl2) (Table 4). The primary sequence alignment is shown below. The secondary structure elements of rBHT(23-594)-HIS and their names are shown above the alignment. β strands are indicated by black arrows, and the α helix structure formed by coils, the exact α turn (TTT letter), the β turn (TT letter), and η refer to a 310-helix random coil. The secondary structure element numbers of the (α / β)-Tim barrel structure are shown above the structural alignment as (α1-α8) and (β1-β8). In the analysis of the unstructured region of HsBglA(23-594)-HIS, the amino acid numbers of HsBglA from the amino terminus to the carboxyl terminus include the deletion signal sequence (residues 1-22) shown by dashed arrows and the unstructured region missing from the crystal structure (residues 23-53) shown by dotted lines. The amino acids were aligned using ClustalO based on % sequence similarity. Identical residues are enclosed in white on a black background, and conservative changes are enclosed in gray boxes. Insertions are highlighted on a purple background. Catalytic acid / base nucleophilic residues are indicated by asterisks. Glycosylation sites found in HsBglA(23-594)-HIS are indicated by triangles. Figure 6A shows the predicted phosphorylation and potential O-glycosylation sites at the N-terminus shown in Figure 1, indicated by squares and circles, respectively. The consensus sequence is shown below the aligned sequence.The images were generated using the ENDscript2.0 web server (http: / / endscript.ibcp.fr / ESPript / ENDscript / )(5), and derived from a comparison of the 3D crystal structure and the protein databank crystal structure based on HsBglA(23-594)-HIS (PDB ID:M6E4) using data obtained from the Dali protein structure comparison server (http: / / ekhidna2.biocenter.helsinki.fi / dali / ) (Holm, 2019). [Figure 6B] This figure shows multiple secondary structure alignments of 6m4e(HsBglA(23-594)-HIS) with structurally homologous proteins of GH1. The four elongation loops A, B, C, and D of HsBglA(23-594)-HIS (PDBID:M6E4) are colored blue, green, yellow, and red, respectively, and are shown as same-colored arrows forming the substrate binding pocket entrances, overlaid on the secondary structure of (A). Generated using PyMOL (https: / / pymol.org / 2 / ). [Figure 6C] This figure shows multiple secondary structure alignments of 6m4e(HsBglA(23-594)-HIS) with structurally homologous GH1 proteins. The degree of conservation of HsBglA(23-594)-HIS (PDBID:M6E4) is represented by a color gradient from red to blue. Deep red indicates more conserved residues, while deeper blue indicates more variable residues. Generated using PyMOL (https: / / pymol.org / 2 / ). [Figure 7] Figure 7A: This figure shows SAXS data for BHT at 1 mg / ml (red) and 4 mg / ml (blue). The SAXS data is shown as a logarithmic plot (left). I(Q) is in arbitrary units. Figure 7B: The P(r) curve calculated from the SAXS data is normalized to a maximum height of 1.0.

Claims

[Claim 1] A functional recombinant β-hexosyl-transferase (rBHT) polypeptide having at least 90% sequence identity with SEQ ID NO: 1, and comprising at least one N-terminal shortening of an amino acid with respect to SEQ ID NO: 1.