Carbohydrate molecule sequencing method based on glycosidase and nanopore and system and application thereof
By combining glycosidase and nanopore technology, and utilizing changes in electrical signals and hydrolysis reactions, the problem of difficult sequence information analysis of carbohydrate compounds has been solved, achieving efficient and accurate sequencing of carbohydrate molecules.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI INSTITUTE OF MATERIA MEDICA CHINESE ACADEMY OF SCIENCES
- Filing Date
- 2025-11-21
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies cannot effectively resolve the sequence information of carbohydrate compounds, especially the composition, linkage, order, and configuration of polysaccharides. Furthermore, traditional methods are cumbersome, time-consuming, and costly. Nanopore technology has not yet achieved sequencing in glycan detection.
A detection system based on glycosidases and nanopores is adopted. The changes in the current signal of sugar molecules are detected by nanopores. Combined with the hydrolysis reaction of glycosidases, specific or non-specific glycosidases are used to hydrolyze sugar molecules at specific sites. The sugar sequence information is then analyzed by comparing the electrical signals.
It enables efficient and accurate sequencing of carbohydrate molecules, allowing for the sequential analysis of composition, linkages, sequence, and conformation, thereby improving sequencing efficiency and accuracy while reducing costs.
Smart Images

Figure CN122038529A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a method, system, and application of carbohydrate molecule sequencing based on glycosidases and nanopores. Background Technology
[0002] Carbohydrates are the third class of biological macromolecules after proteins and nucleic acids, and also the largest family of natural compounds. Different types of monosaccharides can form oligosaccharides or polysaccharides through polymerization or dehydration condensation. Polysaccharides are a class of carbohydrates with complex molecular structures and large volumes; carbohydrates and their derivatives that meet the definition of macromolecules are all called polysaccharides. Currently, there are approximately 200 types of monosaccharides that make up polysaccharides. The different linkage modes of glycosidic bonds and branching structures lead to their structural complexity. Therefore, the sequence information of carbohydrate compounds includes the type of monosaccharide, the functional group modifications of the monosaccharides, the linkage mode between monosaccharides, α- or β-anomers, branching sites, and the aforementioned sequence information of branching sequences. The complexity of carbohydrate compound sequences determines the diversity of their biological functions. Therefore, the analysis of carbohydrate compound sequences is extremely necessary.
[0003] Carbohydrates have structures far more complex than proteins and nucleic acids, making them arguably the most complex biological macromolecules. From a chemical perspective, the complexity of polysaccharide structures presents significant challenges to their structural analysis. Traditional polysaccharide structure analysis typically involves five parts: identification of the basic compound type, determination of the molecular formula, identification of specific functional groups, and deduction of the planar and stereostructure. This requires the use of chemical methods, such as partial acid hydrolysis, complete acid hydrolysis, periodic acid oxidation, Smith degradation, and methylation reactions, as well as biological methods, such as specific glycosidase digestion and immunological methods. It also necessitates the use of costly instruments such as infrared spectroscopy, nuclear magnetic resonance, and mass spectrometry. It is evident that traditional polysaccharide analysis techniques require multidisciplinary collaboration, are cumbersome, time-consuming, consume large quantities of samples, and are expensive. Therefore, the search for and development of new polysaccharide structure analysis techniques is urgently needed.
[0004] Nanopore analysis is a novel single-molecule analytical technique that began to develop in the mid-1990s. In 1996, Damer et al. first reported the detection of single-stranded DNA (ssDNA) and RNA molecules using a natural biological channel formed by Staphylococcus aureus α-hemolysin (αHL), obtaining for the first time the current signal formed by single-stranded DNA (ssDNA) and RNA molecules passing through the αHL channel. This epoch-making research laid the foundation for nanopore analysis, which has since been widely used in many research fields such as chemistry and biology due to its advantages of speed, low cost, and no need for fluorescent labeling. The basic working principle of this technique is that in a cavity filled with electrolyte, an insulating and impermeable membrane with nanoscale pores divides the cavity into two chambers. When a voltage is applied to the electrolyte chamber, ions pass through the pores under the action of an electrochemical gradient, forming a stable and detectable current. Driven by electric field, electroosmotic force, and concentration gradient, analyte molecules enter or pass through nanopores. As the analyte molecules pass through the pore region, they cause fluctuations in the current signal. Different analyte molecules generate current signals with varying timing, amplitude, frequency, and shape. By statistically analyzing a large number of current signals, information such as the molecular structure, composition, size, and charge can be obtained at the single-molecule level. Applying this efficient, specific, and convenient nanopore technology to the detection of structurally complex carbohydrate molecules has broad application prospects, thereby reducing the cumbersome nature of traditional methods for analyzing carbohydrate molecular structures.
[0005] Currently, nanopore detection technology is mainly applied in DNA sequencing, nucleic acid analysis, protein and peptide analysis, metal ion detection, and small organic molecule detection. Recently, researchers have been working to apply nanopore technology to sugar molecule detection and have proposed various strategies to improve its performance in distinguishing minute structural differences in sugar molecules. It has been confirmed that nanopores can accurately distinguish minute structural differences in glycan chains, such as functional groups, diverse building blocks, different glycosidic bonds, variations in length extension, and branching patterns. However, these advances only demonstrate the potential of nanopores in glycan chain detection and have not yet been able to resolve the sequence information of glycan chains. Methods for polysaccharide sequencing using biological nanopores have not yet been reported.
[0006] Compared to nucleic acids and peptides, carbohydrates are composed of hundreds of different monosaccharides linked by various glycosidic bonds, involving multiple spatial conformations. Therefore, even with machine learning, reading out and analyzing glycan-induced nanopore signals remains a significant challenge. Furthermore, unlike nucleic acids and peptides with linear structures and charges, glycans with branched structures and / or low charge densities are difficult to control into fully extended molecules for nanopore sensing. Therefore, how to use nanopore technology to obtain sequence information within carbohydrate molecules—such as the types, positional order, linkages, isomerism, and branching information—remains an unsolved problem. All reported related studies utilize nanopores to identify carbohydrate molecules as a whole, thus falling under the category of carbohydrate detection. For example, in 2018, Hagen Bayley's group reported a method for identifying monosaccharide isomers using nanopores; in 2021, Qing Guangyan and Liang Xinmiao's group reported a method for distinguishing complex polysaccharides such as heparin using nanopores; in 2023, Gao Zhaobing's group reported a method for establishing oligosaccharide fingerprints using nanopores, and in 2024, they reported a method for detecting complex oligosaccharides of different lengths using nanopores. These reported studies all focus on the overall detection of sugar molecules and have not developed methods for sugar sequencing using nanopores. Chinese patent CN118103711A discloses a method for the overall detection of oligosaccharides based on nanopores. This method mainly identifies oligosaccharide molecules as a whole and cannot distinguish the composition, positional order, linkage, or isomerism information of monosaccharides within polysaccharide molecules.
[0007] The high complexity of sugar sequences endows them with a huge information capacity, and the sequence information of sugar molecules cannot be deduced from the genetic code. Furthermore, sugar molecule characterization techniques and efficiency lag far behind those for nucleic acids and proteins. Existing sugar sequence resolution techniques face numerous bottlenecks: First, the basic principle of nuclear magnetic resonance (NMR) technology is to place the sample in a strong magnetic field. Atomic nuclei (such as ^1H, ^13C, ^15N, etc.) align their magnetic moments along the magnetic field direction due to their spin. After applying a radio frequency pulse of a specific frequency, the nuclei undergo energy level transitions (resonance). When the nuclei relax back to equilibrium, they release signals carrying structural information. The structure is inferred by detecting the chemical shift, coupling constant, signal intensity, and relaxation time of these signals. This leads to a fundamental problem: sugar molecules are mainly composed of C, H, and O atoms, with similar chemical environments. In particular, the proton (^1H) has a very narrow chemical shift range (typically concentrated in the 3.0-5.5 ppm range), resulting in significant signal overlap among numerous H2, H3, H4, H5, H6, and H6' residues in different sugar residues, making identification and resolution difficult. The bottleneck lies in the fact that even with the highest field strength NMR (e.g., 1.2 GHz), the signal overlap in the ^1H and ^13C spectra remains extremely severe for structurally similar oligosaccharides or polysaccharides, making sequence determination and residue identification exceptionally difficult. The difference between stereochemistry and the weak signals of anodic linkages lies in the fact that the essence of glycoscience lies in anodic configurations and glycosidic bond positions. Although the chemical shift of the anodic proton (H1) and the ^3J_{H1,H2} coupling constant can distinguish α / β configurations, these values have overlapping ranges and are not absolutely reliable. For linkage sites (e.g., 1-4 vs 1-6), NMR mainly relies on minute changes in chemical shifts (~0.1 ppm), which are difficult to measure and assign precisely in complex structures. The bottleneck is that NMR cannot directly "see" chemical bonds; it can only infer them indirectly through interatomic interactions via bonds or space. This inference is insensitive to minute differences in carbohydrate structures, leading to uncertainty in determining linkage modes and stereochemistry. Conformational dynamics and solution inhomogeneity, along with the considerable flexibility of glycosidic bonds, mean that many sugar chains do not exist in a single rigid conformation in solution, but rather as a set of conformations. The parameters measured by NMR are time-averaged values for all these conformations. The bottleneck lies in the fact that this obscures the true stereochemical information, making precise determination of the three-dimensional structure using data such as NOE complexes, or even impossible, especially for linear or highly flexible sugar chains.
[0008] Secondly, regarding mass spectrometry, it involves converting sample molecules into gaseous charged ions (M+, [M+H]+, [MH]-, etc.) in an ion source, followed by separation and detection in a mass analyzer based on their mass-to-charge ratio. By precisely measuring the mass of molecules (especially high-resolution mass spectrometry, HRMS), their elemental composition can be determined. By fragmenting the parent ion through techniques such as collision-induced dissociation and analyzing the resulting daughter ions, the molecular sequence, branching sites, and modification types can be inferred. The "connection" and "stereochemical" blind spots of mass spectrometry make it impossible to distinguish between stereoisomers and angioisotopes: the core of MS is the mass-to-charge ratio. Two stereoisomers (such as glucose and galactose) or angioisotopes (α / β) have exactly the same mass. Mass spectrometry is inherently "blind" to stereochemistry. MS alone cannot determine which stereoisomer we are measuring. The ambiguity of glycosidic bond breaking and the indirectness of connection site identification: In collision-induced dissociation, the breaking of glycosidic bonds produces characteristic fragment ions (such as B, C, Y, Z ions). However, fragment ions produced by different connection methods (e.g., 1-2, 1-3, 1-4, 1-6) may have the same mass. The fragmentation modes are complex, including internal fragmentation and cross-ring fragmentation, making spectral resolution difficult. For branched sugars, determining the branching position depends on the presence of specific fragment ions, but these ions may be very low in abundance or not present at all. The connection information provided by MS is indirect and speculative, heavily reliant on comparison with standard spectra or complex algorithmic predictions, and is fraught with uncertainty for novel structures. The effect of anomeric configuration on fragmentation energy is weak and unreliable: theoretically, the glycosidic bond stability of α and β anomers differs slightly, potentially leading to minor differences in fragmentation efficiency. However, in complex MS / MS experiments, this difference is easily masked by other factors (e.g., charge position, solvation effect, instrument parameters). Therefore, it cannot serve as a reliable means of distinguishing anomeric configurations.
[0009] In other words, current technologies cannot completely and sequentially interpret the composition, linkages, accurate order, and configuration of complex sugars (such as charged sugars, branched sugars, and complex sugars with more than 10 constituent units); moreover, they cannot meet the needs of sugar function research and industrial development in terms of efficiency, sensitivity, and accuracy. There is an urgent need to develop high-precision sugar sequencing technologies that can identify composition, resolve linkages, interpret order, resolve branches, identify configurations, analyze modifications, and determine sites.
[0010] Currently, many naturally occurring glycosidases (divided into exoglycosidases and endoglycosidases) provide an excellent foundation for nanopore glycan hydrolysis sequencing. For example, exoglycosidases (EXGases) can release specific monosaccharides from the non-reducing ends of glycans, are widely present in bacteria, plants, and mammals, and have been well purified. While maintaining high hydrolytic activity, EXGases exhibit strong specificity for monosaccharide structures, isoform configurations, and glycosidic bonds. The precise cleavage sites and specificity of EXGases provide the advantage of directional and continuous digestion of glycans, facilitating the extraction of detailed structural information. The high sensitivity of nanopore technology enables it to effectively characterize glycan structures. Previous studies have shown that nanopores can detect changes in ion current signals caused by the hydrolysis of glycosaminoglycans by glycosidases, but this only involves the monitoring of the hydrolysis reaction by nanopores, and this report has not yet reported a glycosidase-assisted nanopore glycan sequencing method. The application of glycosidases to nanopore-based sugar sequencing still faces several unresolved challenges: (1) obtaining high-resolution nanopores to acquire identifiable continuous differences before and after hydrolysis; (2) obtaining glycosidases that are sufficient for at least one type of sugar and meet the hydrolysis specificity requirements; (3) developing glycosidase-nanopore arrays to improve the efficiency of the enzyme hydrolysis-detection system; and (4) developing algorithms to automatically identify current signals before and after hydrolysis. Summary of the Invention
[0011] Purpose of the invention: The technical problem to be solved by the present invention is to provide a sequencing method for carbohydrate molecules based on glycosidases and nanopores that can completely and sequentially resolve the composition, linkage, order and configuration of sugars (linear sugars or branched sugars).
[0012] Another technical problem to be solved by this invention is to provide a carbohydrate molecule sequencing system based on glycosidase and nanopores and its preparation method.
[0013] The technical problem to be solved by the present invention is to provide a nanopore mutant protein, a nucleic acid molecule encoding the nanopore mutant protein, an expression cassette containing the nucleic acid molecule, a recombinant vector, a recombinant cell or recombinant strain, a biological nanopore containing the nanopore mutant protein, and a biological nanopore-glycosidase complex containing the mutant protein.
[0014] The technical problem to be solved by the present invention is to provide the application of the nanopore mutant protein, the nucleic acid molecule encoding the nanopore mutant protein, the expression cassette containing the nucleic acid molecule, the recombinant vector, the recombinant cell or recombinant strain, the bio-nanopore containing the nanopore mutant protein, and the bio-nanopore-glycosidase complex containing the mutant protein in carbohydrate molecule sequencing.
[0015] Another technical problem that this invention aims to solve is to provide a nanopore sugar sequencer.
[0016] The final technical problem to be solved by this invention is to provide a data processing method, system and program product to automate glycosidase hydrolysis-assisted nanopore sugar sequencing.
[0017] Technical Solution: To solve the above-mentioned technical problems, a first aspect of the present invention provides a method for sequencing carbohydrate molecules based on nanopores, the method comprising: (1) Provide nanopores; (2) Construct a detection system containing the nanopores and glycosidases; (3) Add the sugar sample to be tested to the detection system; (4) The sugar sample to be tested undergoes a hydrolysis reaction with glycosidase. The sugar sequence information of the sugar sample to be tested can be inferred by comparing the changes in electrical signals before and after the hydrolysis reaction, or by inferring the sugar sequence information of the sugar sample to be tested by the characteristic signals generated by the monosaccharides released by the hydrolysis.
[0018] Wherein, the nanopores in step (1) are channels that are matched with the molecular size of the sugar sample to be tested and can be passed through by collision or displacement of sugar molecules, or not passed through by collision, or not passed through by covalent bonding, or have dissociation characteristics under the influence of external force, thereby producing a change in properties. The nanopores described in this invention have the following characteristics: they can be inserted into a planar lipid bilayer membrane in an electrolyte solution to form a stable open-pore current that can last for several hours; they can generate ionic current signals of sufficient frequency due to the interaction of sugar molecules passing through or not passing through; and they can continuously distinguish the chain length of oligosaccharides based on at least one characteristic parameter of the nanopore ionic current signal, thus achieving monosaccharide resolution.
[0019] Preferably, the channel whose properties change is a channel in which the change in ion current manifests as a change in electrical signal; Preferably, the nanopores include biological nanopores and / or solid nanopores; Preferably, the bio-nanopores include αHL, MspA, or aerolysin; Preferably, the nanopore is an αHL heptameric bio-nanopore with a monomer molecular weight of 35 kDa.
[0020] Preferably, the bio-nanopore comprises a protein complex consisting of at least one wild-type protein monomer and at least one mutant monomer, or a homologous protein complex consisting of homologous mutant monomers; Preferably, the mutant monomer is obtained by mutation at at least one position in NCBI Reference Sequence:WP_343219578.1 of αHL nanoporin.
[0021] Preferably, the αHL nanoporous mutant monomer includes αHL T109AαHL E111A αHL M113F αHL M113R αHL M113V αHL K147N αHL T115A αHL T117A αHL T117C αHL T117G αHL T117S αHL N121A αHL N121D αHL N121Q αHL N123D αHL N123Q αHL N123A αHL N139D αHL N139Q αHL M113H αHL M113K αHL M113D αHL M113E αHL T145R αHL G143R αHL M113R / T145R αHL M113R / G143R αHL M113R / S141A αHL M113R / T145A αHL M113R / T117A αHL M113R / T115A αHL M113R / E111A αHL M113R / K147A .
[0022] Preferably, this includes single-point mutagenesis homoheptamer bio-nanopore αHLM113R based on the narrowest part and vicinity of αHL, and multi-point mutagenesis homoheptamer bio-nanopore αHL based on the αHLβ barrel shape. M113R / T115A αHL M113R / K147A .
[0023] Preferably, the nanopore of the present invention is a nanopore capable of real-time detection of the hydrolysis of the target sugar molecule by exoglycosidases. Since the hydrolysis of sugar molecules by exoglycosidases proceeds stepwise from the non-reducing end, the nanopore of the present invention, in addition to requiring chain length monosaccharide resolution, also possesses the ability to exclude interference from the enzyme's own signal and the signal of the released monosaccharide.
[0024] The detection system for nanopores and glycosidases further includes aptamers, sugar-binding proteins, and / or sugar molecule modification tags; preferably, the aptamers include β-cyclodextrin, α-cyclodextrin, and γ-cyclodextrin; preferably, the sugar-binding proteins include lectins or sugar-capturing proteins.
[0025] The step (2) of constructing the detection system containing the nanopores and glycosidase includes: (i) Provide electrolyte solution: a solution capable of dissolving carbohydrate compounds and driving sugar molecules through electroosmosis to generate detectable ionic currents in the form of passing through or not passing through nanopores; Preferably, the electrolyte solution includes, but is not limited to, solutions of lithium chloride, sodium chloride, potassium chloride, magnesium chloride, and guanidine chloride at certain concentrations and pH values.
[0026] Among them, nanopore detection systems adapted to monosaccharides include, but are not limited to, engineered nanopores that enhance monosaccharide sensing and aptamers that enhance monosaccharide capture (such as β-cyclodextrin, α-cyclodextrin, γ-cyclodextrin), sugar-binding proteins (such as lectins, sugar-capturing proteins), and sugar molecule modification tags.
[0027] Among them, the electrical signal fingerprint spectrum of monosaccharide standards includes, but is not limited to, the size and distribution characteristics of individual electrical signal features caused by monosaccharide molecules in the nanopore detection system, two-dimensional scatter plots (e.g., ΔI1 / I0), and multi-dimensional feature sets.
[0028] The method of comparing and identifying the released monosaccharide electrical signal with the electrical signal fingerprint spectrum of the monosaccharide standard includes manual comparison or automatic comparison.
[0029] Preferably, the electrolyte solution includes KCl, NaCl, MgCl2 ion solutions of different concentrations or supersaturated solutions of the above ions; (ii) Providing an insulating layer: The insulating layer comprises a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores; preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts; (iii) Inserting and penetrating the nanoporous protein into the insulating layer to form a nanopore, and connecting the positive and negative terminals of the power supply on both sides of the insulating layer; preferably, the nanopore is located in the insulating layer and connects the two parts of the electrolyte solution; Preferably, the glycosidase is free in the electrolyte solution or is attached to the nanopore by chemical linkage or fusion expression to form a nanopore-glycosidase complex. Preferably, the nanopore-glycosidase complex simultaneously possesses the function of cleaving the sugar molecules to be tested and the function of generating electrical signals by the sugar molecules.
[0030] Preferably, the potential difference across the nanopore is typically between tens of mV and hundreds of mV; if it exceeds the upper limit, the pore is unstable and cannot be used for detection.
[0031] The detection system described in step (2) further includes the magnitude and direction of the potential applied across the nanopore, the type and concentration of the electrolyte, the pH of the electrolyte solution, the type and concentration of the buffer salt, and the direction of sugar addition. The electric field force on the sugar molecules is controlled by adjusting the magnitude and direction of the potential applied across the solution, and the magnitude and direction of the electroosmotic force on the sugar molecules are adjusted by controlling the magnitude and direction of the potential applied across the nanopore, the electrolyte salt concentration and pH, and the direction of sugar addition. The concentration-driven force on the sugar molecules is controlled by controlling the concentration of the sugar molecules in the electrolyte solution. Under the combined action of the electric field force, the concentration-driven force, and the electroosmotic force, the sugar molecules can generate an electrical signal, and this electrical signal is characteristic.
[0032] The glycosidase mentioned in step (2) includes an enzyme that recognizes the monosaccharide sequence and / or glycosidic bond of a sugar molecule and hydrolyzes and cleaves it from the glycosidic bond site of the sugar molecule. Preferably, the glycosidase includes a specific glycosidase or a non-specific glycosidase; Preferably, the glycosidase can be a single glycosidase or an enzyme array composed of multiple glycosidases; the enzyme array includes a variety of different arrangements and combinations. Preferably, the specific glycosidase includes an exoglycosidase that specifically recognizes the non-reducing end monosaccharide and glycosidic bond of a sugar molecule and hydrolyzes it to release the monosaccharide, or an endoglycosidase that specifically recognizes the internal sequence and glycosidic bond of a sugar molecule and hydrolyzes it to release disaccharide or other oligosaccharide fragments. Preferably, the nonspecific glycosidase includes a nonspecific exoglycosidase that hydrolyzes sugar molecules one by one from the ends or a nonspecific endoglycosidase that hydrolyzes sugar molecules from the inside. Preferably, the glycosidase comprises a specific exoglycosidase array suitable for nanopore sugar sequencing. The provided specific glycosidase can efficiently hydrolyze and release monosaccharides from the non-reducing ends of the target sugar molecule, one by one, in the provided detection system or other solutions. The exoglycosidase can be directly added to the above-mentioned detection system or other solutions to hydrolyze the sugar molecules, or it can be attached to the sugar molecule inlet of the nanopore provided in the first aspect to hydrolyze the sugar molecules entering the pore. The exoglycosidase array is a combination of all exoglycosidases that can potentially hydrolyze a certain class of sugar compounds.
[0033] Preferably, the glycosidase is located on the same side of the sugar sample to be tested.
[0034] In step (3), the sugar sample to be tested includes naturally occurring oligosaccharides, polysaccharides or their derivatives; Preferably, the sugar sample to be tested is a sugar with a chain length of n, where n is any integer; preferably, n is a natural number from 2 to 100; preferably, n is from 2 to 50; preferably, n is from 2 to 10; preferably, the sugar sample to be tested includes one or more of the following: decasaccharides, nonasaccharides, octasaccharides, heptasaccharides, hexasaccharides, pentasaccharides, tetrasaccharides, trisaccharides, or disaccharides. Preferably, the sugar sample to be tested is located in either side of the chamber.
[0035] The hydrolysis reaction in step (4) includes specific external hydrolysis, specific internal hydrolysis, non-specific external hydrolysis, or non-specific internal hydrolysis.
[0036] In step (4), the changes in electrical signals before and after the hydrolysis reaction are compared to infer whether the sequencing is forward or reverse. The forward sequencing is based on comparing the electrical signals caused by the sugar molecules released by glycosidase hydrolysis with the above-mentioned electrical signal fingerprint of sugar molecules to infer the sequence structure of sugar molecules; Specifically, regarding forward sequencing, firstly, an electrical signal fingerprint spectrum of monosaccharide standards is established using a monosaccharide-adapted nanopore detection system; then, monosaccharides are released one by one from the non-reducing ends of the target sugar molecules using enzymatic hydrolysis; the released monosaccharides are captured by the nanopores to generate electrical signals; the released monosaccharide electrical signals are compared and identified with the electrical signal fingerprint spectrum of the monosaccharide standards to obtain monosaccharide composition information.
[0037] Among them, monosaccharide-adapted nanopore detection systems include, but are not limited to, engineered nanopores that enhance monosaccharide sensing and aptamers that enhance monosaccharide capture (such as β-cyclodextrin (β-CD), α-cyclodextrin, γ-cyclodextrin), sugar-binding proteins (such as lectins, sugar-capturing proteins), and sugar molecule modification tags.
[0038] Among them, the electrical signal fingerprint spectrum of monosaccharide standards includes, but is not limited to, the size and distribution characteristics of individual electrical signal features caused by monosaccharide molecules in the nanopore detection system, two-dimensional scatter plots (e.g., ΔI1 / I0), and multi-dimensional feature sets.
[0039] The method of comparing and identifying the released monosaccharide electrical signal with the electrical signal fingerprint spectrum of the monosaccharide standard includes manual comparison or automatic comparison.
[0040] The reverse sequencing is based on comparing the electrical signals caused by carbohydrate compounds with the electrical signals caused by the residual oligosaccharides after glycosidase hydrolysis of sugar molecules to determine whether the glycosidase has successfully hydrolyzed the sugar molecules. Then, the sequence information of the sugar molecules is inferred based on the specificity of the glycosidases involved in the hydrolysis. Specifically, regarding the reverse sequencing method, this invention determines the occurrence of the hydrolysis reaction based on the difference between the electrical signal of the remaining sugar chain after hydrolysis and the electrical signal of the original substrate sugar molecule. Then, based on the specificity of the glycosidases involved in the hydrolysis reaction, it determines the monosaccharide composition, linkage structure, and stereoisomer information at that position.
[0041] Preferably, the comparison of changes in electrical signals before and after the hydrolysis reaction includes changes in the amplitude of the comparison ion current, changes in dwell time, changes in open frequency, Gaussian distribution fitting value or exponential distribution fitting value of the standard deviation, or parameters for comparing one or more characteristic distributions. Preferably, the parameters include KL divergence, JS divergence, Earth mover's distance (EMD distance), overlap coefficient, and Bhattacharyya distance. Preferably, the signal comparison is performed manually after statistical analysis or automatically by machine learning to identify signal differences.
[0042] The sugar sequence information mentioned in step (4) includes, but is not limited to, one or more of the following: monosaccharide unit type, monosaccharide unit arrangement order, connection mode between monosaccharide units (including α / β configuration and connection site), D / L configuration, modification rest (including but not limited to modification type such as sulfonation and acetylation, modification location, modification quantity, etc.), and branching information (including branching point location, number of branches, and the above sequence information of the branches).
[0043] A second aspect of the present invention provides a carbohydrate molecule sequencing system based on glycosidases and nanopores, the carbohydrate molecule sequencing system comprising: (1) Nanopore: The nanopore is a channel that has a size that matches the molecular size of the sugar sample to be tested and can be passed through by collision or displacement of sugar molecules, or not passed through by collision, or not passed through by covalent bonding, or has dissociation characteristics under the influence of external force, thereby producing a change in properties; (2) Glycosidase: Glycosidases that exist naturally or have been modified by enzyme engineering and may or may not have specificity; (3) Electrolyte solution: A solution that can dissolve carbohydrate compounds and drive sugar molecules to generate detectable ionic currents by electroosmosis in the form of passing through or not passing through nanopores; (4) Insulating layer: The insulating layer includes a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores; Preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts.
[0044] The nanopores are located in the insulating layer and are connected to the electrolyte solution. Preferably, the channel whose properties change refers to a channel in which the change in ion current manifests as a change in electrical signal; Preferably, the nanopores include biological nanopores and / or solid nanopores; Preferably, the bio-nanopores include α-hemolysin (abbreviated as αHL), MspA, or aerolysin; Preferably, the bio-nanopore comprises a protein complex consisting of at least one wild-type protein monomer and at least one mutant monomer, or a homologous protein complex consisting of homologous mutant monomers; The mutant monomer is obtained by mutation at at least one position in the NCBI Reference Sequence: WP_343219578.1 of αHL nanoporin; Preferably, the αHL nanoporous mutant monomer includes αHL T109A αHL E111A αHL M113F αHL M113R αHL M113V αHL K147N αHL T115A αHL T117A αHL T117C αHL T117G αHL T117S αHL N121A αHL N121D αHL N121Q αHL N123D αHL N123Q αHL N123A αHL N139D αHL N139Q αHL M113H αHL M113K αHL M113D αHL M113E αHL T145R αHL G143R αHL M113R / T145R αHL M113R / G143R αHL M113R / S141A αHL M113R / T145A αHL M113R / T117A αHL M113R / T115A αHL M113R / E111A αHL M113R / K147A .
[0045] The glycosidase is either free in the electrolyte solution or connected to the nanopore through chemical linkage or fusion expression to form a nanopore-glycosidase complex.
[0046] The glycosidase includes an enzyme that recognizes the monosaccharide sequence and / or glycosidic bond of a sugar molecule and hydrolyzes it from the glycosidic bond site of the sugar molecule. Preferably, the glycosidase includes a specific glycosidase or a non-specific glycosidase; Preferably, the specific glycosidase includes an exoglycosidase that specifically recognizes the non-reducing end monosaccharide and glycosidic bond of a sugar molecule and hydrolyzes it to release the monosaccharide, or an endoglycosidase that specifically recognizes the internal sequence and glycosidic bond of a sugar molecule and hydrolyzes it to release disaccharide or other oligosaccharide fragments. Preferably, the nonspecific glycosidase includes a nonspecific exoglycosidase that hydrolyzes sugar molecules one by one from the ends or a nonspecific endoglycosidase that hydrolyzes sugar molecules from the inside.
[0047] The electrolyte solution includes ion solutions of different concentrations such as KCl, NaCl, LiCl, and MgCl2, or supersaturated solutions of the above ions. The insulating layer includes a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores. Preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts to form two side chambers.
[0048] The bio-nanopore-based sugar sequencing system further includes an electrical signal monitoring system and a data processing system. The electrical signal monitoring system is used to monitor changes in the electrical signal of the nanopore caused by sugar molecules, and the data processing system is used to infer the sugar sequence information of the sugar sample to be tested by comparing the changes in electrical signals before and after the hydrolysis reaction.
[0049] The present invention includes, but is not limited to, engineered nanopores for enhanced monosaccharide sensing and aptamers for enhanced monosaccharide capture (such as β-cyclodextrin, α-cyclodextrin, γ-cyclodextrin), sugar-binding proteins (such as lectins, sugar-capturing proteins), and sugar molecule modification tags. A third aspect of the present invention provides a method for preparing a carbohydrate molecule sequencing system based on glycosidases and nanopores, comprising the following steps: (1) Providing nanopores: The nanopores are channels that are matched with the molecular size of the sugar sample to be tested and can be passed through by collisions or displacement of sugar molecules, or not passed through by collisions, or not passed through by covalent bonds, or have dissociation characteristics under the influence of external forces, thereby producing changes in properties. (2) Construct a detection system containing the nanopores and glycosidases.
[0050] The detection system for nanopores and glycosidases further includes aptamers, sugar-binding proteins, and / or sugar molecule modification tags; preferably, the aptamers include β-cyclodextrin, α-cyclodextrin, and γ-cyclodextrin; preferably, the sugar-binding proteins include lectins or sugar-capturing proteins.
[0051] The construction of the detection system containing the nanopores and glycosidase includes: (i) Provide electrolyte solution: a solution capable of dissolving carbohydrate compounds and driving sugar molecules through electroosmosis to generate detectable ionic currents in the form of passing through or not passing through nanopores; Preferably, the electrolyte solution includes ion solutions of different concentrations such as KCl, LiCl, NaCl, and MgCl2, or supersaturated solutions of the above-mentioned ions. (ii) Providing an insulating layer: The insulating layer comprises a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores; preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts; (iii) Inserting and penetrating the nanoporous protein into the insulating layer to form a nanopore, and connecting the positive and negative terminals of the power supply on both sides of the insulating layer; preferably, the nanopore is located in the insulating layer and connects the two parts of the electrolyte solution; Preferably, the glycosidase is free in the electrolyte solution or is attached to the nanopore by chemical linkage or fusion expression to form a nanopore-glycosidase complex. Preferably, the nanopore-glycosidase complex simultaneously possesses the function of cleaving the sugar molecules to be tested and the function of generating electrical signals by the sugar molecules.
[0052] In a fourth aspect, the present invention provides a nanopore mutant protein, said nanopore mutant protein comprising an amino acid mutation at at least one position in NCBI Reference Sequence: WP_343219578.1.
[0053] The nanopore mutant protein includes the mutant αHL. T109A αHL E111A αHL M113F αHL M113R αHL M113V αHL K147N αHL T115A αHL T117A αHL T117C αHL T117G αHL T117S αHL N121A αHL N121D αHL N121Q αHL N123DαHL N123Q αHL N123A αHL N139D αHL N139Q αHL M113H αHL M113K αHL M113D αHL M113E αHL T145R αHL G143R αHL M113R / T145R αHL M113R / G143R αHL M113R / S141A αHL M113R / T145A αHL M113R / T117A αHL M113R / T115A αHL M113R / E111A αHL M113R / K147A .
[0054] A fifth aspect of the invention provides a nucleic acid molecule encoding the aforementioned nanopore mutant protein.
[0055] In a sixth aspect, the present invention provides an expression cassette, a recombinant vector, a recombinant bacterial strain or a recombinant cell, comprising the aforementioned nucleic acid molecule.
[0056] A seventh aspect of the present invention provides a biological nanopore comprising the aforementioned nanopore mutant protein.
[0057] An eighth aspect of the present invention provides a nanopore-glycosidase complex comprising the aforementioned nanopore mutant protein.
[0058] The glycosidase is attached to the nanopore through chemical linkage or fusion expression to form a nanopore-glycosidase complex.
[0059] A ninth aspect of the present invention provides the application of the nanopore mutant protein, the nucleic acid molecule, the expression cassette, the recombinant vector, the recombinant strain or recombinant cell, the biological nanopore, and the nanopore-glycosidase complex in carbohydrate molecule sequencing.
[0060] In a tenth aspect, the present invention provides a nanopore sugar sequencer, the nanopore sugar sequencer comprising the aforementioned nanopore mutant protein.
[0061] The nanopore sugar sequencer further includes a glycosidase hydrolysis system, a nanopore glycan electrical signal monitoring system, and / or a data processing system. The glycosidase hydrolysis system is used to hydrolyze sugar molecules with glycosidases, the nanopore glycan electrical signal monitoring system is used to monitor changes in the electrical signal of the nanopore caused by sugar molecules, and the data processing system is used to infer the sugar sequence information of the sugar sample to be tested by comparing the changes in the electrical signal before and after the hydrolysis reaction.
[0062] In an eleventh aspect, the present invention provides a data processing method for glycosidase-assisted nanopore sugar sequencing, comprising the following steps: Obtain the electrical signals before and after the reaction of glycosidase hydrolyzing sugar molecules; Data analysis is performed on the acquired electrical signals. Based on the comparison between the first electrical signal before the hydrolysis reaction and the second electrical signal after the hydrolysis reaction, it is determined whether the glycosidase has successfully hydrolyzed the sugar molecules. Then, based on the characteristics of the glycosidase involved in the hydrolysis, the sugar molecules that are released step by step are inferred, thereby inferring the sequence structure of the sugar chain in reverse.
[0063] In some embodiments, the signal comparison includes comparing the amplitude variation, dwell time variation, open frequency variation, Gaussian distribution fitting value or exponential distribution fitting value of the standard deviation of the current signal, or comparing parameters of one or more feature distributions, including KL divergence, JS divergence, EMD distance, overlap coefficient and Bartholomew's distance.
[0064] In some embodiments, determining whether the glycosidase has successfully hydrolyzed sugar molecules based on comparing a first electrical signal before the hydrolysis reaction with a second electrical signal after the hydrolysis reaction includes: Extract the first electrical signal features and the second electrical signal features, calculate the difference between the first electrical signal features and the second electrical signal features, input the difference into a trained classification model, and obtain a classification result on whether the first electrical signal features and the second electrical signal features are the same; or, extract the first electrical signal features and the second electrical signal features, input the first electrical signal features and the second electrical signal features into a trained classification model, and obtain a classification result on whether the first electrical signal features and the second electrical signal features are the same. If the first electrical signal characteristic is different from the second electrical signal characteristic, it is determined that the glycosidase has successfully hydrolyzed the sugar molecule.
[0065] In some embodiments, for an unknown glycan sequence, the process of inferring the complete glycan sequence by reverse deduction includes: S1, obtain the initial current signal feature Data_m, where m represents the number of successful enzymatic hydrolysis and the position of free monosaccharides in the original glycan chain, with an initial value of 1; S2, obtain the current signal characteristics of each glycosidase hydrolysis product in the glycan and enzyme array, Data_m_f, where f represents the sample number after enzymatic hydrolysis of different types of enzymes; S3. Compare all Data_m_f with Data_m to determine whether hydrolysis has occurred. For samples that have been determined to have hydrolyzed, output f and m. The corresponding Data_m_f is used as the reference data Former_data for the next round of hydrolysis comparison. This sample is used as the initial sample Former_sample for subsequent hydrolysis. Update the successful hydrolysis count m=m+1. S4, repeat S2-S3 until the complete sugar chain sequence is determined.
[0066] A twelfth aspect of the present invention also provides a data processing method for glycosidase-assisted nanopore sugar sequencing, comprising the following steps: Obtain the electrical signals caused by the release of sugar molecules from glycosidase hydrolysis; Based on the comparison of electrical signals and fingerprint patterns caused by the release of sugar molecules by glycosidase hydrolysis, the gradually released sugar molecules are inferred, thereby leading to the forward inference of the sequence structure of the sugar chain.
[0067] In some embodiments, the fingerprint spectrum includes a monosaccharide fingerprint spectrum library, a disaccharide fingerprint spectrum library, and other oligosaccharide fingerprint spectrum libraries. The fingerprint spectrum is used to assign signals to electrical signals, obtain the types of terminal oligosaccharide fragments during hydrolysis, and deduce the compositional order of oligosaccharide fragments based on the order of occurrence of electrical signal events and / or the order of signal abundance.
[0068] In some embodiments, the fingerprint is established based on electrical signal features, which include one or more of the following: amplitude of current signal, dwell time, open frequency, and standard deviation.
[0069] In a thirteenth aspect of the present invention, a data processing system for glycosidase-assisted nanopore sugar sequencing is also provided, comprising: The reverse sequencing input module is used to acquire electrical signals before and after the glycosidase hydrolysis of sugar molecules. The reverse sequencing information processing module is used to analyze the acquired electrical signals. Based on the comparison of the first electrical signal before the hydrolysis reaction and the second electrical signal after the hydrolysis reaction, it determines whether the glycosidase has successfully hydrolyzed the sugar molecules. Then, based on the characteristics of the glycosidase involved in the hydrolysis, it infers the sugar molecules that are released step by step, thereby inferring the sequence structure of the sugar chain in reverse.
[0070] A fourteenth aspect of the present invention also provides a data processing system for glycosidase-assisted nanopore sugar sequencing, comprising: The forward sequencing input module is used to acquire the electrical signals caused by the release of sugar molecules by glycosidase hydrolysis; The forward sequencing information processing module infers the progressively released sugar molecules based on the electrical signals and fingerprint patterns caused by the hydrolysis of sugar molecules by glycosidases, thereby forward inferring the sequence structure of the sugar chain.
[0071] According to the fifteenth aspect of the present invention, a computer system is also provided, including a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor, characterized in that, when the computer program / instructions are executed by the processor, they implement the steps of the data processing method for glycosidase hydrolysis-assisted nanoporous sugar sequencing as described in either the eleventh or twelfth aspect.
[0072] In a sixteenth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a data processing method for glycosidase hydrolysis-assisted nanoporous sugar sequencing as described in either the eleventh or twelfth aspect.
[0073] In a seventeenth aspect of the present invention, a computer program product is also provided, comprising a computer program / instructions, characterized in that, when the computer program / instructions are executed by a processor, they implement the steps of a data processing method for glycosidase hydrolysis-assisted nanoporous sugar sequencing as described in any one of the eleventh or twelfth aspects.
[0074] An eighteenth aspect of the present invention also provides a nanopore sequencer based on glycosidase hydrolysis, comprising a glycosidase hydrolysis system, a nanopore glycan electrical signal monitoring system, and a data processing system; wherein the glycosidase hydrolysis system is used for glycosidase hydrolysis of sugar molecules, and the nanopore glycan electrical signal monitoring system is used for monitoring changes in nanopore electrical signals caused by sugar molecules; the data processing system is a glycosidase hydrolysis-assisted nanopore glycan sequencing data processing system as described in either the thirteenth or fourteenth aspect; or, the data processing system automatically generates the sequence structure of glycans based on the computer program product described in the seventeenth aspect.
[0075] A nineteenth aspect of the present invention also provides a nanopore glycan electrical signal monitoring system, comprising two chambers separated by an insulating layer, a nanopore formed on the insulating layer, electrodes placed in the two chambers for measuring changes in electrical signals, and an instrument for acquiring changes in electrical signals; wherein the nanopore is a channel having a size matching the molecular size of the sugar sample to be tested and being able to pass through by collisions or displacement of sugar molecules, or not pass through by collisions, or not pass through by covalently bonded molecules, or pass through or not pass through under the influence of external forces, thereby causing a change in properties; Preferably, the channel whose properties change refers to a channel in which the change in ion current manifests as a change in electrical signal; Preferably, the nanopores include biological nanopores and / or solid nanopores; Preferably, the bio-nanopores include αHL, MspA, or aerolysin; Preferably, the bio-nanopore comprises a protein complex consisting of at least one wild-type protein monomer and at least one mutant monomer, or a homologous protein complex consisting of homologous mutant monomers; The mutant monomer is obtained by mutation at at least one position in αHL nanopore NCBI Reference Sequence: WP_343219578.1.
[0076] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: 1) This invention provides a novel sugar sequencing method and system that utilizes the specificity and efficiency of glycosidases to rapidly release sugar units, and leverages the high sensitivity of nanopores to detect the structural information of the released monosaccharides or remaining sugar chains. This not only retains the high sensitivity, high resolution, and high throughput characteristics of nanopore technology, but also helps expose the internal structural details of the sugar chains for complete recognition by the nanopores, avoiding the problem of structural extension. Compared with decaglycolic chains and N-saccharides with branched sugars of similar complexity, this invention achieves an analysis accuracy exceeding 98%, a detection sensitivity of 0.1 μM, representing a sensitivity improvement of more than 50 times and a time saving of more than 5 times.
[0077] 2) The sugar sequencing method of this invention can employ both forward and reverse sequencing methods. The reverse sequencing method, by monitoring any changes in electrical signals before and after hydrolysis, can directly determine which exoglycosidase is involved in the hydrolysis reaction. This allows the invention to reverse-engineer the terminal monosaccharides and glycosidic bonds of the glycan chain; repeating this process continuously will achieve de novo sequencing. The reverse sequencing method avoids dependence on electrofinite fingerprint libraries of monosaccharides and more complex glycan chains, overcoming a key limitation in the field caused by the lack of high-resolution and comprehensive glycan fingerprint databases. This is unprecedented in the fields of sugar structure analysis and nanopore-based nucleotide and protein sequencing.
[0078] 3) Based on the data processing method, system, and program products provided by this invention, automated result interpretation can be achieved, eliminating the need for manual post-analysis of electrical signals and reducing the professional knowledge requirements of operators. As long as the hydrolysis reaction proceeds effectively, the sugar chain composition can be continuously read with high accuracy and speed, meeting a wider range of analytical needs. Attached Figure Description
[0079] Figure 1 This is a structural diagram of the carbohydrate molecule sequencing system based on glycosidase and nanopores of the present invention.
[0080] Figure 2 This is a flowchart of a nanoporous sugar sequencing method based on glycosidase hydrolysis.
[0081] Figure 3 This is a schematic diagram of the glycosidase, glycan chain, nanopore unit, and single-channel or array-based network detection device of the present invention.
[0082] Figure 4 This is a purified gel image of nanoporin αHL (M113R / T115A) (abbreviated as nanoporous A).
[0083] Figure 5This is a schematic diagram of the nanopore detection system and the property identification of the nanopores in Example 1. A. Nanopore detection system; B. Single-channel current of αHL (M113R / T115A); C. Current property analysis of αHL (M113R / T115A), including current-voltage curves and Gaussian fitting of the opening current.
[0084] Figure 6 This is a current signal event (A) and its characteristic description diagram (B). Different levels of current blockage are named In, such as the first level of blockage being named I1, the second level of blockage being named I2, and each current blockage is a signal event.
[0085] Figure 7 This is a schematic diagram of the sugar chain, detection device, and electrical signals under different detection conditions in Example 2. A. Structure of Poly-LacNAc monosaccharide to decasaccharide. B. Under the first detection condition (Trans side +100mV, Cis sugar to be tested added). C. Schematic diagram showing no characteristic signal generated under the second detection condition (Trans side -40mV, Trans side sugar to be tested added).
[0086] Figure 8 These are the electrical signal characteristic diagrams of the sugar chain standards involved in Example 2. A. Level 1 events for poly-LacNAc decasaccharides to trisaccharides. B. Level 2 events for disaccharides.
[0087] Figure 9 These are the electrical signal analysis graphs of the glycan standards involved in Example 2. A. Scatter plot of Dwell time versus ΔI1 / I0 for decasaccharide to trisaccharide signals. B. Gaussian distribution fitting of ΔI1 / I0 for decasaccharide to trisaccharide signals. C. Exponential fitting of Dwell time for decasaccharide to trisaccharide signals. D. Scatter plot of Dwell time versus ΔI2 / I1 for disaccharide signals. E. Gaussian distribution fitting of ΔI2 / I for disaccharide signals. F. Exponential distribution fitting of Dwell time for disaccharide signals.
[0088] Figure 10 This is a schematic diagram showing that the two monosaccharides Gal and GlcNAc, the standards of Example 2, do not produce characteristic signals under either the first detection condition (Trans side +100mV, Cis added to the test sugar) or the second detection condition (Trans side -40mV, Trans side added to the test sugar).
[0089] Figure 11 This is a graph showing the results of Coomassie Brilliant Blue staining and Western blot identification of purified β-N-acetylglucosidase (GlcNAcH) and β1-4-galactosidase (GalH) in Example 3.
[0090] Figure 12 In Example 3, the final concentrations of GlcNAcH(A) or Gal(B) glycosidase were set to 0.005 mg / mL, 0.05 mg / mL, and 0.5 mg / mL, respectively, and the final concentration of the substrate nonaglycone was set to 10 mM. The reaction system was reacted at 37°C, and samples were taken at different time points, including 1, 2, 5, and 10 min. The results were analyzed by TLC.
[0091] Figure 13 This is a flowchart of the hydrolysis-assisted nanopore "reverse sugar sequencing" strategy in Example 4.
[0092] Figure 14 This is a schematic diagram of the stepwise hydrolysis of the decasaccharide 1X-2X-3X-4X-5X-6X-7X-8X-9X-10X in Example 4.
[0093] Figure 15 This is a real-time detection graph of the stepwise hydrolysis system of the decasaccharide based on nanopores in Example 4.
[0094] Figure 16 These are the characteristic currents and analysis graphs of the sugar and its hydrolysis products tested in Example 4. A. Changes in current signal during stepwise hydrolysis; B. Two-dimensional scatter plot of signal characteristics of the stepwise hydrolyzed sample; C. Gaussian distribution fitting of ΔI1 / I0 or I2 / I0 for the stepwise hydrolyzed sample.
[0095] Figure 17 The sequence diagram of the deca-saccharide to be tested, derived from the results of the nine-step hydrolysis in Example 4, is as follows: Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc.
[0096] Figure 18 These are the liquid chromatography verification chromatogram (A) and mass spectrometry verification chromatogram (B) of the stepwise hydrolysis sample of Example 4.
[0097] Figure 19 This is a flowchart of the automated nanopore sugar sequencing program based on machine learning in Example 5.
[0098] Figure 20 This is the automated sequencing flowchart of Example 5.
[0099] Figure 21 This is a diagram showing the automated sugar sequencing results obtained using the machine learning model and automated sequencing process established in Example 5.
[0100] Figure 22 This is a graph showing the machine learning model and prediction accuracy results in a real-world scenario, as described in Example 5.
[0101] Figure 23 This is a schematic diagram of the forward sequencing strategy in Example 6, where nanopores capture and collect terminal oligosaccharide fragments released by hydrolysis.
[0102] Figure 24 These are the purification results and detection system diagrams for nanoporous B in Example 6. A. Coomassie brilliant blue staining to identify the purification results of nanoporous B; B. Nanoporous B detection system.
[0103] Figure 25 This is a schematic diagram of the characteristic parameters of the Level 2 signal of sugar molecules collected by the nanopore B detection system in Example 6.
[0104] Figure 26 It is the fingerprint spectrum of the three monosaccharides sialic acid (Neu5Ac), galactose (Gal), and glucosamine (GalNAc) in Example 6.
[0105] Figure 27 This is a two-dimensional scatter plot of the stepwise hydrolysis of the trisaccharide during forward sequencing in Example 6.
[0106] Figure 28 This is a diagram illustrating the stepwise hydrolysis process of the branched sugar N-sugar H4 in Example 5.
[0107] Figure 29 The diagram shows the enzyme array hydrolysis-sequencing system and sequencing derivation results of branched thymol N-sugar in Example 5.
[0108] Figure 30 The images show the signal spectra collected by the nanopore detection system during the stepwise hydrolysis of N-sugar H4 in Example 5. A is a scatter plot of H4; B is a scatter plot of H3A; C is a scatter plot of H2B; D is a scatter plot of H1C. Figure 31 This refers to the monosaccharides and remaining sugar chains released by the stepwise cleavage of N-glycan H4 by glycosidase in Example 5.
[0109] Figure 32 shows the reverse sequencing process of HMO sugars in Example 5. A is a diagram of the nanopore C detection system; B is a schematic diagram of the hydrolysis of HMO sugar molecules by FucH and NanA; C and D are the current signal changes of the nanopore protein C detection system. Detailed Implementation
[0110] In this invention, we developed a glycosylation sequencing strategy combining glycosidase and nanopore detection technology. A specific exoglycosidase is used to hydrolyze and cleave the non-reducing ends of the target glycan molecule, with each hydrolysis reaction generating a released monosaccharide and a residual oligosaccharide. High-precision and high-sensitivity nanopore technology is used to detect the hydrolysis system of the glycosidase and the target glycan molecule. Then, by analyzing the characteristic current changes of the glycosidase hydrolysis products generated by the nanopore detection system, the monosaccharide unit composition, sequence arrangement, and linkage of the target glycan molecule can be read and inferred in real-time and stepwise. In this process, an automated analysis and judgment algorithm was developed by incorporating machine learning algorithms, thereby achieving real-time automatic and rapid identification of different current patterns. Finally, an automated nanopore process assisted by stepwise hydrolysis using a multi-exoglycosidase array was developed, achieving efficient automated nanopore glycan sequencing. The nanopore glycan sequencing method mainly involves the following steps: Step 1: Use exoglycosidase to stepwise hydrolyze the sugar chains of the sugar sample to be tested, and obtain the stepwise hydrolysis products of the sugar chains; Step 2: Use a nanopore sensor to obtain the current signal generated by the sugar chain to be tested and its stepwise hydrolysis products in the nanopore detection system, perform noise reduction processing on the obtained current signal, and obtain the current signal data of the sugar molecule to be tested and its hydrolysis products. Step 3: Glycan sequencing using nanopores and glycosidases: Extract the parameter distribution characteristics of the current signals obtained before and after hydrolysis; use the developed machine learning method to assign the current signals (for forward sequencing) or evaluate the differences (for reverse sequencing) to realize the exonuclease hydrolysis reaction at each step; Step 4: Based on the signal attribution or differential assessment results of the stepwise hydrolysis in Step 3, and the specific characteristics of the enzymes used, the sequence information of sugar molecules is automatically generated using the developed automated sequencing program, including the types of monosaccharide units, the unit order, and the connection methods between them.
[0111] This invention relates to nanopores with monosaccharide resolution. The nanopores are αHL heptamer bio-nanopores with a monomer molecular weight of 35 kDa. This includes single-point mutagenesis of homoheptamer bio-nanopores αHL based on the narrowest part and vicinity of the αHL. M113R Multi-point mutant homoheptamer bionanopores based on αHLβ barrel structure αHL M113R / T115A αHL M113R / K147A The provided nanopores possess the following characteristics: they can be inserted into a planar lipid bilayer in an electrolyte solution to form a stable open-pore current that can last for several hours; they can generate ionic current signals of sufficient frequency due to the interaction of sugar molecules passing through or not passing through; and they can continuously distinguish the chain length of oligosaccharides based on at least one characteristic parameter of the nanopore ionic current signal, achieving monosaccharide resolution.
[0112] This invention also relates to a detection system that can drive analyte sugar molecules to generate characteristic electrical signals by passing through or not passing through nanopores. This detection system includes the magnitude and direction of the potential applied across the nanopore, the type and concentration of the electrolyte, the pH of the electrolyte solution, the type and concentration of the buffer salt, and the direction of sugar addition. The electric field force received by the sugar molecules is controlled by controlling the magnitude and direction of the potential applied across the solution. The magnitude and direction of the electroosmotic force on the sugar molecules are adjusted by controlling the magnitude and direction of the potential applied across the nanopore, the electrolyte salt concentration and pH, and the direction of sugar addition. The concentration-driven force on the sugar molecules is controlled by controlling the concentration of the analyte sugar molecules in the electrolyte solution. Under the combined action of the electric field force, the concentration-driven force, and the electroosmotic force, the analyte sugar molecules can generate an electrical signal, and this electrical signal is characteristic.
[0113] This invention also relates to a specific exoglycosidase array for use in nanopore sugar sequencing. The provided specific glycosidases can efficiently hydrolyze and release monosaccharides from the non-reducing ends of the target sugar molecules one by one, either in the detection system or in other solutions. The exoglycosidases can be added directly to the detection system or other solutions to hydrolyze the sugar molecules, or they can be attached to the sugar molecule inlet of the nanopore to hydrolyze the sugar molecules entering the pore. The exoglycosidase array is a combination of all exoglycosidases that can potentially hydrolyze a specific class of sugar compounds.
[0114] This invention also relates to nanopores capable of real-time detection of the hydrolysis of analyte sugar molecules by exoglycosidases. The hydrolysis of sugar molecules by exoglycosidases proceeds stepwise from the non-reducing end. Therefore, the nanopores of this invention, in addition to possessing long-chain monosaccharide resolution, also have the ability to eliminate interference from the enzyme's own signal and the signal of the released monosaccharide.
[0115] This invention also relates to a procedure for automating nanopore glycolysis sequencing.
[0116] The sequencing principle of this invention: This invention uses a specific exoglycosidase to hydrolyze and cleave the non-reducing ends of the target glycan molecule, with each hydrolysis reaction generating a released monosaccharide and a residual oligosaccharide. High-precision, high-sensitivity nanopore technology is used to detect the hydrolysis system of the glycosidase and the target glycan molecule. Then, by analyzing the characteristic current changes of the glycosidase hydrolysis products generated by the nanopore detection system, the monosaccharide unit composition, sequence arrangement, and linkage of the target glycan molecule can be read and inferred in real-time and stepwise. In this process, an automated analysis and judgment algorithm was developed by combining machine learning algorithms, thereby achieving real-time automatic and rapid identification of different current patterns. Finally, an automated nanopore process assisted by stepwise hydrolysis of multiple exoglycosidase arrays was developed, achieving efficient automated nanopore glycan sequencing. The basic principle of the sequencing system in this invention is based on the natural specificity of glycosidases, combined with the sensitivity of nanopores to subtle structural changes in the analyte. Specifically, it includes: 1) The exoglycosidases tested in this invention currently possess specificity for recognizing sugar units and glycosidic bonds. They can sequentially and precisely recognize the monosaccharides at the non-reducing ends of the target sugar chain and their linkage patterns, and then hydrolyze and cleave them, thereby generating two hydrolysis products: monosaccharides released from the ends and residual oligosaccharides. This is different from the other two types of biomolecules (nucleic acids and proteins) currently being tested.
[0117] 2) The detection system used in this invention: The nanopore sequencing system of this invention distinguishes, identifies, and identifies the structural differences of analytes by detecting the current disturbance generated after the analyte enters or passes through the nanopore. This is completely different from previously reported glycan structure detection, identification, and characterization systems. Therefore, in the sequencing system of this invention, the current changes generated by the release of monosaccharides or residual oligosaccharides in each step of the exoglycosidase hydrolysis reaction are captured by the nanopore detection system, thereby sequentially interpreting the structural information of the glycans. Based on the capture of monosaccharides or residual oligosaccharides by the nanopore detection system, this sequencing method is divided into "forward sequencing" and "reverse sequencing".
[0118] Forward sequencing refers to the process where a nanopore detection system captures monosaccharides released by hydrolysis and generates characteristic electrical signals. Based on a pre-established monosaccharide fingerprint, the electrical signals are assigned to determine the types of monosaccharide units in the target sugar molecule. The positional order of the monosaccharide units in the target sugar molecule is then inferred based on the order of monosaccharide release during hydrolysis.
[0119] Reverse sequencing refers to the nanopore detection system capturing the residual sugar chains after monosaccharide release during hydrolysis. Since the residual sugar chains generated after each glycosidase hydrolysis step differ from the sugar chains before that step by a single monosaccharide, the nanopore detection system with monosaccharide resolution can generate different electrical signals before and after the hydrolysis reaction. Machine learning models are used to identify these differences. If the machine learning results show a difference in the electrical signals before and after a certain hydrolysis step, it indicates that the exoglycosidase hydrolysis reaction at that step occurred successfully. Based on the specificity of the exoglycosidase at that step, the type of terminal monosaccharide and the linkage type can be deduced. Based on multi-step enzymatic digestion reactions, the continuous sequence information of the target sugar molecule can be completely inferred.
[0120] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0121] The partial glycosidases used for hydrolyzing nanopore sugars in the embodiments of the present invention are shown in Table 1.
[0122] Table 1: Glycosidases used for hydrolysis-assisted nanopore sugar sequencing (partial list)
[0123] Example 1: Preparation of a nanopore-based sequencing system and construction of a sequencing method Using poly-LacNAc decaglucan as a template glycan, the sensitivity of the nanopore to glycans of different lengths and its ability to achieve monosaccharide resolution were confirmed.
[0124] The specific experimental method is as follows: 1. Preparation of nanopores with monosaccharide resolution: 1.1 Construction of expression vector: The α-hemolysin (α-HL) gene sequence with a 6×His tag at the C-terminus was inserted into the pEASY vector (purchased from BGI Sequencing) through NdeI and HindⅢ restriction sites to form a plasmid that can express α-hemolysin.
[0125] Gene sequence of α-hemolysin: NCBI Reference Sequence: WP_343219578.1.
[0126] 5'' 1.2 Mutation of the α-hemolysin gene: The α-hemolysin (α-HL) gene sequence with a 6×His tag at the C-terminus was inserted into the pEASY vector (purchased from BGI Sequencing) through NdeI and HindⅢ restriction sites to form a plasmid that can express α-hemolysin.
[0127] Taking α-HL (M113R / T115A) as an example, we carried out genetic engineering modification of α-hemolysin nanopores. Based on the α-hemolysin gene sequence, we designed and synthesized point mutation primers for M113R and T115A, and used the point mutation primers to perform point mutations on plasmids expressing α-hemolysin.
[0128] PrimeSTAR Max DNA polymerase (purchased from Novizan) was used with synthesized point mutation primers. Using the pEASY plasmid containing the α-hemolysin gene as a template, KeyPo Master Mix (purchased from Novizan, catalog number PK511) was used for PCR to perform point mutation and amplification of the mutant plasmid. The PCR amplification system and program settings were followed according to the KeyPo Master Mix instruction manual. PCR amplification yielded a linear plasmid with a site-directed mutation. Dpn I restriction endonuclease was added to the system to digest and destroy the unmutated plasmid template. The mutant plasmid was transformed into Trans5α competent *E. coli* (transformation steps were followed according to the *Trans5α competent *E. coli* instruction manual of Beijing TransGen Biotech Co., Ltd.). The transformation steps were as described in the antibiotic plate coating instructions. After selecting single clones and confirming their sequence by BGI Genomics, amplification was performed. The plasmid was then extracted using a small-scale extraction kit (purchased from Yisheng Biotechnology). This plasmid was the mutant plasmid containing the α-HL (M113R / T115A) gene sequence. The plasmid was stored at -20℃ before use.
[0129] Point mutation primers for α-HL (M113R): Upstream primer 5'-GATACAAAAGAGTATCGTAGTACTTTAACTTATG-3' Downstream primer: 5'-CATAAGTTAAAGTACTACGATACTCTTTTGTATC-3' Mutant primers for α-HL (T115A): Upstream primer: 5'-GAGTATCGTAGTgccTTAACTTATG-3' Downstream primer: 5'-CATAAGTTAAggcACTACGATACTC-3' 1.3 The mutant plasmid from 1.2 was transformed into competent *E. coli* BL21(DE3)pLysS, amplified, and induced to express nanoporin α-HL (M113R / T115A). The protein was purified by Ni-NTA affinity resin column purification. The purified α-hemolysin (M113R / T115A) protein was stored in buffer (10 mM Tris-HCl, pH 8.0 and 50 mM NaCl). In the following description, nanoporin α-HL (M113R / T115A) will be referred to as nanoporin A. Figure 4 ).
[0130] 2. Construction of the nanopore detection system 2.1 Construction of planar lipid bilayer: Holes with a diameter of 30~120 μm were electrically spark-drilled on a polytetrafluoroethylene membrane, and an artificial lipid bilayer membrane was constructed on the micropores using diaphytophosphatidylcholine (DPhPC) to divide the chamber of the measuring device into two sides.
[0131] 2.2 Single-channel nanopore insertion into lipid bilayer membrane: 260 μL of electrolyte solution was added to both chambers described in 2.1, and a certain potential difference was applied across the lipid bilayer membrane ( Figure 5 (A) After dilution of nanoporin α-HL (M113R / T115A) and addition to either chamber, the nanoporin oligomerizes on the lipid bilayer membrane described in 2.1 to form nanopores. When a stable single-channel current is recorded ( I When the value is 0, it indicates that a single nanoporous protein has successfully inserted into the lipid bilayer membrane. Figure 5 (B in the text). During this process, we measured the current properties of the engineered nanopore A. Figure 5 C in the text). The addition of protein to the side is defined as cis( Cis ) chamber, defining the opposite chamber as inverse ( Trans ) chamber.
[0132] The nanopores used in this invention are not limited to the mutant nanopore; other nanopores can also be used to build the nanopore detection system described above.
[0133] The nanopore detection system may also include glycosidases or glycosidase arrays.
[0134] 3. Using a nanopore detection system to identify and acquire signals from different sugar chains. The solution of the sugar molecules to be tested is added one by one to either side of the nanopore detection system described in step 2 at a certain initial concentration and volume to form a certain final concentration (including but not limited to final concentrations at the pM, μM, and mM levels; in this embodiment, 0.1 μM is preferred). The sugar molecules to be tested will enter the nanopore channels or collide with the nanopore channels under the driving force of concentration and / or electrophoresis and / or electroosmotic flow, causing changes in the opening current of the nanopore channels. These changes in opening current are generally, but not limited to, current blockage. Different levels of current blockage are named... I n For example, the first layer of blockage is named I 1. The second layer of blockage is named I 2. Each current blockage becomes a signal event. Figure 6 A). Signal events were recorded and collected using a nanopore instrument, Cube-D2, and then denoised. In this embodiment, a 5 kHz low-pass filter and a 50 kHz sampling frequency were used.
[0135] 4. Extract the signal event features of sugar molecules as signal data and perform statistical analysis. All event current traces were processed using PyNanoLab software. Characteristic parameters of the signal events of different chain length sugar molecules generated by the nanopore monitoring system in step 3 were extracted using PyNanoLab software, including but not limited to the dwell time and amplitude (ΔI) of the signal events. n ), ΔI n / I0, standard deviation (Std). In this embodiment, signals with Dwell time ≤ 0.2 ms or ΔI1 / I0 ≤ 0.15 are ignored for further noise reduction. The features of the extracted sugar molecule signal events are referred to as the signal data. Further statistical analysis and plotting are performed using GraphPad Prism 9 and Origin 2022 software to describe the resolution of sugar molecules of different chain lengths by the nanopore. In this embodiment, a scatter plot of Dwell timeversus ΔI1 / I0 is used to visualize the degree of distinction of different sugar chains by the nanopore, and the separation degree (S) between the Gaussian distributions of the individual parameter ΔI1 / I0 in the signal data of different sugar molecules is used to quantitatively describe the degree of distinction of different sugar chains by the nanopore. Figure 6 B). Resolution (S) is calculated using the following formula: S = |Ps-Pm| / (Ws+Wm), where Ps and Pm represent the peak values of the Gaussian fitting of the parameters ΔI1 / I0 for different sugar molecules, and Ws and Wm represent the half-width of the Gaussian fitting of the parameters ΔI1 / I0 for different sugar molecules.
[0136] 5. Verify the monosaccharide resolution of nanopores for sugar molecules of different chain lengths. Nanopore detection systems can be used to verify the monosaccharide resolution of nanopores by detecting decasaccharides down to monosaccharides. When the detection system also includes glycosidases or glycosidase arrays, it can also be used for sequencing actual sugar samples.
[0137] Example 2: Sequencing method for decasaccharide to monosaccharide molecules based on a nanopore sequencing system
[0138] In this embodiment, poly-LacNAc series sugar molecules of different chain lengths, composed of galactose (Gal) and acetylglucosamine (GlcNAc) repeating units (LacNAc), were synthesized as model sugar chains. LacNAc polysaccharides were used as model sugar chains to test the resolution of the nanopore. Figure 7 A).
[0139] The decasaccharide in this embodiment: Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Nonose: GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Octose: Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Heptasaccharide: GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Hexasaccharide: Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Five sugars: GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Tetrasaccharide: Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc Trisaccharides: GlcNAcβ1-3Galβ1-4GlcNAc Disaccharide: Galβ1-4GlcNAc Monosaccharides: Galβ; GlcNAc; The specific experimental steps and results of this embodiment are as follows:
[0140] (1) α-HL (M113R / T115A), abbreviated as nanoporous protein A, was prepared according to step 1 of Example 1. Figure 4 ).
[0141] (2) A nanopore detection system was constructed using purified nanopore protein A according to step 2 of Example 1.
[0142] Specific experimental steps and conditions: Add 260 μL of 3M KCI buffer (containing 10 mM citric acid, pH 5.0) to both the Cis and Trans chambers. Apply a +100 mV voltage to the Trans side. After constructing a phospholipid bilayer using diaphytylphosphatidylcholine (DPhPC), add nanopore A protein dilution buffer (diluted to mg / mL levels, e.g., 0.1 mg / mL, using TBS buffer) to the Cis side. The dilution buffer is added to the nanopore detection system to a final concentration of μg / mL, e.g., 0.1 μg / mL. Nanopore A exhibits a Gaussian fitting value of 269.50 ± 3.23 pA (I) at +100 mV. 0 The single-channel orifice current and conductivity (G) of 2458.21 ± 16.96 pS ( Figure 5 ).
[0143] (3) Following step 3 of Example 1, the current signal events of decasaccharides, nonasaccharides, octasaccharides, heptasaccharides, hexasaccharides, pentasaccharides, tetrasaccharides, trisaccharides, disaccharides, and monosaccharides were detected and collected sequentially using the nanopore A detection system. Figure 7 A). Each sugar standard sample was added to the nanopore detection system at a final concentration of 0.1 μM. During the detection process, all analyte sugar chains were first added to... Cis On the side, signal acquisition is performed. Figure 7 B); as in Cis If no characteristic signal is detected by side detection, the sugar molecule to be tested is added. Trans Detection is performed on the side ( Figure 7 C). Furthermore, during this process, if no characteristic signal is observed at the current detection voltage, the electroosmotic flow will be altered by adjusting the detection voltage, thereby increasing the signal capture of the glycan chains by the nanopores. The results are as follows... Figure 8 As shown in Figure A, the decasaccharides, nonasaccharides, octasaccharides, heptasaccharides, hexasaccharides, pentasaccharides, tetrasaccharides, and trisaccharides all generated reversible and continuous first-level current blocking signal events, named Level 1 signal events. I 1). Regarding disaccharides, the results are as follows: Figure 8 B shows that only when disaccharide is added... TransOn the side, and with the detection voltage adjusted to -40mV, a characteristic current signal can be observed generated by the disaccharide in the nanopore A detection system. The characteristic of this signal is that a second level of current blockage signal event is formed on the basis of the first level, named the Level 2 signal event. I 2).
[0144] (4) Following step 4 of Example 1, extract the signal event features of different sugar molecules as their respective signal data and perform statistical analysis. In this example, the Amplitude (I1) and Dwell time of Level 1 signal events for decasaccharides to trisaccharides are extracted, and the Amplitude (I2) and Dwell time of Level 2 signal events for disaccharides are extracted. The Dwell time versus ΔI1 / I0 scatter plot of decasaccharide to trisaccharide signals clusters a single event group ( Figure 9 A). The Gaussian distribution fitting values of ΔI1 / I0 for decasaccharides to trisaccharides are 0.73, 0.63, 0.47, 0.39, 0.34, 0.37, 0.22, and 0.16, respectively. Figure 9 B), with an average separation ratio of 1.78. The retention time index fitted values from decasaccharides to trisaccharides were 0.12 ms, 0.13 ms, 0.14 ms, 0.16 ms, 0.15 ms, 0.25 ms, 0.14 ms, and 0.17 ms, respectively. Figure 9 C). Furthermore, the disaccharide's Dwell time versus ΔI2 / I1 scatter plot clusters into a single event group ( Figure 9 The mean of the Gaussian distribution fit for ΔI² / I is 0.54 (D). Figure 9 E), the fitted τ value of the exponential distribution of the residence time is 10.71 ms ( Figure 9 F). In the current nanopore A detection system, the two monosaccharides Gal and GlcNAc, as standards, did not produce characteristic signals under any detection conditions. Figure 10 A, Figure 10 (B) Regarding the fact that monosaccharides do not generate characteristic signals in this nanopore A detection, this embodiment offers two advantages: First, during the sequential hydrolysis process using exoglycosides, the resulting monosaccharide units do not interfere with the signal attribution and differential identification of the hydrolysis products; second, during the sequential hydrolysis from disaccharide to monosaccharide, if the characteristic signal events of the disaccharide decrease or disappear, it can be directly distinguished from the disaccharide. Thus, combining these two detection conditions, the nanopore A detection system of this invention can satisfy the characteristic signal acquisition from decasaccharides to monosaccharides; in summary, the nanopore detection system of this invention has high-precision monosaccharide resolution for glycan chains from decasaccharides to monosaccharides.
[0145] Example 3: Optimization of a reaction system for stepwise hydrolysis of glycan chains by glycosidases for nanopore glycan sequencing
[0146] This embodiment first uses the poly-LacNAc glycan nonasugar described in Example 2 as a template glycan and verifies the hydrolytic reactivity of the glycosidase by using a glycosidase that hydrolyzes the non-reducing end of the glycan.
[0147] The specific experimental method is as follows: 1. Preparation of the specific exoglycosidases involved in the test sugar molecule. Specific exoglycosidases include, but are not limited to, the enzymes described in Table 1. This example uses a nonaglycone belonging to the poly-LacNAc series of sugar chains as the model sugar chain. The poly-LacNAc series of sugar chains involves two exoglycosidases, including β-N-acetylglucosidase (GlcNAcH) and β1-4-galactosidase (GalH), which specifically recognize and hydrolyze the non-reducing ends of the test sugar molecule to release β1,3-linked N-acetylglucosamine (GlcNAc) and β1,4-linked galactose (Gal), respectively. The genes for glycosidases GalH and GlcNAcH were cloned into the pET-28a vector through the restriction sites NdeⅠ and XhoI, respectively, with a 6×His tag at the N-terminus or C-terminus. The sequenced construct was then transformed into *E. coli* BL21(DE3) strain, amplified, and induced to express protein, followed by purification via Ni-NTA agarose column. Prior to purification, the Ni-NTA column was equilibrated with lysis buffer (50 mM Tris-HCl, 300 mM NaCl, 10 mM imidazole; pH 7.5). Ni-NTA was eluted with two column volumes of lysis buffer, followed by elution with elution buffer (50 mM Tris-HCl, 300 mM NaCl, 300 mM imidazole; pH 7.5). The purified enzyme protein was identified using Coomassie brilliant blue staining and Western blot. Figure 11 ).
[0148] Amino acid sequence of GlcNAcH protein: NCBI Reference Sequence: WP_010992686.1 MRNLFKIAGLLALTGFISSCNDKETTANYQVIPLPQEITTAQSQPFTLNGSVKIIYPEGNEKMQRNAQFLADYLKKATGKDYAVEAGTEGKGAILLKLGMESENPEAYQLSVNADGVTIAAPTEAGVFYGIQTLRKSIPVAIGTTPSLPAVEISDYPRFSYRGAHFDVGRHFFTVDEVKTYIDMMALHNMNRLHWHLTEDQGWRLEIKKYPKLTEIGSKRSETVIGRNSGEYDGKPYSGFYTQEEAREIVAYAADRYITVIPEIDLPGHMQGALAAYPHLGCTGGPYEVWKIWGVSDQVLCAGNDSVLTFIDDVLTEVMDIFPSEYIHVGGDECPKTEWAKCPKCQARIKALGIKSDAKHSKEEYLQSFVINHAEKFLNEHGRQIIGWDEILEGGLAPNATVMSWRGEGGGIEAAKQKHDVIMTPNTYLYFDYYQTKDTENEPLAIGGYVPLERVYGYEPMPSSLTPEEQKHIIGVQANLWTEYIPTFSQAQYMVLPRWAALAEVQWSNPEKKNYDNFLSRLPQLINIYDAEGYNYAKHVFDVKSEFVANSATGAVDVVMTTIDGAPIHYTLDGTEPTAASPVCDSILTIKESCTLKAVAVRPTGNSKMLTEQIVFSKSTSKPIKANQPVNKQYEFGGVSTLVDGLKGNGNYKTGRWIAFYKNDMDVTIDLQQPTEISSVAITTCVEKGDWVFDARSFSIEVSDDDKTFTKVASEAYPEMKETDRNGLYEHKLTFDPVKTRYVKVIATSEHSIPAWHGGKGNPGFLFVDEITLN Amino acid sequence of GalH protein: NCBI Reference Sequence: WP_122128572.1 MNKKIKIAFASMLAVPLLACAQVRTEQTFEKGWKFTREDSKDFSNSTYDDAKWQSVTVPHDWAIYGPFSINNDKQNVAISQDGQKEAMEHAGRTGGLPFVGVGWYRLNFDAPSFSKGKKATLVFDGAMSHAHVYINGQEAGYWPYGYNSFYVDATPYLKPGEKNTLAVRLENENESSRWYPGAGLYRNVHLVVNEDAHIPTWGTQLTTPVVKDEFAKVNLKTKLDVPAGKAFEGYRIVTELKDKDGKVVAANEKKGGPFDDNVFEQDFVVTSPALWTPDTPHLYSAVSKVYEGNTLKDEYTTSFGIRSIEIIPNKGFFLNGKKTMFKGVCNHHDLGPLGGIANDAGIRRQIRILKDMGCNAIRTSHNMPAPELIKACDEMGMMIMAESFDEWKAAKVQNGYHKVFDEWVEKDLVNLIHQYRNNPSVVMWCIGNEVPDQWNGDRGPKLSRFLQDICHREDPTRPVTQGMDAPDAVVNNNMAAVMDVAGFNYRPHKYQENYKKLPQQIILGSETASTVSSRGVYKFPVVRRAMQKYDDHQSSSYDVEHCGWSNLPEDDFIQHEDLPYCIGEFVWTGFDYLGEPTPYYTDWPSHSSLFGIIDLAGLPKDRYYLYRSHWNKDKETLHILPHWNWEGREGEVTPVFVYTNYPSAELFINGKSQGKRTKDLSVTVNNSGDSTSVANFKRQQRYRLMWMDTKYEPGTVKVVAYDKDGKAVAEKEIHTAGKPDHIELVADRSVIDANGKDLSFVTVKVVDKEGNLCPLADNEISFKVKGAGTY Gene sequence of GalH protein: Gene ID: 60368346 1atgaacaaga aaatcaaaat tgcatttgct tcgatgctcg ctgtgccgct gttggcgtgt 61 gcgcaagtcc gtacggaaca aacctttgag aagggatgga agttcactcg tgaggatagt 121 aaagacttta gtaactctac gtatgatgat gcgaagtggc aatctgtgac cgttccgcat 181 gattgggcta tttacggacc attcagtatt aatatgata aacagaacgt agccatttct 241 caggatggac agaaagaggc gatggagcat gccggacgta ccgggggact tccttttgtc 301 ggtgtgggat ggtataggct caattttgat gctccttcat tcagtaaggg caaaaaggct 361 actctggttt tcgatggagc catgagccac gcacatgtct atatcaacgg acaagaggcc 421 ggttattggc cttacggtta caatagtttt tatgtggatg ccactcctta tttgaaaccg 481 ggtgaaaaga acacattggc ggttcggctt gaaaacgaga acgaatcttc acgttggtat 541 ccgggtgccg gattatatcg caatgtacat ctggtagtga acgaagatgc ccatattcct 601 acttggggta ctcagctgac tactcctgta gtgaaagatg aatttgcaaa ggtaaacctg 661 aaaaccaaac tcgatgtacc agccggaaaa gcattcgaag gatatcgtat tgttaccgaa 721 ctgaaggata aggatggtaa agtggtggca gccaacgaaa agaagggagg tccgttcgat 781 gacaatgttt tcgaacaaga tttcgtagtg acttctccgg ctttgtggac tccggatact 841 ccgcatcttt atagtgccgt ttctaaagta tatgagggga atactctgaa ggatgaatat 901 actacttcat tcggtattcg ttccattgag attattccga ataaaggttt cttcctgaat 961 ggaaagaaaa ccatgttcaa aggtgtgtgc aatcaccatg atctcggacc gttgggaggt 1021 attgccaatg atgccggtat tcgccgtcag attcgtatcc tgaaagacatgggctgcaat 1081 gccatccgta cttcacataa tatgcctgct cccgaattga ttaaagcttgtgatgaaatg 1141 ggcatgatga taatggccga gtcatttgac gagtggaaag cagccaaagtacagaatggt 1201 tatcacaaag tgtttgatga atgggtagag aaagatctgg tgaatctgattcatcagtat 1261 cgcaataatc cgagcgtggt tatgtggtgt atcggtaatg aggtgcccgaccagtggaac 1321 ggtgaccgtg gaccgaaact gtcacgcttc ttgcaggata tctgccatcgtgaagatcct 1381 acacgtcctg ttactcaggg tatggatgct cccgatgcag tagtcaataataatatggcg 1441 gctgtgatgg acgtagccgg ttttaattac cgtcctcaca aatatcaggaaaactacaaa 1501 aaacttccgc aacagatcat cctgggtagc gagacagctt ctaccgttagttcacgtggc 1561 gtttataaat tccctgttgt ccgtcgtgcg atgcagaagt atgacgatcatcagtcttca 1621 tcatatgatg tagagcattg cggttggtcc aatcttcctg aagacgattttatccagcat 1681 gaggatttgc cttattgcat tggtgagttc gtatggaccg gattcgattatttaggcgaa 1741 ccgactccgt actataccga ttggcccagc cactcttcgt tgtttggaattatagacctt 1801 gccggacttc cgaaagaccg ttattatctc tatcgcagtc attggaataaagataaggaa 1861 acgctgcata tcttgcctca ctggaattgg gaaggacgtg agggtgaagttaccccggta 1921 tttgtatata ccaattatcc ttcggctgaa ctgtttatta acggcaaaagtcagggtaag 1981 cgtacaaagg atctctccgt tacggtaaac aatagcggtg actctacttctgtggccaac 2041 ttcaaacgtc agcaacgtta tcgcctgatg tggatggata cgaaatacgagccgggtact 2101 gtaaaagtgg ttgcttacga taaagatggc aaagccgttg ctgaaaaggaaattcataca 2161 gccggaaagc ccgatcacat tgagttggtg gcagaccgta gcgtgattgacgccaatggt 2221 aaagatctct catttgtgac agtgaaggta gtcgataaag agggtaatctttgtccgctg 2281 gcagacaatg agatcagttt caaggtgaag ggagcaggta cttatcgtgcaggagccaat 2341 ggtaatccgg cttctcttga gtcattccag accccgaaaa tgaaagtattcagcggtatg 2401 atgacagcga ttgtgcagtc tactgaaaaa gcaggaaaga tcacacttgaagcaacagga 2461 aaaggtctga aaaaaggtac attgctgatc gaaagcaaat aa Gene sequence of GlcNAcH glycosidase protein, NCBI Reference Sequence: NC_003228.3 1 atgagaaatc tttttaaaat tgccggttta ctggcgctca cgggatttat ttcttcgtgt 61 aacgacaaag aaactacagc taattaccag gtaattcctt tacctcagga aattacgact 121 gctcaaagcc aaccatttac attaaatggt agtgtaaaaa tcatctatcc ggaaggaaat` 181 gaaaagatgc aacgcaatgc ccaattcctg gctgattatc tgaagaaagc cacaggtaag 241 gattatgccg ttgaagcagg tactgagggc aaaggtgcta ttctgcttaa gttgggtatg 301 gaatcagaaa accccgaagc ttatcagttg agtgtgaacg cagatggtgt aactattgct[[ID=二十一]] [[ID=二十二]]361 gctcctactg aagctggtgt attctatggt atccagactt tacgtaagtc tattccggta[[ID=二十三]] 421 gctatcggta caactccttc attgccggct gtcgagatta gcgactatcc tcgttttagt 481 tatcgtggtg cccattttga tgtaggtcgc catttcttca cagtagacga agtaaagaca 541 tacatcgata tgatggcttt gcataacatg aaccgtttac actggcactt aaccgaagat 601 cagggatgga gacttgaaat caagaagtat cctaaactga ctgagatcgg ttcaaaaaga 661 tcggaaactg taatcggacg taattccggt gaatatgacg gaaaacctta cagtggtttc 721 tatactcagg aagaagccag agaaatcgta gcgtatgctg cagatcgtta tatcactgtt 781 attcctgaaa tcgaccttcc gggacatatg cagggagcat tggccgctta tccgcatttg 841 ggatgtacag gtggaccgta tgaggtttgg aaaatatggg gagtatctga tcaggtgctt 901 tgtgccggaa acgacagtgt gctgactttt attgatgatg tcctgacgga agtaatggat 961 atcttccctt ctgaatatat ccatgtcgga ggtgacgaat gtccgaagac tgaatgggcg 1021 aaatgtccga agtgccaggc tcgtatcaag gccttgggaa tcaagagtga cgctaaacat 1081 tcaaaagaag agtatttgca gagttttgtt atcaatcatg cagaaaaatt cctgaacgaa 1141 catggacgtc agattattgg ctgggacgaa atcctggaag gtggactggc cccgaatgct 1201 acagtgatgt catggcgcgg cgaaggtggc ggtatcgaag ctgccaagca gaaacatgat 1261 gttatcatga ccccgaacac ttatctttac tttgactact atcagaccaa ggatacagag 1321 aatgaacctt tggctatcgg tggttatgtg cctttggaaa gagtttatgg ttacgaaccg 1381 atgccttctt cattgacacc ggaagaacaa aaacatatta ttggcgtaca ggctaatctt 1441 tggacggaat acattcctac tttctctcag gctcagtata tggtgttgcc gcgttgggct 1501 gctttagcag aagttcaatg gtctaatcct gaaaagaaga attatgataa cttcctgagc 1561 cgtttgccgc agttgattaa tatttatgat gctgaaggat ataattatgc caaacatgta 2. Determine the reaction conditions for specific exoglycosidases. After purifying the glycosidase, the reaction system for the glycosidase was determined and its activity was tested, including the concentration of glycosidase, the reaction system, the concentration of substrate sugar, and the reaction time. In this example, different concentrations of glycosidase GlcNAcH or GalH were reacted with the substrate nonaglycone to test the hydrolytic activity of glycosidase GlcNAcH or GalH. Thin-layer chromatography (TLC) was used as the detection system for the reaction conditions during the hydrolysis process to determine the reaction time. Specifically: gradient concentrations of glycosidase GlcNAcH or glycosidase GalH were mixed with the substrate nonaglycone in aqueous solution. The final concentration of the enzyme system was set to 0.005, 0.05, and 0.5 mg / mL, and the final concentration of the substrate nonaglycone was set to 10 mM. The reaction system was reacted at 37°C, and samples were taken at different time points, including 1, 2, 5, and 10 min. The samples were detected by TLC. Figure 12 ).
[0149] Specific experimental procedures and results:
[0150] (1) Prepare two enzymes, GlcNAcH and GalH, according to step 1.
[0151] (2) Following step 2, both exoglycosidases at a concentration of 0.5 mg / mL can completely hydrolyze the 10 mM substrate within 10 minutes. GlcNAcH at a concentration of 0.5 mg / mL only requires 1 minute to achieve complete hydrolysis. Therefore, to ensure the enzyme reaction is as complete as possible, in this embodiment, when the sugar concentration is 10 mM, the enzyme concentration is set to 0.5 mg / mL, and the reaction time is set to 10 minutes; if other sugar concentrations are used, the enzyme reaction time is shortened or reduced proportionally to the concentration. Figure 12 The method can detect sugar concentrations down to 0.1 μM, achieving a sensitivity of 0.1 μM and an accuracy of 98%.
[0152] Example 4: A nanopore "reverse sugar sequencing" method based on glycosidase hydrolysis assistance
[0153] The hydrolysis-assisted nanopore "reverse sugar sequencing" method proposed in this embodiment is achieved by combining the nanopore detection system and method described in Example 1, the sequencing method described in Example 2, and the sugar chain hydrolysis reaction system described in Example 3.
[0154] Specifically: First, an array of hydrolysis chips or kits composed of various glycosidase combinations is used to hydrolyze unknown or test sugar chains. Nanopores are used for real-time monitoring of the products formed in the arrayed hydrolysis reaction system. Figure 13 By monitoring the differences in electrical signals before and after the reaction, and combining this with the specific characteristics of exoglycosidases, the glycan sequence information, including monosaccharide composition and linkage, is inferred in reverse. The signal difference before and after hydrolysis refers to the difference between the electrical signal of the remaining glycan after each step of the exoglycosidase hydrolyzes and cleaves the target glycan molecule to release monosaccharides and the current signal of the substrate sugar molecule before the hydrolysis. The significance of this signal difference before and after hydrolysis can determine whether the hydrolysis reaction has occurred. If the signal difference is significant, it indicates that the exoglycosidase in that step has performed the hydrolytic function. Then, based on the specificity of exoglycosidases in recognizing terminal monosaccharides and their linkages, the monosaccharide unit composition and linkage of the sequence information at that position are inferred in reverse. Based on continuous stepwise specific exoglycosidase array enzymatic hydrolysis reactions and nanopore detection, the complete sequence information of the target glycan molecule is continuously obtained in reverse. The complete sequence information includes the types, quantities, and positional order of all monosaccharide units in the glycan chain. Figure 13 ).
[0155] The specific experimental method is as follows: 1. As described in Examples 1 and 2, a nanopore detection system was constructed. In this example, a detection system based on nanopore protein A with sugar resolution was used. First, the raw electrical signal characteristics of the decaglycosylation to be tested (named S10) were detected and collected.
[0156] 2. Establish an array reaction system for stepwise hydrolysis of glycans using glycosidases for nanopore sugar sequencing. The specific exoglycosidase arrays involved in the poly-LacNAc series of sugars consist of two enzymes: GalH and GlcNAc. Starting with decaglycosaccharides, each step of the glycosidase hydrolysis reaction involves hydrolysis by an array of these two enzymes. Due to enzyme specificity, one enzyme specifically recognizes and hydrolyzes non-reducing terminal monosaccharides and glycosidic bonds to generate monosaccharides and a residual glycan chain with one less monosaccharide in length. The other enzyme, unable to recognize non-specific terminals, does not hydrolyze these terminals. Before each subsequent hydrolysis step, the glycosidase from the previous step is removed before the next glycosidase is added. Stepwise hydrolysis using the glycosidase array continues until all enzymes in the array can no longer hydrolyze the sugars. In this example, a dual-enzyme array of GalH and GlcNAcH is used to alternately hydrolyze decaglycosaccharides until monosaccharides are reached. In the dual-enzyme array, GalH is designated as enzyme A, and GlcNAcH as enzyme B.
[0157] 3. Stepwise hydrolysis of the test sugar chain and monitoring of electrical signals. In this embodiment, a dual-enzyme array including enzyme A and enzyme B is used to hydrolyze the test decasaccharide separately. The enzyme reaction conditions determined in step 4 of Example 3 are set as follows: the final sugar concentration is 10 mM, the final enzyme concentration is 0.5 mg / mL, and the reaction time at room temperature is set to 10 minutes. One of the enzymes will hydrolyze the non-reducing terminal 1X of the test decasaccharide 1X-2X-3X-4X-5X-6X-7X-8X-9X-10X. Figure 14 The first step involves reacting the decaglycoside S10 to be tested with either enzyme A or enzyme B. A nanopore is used to monitor the current signal during the reaction. To avoid interference from the enzymes on the electrophysiological system during hydrolysis, the system employs methods to inactivate or remove the coupled enzymes. Methods for removing or inactivating glycosidases include, but are not limited to, denaturing the enzyme through heating, attaching the enzyme to a solid support, and direct ultracentrifugation. In this embodiment, the heating system built into the detection system is used. After S10 reacts with enzyme A or enzyme B for 10 minutes, it is heated at 50°C for 10 minutes. After the first step of hydrolysis, the products of the glycan S10 are named HS9 (HS9A and HS9B). The current signal events of the hydrolyzed samples HS9A and HS9B are detected and collected using the nanopore detection system described in Examples 1 and 2.
[0158] 4. Stepwise cyclic hydrolysis of glycosidase array and nanopore detection. Repeat step 3, i.e., after the previous glycosidase array hydrolysis reaction is completed and a signal event is collected based on the nanopore detection system, remove the glycosidase from the reaction system. Then proceed to the next glycosidase array reaction, similarly performing signal event collection and glycosidase removal. Each step of the enzyme array hydrolysis involves one glycosidase. This cycle continues until 10X is released during hydrolysis. Thus, we can obtain a stepwise hydrolysis reaction system from S10, HS9A, HS9B, HS8A, HS8B, HS7A, HS7B, HS6A, HS6B, HS5A, HS6B, HS4A, HS4B, HS3A, HS3N, HS2A, HS2B, HS1A, HS1B. Figure 15 When the sugar chain shortens to the point where the conditions in the nanopore detection system cannot produce a response, the detection conditions can be switched, including but not limited to voltage sign and magnitude, electrolyte solution in the chamber, nanopore type, or mutation site. In this embodiment, all products after stepwise hydrolysis were added to the system as described in Examples 1 and 2. Cis If no characteristic signal appears on the side, then add to Trans The side, and by changing the voltage, increases its capture.
[0159] 5. Based on the difference in current signals in the stepwise glycosidase array hydrolysis system, the reverse sequence of the analyte molecule is deduced according to the enzyme specificity. In this embodiment, the current signal time characteristics of the above-mentioned analyte sugar molecule and the stepwise hydrolysis reaction system S10, HS9A, HS9B, HS8A, HS8B, HS7A, HS7B, HS6A, HS6B, HS5A, HS6B, HS4A, HS4B, HS3A, HS3N, HS2A, HS2B, HS1A, and HS1B are extracted to form signal data and statistical analysis is performed. Based on the statistical analysis results of one or more features, it is determined whether the signal characteristics before and after each hydrolysis step have changed significantly. If one or more signal characteristics have changed significantly, it is determined that there is a significant signal difference before and after hydrolysis. If there is a significant signal difference before and after hydrolysis, the sequence information of the analyte decaglycoside is deduced based on the specificity of exoglycosidase A or B.
[0160] Experimental results: According to the above experimental steps, signal characteristic data of the decasaccharide standard sample (S10) and its nine-step hydrolysis reaction system samples S10, HS9A, HS9B, HS8A, HS8B, HS7A, HS7B, HS6A, HS6B, HS5A, HS5B, HS4A, HS4B, HS3A, HS3B, HS2A, HS2B, HS1A, and HS1B were collected and statistically analyzed. In this embodiment, only the statistical analysis graph of the changes in electrical signals before and after the reaction is shown. Figure 16A). During the hydrolysis and monitoring process, when the cycle reached step 7, i.e., products HS2A and HS2B no longer produced signal shift in the first detection system (Trans applied a voltage of +100mV, sample added to the Cis side), we changed the sample addition side to the Trans side and adjusted the voltage from +100 mV to -40 mV.
[0161] A two-dimensional scatter plot of Dwell time versus ΔI1 / I0 was created based on the signal characteristic data of each test sample.
[0162] like Figure 16 As shown in Figure B, the scatter points of each sample can be clustered well and exhibit significant differences. Specifically, the center values of ΔI1 / I0 for the scatter points S10, HS9A, HS8B, HS7A, HS6B, HS5A, HS4B, HS3A, HS2B, and HS1A show significant shifts. To quantify the changes in the ΔI1 / I0 feature, Gaussian distribution fitting was performed on the signal feature ΔI1 / I0 of all samples, resulting in mean ΔI1 / I0 values of 0.73, 0.64, 0.47, 0.37, 0.34, 0.36, 0.22, and 0.15 for S10, HS9A, HS8B, HS7A, HS6B, HS5A, HS4B, and HS3A, respectively. Figure 16 (C) According to the aforementioned formula, the signal characteristic ΔI1 / I0 separation ratios before and after each hydrolysis step are calculated to be 1.25, 3.11, 1.99, 0.62, 0.46, 4.11, and 2.05, respectively. This means that the stepwise hydrolysis signals collected by the nanopore detection system in this invention can be accurately identified. Furthermore, the further hydrolysis of HS2B can be rapidly interpreted by observing the phenomenon of Level 2 signal transition from presence to absence (HS1A) to determine the hydrolysis process and product signals.
[0163] Since the signal characteristics of HS9A, generated by the hydrolysis of S10 by enzyme A (GalH), are significantly different from those of S10, it is determined that the first step of the hydrolysis reaction proceeded smoothly. Based on the specificity of GalH, it is deduced that the terminal 1X of the analyte molecule is Gal linked by a β1,4 glycosidic bond, and similarly, 2X is identified as GlcNAc linked by a β1,3 glycosidic bond. Based on the results of the nine-step hydrolysis, the sequence of the analyte decaglycose is derived as: Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAcβ1-3Galβ1-4GlcNAc (…). Figure 17 This is consistent with the structure identified by its mass spectrometry and nuclear magnetic resonance (NMR). Figure 18It is worth noting that the nanopore "reverse sugar sequencing" method based on glycosidase hydrolysis can detect the sugar concentration as low as 0.1 μM, i.e., the sensitivity is 0.1 μM. It can reduce the amount of sugar used from 1~10 mg / time to 20 μg / time, achieve an accuracy of 98%, and reduce the detection time in sequence analysis from 12h to 1h.
[0164] Example 5: A glycosidase-assisted nanopore sequencing method based on a "reverse sugar sequencing" strategy applied to branching sugar sequencing.
[0165] This embodiment first uses N-sugars as an example of branched sugars. N-sugars consist of a core pentasaccharide structure and a branched structure. Exoglycosidases can progressively cleave the terminal monosaccharides starting from the non-reducing end of the branched structure of the N-sugar. Because the exoglycosidases used have strong specificity, they can accurately identify the monosaccharide units and glycosidic bonds at the ends of the sugar chain, thus one specific exoglycosidase corresponds to one type of monosaccharide and glycosidic bond cleaved. N-sugars are sequenced according to the "reverse sugar sequencing" strategy described in Example 4, and automated N-sugar sequencing is achieved according to the machine learning model and automated process described in Example 4. Specifically: First, an array of hydrolysis chips or kits composed of various glycosidase combinations is used to hydrolyze the unknown or test N-sugar. The products formed in the arrayed hydrolysis reaction system are monitored in real time using nanopores. By monitoring the difference in electrical signals before and after the reaction, combined with the specific characteristics of the exoglycosidase, the N-sugar sequence information, including monosaccharide composition and linkage mode, is inferred in reverse. The signal difference before and after hydrolysis refers to the difference between the electrical signal of the residual sugar chain after the exoglycosidase hydrolyzes and cleaves the N-sugar to release monosaccharides at each step and the current signal of the substrate sugar molecule before the hydrolysis. The significance of this signal difference before and after hydrolysis can determine whether the hydrolysis reaction has occurred. If the signal difference is significant, it indicates that the exoglycosidase in that step has performed the hydrolytic function. Then, based on the specificity of the exoglycosidase in recognizing the terminal monosaccharide and its linkage of the N-sugar, the monosaccharide unit composition and linkage mode of the sequence information at that position are deduced. Based on continuous stepwise specific exoglycosidase array enzymatic hydrolysis and nanopore detection, the complete sequence information of the N-sugar molecule to be tested is continuously deduced. The frequency of current blockage events generated by 10 μM N-sugar H4 in the nanopore system reaches 8 s. -1 This technology is already superior to existing technologies for detecting N-glycans.
[0166] This embodiment uses a bifurcated N-sugar (referred to as H4) as an example, with the following structural formula: Figure 28As shown, X represents the sequence information at a position. The exoglycosidases involved in this N-sugar are NanA, BgaA, and H1811, as listed in Table 1. NanA hydrolyzes sialic acid linked to the non-reducing end at α2,3, BgaA hydrolyzes galactose linked to the non-reducing end at β1,4, and H1811 hydrolyzes acetylglucosamine linked to β1,2. A three-enzyme array was constructed using NanA, BgaA, and H1811. The N-sugar H4 to be tested was hydrolyzed using this three-enzyme array and then sequenced. The sequencing method and results are as follows: 1. Establishment of a nanopore detection system. A nanopore detection system with good capture rate and signal response for branched N-glycans was developed. In this embodiment, αHL (M113R / K147R), abbreviated as nanopore protein C, with good capture rate for N-glycans, was used. The αHL gene was subjected to a double-point mutation M113R / K147R according to the method described in Example 1, and nanopore αHL (M113R / K147R) was obtained by expression and purification using an E. coli expression system. A nanopore C detection system was established using nanopore C.
[0167] 2. Establish an array reaction system for the stepwise hydrolysis of N-glycans by glycosidases for nanopore N-glycan sequencing. In this embodiment, the specific exoglycosidase array involved in the N-glycan is composed of three enzymes: NanA, BgaA, and H1811. Starting with N-glycan H4, each step of the glycosidase hydrolysis reaction involves hydrolysis by the enzyme array composed of these three enzymes. Due to enzyme specificity, one enzyme specifically recognizes and hydrolyzes non-reducing terminal monosaccharides and glycosidic bonds to generate monosaccharides and residual sugar chains with a chain length reduced by one monosaccharide. The other enzyme, unable to recognize non-specific terminals, does not perform enzymatic hydrolysis. Before the next hydrolysis step, the glycosidase from the previous step is removed before proceeding to the next three-enzyme array hydrolysis reaction. Stepwise glycosidase array hydrolysis is performed until all enzymes in the glycosidase array can no longer hydrolyze the N-glycans, at which point the process stops. In this embodiment, the NanA, BgaA, and H1811 three-enzyme array is used to alternately hydrolyze decasaccharides until the core pentasaccharide structure is reached. In the dual-enzyme array, NanA is named enzyme A, BgaA is named enzyme B, and H1811 is named enzyme C.
[0168] 3. Stepwise hydrolysis of the N-sugar to be tested and monitoring of electrical signals. In this embodiment, a three-enzyme array including enzyme A, enzyme B, and enzyme C is used to hydrolyze the N-sugar to be tested, H4. The enzyme reaction conditions determined in step 4 of embodiment 3 are set as follows: the final sugar concentration is 10 mM, the final enzyme concentration is 0.5 mg / mL, and the reaction time is set to 10 minutes at room temperature. One of the enzymes will hydrolyze the non-reducing terminal monosaccharide of the test H4 ( Figure 28The first step involves reacting the H4 sample with each enzyme in the three-enzyme array, using a nanopore C detection system to collect the current signals during the reaction. To avoid interference from the enzymes to the electrophysiological system during hydrolysis, the system employs methods to inactivate or remove the coupled enzymes. Methods for removing or inactivating glycosidases include, but are not limited to, denaturing the enzymes through heating, attaching the enzymes to a solid support, and direct ultracentrifugation. In this embodiment, the detection system's built-in heating system is used; after H4 reacts with enzyme A, B, or C for 10 minutes, it is heated at 50°C for 10 minutes. After the first step of hydrolysis, the product of the H4 sample is named H3. Based on the different enzymes in the three-enzyme array, the three H3 samples are named H3A, H3B, and H3C, respectively. The nanopore detection system described in Example 1 is used to detect and collect the current signal events of the hydrolyzed samples H3A, H3B, and H3C.
[0169] 4. Stepwise cyclic hydrolysis of glycosidase array and nanopore detection. Step 3 is repeated stepwise, i.e., after the previous glycosidase array hydrolysis reaction is completed and a signal event is collected based on the nanopore detection system, the glycosidase in the reaction system is removed. Then, the next glycosidase array reaction is performed, and the same steps of signal event collection and glycosidase removal are repeated. Each step of the enzyme array hydrolysis involves one glycosidase. This cycle continues until no signal is generated after the hydrolysis of any enzyme in the enzyme array. Thus, a stepwise hydrolysis reaction system of H4, H3A, H3B, H3C, H2A, H2B, H2C, H1A, H1B, and H1C can be obtained. Figure 29 ).
[0170] 5. Based on the difference in current signals in the stepwise glycosidase array hydrolysis system, the reverse sequence of the analyte molecule is deduced according to the enzyme specificity. In this embodiment, the characteristics of the current signal time of the above-mentioned analyte sugar molecule and the stepwise hydrolysis reaction system H4, H3A, H3B, H3C, H2A, H2B, H2C, H1A, H1B, and H1C are extracted to form signal data and statistical analysis is performed. Based on the statistical analysis results of one or more features, it is determined whether the signal characteristics before and after each step of the hydrolysis reaction have changed significantly. If one or more signal characteristics have changed significantly, it is determined that there is a significant signal difference before and after hydrolysis. If there is a significant signal difference before and after hydrolysis, the sequence information of the analyte H4 is deduced based on the specificity of exoglycosidase A, B, or C.
[0171] Experimental results: 1. In this embodiment, the nanopore is nanopore C, the electrolyte buffer is 3M KCl buffer (5mM citric acid, pH 5.0), the sample is added to the Cis side chamber of the nanopore, and a detection voltage of +100mV is applied to the Trans side. This detection system has good capture ability for N-glycans.
[0172] 2. Based on the above experimental steps, signal characteristic data of the N-sugar sample (H4) to be tested and its three-step hydrolysis reaction system samples H4, H3A, H3B, H3C, H2A, H2B, H2C, H1A, H1B, and H1C were collected and statistically analyzed. In this embodiment, only the statistical analysis graph of the changes in electrical signals before and after the reaction is shown.
[0173] A two-dimensional scatter plot of Dwell time versus ΔI1 / I0 was created based on the signal characteristic data of each test sample.
[0174] like Figure 30 As shown, the scatter plots of each sample can be clustered well and exhibit significant differences, with a significant shift in the ΔI1 / I0 center values of the H4, H3A, H2B, and H1C scatter plots. This indicates that the stepwise hydrolysis signal collected by the nanopore detection system in this invention can be accurately identified.
[0175] 3. Based on the difference in current signals from the stepwise glycosidase array hydrolysis system, the reverse sequence of the analyte molecule is deduced according to the enzyme specificity. Since the signal characteristics of H3A, generated by the hydrolysis of H4 by enzyme A (NanA), are significantly different from those of S10, it is determined that the first step of the NanA glycosidase hydrolysis reaction proceeded successfully. Therefore, based on the specificity of NanA, it is deduced that the first terminal position of the analyte molecule is sialic acid (Neu5Ac) linked by an α2,6 glycosidic bond. Similarly, it is deduced that the second terminal position of H4 is galactose (Gal) linked by a β1,4 glycosidic bond, and the third terminal position of H4 is acetylglucosamine (GlcNAc) linked by a β1,2 glycosidic bond. The structure of the analyte H4 is then deduced based on the results of the three-step hydrolysis. Figure 31 The nanopore "reverse sugar sequencing" method based on glycosidase hydrolysis reduces the amount of sugar used from 1-10 mg to 10 μg, achieves an accuracy of 99%, and reduces the detection time in sequence analysis from 12 h to 1 h.
[0176] In addition, this embodiment also used sugar molecules belonging to human milk oligosaccharides (HMOs) as the sugar molecules to be tested (reference: 10.1073 / pnas.1701785114) to verify the feasibility of reverse sequencing based on nanopores.
[0177] The nanopore detection system used in this embodiment is the nanopore protein C detection system constructed above.
[0178] This embodiment uses a hydrolase array consisting of two hydrolases: FucH (purchased from New England Biolabs, catalog number P0769S) and NanA (purchased from New England Biolabs, catalog number P0722S). FucH can specifically hydrolyze α-1,3-linked fucose, and NanA can specifically hydrolyze α-2,3-linked sialic acid. Figure 32 AB).
[0179] Reverse sequencing of human milk oligosaccharides (HMOs) was performed using a nanoporous protein C (NPC) detection system and a hydrolytic enzyme array composed of FucH and NanA. First, the NPC detection system was used to detect the target HMO sugar molecules, generating an initial signal. Then, 5 mM of the target HMO molecules were mixed with 0.5 mg / mL of FucH or NanA enzyme and reacted for 10 min. The enzyme was then inactivated by boiling at 100°C, followed by centrifugation and retention of the supernatant. 0.26 μL of the supernatant was added to the NPC detection system to obtain the signal after the first step of the reaction. The initial signal and the signal after the first step of the reaction were compared. If a significant difference was found, the supernatant was mixed with 0.5 mg / mL of FucH or NanA enzyme and reacted for 10 min. This inactivation and detection process was repeated to obtain the signal after the second step of the reaction. The signal changes before and after the second step of the enzyme reaction were observed to see if there was a significant difference. If a significant difference was found, the enzyme was determined to be capable of hydrolysis. Based on the enzyme's specificity, the type of monosaccharide and glycosidic bond corresponding to that position was then determined.
[0180] Experimental results: The HMO sugar molecules to be tested were hydrolyzed using a hydrolytic enzyme array composed of FucH and NanA, and observations were made. Figure 32 CD), the HMO sugar molecules to be tested generated an initial current signal in the nanoporin C detection system. In the first step of hydrolysis, FucH hydrolysis significantly altered the signal of the HMO sugar molecule being tested, with the ΔI1 / I0 value changing from 0.2681 to 0.2469. In the second step, NanA hydrolysis significantly altered the electrical signal of the product from the first step, with the ΔI1 / I0 value changing from 0.2469 to 0.2893. These characteristic changes in the current signal reflect the successful occurrence of enzymatic hydrolysis. Therefore, based on the specificity and hydrolysis sequence of FucH and NanA enzymes, the type of monosaccharide and glycosidic bond at this position in the HMO molecule being tested can be inferred: the enzyme hydrolyzing in the first step is FucH, and the specificity of ΔI1 / I0 indicates hydrolysis of α-1,3-linked fucose, thus it is inferred that this position contains α-1,3-linked fucose; similarly, the enzyme hydrolyzing in the first step is NanA, and the specificity of NanA indicates hydrolysis of α-2,3-linked sialic acid, thus it is inferred that this position contains α-2,3-linked sialic acid.
[0181] Example 6: A data processing method for glycosidase-assisted nanoporous sugar sequencing
[0182] This invention discloses a data processing method for glycosidase-assisted nanopore sugar sequencing, which employs a reverse sequencing strategy. The method includes: acquiring electrical signals before and after the glycosidase hydrolysis of sugar molecules; performing data analysis on the acquired electrical signals; determining whether the glycosidase successfully hydrolyzed the sugar molecules based on a comparison of the first electrical signal before the hydrolysis reaction and the second electrical signal after the hydrolysis reaction; and then, based on the characteristics of the glycosidase involved in the hydrolysis, inferring the gradually released sugar molecules, thereby inferring the sequence structure of the sugar chain in reverse.
[0183] The signal comparison includes comparing the amplitude variation, dwell time variation, open frequency variation, Gaussian distribution fitting value or exponential distribution fitting value of the standard deviation of the current signal, or comparing the parameters of one or more feature distributions, including KL divergence, JS divergence, EMD distance, overlap coefficient and Bartholomew distance.
[0184] Signal comparison can be performed manually through statistical analysis or automatically by machine learning to identify signal differences.
[0185] Preferably, machine learning is used to help compare the first electrical signal before the hydrolysis reaction with the second electrical signal after the hydrolysis reaction to determine whether the glycosidase has successfully hydrolyzed the sugar molecules.
[0186] Specifically, in some embodiments, by extracting first electrical signal features and second electrical signal features, calculating the difference between the first and second electrical signal features, and inputting the difference into a trained classification model, a classification result is obtained as to whether the first and second electrical signal features are the same. The difference can be measured by one or more of the following: KL divergence, JS divergence, EMD distance, overlap coefficient, and Bach distance.
[0187] In other embodiments, the first electrical signal features and the second electrical signal features can be extracted and directly input into the trained classification model to obtain the classification result of whether the first electrical signal features and the second electrical signal features are the same.
[0188] If the first electrical signal characteristic is different from the second electrical signal characteristic, it is determined that the glycosidase has successfully hydrolyzed the sugar molecule.
[0189] In practice, the data processing method is implemented through software programs to automate the process. Specifically, it mainly includes the following program modules: The signal acquisition module automatically collects current signal events before and after the hydrolysis reaction, which can be implemented, for example, based on the PyNanolab software program.
[0190] The data analysis module uses machine learning to assist in the analysis of electrical signal data before and after glycosidase hydrolysis. The difference discrimination module can automatically identify and determine whether the hydrolysis reaction of the test sugar molecule by glycosidase has occurred; The automatic sequence generation module cyclically identifies whether a glycosidase hydrolysis reaction has occurred, and automatically generates glycan structures based on whether a stepwise glycosidase hydrolysis reaction has occurred and the specificity of the glycosidases involved in the hydrolysis.
[0191] The program processing flow is as follows Figure 19 As shown.
[0192] The automated processing flow of this embodiment is described below using the poly-LacNAc series of decaglycosides to be tested in Example 4 as an example.
[0193] Experimental methods: Step S51. Construct a machine learning training set. Collect a large amount of electrical signal data of different sugars in the nanopore detection system to construct a training set for use by the machine learning model in this embodiment.
[0194] Step S52. Model building and evaluation.
[0195] The first step is to use the training set provided in step 1 to calculate the two-dimensional frequency distribution of each data point based on the current signal event feature parameters ΔI1 / I0 and lg(Dwelltime).
[0196] The second step is to use the following five parameters to describe the difference between the two distributions: KL divergence, JS divergence, Earth mover's distance, overlap coefficient, and Bhattacharyya distance.
[0197] The third step is to combine these parameters into a single dataset for machine learning. This involves comparing the frequency distributions of the various feature parameters of the telecommunications signal events provided in the training set to see if there are any discrepancies.
[0198] The fourth step involves dividing the events in the dataset into two categories: "Yes" (same electrical signal characteristics) and "No" (different electrical signal characteristics). To mitigate the problem caused by data imbalance, the SMOTE algorithm is used to increase the number of minority class samples, balancing them with the majority class samples. The input data is randomly divided into a training set (80% of the labeled dataset) and a test set (20%) for model training and testing. The training set data is first standardized and then used to train five classification models, including Support Vector Machine (SVM), KNN, Random Forest, Naïvebayes, and Multilayer Perceptron. The model with the highest accuracy is selected based on 10-fold cross-validation accuracy. A confusion matrix is generated using the test set for model evaluation.
[0199] Step S53. Establish an automated interpretation process. In this embodiment, the model from step 2 is used to establish an automated interpretation process. Specifically, as follows: Figure 20 As shown, it includes the following steps: Step S531. The unknown glycan sequence (G) is detected by the nanopore detection system described in Example 1, generating an initial current signal feature (Data_m), where m represents the number of successful enzymatic digestions and the position of free monosaccharides in the original glycan, with an initial value of 1.
[0200] Step S532. Next, the glycan chain is enzymatically digested with an enzyme array composed of multiple specific exoglycosidases. Assuming there are n glycosidases, they are named E1, E2, E3 to En, respectively. The current signal characteristics (Data_m_f) of the hydrolysis products are collected using a nanopore detection system. Here, f represents the sample number after enzymatic digestion of different types of enzymes, starting from 1 and ranging from 1 to n.
[0201] Step S533. Compare all data_m_f with data_m, extract signal features, and input them into the best classification model (random forest model in this example). The model will provide prediction results for three parallel hydrolysis experiments and calculate the probability of hydrolysis occurring. If the probability exceeds a threshold (set to 60%), hydrolysis is confirmed. For samples confirmed to have undergone hydrolysis, output the corresponding data name containing m and f (representing the position and type of monosaccharide), and this data will be used as reference data (Former_data) for the next round of hydrolysis comparison. This sample will also serve as the initial sample (Former_sample) for subsequent hydrolysis, and the successful hydrolysis count will be updated (m=m+1).
[0202] Step S534. Repeat steps S532-S533 until the complete sugar chain sequence is determined.
[0203] Since disaccharides and monosaccharides show almost no signal under the current experimental conditions, the difference is very small, indicating that hydrolysis has reached the disaccharide stage. At this point, the experiment switches to the second set of experimental conditions to detect the Former_sample and obtain New_data. New_data will serve as reference data for the next round of hydrolysis comparison. Then, following the previous loop, after identifying the type and connection method of the second-to-last sugar unit in the glycan chain, the loop ends, outputting the complete glycan sequence.
[0204] Step S54. Apply the established machine learning model to automatically analyze the differences in electrical signals before and after the hydrolysis of the sugar to be tested.
[0205] Experimental results: In this embodiment, a nanopore detection system is used to detect and collect current signal events of sugar standards with different sugar chain lengths, including current signal events of deca-saccharides, nona-saccharides, octa-saccharides, heptasaccharides, hexa-saccharides, penta-saccharides, tetra-saccharides, tri-saccharides, disaccharides, and monosaccharides as a training set.
[0206] Following step S52, a two-dimensional probability distribution for each data point is first calculated for the ΔI1 / I0 and residence time features of the electrical signal characteristic events of the glycan samples in the training set. The difference between each pair of distributions is described by five parameters: KL divergence (KL), JS divergence (JS), Earth mover's distance (EMD), Overlap coefficient (OC), and Bhattacharyya distance (BD). The difference between three parallel data points of the same type of glycan is used to simulate the case where hydrolysis has not occurred, and the difference between adjacent glycan length datasets is used to simulate the case where hydrolysis has occurred. The dataset contains 114 events, divided into two classes: "Yes" (same electrical signal features) and "No" (different electrical signal features). The input data is randomly divided into a training set (80% labeled dataset) and a test set (20%) for model training and testing. The training set data is first standardized and then used to train five models, including SVM, KNN, Random forest, Naïve Bayes, and Multilayer perceptron. The Random Forest model, which achieved the highest accuracy (98.4%), was selected. In the Random Forest model, the feature importance values for KL, JS, EMD, OC, and BD were 18.01, 16.69, 22.33, 22.41, and 20.55, respectively. A confusion matrix was generated using the test set for model evaluation, showing prediction accuracies of 100% for "Yes" and 94% for "No." Figure 22 .
[0207] The interpretation software developed in step S53 was used for automated interpretation of the hydrolyzed sugar sequencing data. The data in this embodiment came from the electrochemical signals of the glycans before and after hydrolysis in Example 4. The results showed that in this automated interpretation cycle, the monosaccharide constituent units, linkage methods, and positional order of the glycans to be tested could be interpreted sequentially with an accuracy exceeding 98%. The successful implementation of this embodiment represents the first case of continuous glycan unit sequencing using nanopores, and the first time the read length has exceeded 10 monosaccharides. Figure 21 .
[0208] Example 7 A hydrolysis-assisted nanopore "forward sugar sequencing" strategy
[0209] Forward sequencing refers to the process of deriving the sequence of a target sugar molecule from the current signal characteristics of oligosaccharides (including monosaccharides, disaccharides, or other oligosaccharides, unless otherwise specified) released by glycosidase hydrolysis. Compared to reverse sequencing, the forward sequencing strategy described here utilizes a nanopore detection system to collect terminal monosaccharide or oligosaccharide fragments released by the hydrolytic enzyme, rather than the remaining sugar chains after hydrolysis. First, a high-resolution nanopore is developed to construct the nanopore detection system. The constructed nanopore detection system is used to detect and collect current signal events of monosaccharides, disaccharides, or other oligosaccharide fragments that may be released by hydrolysis, and the signal event characteristics of various sugar molecules are extracted according to the methods described in Examples 1 and 2, thereby establishing an oligosaccharide current signal fingerprint library, including but not limited to monosaccharide fingerprint libraries, disaccharide fingerprint libraries, and other oligosaccharide fingerprint libraries. Then, exonucleases, endonucleases, or combinations of glycosidases (hereinafter collectively referred to as glycosidases unless otherwise specified) are used to hydrolyze the sugar molecule, releasing terminal monosaccharide, disaccharide, or other oligosaccharide fragments. Glycosidases can be added to the same side of the target sugar molecule in the nanopore detection system, or they can be expressed by chemical linker linkage or fusion with the nanopore. The nanopore detection system captures and collects the terminal oligosaccharide fragments of the target sugar chain released by glycosidase hydrolysis. Based on the oligosaccharide electro-signal fingerprinting, characteristic electrical signals are assigned to determine the types of terminal components of the sugar molecule. The compositional order of the oligosaccharide fragments is deduced based on the order of occurrence and signal abundance of characteristic electrical signal events during hydrolysis, thus obtaining their complete sequence information. Figure 23 ).
[0210] Accordingly, the embodiments relate to a data processing method for glycosidase hydrolysis-assisted nanopore sugar sequencing using a forward sequencing strategy, comprising: acquiring the electrical signals caused by sugar molecules released by glycosidase hydrolysis; inferring the progressively released sugar molecules based on the comparison of the electrical signals and fingerprints of the sugar molecules released by glycosidase hydrolysis, thereby forward inferring the sequence structure of the glycan chain. The fingerprint is established based on electrical signal characteristics, including but not limited to one or more of the following: amplitude, dwell time, open frequency, and standard deviation of the current signal. The fingerprint includes monosaccharide fingerprint libraries, disaccharide fingerprint libraries, and other oligosaccharide fingerprint libraries. The electrical signals are assigned using the fingerprints to obtain the types of terminal oligosaccharide fragments during hydrolysis. The compositional order of the oligosaccharide fragments is deduced based on the order of occurrence of electrical signal events and / or the abundance of signals (the abundance of signals from sugars hydrolyzed earlier is higher, and the abundance of signals from sugars hydrolyzed later is lower).
[0211] In this embodiment, the sugar molecule to be tested is a trisaccharide molecule (trisaccharide: trisaccharide 8, reference: 10.1021 / jacs.9b08964). An exonuclease kit consisting of two exonucleases, α-2,6 neuraminidase (NanA) (purchased from New England Biolabs, catalog number P0722S) and β-1-4 galactosidase (BgaA) (purchased from New England Biolabs, catalog number P0777S), was used to validate the methodology of the nanopore-based glycolysis sequencing method based on the forward sequencing strategy.
[0212] The unknown sequence of the trisaccharide to be tested is represented as 1X-2X-3X (the numbers represent the unit positions from the non-reducing end to the reducing end). In this embodiment, a monosaccharide fingerprint spectrum including the current signal fingerprint spectrum of galactose, acetylglucosamine, and sialic acid was first established using a nanopore detection system. The glycans to be tested in the nanopore detection system were hydrolyzed using the exoglycosidase kit, with the exoglycosidase performing stepwise hydrolysis from the non-reducing end of the glycan. The enzyme reaction rate was controlled by adjusting system conditions, including but not limited to enzyme concentration, salt concentration, and pH, so that the nanopore detection system could capture the terminal monosaccharides released stepwise by enzymatic decomposition. By comparing the monosaccharide signals collected by the nanopore detection system with the established monosaccharide fingerprint spectrum, the type of monosaccharide unit of the glycan to be tested can be identified. Furthermore, the positional order of the monosaccharide types can be deduced based on the order of the released monosaccharide signals and / or the order of their abundance.
[0213] The experimental methods and results are described below.
[0214] Experimental methods: Step S61. Construct a nanopore detection system based on high-resolution nanopores. In this example, the purification method described in Example 1 is used to obtain a nanopore αHL(M113R) with high resolution for the trisaccharide to be tested, referred to as nanopore B ( Figure 24 A). In this example, to generate a detectable characteristic current signal for monosaccharides, based on the nanopore detection system in Example 1, the electrolyte solution was 1.5 M NaCl (10 mM NaH2PO4, pH=3.0), and the Trans-side voltage was set to -100 mV. β-cyclodextrin (abbreviated as β-CD) (purchased from Ambeed, catalog number A248973) was added to the Trans side to a concentration of 20 μM. The nanopore B detection system was thus formed based on the conditions described for nanopore B.
[0215] Step S62. Establish an oligosaccharide current signal fingerprint library. Based on the nanopore detection system built in step S61, establish a fingerprint library of terminal oligosaccharide fragments, including a monosaccharide fingerprint library, a disaccharide fingerprint library, and other oligosaccharide fingerprint libraries. In this example, using the nanopore B detection system and the electrical signal acquisition method and signal feature extraction method described in Example 1, establish electrical signal fingerprint libraries for sialic acid (Neu5Ac), galactose (Gal), and glucosamine acetylglucosamine (GlcNAc).
[0216] Step S63. Hydrolyze the analyte sugar chain using the glycosidase kit. The glycosidase kit is a combination of glycosidases that can hydrolyze and cleave any glycosidic bonds that may be present in the analyte sugar chain. Glycosidases can be added to the same side of the analyte sugar in the nanopore detection system, or they can be linked to the nanopore via a chemical linker or fusion expression. Glycosidases hydrolyze sugar molecules, releasing terminal monosaccharide, disaccharide, or other oligosaccharide fragments. In this example, an exoglycosidase kit (Kit) is composed of two exoglycosidases: α-2,6-neuraminidase (NanA) and β-1-4-galactosidase (BgaA). This kit is added to the same side of the nanopore detection system as the analyte trisaccharide to hydrolyze the trisaccharide.
[0217] Step S64. Acquire signal events of the glycan hydrolysis process to be tested. Using the signal acquisition and feature extraction method described in Example 1, acquire signal events of the glycan hydrolysis process in the nanopore detection system in real time and extract signal feature parameters, including but not limited to the signal amplitude at the nth level (In). n The data collected included parameters such as Dwell time and standard deviation (Std). In this example, the amplitude (ΔI2) and Dwell time of the second-layer horizontal signal events occurring during the hydrolysis process at different time periods were collected.
[0218] Step S65. Proactively deduce the sequence information of the glycan to be tested based on the signal events of oligosaccharide fragment release during the hydrolysis of the glycan to be tested. Using the signal event feature parameters of the hydrolysis process of the glycan to be tested extracted in step S64 at different time periods, create 2D or multi-dimensional maps for different hydrolysis time periods. Compare these maps with the previously established monosaccharide fingerprint maps to identify the types of monosaccharide units in the glycan to be tested. Deduce the positional order of monosaccharide units in the glycan to be tested based on the chronological order of the monosaccharide maps at different time periods. In this example, the amplitude (ΔI2) of the second-level signal event obtained in step S64 and its Dwell time signal parameters are plotted as a ΔI2 versus Dwell time 2D scatter plot and compared with the previously established fingerprint maps of glucose, galactose, and glucosamine.
[0219] Experimental results: Following step S61, α-HL (M113R), abbreviated as nanoporin B, is prepared. Figure 24 A).
[0220] A nanopore detection system was constructed using purified nanoporin B and β-CD ligand. Experimental conditions: conductivity buffer was 1.5 M NaCl (10 mM NaH2PO4, pH=3.0); a clamping voltage of -100 mV was applied to the Trans side; the membrane phospholipid was diaphytylphosphatidylcholine (DPhPC); nanopore B exhibited an opening current with a Gaussian fitting value of -153 pA (I0) at -100 mV. Without β-CD, adding monosaccharides to the Trans side of the nanopore did not produce a significant signal. After adding β-CD to the Trans side, a first layer of blockage signal (Level 1) was generated, which was caused by β-CD. Then, the sugar molecule sample to be tested was added to the Trans side, and the sugar molecules generated a second layer of blockage (Level 2) on top of the first layer of blockage by β-CD. Based on the above conditions of nanoporin B, a nanopore B detection system was formed and detection was performed using β-CD as an aptamer. Figure 24 B).
[0221] Following step S62, a nanopore detection system was constructed using nanopore protein B to establish a fingerprint library of the constituent units of the target sugar molecules, mainly including the electrical signal fingerprint libraries of sialic acid (Neu5Ac), galactose (Gal), and glucosamine (GlcNAc). Each sugar standard sample was added to the nanopore detection system at a final concentration of 500 μM (each sugar standard sample was detected independently). During the detection process, all target sugar chains were added to the Trans side, with a Trans side voltage of -100 mV. All three monosaccharides—sialic acid (Neu5Ac), galactose (Gal), and glucosamine (GlcNAc)—generated a second-level blocking event based on the first-level current blocking, named the Level 2 signal event. I 2) The Level 2 events are reversible and repetitive. The Level 2 signal event characteristics of different sugar molecules are extracted as their respective signal data and statistically analyzed. In this example, the Amplitude (ΔI2) and Dwell time (ΔI2) of the Level 2 signal events of three monosaccharides are extracted. Figure 25 Statistical analysis was performed on the signal data of the three monosaccharides to visualize the fingerprint spectrum. In this example, a two-dimensional fingerprint spectrum of ΔI²Versus Dwell time was generated, showing that the scatter points of the three trisaccharides can be clustered and distinguished from each other quite well. Figure 26 ).
[0222] Following step S63, an exoglycosidase kit (Kit) is constructed using α-2,6 neuraminidase (NanA) and β-1-4 galactosidase (BgaA). The kit is then added to the same side of the nanopore detection system as the trisaccharide to be tested, and the trisaccharide to be tested is hydrolyzed.
[0223] Following step S64, real-time electrical signal monitoring is performed on the hydrolysis products from step S63. In this embodiment, the amplitude (ΔI2) and drain time of the second-level signal events occurring during hydrolysis at different time periods were collected. The obtained amplitude (ΔI2) and drain time signal parameters of the second-level signal events were plotted as a ΔI2 versus drain time two-dimensional scatter plot and compared with the previously established fingerprint spectra of glucose, galactose, and acetylglucosamine to deduce the sequence of the trisaccharide to be tested as Neu5Ac-Gal-GlcNAc (…). Figure 27 ). Figure 27 Figure A shows the signal of the trisaccharide itself, while Figure B shows the signal generated after the first enzyme digestion. By comparing it with the monosaccharide fingerprint, sialic acid can be identified as being cleaved. Figure C shows the result of the second enzyme digestion based on Figure B, which generates three monosaccharides, corresponding to the fingerprint patterns of these three monosaccharides.
[0224] Example 8: A data processing system for glycosidase-assisted nanopore sugar sequencing
[0225] This embodiment discloses a data processing system for glycosidase hydrolysis-assisted nanopore sugar sequencing, which, like the data processing method for glycosidase hydrolysis-assisted nanopore sugar sequencing disclosed in Embodiment 6, is based on the same inventive concept. The system includes: a reverse sequencing input module for acquiring electrical signals before and after the glycosidase hydrolysis of sugar molecules; and a reverse sequencing information processing module for analyzing the acquired electrical signals. Based on a comparison of the first electrical signal before the hydrolysis reaction and the second electrical signal after the hydrolysis reaction, it determines whether the glycosidase has successfully hydrolyzed the sugar molecules. Then, based on the characteristics of the glycosidase involved in the hydrolysis, it infers the gradually released sugar molecules, thereby inferring the sequence structure of the sugar chain. Specific implementations of each module can be found in Embodiment 6 above, and will not be repeated here.
[0226] Example 9: A data processing system for glycosidase-assisted nanoporous sugar sequencing
[0227] This embodiment discloses a data processing system for glycosidase hydrolysis-assisted nanopore sugar sequencing, which, like the data processing method for glycosidase hydrolysis-assisted nanopore sugar sequencing disclosed in Embodiment 7, is based on the same inventive concept. The system includes: a forward sequencing input module for acquiring the electrical signals caused by sugar molecules released by glycosidase hydrolysis; and a forward sequencing information processing module for inferring the progressively released sugar molecules based on comparing the electrical signals and fingerprint patterns caused by the sugar molecules released by glycosidase hydrolysis, thereby forward inferring the sequence structure of the sugar chain. Specific implementations of each module can be found in Embodiment 7 above and will not be repeated here.
[0228] Example 10 A computer system
[0229] An embodiment of the present invention discloses a computer system, including a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor. When the computer program / instructions are executed by the processor, they implement the steps of a data processing method for glycosidase hydrolysis-assisted nanopore sugar sequencing disclosed in Embodiment 6 or Embodiment 7.
[0230] Example 11 A computer-readable storage medium
[0231] This invention discloses a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of a data processing method for glycosidase hydrolysis-assisted nanopore sugar sequencing disclosed in Embodiment 6 or Embodiment 7.
[0232] Example 12 A computer program product
[0233] The present invention discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of a data processing method for glycosidase hydrolysis-assisted nanopore sugar sequencing disclosed in Embodiment 6 or Embodiment 7.
[0234] Example 13 A nanopore sequencer based on glycosidase hydrolysis
[0235] This invention discloses a nanopore sequencer based on glycosidase hydrolysis, comprising a glycosidase hydrolysis system, a nanopore glycan electrical signal monitoring system, and a data processing system. The glycosidase hydrolysis system is used to hydrolyze sugar molecules with glycosidases; the nanopore glycan electrical signal monitoring system is used to monitor changes in the nanopore electrical signal caused by sugar molecules; the data processing system is a glycosidase hydrolysis-assisted nanopore glycan sequencing data processing system disclosed in Embodiment 6 or Embodiment 7; or, the data processing system automatically generates the glycan sequence structure based on the computer program product disclosed in Embodiment 12.
[0236] Example 14 A Nanoporous Glycan Chain Electrical Signal Monitoring System
[0237] This invention discloses a nanopore glycan electrical signal monitoring system, comprising two chambers separated by an insulating layer, a nanopore formed on the insulating layer, electrodes placed in the two chambers for measuring changes in electrical signals, and an instrument for acquiring these changes. The nanopore can be a conventional nanopore or an improved nanopore disclosed in Embodiment 1.
Claims
1. A sequencing method for carbohydrate molecules based on glycosidases and nanopores, characterized in that, The method includes: (1) Provide nanopores; (2) Construct a detection system containing the nanopores and glycosidases; (3) Add the sugar sample to be tested to the detection system; (4) The sugar sample to be tested undergoes a hydrolysis reaction with glycosidase. The sugar sequence information of the sugar sample to be tested can be inferred by comparing the changes in electrical signals before and after the hydrolysis reaction, or by inferring the sugar sequence information of the sugar sample to be tested by the characteristic signals generated by the monosaccharides released by the hydrolysis.
2. The sequencing method for carbohydrate molecules based on glycosidases and nanopores according to claim 1, characterized in that, The nanopores mentioned in step (1) are channels that match the molecular size of the sugar sample to be tested and can be passed through by collisions or displacement of sugar molecules, or not passed through by collisions, or not passed through by covalently bonded molecules, or have dissociation characteristics under the influence of external forces, thereby causing changes in properties; preferably, the channels that change in properties are channels in which changes in ion current manifest as changes in electrical signals; preferably, the nanopores include biological nanopores and / or solid nanopores; Preferably, the bio-nanopore comprises αHL, MspA, or aerolysin; preferably, the nanopore is an αHL heptameric bio-nanopore with a monomer molecular weight of 35 kDa; preferably, the bio-nanopore comprises a protein complex consisting of at least one wild-type protein monomer and at least one mutant monomer, or a homologous protein complex consisting of homologous mutant monomers; preferably, the mutant monomer comprises a mutation at at least one position in the NCBI Reference Sequence: WP_343219578.1 of the αHL nanopore protein; preferably, the αHL nanopore mutant monomer comprises αHL... T109A αHL E111A αHL M113F αHL M113R αHL M113V αHL K147N αHL T115A αHL T117A αHL T117C αHL T117G αHL T117S αHL N121A αHL N121D αHL N121Q αHL N123D αHL N123Q αHL N123A αHL N139D αHL N139Q αHL M113H αHL M113K αHL M113D αHL M113E αHL T145R αHL G143R αHL M113R / T145R αHL M113R / G143R αHL M113R / S141A αHL M113R / T145A αHL M113R / T117A αHL M113R / T115A αHL M113R / E111A αHL M113R / K147A Preferably, the bio-nanopore αHLM113R is a single-point mutant homoheptamer based on the narrowest part and vicinity of αHL, and the bio-nanopore αHLM113R is a multi-point mutant homoheptamer based on the barrel shape of αHLβ. M113R / T115A αHL M113R / K147A .
3. The sequencing method for carbohydrate molecules based on glycosidases and nanopores according to claim 1, characterized in that, Step (2) involves constructing a detection system containing the nanopores and glycosidase, including: (i) Provide an electrolyte solution: a solution capable of dissolving carbohydrate compounds and driving sugar molecules to generate detectable ionic currents by electroosmotic flow in the form of passing through or not passing through nanopores; preferably, the electrolyte solution includes, but is not limited to, lithium chloride, sodium chloride, potassium chloride, magnesium chloride, and guanidine chloride solutions of a certain concentration and a certain pH. (ii) Providing an insulating layer: The insulating layer comprises a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores; preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts; (iii) A nanoporous protein is inserted into and penetrates an insulating layer to form a nanopore, and a power supply is connected to the positive and negative terminals on both sides of the insulating layer; preferably, the nanopore is located in the insulating layer and connects to two portions of the electrolyte solution; preferably, the glycosidase is free in the electrolyte solution or is connected to the nanopore through chemical linkage or fusion expression to form a nanopore-glycosidase complex; preferably, the nanopore-glycosidase complex has the function of cleaving the sugar molecule to be tested and the function of generating an electrical signal by the sugar molecule; preferably, the potential difference on both sides of the nanopore is generally from tens of mV to hundreds of mV, and if it exceeds the upper limit, the pore is unstable and cannot be used for detection; Preferably, the detection system of the nanopore and glycosidase further includes aptamers, sugar-binding proteins and / or sugar molecule modification tags; preferably, the aptamers include β-cyclodextrin, α-cyclodextrin, and γ-cyclodextrin; preferably, the sugar-binding proteins include lectins or sugar-capturing proteins.
4. The sequencing method for carbohydrate molecules based on glycosidases and nanopores according to claim 1, characterized in that, The glycosidase mentioned in step (2) includes an enzyme that recognizes the monosaccharide sequence and / or glycosidic bond of a sugar molecule and hydrolyzes it from the glycosidic bond site of the sugar molecule; preferably, the glycosidase includes a specific glycosidase or a non-specific glycosidase. Preferably, the glycosidase can be a single glycosidase or an enzyme array composed of multiple glycosidases; the enzyme array includes various different arrangements and combinations; preferably, the specific glycosidase includes an exoglycosidase that specifically recognizes the non-reducing end monosaccharide and glycosidic bond of a sugar molecule and hydrolyzes it to release the monosaccharide, or an endoglycosidase that specifically recognizes the internal sequence and glycosidic bond of a sugar molecule and hydrolyzes it to release disaccharide or other oligosaccharide fragments; preferably, the non-specific glycosidase includes a non-specific exoglycosidase that hydrolyzes non-specifically from the ends of sugar molecules one by one, or a non-specific endoglycosidase that hydrolyzes non-specifically from the inside of sugar molecules; preferably, the glycosidase includes a specific exoglycosidase array that can be used for nanopore sugar sequencing. The provided specific glycosidase can efficiently hydrolyze and release monosaccharides from the non-reducing ends of the target sugar molecule one by one in the provided detection system or in other solutions. The exoglycosidase can be added directly to the above detection system or other solutions to hydrolyze the sugar molecules, or it can be connected to the sugar molecule inlet of the nanopore provided in the first aspect to hydrolyze the sugar molecules entering the pore. An exoglycosidase array is a combination of all exoglycosidases that can potentially hydrolyze a particular class of carbohydrates. Preferably, the glycosidases are located on the same side of the sugar sample to be tested.
5. The sequencing method for carbohydrate molecules based on glycosidases and nanopores according to claim 1, characterized in that, In step (3), the sugar sample to be tested includes naturally occurring oligosaccharides, polysaccharides, or their derivatives; preferably, the sugar has a chain length of n, where n is any integer; preferably, n is a natural number from 2 to 100; preferably, n is from 2 to 50; preferably, n is from 2 to 10; preferably, the sugar sample to be tested includes one or more of decasaccharides, nonasaccharides, octasaccharides, heptasaccharides, hexasaccharides, pentasaccharides, tetrasaccharides, trisaccharides, or disaccharides; preferably, the sugar sample to be tested is located in any one of the chambers.
6. The sequencing method for carbohydrate molecules based on glycosidases and nanopores according to claim 1, characterized in that, The hydrolysis reaction described in step (4) includes specific external hydrolysis, specific internal hydrolysis, non-specific external hydrolysis, or non-specific internal hydrolysis; preferably, the inference in step (4) by comparing the changes in electrical signals before and after the hydrolysis reaction includes forward sequencing or reverse sequencing.
7. A carbohydrate molecule sequencing system based on glycosidases and nanopores, characterized in that, The sugar sequencing system includes: (1) Nanopore: The nanopore is a channel that has a size that matches the molecular size of the sugar sample to be tested and can be passed through by collision or displacement of sugar molecules, or not passed through by collision, or not passed through by covalent bonding, or has dissociation characteristics under the influence of external force, thereby producing a change in properties; (2) Glycosidase: Glycosidases that exist naturally or have been modified by enzyme engineering and may or may not have specificity; (3) Electrolyte solution: A solution that can dissolve carbohydrate compounds and drive sugar molecules to generate detectable ionic currents by electroosmosis in the form of passing through or not passing through nanopores; (4) Insulating layer: The insulating layer includes a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores.
8. The carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 7, characterized in that, The nanopores are located in the insulating layer and are connected to the electrolyte solution; the electrolyte solution includes ion solutions of different concentrations such as KCl, NaCl, LiCl, MgCl2, etc., or supersaturated solutions of the above ions; preferably, the channel with changing properties refers to a channel in which the change of ion current manifests as a change in electrical signal.
9. The carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 7, characterized in that, The nanopores include biological nanopores and / or solid nanopores; preferably, the biological nanopores include α-hemolysin, MspA, or aerolysin; preferably, the biological nanopores include a protein complex consisting of at least one wild-type protein monomer and at least one mutant monomer, or a homologous protein complex consisting of homologous mutant monomers; the mutant monomers are obtained by mutation at at least one position in the NCBI Reference Sequence: WP_343219578.1 of the αHL nanopore protein; preferably, the αHL nanopore mutant monomers include αHL... T109A αHL E111A αHL M113F αHL M113R αHL M113V αHL K147N αHL T115A αHL T117A αHL T117C αHL T117G αHL T117S αHL N121A αHL N121D αHL N121Q αHL N123D αHL N123Q αHL N123A αHL N139D αHL N139Q αHL M113H αHL M113K αHL M113D αHL M113E αHL T145R αHL G143R αHL M113R / T145R αHL M113R / G143R αHL M113R / S141A αHL M113R / T145A αHL M113R / T117A αHL M113R / T115A αHL M113R / E111A αHL M113R / K147A .
10. The carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 7, characterized in that, The sugar sequencing system further includes aptamers, sugar-binding proteins, and / or sugar molecule modification tags; preferably, the aptamers include β-cyclodextrin, α-cyclodextrin, and γ-cyclodextrin; preferably, the sugar-binding proteins include lectins or sugar-capturing proteins.
11. The carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 7, characterized in that, The glycosidase is free in the electrolyte solution or linked to the nanopore through chemical linkage or fusion expression to form a nanopore-glycosidase complex; the glycosidase includes enzymes that recognize monosaccharide sequences and / or glycosidic bonds of sugar molecules and hydrolyze them from the glycosidic bond sites of sugar molecules; preferably, the glycosidase includes specific glycosidases or non-specific glycosidases; preferably, the specific glycosidase includes exoglycosidases that specifically recognize the non-reducing terminal monosaccharide and glycosidic bond of sugar molecules and hydrolyze them to release monosaccharides, or endoglycosidases that specifically recognize the internal sequence and glycosidic bond of sugar molecules and hydrolyze them to release disaccharides or other oligosaccharide fragments; preferably, the non-specific glycosidase includes non-specific exoglycosidases that hydrolyze sugar molecules one by one from the ends, or non-specific endoglycosidases that hydrolyze sugar molecules from the inside.
12. The carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 7, characterized in that, The insulating layer includes a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores; preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts to form two side chambers.
13. The carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 7, characterized in that, The sugar sequencing system also includes an electrical signal monitoring system and a data processing system; the electrical signal monitoring system is used to monitor changes in the electrical signal of the nanopore caused by sugar molecules, and the data processing system is used to infer the sugar sequence information of the sugar sample to be tested by comparing the changes in electrical signals before and after the hydrolysis reaction.
14. A method for preparing a carbohydrate molecule sequencing system based on glycosidases and nanopores, characterized in that, Includes the following steps: (1) Providing nanopores: The nanopores are channels that are matched with the molecular size of the sugar sample to be tested and can be passed through by collisions or displacement of sugar molecules, or not passed through by collisions, or not passed through by covalent bonds, or have dissociation characteristics under the influence of external forces, thereby producing changes in properties. (2) Construct a detection system containing the nanopores and glycosidases.
15. The method for preparing a carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 14, characterized in that, The detection system of nanopores and glycosidases further includes aptamers, sugar-binding proteins, and / or sugar molecule modification tags; preferably, the aptamers include β-cyclodextrin, α-cyclodextrin, and γ-cyclodextrin; preferably, the sugar-binding proteins include lectins or sugar-capturing proteins.
16. The method for preparing a carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 14, characterized in that, The construction of the detection system containing the nanopores and glycosidase includes: (i) Provide an electrolyte solution: a solution capable of dissolving carbohydrate compounds and driving sugar molecules to generate detectable ionic currents by electroosmotic flow in the form of passing through or not passing through nanopores; preferably, the electrolyte solution comprises ionic solutions of different concentrations such as KCl, LiCl, NaCl, MgCl2, etc., or supersaturated ionic solutions of the above. (ii) Providing an insulating layer: The insulating layer comprises a phospholipid bilayer, a thin film made of other materials, or a thin film material for preparing solid nanopores; preferably, the insulating layer is located in the middle of the electrolyte solution, dividing the electrolyte solution into two parts; (iii) Inserting nanoporous proteins into and penetrating the insulating layer to form nanopores, and connecting the positive and negative terminals of the power supply on both sides of the insulating layer; preferably, the nanopores are located in the insulating layer and are connected to the two portions of the electrolyte solution.
17. The method for preparing a carbohydrate molecule sequencing system based on glycosidase and nanopores according to claim 14, characterized in that, The glycosidase is either free in the electrolyte solution or connected to the nanopore through chemical linkage or fusion expression to form a nanopore-glycosidase complex; preferably, the nanopore-glycosidase complex has the function of cleaving the sugar molecule to be tested and the function of being induced by the sugar molecule to generate an electrical signal.
18. A nanoporous mutant protein, characterized in that, The nanopore mutant protein is obtained by mutating an amino acid at at least one position in the NCBI Reference Sequence: WP_343219578.1 of the αHL nanopore mutant monomer sequence.
19. The nanoporous mutant protein according to claim 18, characterized in that, The nanopore mutant protein includes mutant αHL T109A , αHL E111A , αHL M113F , αHL M113R , αHL M113V , αHL K147N , αHL T115A , αHL T117A , αHL T117C , αHL T117G , αHL T117S , αHL N121A , αHL N121D , αHL N121Q , αHL N123D , αHL N123Q , αHL N123A , αHL N139D , αHL N139Q , αHL M113H , αHL M113K , αHL M113D , αHL M113E , αHL T145R , αHL G143R , αHL M113R / T145R , αHL M113R / G143R , αHL M113R / S141A , αHL M113R / T145A , αHL M113R / T117A , αHL M113R / T115A , αHL M113R / E111A , αHL M113R / K147A .
20. A nucleic acid molecule encoding the nanopore mutant protein of claim 18 or 19.
21. An expression cassette, recombinant vector, recombinant bacterial strain, or recombinant cell comprising the nucleic acid molecule of claim 20.
22. A biological nanopore or nanopore-glycosidase complex, characterized in that, The bio-nanopore or nanopore-glycosidase complex includes the nanopore mutant protein of claim 18 or 19.
23. The application of the nanopore mutant protein of claim 18 or 19, the nucleic acid molecule of claim 20, the expression cassette, recombinant vector, recombinant strain or recombinant cell of claim 21, the biological nanopore of claim 22, or the nanopore-glycosidase complex of claim 22 in carbohydrate molecule sequencing.
24. A nanopore sugar sequencer, characterized in that, The nanopore sugar sequencer includes the nanopore mutant protein as described in claim 18 or 19.
25. A data processing method for glycosidase-assisted nanoporous sugar sequencing, characterized in that, Includes the following steps: Obtain the electrical signals before and after the reaction of glycosidase hydrolyzing sugar molecules; Data analysis is performed on the acquired electrical signals. Based on the comparison between the first electrical signal before the hydrolysis reaction and the second electrical signal after the hydrolysis reaction, it is determined whether the glycosidase has successfully hydrolyzed the sugar molecules. Then, based on the characteristics of the glycosidase involved in the hydrolysis, the sugar molecules that are released step by step are inferred, thereby inferring the sequence structure of the sugar chain in reverse.
26. The data processing method for glycosidase-assisted nanoporous sugar sequencing according to claim 25, characterized in that, The signal comparison includes comparing the amplitude variation, dwell time variation, open frequency variation, Gaussian distribution fitting value or exponential distribution fitting value of the standard deviation of the current signal, or comparing the parameters of one or more feature distributions, including KL divergence, JS divergence, EMD distance, overlap coefficient and Bartholomew distance.
27. The data processing method for glycosidase-assisted nanoporous sugar sequencing according to claim 25, characterized in that, The method of determining whether the glycosidase has successfully hydrolyzed sugar molecules based on comparing the first electrical signal before the hydrolysis reaction with the second electrical signal after the hydrolysis reaction includes: Extract the first electrical signal features and the second electrical signal features, calculate the difference between the first electrical signal features and the second electrical signal features, and input the difference into a trained classification model to obtain a classification result on whether the first electrical signal features and the second electrical signal features are the same; or, extract the first electrical signal features and the second electrical signal features, input the first electrical signal features and the second electrical signal features into a trained classification model to obtain a classification result on whether the first electrical signal features and the second electrical signal features are the same.
28. The data processing method for glycosidase-assisted nanoporous sugar sequencing according to claim 25, characterized in that, For an unknown glycan sequence, the process of inferring the complete glycan sequence from the reverse includes: S1, obtain the initial current signal feature Data_m, where m represents the number of successful enzymatic hydrolysis and the position of free monosaccharides in the original glycan chain, with an initial value of 1; S2, obtain the current signal characteristics of each glycosidase hydrolysis product in the glycan and enzyme array, Data_m_f, where f represents the sample number after enzymatic hydrolysis of different types of enzymes; S3. Compare all Data_m_f with Data_m to determine whether hydrolysis has occurred. For samples that have been determined to have hydrolyzed, output f and m. The corresponding Data_m_f is used as the reference data Former_data for the next round of hydrolysis comparison. This sample is used as the initial sample Former_sample for subsequent hydrolysis. Update the successful hydrolysis count m=m+1. S4, repeat S2-S3 until the complete sugar chain sequence is determined.
29. A data processing method for glycosidase-assisted nanoporous sugar sequencing, characterized in that, The steps include: obtaining the electrical signal caused by the release of sugar molecules by glycosidase hydrolysis; Based on the comparison of electrical signals and fingerprint patterns caused by the release of sugar molecules by glycosidase hydrolysis, the gradually released sugar molecules are inferred, thereby leading to the forward inference of the sequence structure of the sugar chain.
30. The data processing method for glycosidase-assisted nanoporous sugar sequencing according to claim 29, characterized in that, The fingerprint spectrum includes a monosaccharide fingerprint spectrum library, a disaccharide fingerprint spectrum library, and other oligosaccharide fingerprint spectrum libraries. The fingerprint spectrum is used to assign signals to electrical signals, obtain the types of terminal oligosaccharide fragments during hydrolysis, and deduce the compositional order of oligosaccharide fragments based on the order of occurrence of electrical signal events and / or the order of signal abundance.
31. The data processing method for glycosidase-assisted nanoporous sugar sequencing according to claim 29, characterized in that, The fingerprint spectrum is established based on electrical signal features, which include one or more of the following: amplitude of current signal, dwell time, open frequency, and standard deviation.
32. A data processing system for glycosidase-assisted nanopore sugar sequencing, used to implement the data processing method according to any one of claims 29 to 31, characterized in that, include: The reverse sequencing input module is used to acquire electrical signals before and after the glycosidase hydrolysis of sugar molecules. The reverse sequencing information processing module is used to analyze the acquired electrical signals. Based on the comparison of the first electrical signal before the hydrolysis reaction and the second electrical signal after the hydrolysis reaction, it determines whether the glycosidase has successfully hydrolyzed the sugar molecules. Then, based on the characteristics of the glycosidase involved in the hydrolysis, it infers the sugar molecules that are released step by step, thereby inferring the sequence structure of the sugar chain in reverse.
33. A data processing system for glycosidase-assisted nanopore sugar sequencing, used to implement the data processing method according to any one of claims 29 to 31, characterized in that, include: The forward sequencing input module is used to acquire the electrical signals caused by the release of sugar molecules by glycosidase hydrolysis; The forward sequencing information processing module infers the progressively released sugar molecules based on the electrical signals and fingerprint patterns caused by the hydrolysis of sugar molecules by glycosidases, thereby forward inferring the sequence structure of the sugar chain.
34. A computer system comprising a memory, a processor, and computer programs / instructions stored in the memory and executable on the processor, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 29 to 31.
35. A computer-readable storage medium storing a computer program, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 29 to 31.
36. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 29 to 31.
37. A nanopore sequencer based on glycosidase hydrolysis, characterized in that, The system includes a glycosidase hydrolysis system, a nanopore glycan electrical signal monitoring system, and a data processing system; the glycosidase hydrolysis system is used for glycosidase hydrolysis of sugar molecules, and the nanopore glycan electrical signal monitoring system is used for monitoring changes in the electrical signal of the nanopore caused by sugar molecules; the data processing system is the data processing system according to claim 32 or 33; or, the data processing system is based on the computer program product according to claim 34 to automatically generate the sequence structure of the glycan.
38. A nanoporous glycan electrical signal monitoring system, characterized in that, It includes two chambers separated by an insulating layer, nanopores formed on the insulating layer, electrodes placed in the two chambers for measuring changes in electrical signals, and an instrument for acquiring changes in electrical signals; the nanopores are channels that match the molecular size of the sugar sample to be tested and can be passed through by sugar molecules through collisions or displacements, or not through collisions, or not through covalently bonded molecules, or have dissociation characteristics under the influence of external forces, thereby causing changes in properties.
39. The nanoporous glycan electrical signal monitoring system according to claim 38, characterized in that, The channel whose properties change refers to the channel through which the change in ion current manifests as a change in electrical signal.
40. The nanoporous glycan electrical signal monitoring system according to claim 39, characterized in that, The nanopores include biological nanopores and / or solid nanopores; preferably, the biological nanopores include αHL, MspA, or aerolysin; preferably, the biological nanopores include a protein complex consisting of at least one wild-type protein monomer and at least one mutant monomer, or a homologous protein complex consisting of homologous mutant monomers; the mutant monomers are obtained by mutation at at least one position in the αHL nanopore NCBI Reference Sequence: WP_343219578.1.