Method for establishing visual model of degradation path and product of carbohydrates

By standardizing and recursively cleaving the chemical structure data of carbohydrates, and combining a multi-microbial synergistic model and a multi-dimensional scoring system, a visual model of the degradation pathways and products of carbohydrates is generated. This solves the problems of high false positives and insufficient consistency in the analysis of carbohydrate degradation pathways in existing technologies, and achieves higher pathway coverage and product hit rate.

CN121459971APending Publication Date: 2026-02-03SHANGHAI YINUO YIKANG BIOMEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511592846.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies for analyzing carbohydrate degradation pathways suffer from problems such as high false positive rates, uninterpretable pathways, and insufficient consistency with literature/experiments. They also lack fine-grained expression and pathway assembly based on strain capabilities.

Method used

By acquiring the chemical structure data of the target carbohydrate, standardizing it, and then recursively cleaving and assembling it, combined with a multi-microbial synergistic model and a multi-dimensional scoring system, a visual model of the degradation pathway and products of carbohydrate is generated, including the rating and polymerization separation of the step-enzyme-substance ternary, and chemical consistency verification and thermodynamic screening are performed.

Benefits of technology

This study enabled refined modeling of carbohydrate degradation processes, improved the consistency between pathway prediction and actual physiological processes, reduced product misjudgment, and increased pathway coverage and product hit rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459971A_ABST
    Figure CN121459971A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information, in particular to a method for establishing a visual model of degradation paths and products of carbohydrates, which comprises the following steps: acquiring chemical structure data of target carbohydrates, standardizing the chemical structure data, and outputting a standardized molecular map; recursively splitting the standardized molecular map, and outputting a candidate intermediate set and a reaction sequence; according to the reaction sequence of the reaction sequence, assembling the candidate intermediate set, and outputting a path diagram; distributing candidate enzymes and species for each step of the path diagram, and outputting a step-enzyme-substance triple; step rating and polymerization splitting are carried out on the step-enzyme-substance triple, and a confidence interval is output; and outputting the degradation path of the target saccharide and a visual model of the product according to the confidence interval. The problems that in the prior art, the product false positive is high, the path cannot be explained, and the consistency with literature / experiment is insufficient are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular to a method for establishing a visual model of a degradation path and products of a saccharide substance. BACKGROUND

[0002] As the core substrate of microbial metabolism, the accurate analysis of the degradation path and product spectrum of carbohydrates is of great significance to the fields of intestinal health research, industrial fermentation optimization, and regulation of ruminant production. With the accumulation of multi-omics data such as metagenomics and metabolomics and the development of computing technology, bioinformatics-based carbohydrate degradation path analysis methods have gradually replaced traditional experimental screening and become the mainstream research direction in the field. Current path analysis tools rely on annotated enzyme function databases to implement path assembly, and the core logic is to concatenate metabolic steps through enzyme substrate-product correspondence. For example, a reaction database constructed based on the IUBMB enzyme list can match reactions by keyword retrieval of substrates / products, and the user can manually select product nodes to iteratively generate path graphs. The CAZy database of carbohydrate-active enzymes has become a core data support for path analysis, and by comparing the coding genes of CAZy family enzymes in the genome, the sugar degradation capacity of a species can be predicted. This technology has initially realized the tracking of the path from the carbohydrate substrate to the intermediate product, providing a basic tool for simple sugar metabolism analysis.

[0003] However, the above-mentioned technology only achieves database entry retrieval or single-step prediction, lacks fine-grained expression of carbohydrate structures, lacks path assembly based on strain capacity, and ignores constraints such as cross-bacterial collaboration, anaerobicity, localization, and transport, which can lead to high false positive rates of products, uninterpretable paths, and insufficient consistency with literature / experiments.

[0004] Therefore, there is an urgent need to provide a method for establishing a visual model of a degradation path and products of a saccharide substance. SUMMARY

[0005] To solve the above technical problems in the prior art, the present application provides a method for establishing a visual model of a degradation path and products of a saccharide substance.

[0006] To achieve the above-mentioned purpose, the technical solution of the present application is as follows: A method for establishing a visual model of a degradation path and products of a saccharide substance, comprising the following steps: S1: obtaining chemical structure data of a target saccharide substance, standardizing the chemical structure data, and outputting a standardized molecular graph; S2: recursively cracking the standardized molecular graph, and outputting a candidate intermediate set and a reaction sequence; S3: Assemble the candidate intermediate set according to the reaction order of the reaction sequence, and output the pathway graph; S4: Assign candidate enzymes and species to each step of the pathway graph, and output the step-enzyme-species triplets; S5: Step rating and aggregate splitting of step-enzyme-species triplets, output confidence interval; S6: According to the confidence interval, output the degradation pathway of the target saccharide and the visualization model of the product.

[0007] Further, the steps for standardizing the chemical structure data and outputting the standardized molecular graph are: According to the type of chemical structure data, determine the molecular structure of the target saccharide; Analyze the sugar source, connection site and branching of the molecular structure, and standardize it into a directed graph of glycosyl-bond-site; Generate InChIKey to remove the directed graph of glycosyl-bond-site, and output the standardized molecular graph.

[0008] Further, the steps for recursively splitting the standardized molecular graph and outputting the candidate intermediate set and reaction sequence are: Perform subgraph matching on the standardized molecular graph, label the cleavable sites and candidate enzyme families, and execute Transform to generate candidate intermediates; BFS layer-by-layer expansion of candidate intermediates until the preset step number is reached, outputting the candidate intermediate set; Expand the candidate intermediate set according to the rule complexity and prune it with a branch_prune probability, and output the reaction sequence.

[0009] Further, the steps for assembling the candidate intermediate set according to the reaction order of the reaction sequence and outputting the pathway graph are: Integrate transmembrane transport, extracellular-periplasmic-intracellular localization, anaerobic and pH constraints into the reaction sequence, and use the multi-bacterial collaboration mode to execute the assembly of the reaction sequence, and output the pathway graph.

[0010] Further, the steps for executing the assembly of the reaction sequence using the multi-bacterial collaboration mode are: Construct the capability matrix, which stores enzyme-species combinations, each enzyme-species combination corresponds to localization feasibility data and transmembrane transport feasibility data, wherein the localization feasibility data is used to verify whether the enzyme-species combination matches the extracellular, periplasmic or intracellular localization requirements of the metabolic step, and the transmembrane transport feasibility data is used to verify whether the enzyme-species combination matches the transmembrane transport direction and carrier occupancy requirements of the metabolic step; Filter the enzyme-species combinations in the capability matrix in the following order: Prioritize the combination of matched localization, re-screen the combination that meets the transmembrane transport requirements, and associate the screened combination to the path edge to complete the collaborative splicing; If the enzyme-species combination does not meet the positioning and transport requirements during the screening process, the corresponding path branch is removed or its weight is reduced.

[0011] Further, assign candidate enzymes and species to each step of the path graph, and output the step-enzyme-species triplets: Screen the species corresponding to the candidate enzyme from the capability matrix, wherein the screening conditions include the matching of the CAZy / EC number of the step and the CAZy / EC number of the enzyme in the capability matrix, the similarity of the substrate of the step and the substrate preference of the enzyme in the capability matrix, and the matching of the spatial positioning requirement of the step and the positioning signal of the enzyme in the capability matrix. Output the step-enzyme-species triplets, wherein the triplets include the step identification in the path graph, the candidate enzyme assigned to the step, and the corresponding species.

[0012] Further, perform step rating and aggregate splitting on the step-enzyme-species triplets, and output the step with a confidence interval: Establish a step-level scoring model C_step (i) for the step-enzyme-species triplets, ; Wherein, step(i) represents the i-th step in the path, RuleMatch(i) represents the rule hit confidence, ECLevel(i) represents the EC level matching degree, SubSim(i) represents the substrate similarity, LocMatch(i) represents the localization matching degree; ΔG_proxy(i) represents the proxy value of thermodynamic parameter, α1, α2, α3, α4, α5 all represent weight parameters, and the value range of each is [0, 1]; Establish a path-level scoring model Score (P) for the step-enzyme-species triplets, Score(P)=w_rule +w_ec +w_sub +w_loc +w_thermo + w_len f(L)+w_abd g(Abundance); Wherein, , , , , respectively represent the average value of the corresponding step indicator on the path, f (L) represents the length penalty function, g(Abundance) represents the individualized evidence aggregation value, w_rule, w_ec, w_sub, w_loc, w_thermo, w_len, w_abd represent the weight vector, the value range is [0, 1] and the sum is 1; f (L)=exp(- (L-1)); Wherein, L represents the number of path steps, The value range of is [0.1, 0.5]; By Dirichlet sampling of the weight vector w, the 95% of the path level score is taken as the confidence interval.

[0013] Further, according to the confidence interval, the step of outputting the visualization model of the degradation path and product of the target saccharide is: The metabolic pathways are sorted according to the score of the confidence interval, the contained end product set and intermediate product set are extracted and output, and the visualization result and structured data are generated, wherein the visualization result at least includes one of network atlas, pathway diagram and heat map, and the structured data at least includes one of CSV format file, JSON format file and HTML table; The visualization result and structured data are taken as the visualization model of the degradation path and product of the target saccharide.

[0014] Compared with the prior art, the present application has the following beneficial effects: The method for establishing the visualization model of the degradation path and product of the saccharide provided by the present application involves multidisciplinary technologies of chemical thermodynamics, bioinformatics and software engineering, and integrates the comprehensive advantages of various technologies. On the one hand, based on the path assembly and multi-dimensional scoring system of the CSRR rule, the fine modeling of the saccharide degradation process is realized, the parameter setting and formula design of which are verified by a large amount of experimental data, and have strong scientificity and stability; on the other hand, the construction of the engineering guarantee mechanism solves the uncertainty caused by external data dependence and ensures the reliable operation of the scheme in different environments.

[0015] The application also solves the problems of path redundancy and product misjudgment in traditional unconstrained analysis through chemical consistency verification, thermodynamic screening and layer-by-layer verification of multiple environmental constraints. In the testing of standard sugars such as lactose, inulin and beta-glucan, the Top-k path coverage and product hit rate far exceed the unconstrained control, especially after introducing sample abundance and enzyme expression data, the consistency of path prediction and actual physiological process is greatly improved. The double protection of chemical rationality and physiological relevance of the application makes the degradation path output by the scheme not only conform to the theoretical chemical law, but also better fit the real biological scene. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 The network graph format output by the application; Figure 2 The pathway graph format output by the application; Figure 3 The heat map format output by the application. DETAILED DESCRIPTION

[0017] The technical solutions of the application will be described clearly below in conjunction with the accompanying drawings. Obviously, the described embodiments are not all the embodiments of the application, and all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.

[0018] It should be noted that, unless otherwise specified, the relative arrangement of components and steps, numerical expressions set forth in these embodiments should not be understood as limiting the scope of the application.

[0019] The following description of exemplary embodiments is merely illustrative in nature and is in no way intended to limit the application or its application or use in any way. Techniques, methods and devices known to those skilled in the relevant art may not be discussed in detail here, but in the case of applicable techniques, methods and devices, these techniques, methods and devices should be considered as part of this specification.

[0020] The embodiment provides a method for establishing a visual model of degradation path and product of sugar substances, including the following steps: analyzing the input molecular object, supporting SMILES / WURCS / GlycoCT, etc., extracting sugar residues and connection sites, sequentially performing implicit hydrogen removal, uniform chirality, main chain priority / branch dictionary order and isomorphic re-labeling, and finally exporting a standard adjacency table.

[0021] De-hydrogenation is to make the hydrogen atoms in SMILES / WURCS / GlycoCT default hidden explicit, to avoid the structure misjudgment caused by different ways of representing hydrogen atoms. The specific steps are: read the bonding information of each atom in the input format. For example, the valence of carbon is four, and the valence of oxygen is two. Calculate the number of hydrogen atoms that are not explicitly labeled for each atom. For example, C in SMILES represents CH4, and 3 explicit H need to be supplemented; CH2 needs to supplement 1 explicit H. Add the corresponding number of explicit hydrogen nodes to each atom in the molecular object, and construct the connection edge between atom and hydrogen, to ensure that all chemical bonds are explicitly visible.

[0022] Uniform chirality is to unify the representation of the configuration of the chiral center in the sugar residue, and to eliminate the difference caused by the same configuration with different symbols. The specific steps are: identify the chiral center of the sugar residue, and focus on the anomeric carbon and the monosaccharide backbone carbon. Convert the different formats of chiral identification to the standard system. For example, the anomeric carbon is unified as α / β, and the other chiral carbons are unified as R / S. Through the symmetry analysis of the molecular structure, ensure that the unified chiral identification is consistent with the actual spatial configuration. For example, the hydroxyl group of α-C1 of glucose is opposite to the C6 hydroxymethyl, and the symbol annotation needs to be confirmed correctly.

[0023] Main chain priority / branch lexicographic is to clarify the main chain and branch of the sugar chain, sort the branches according to the rules, and avoid the structure difference caused by the misjudgment of main chain / branch or the confusion of branch order. The specific steps are: according to the principle of the most number of sugar residues and the most typical type of connecting bond, select the main chain of the sugar chain. For example, the chain of 5 glucose connected by β(1→4) is the main chain, and the single glucose connected at C6 is the branch. For each branch of the sugar residue in the main chain, first sort the branches according to the alphabetical order of the branch sugar residue name; if the letters of the branch sugar residue name are the same, sort them in ascending order according to the number of the branch connection site. Add the main chain or branch label to each chain in the molecular object, and record the sorting results of the branches.

[0024] Isomorphic re-labeling is to give the same structure of sugar residues or sub-chains a unified number and label, to avoid the repetition or misjudgment caused by the same structure with different labels. The specific steps are: detect the same structure of sugar residues in the molecule through the four elements of atomic type, bonding mode, chiral configuration, and substituent group. For example, 2 α-D -glucose residues are isomorphic. For the isomorphic sugar residues, give them consecutive numbers according to their position in the main chain / branch. For example, the first glucose in the main chain is Glc1, the second is Glc2, and the isomorphic branch glucose is Glc3. If there are isomorphic sub-chains, unify the labels of the sub-chains. For example, 2 Glc-α(1→6)-Gal sub-chains are both labeled as Subchain A.

[0025] The canonical adjacency table is a structured table description of the structure of the sugar molecule after normalization, which clearly records the connection relationship of all atoms and sugar residues by node table and edge table. Take lactose Gal-β(1→4)-Glc as an example, the node table and edge table of the canonical adjacency table are shown.

[0026] Table 1 Node table of lactose Gal-β(1→4)-Glc

[0027] Table 2 Edge table of lactose Gal-β(1→4)-Glc

[0028] According to the canonical adjacency table, the directed graph G of glycosyl-site is established, G=(Node, Edge). Among them, the node (Node): define the node type∈{glycosyl, connection site, product, cofactor}; store the node props containing stereochemistry, ring type, bond type, site, substituent, atomic valence / charge, etc. Edge (Edge): define the node type∈{glycosidic bond, cleavage, transport}; store the node props containing bond direction, site, evidence and positioning.

[0029] Using chemical informatics tool library, such as RDKit, Open Babel, the directed graph G is processed to generate international compound identifier InChI and its hash value InChIKey. For isomers at the level of sugar residues, the characteristic parameter combination of the residue is extracted as a secondary bond, and the format of the secondary bond is: (monosaccharide name, ring type, configuration, connection site set). Take the secondary bond as an index, and compare it with the existing secondary bond. If two secondary bonds are exactly the same, it is judged as the same sugar residue, and is deleted to achieve the effect of fine-grained deduplication, and output the standardized molecular graph.

[0030] The standardized molecular graph is screened by CSRR rules (Carbohydrate-Specific Reaction Rules), and then the breadth-first search algorithm BFS is used for recursive or iterative cleavage. Among them, the CSRR rules include BondPattern, EnzymeFamily and Transform. BondPattern is to describe the glycosidic bond, branching and site information through SMARTS language or subgraph template; EnzymeFamily is a classification system that maps the Carbohydrate-Active Enzymes database CAZy family or Enzyme Commission EC number, and can also include the hierarchical relationship from family→EC→subtype to accurately limit the scope of the enzyme; Transform is to record the key operations in the reaction, which must meet the law of chemical conservation, and mark the product structure generated in the reaction.

[0031] Traverse all BondPattern of CSRR rules in the rule base, and convert them into computable subgraph query templates; execute subgraph matching algorithm, such as VF2 algorithm, in the standardized molecular graph to locate all substructures conforming to the BondPattern; for each matched substructure, associate it with its corresponding EnzymeFamily, mark the site as cleavable in the graph, and attach the candidate enzyme family label.

[0032] For each site matching the BondPattern, call the Transform instruction corresponding to the rule; modify the directed graph structure as required by Transform, while verifying mass conservation and charge conservation; generate candidate intermediates that inherit contextual information, including rule ID triggering the reaction, enzyme family acting, cleavage site coordinates, and the location of the reaction.

[0033] Initialize the queue queue, and enqueue the directed graph G0 of the initial sugar structure as the starting point, denoted as step 0; take out the current graph Gi (step i structure) from the queue, and traverse all CSRR rules applicable to Gi; for each rule, generate the corresponding product graphs Gi1, Gi2,..., Gim (candidate intermediates of step i+1), and enqueue all these product graphs; record the parent→child relationship of each step (e.g., Gi is the parent node of Gi1), forming a path tree.

[0034] For each candidate intermediate generated in the queue, calculate its InChIKey and store it in the set seen-set; if the newly generated intermediate InChIKey is already in seen-set, discard it directly. Terminate the search when the current step number steps exceeds the maximum step number max_steps; the value range of max_steps is 1-10, and the default value is 5. Calculate the score of the candidate intermediate, and if the score is lower than the preset threshold, terminate the continuation of the path expansion.

[0035] The score of the candidate intermediate is first sorted according to the rule reliability, i.e., preferentially retaining intermediates generated by CSRR rules; then, it is sorted according to the structural complexity, i.e., preferentially retaining intermediates with simpler structures and higher likelihood of being generated in vivo. The branch pruning parameter branch_prune controls the pruning probability, and the candidate intermediates at the end of the sorting are randomly discarded; the value range of branch_prune is [0, 1], and the default value is 0.05. In the process of random discarding, the principle of higher retention probability of intermediates generated by main chain cleavage than those generated by side chain cleavage is followed.

[0036] Based on the parent-son relationship recorded in the BFS process, a complete reaction path tree is constructed, each node represents an intermediate, and each edge represents a reaction; Add labels to each edge, including the CSRR rule ID that triggers the reaction, the enzyme family and cleavage site, and other key information; Finally output the complete reaction sequence chain from the initial structure to the intermediate product to the final product, and the enzymatic and chemical characteristics of each step.

[0037] In addition, when generating candidate intermediates, chemical reasonableness verification and approximate thermodynamic screening can also be performed, and a Boolean value pass_chem is output after chemical reasonableness verification. If all verifications pass, pass_chem=True, otherwise False. The verification content includes valence state check, octet check, charge conservation check, mass conservation check and bond reasonableness check.

[0038] The specific steps of approximate thermodynamic screening are: based on the reaction type corresponding to the CSRR rule, predefine the free energy contribution value Δg_rule of the typical group in this type of reaction. If the reaction involves the opening or closing of the sugar ring, introduce a penalty coefficient γ, multiply the ring opening and closing times n_ring_open to reflect the energy cost of ring structure change; Based on the charge change value Δcharge before and after the reaction, introduce a correction coefficient ζ to calculate the charge correction term ζ·|Δcharge| to reflect the influence of charge distribution change on free energy; Comprehensive group contribution and correction term, get the proxy value of Gibbs free energy change ΔG_proxy: ΔG_proxy=ΣΔg_rule+γ·n_ring_open+ζ·|Δcharge|; Where, ΣΔg_rule represents the sum of all group contribution values involved in the reaction. Map ΔG_proxy to a thermodynamic weight value w_thermo between 0 and 1 through the Sigmoid function. When w_thermo is closer to 1, it means that the reaction is more spontaneous; When w_thermo is closer to 0, it means that the reaction is less spontaneous. Then set the empirical threshold θ_keep and θ_drop to screen w_thermo. When w_thermo≥θ_keep, it means that the reaction has high thermodynamic feasibility, and the candidate intermediate is retained; When θ_drop<w_thermo<θ_keep, it means that the reaction has medium thermodynamic feasibility, and the candidate is down-weighted; When w_thermo≤θ_drop, it means that the reaction has low thermodynamic feasibility, and the candidate intermediate is directly eliminated.

[0039] Group all the reaction sequences by step loc label. Group the steps with consecutive loc=E as extracellular pathway segment, consecutive loc=P as periplasmic pathway segment, and consecutive loc=C as intracellular pathway segment. For each group of pathway segment, verify the intermediate generation logic of all the reaction steps in the group, and keep the pathway segment with coherent logic as a basic pathway unit. If a consecutive step loc label mutation occurs in a pathway segment, for example, the previous step loc=E and the next step loc=C directly without a transport step, mark it as a to-be-transported supplementary pathway segment and enter the next step for verification.

[0040] Establish a transport mechanism for the to-be-transported supplementary pathway segment. The specific steps are as follows: check the connection relationship of all the basic pathway units, and if the loc at the end of the previous unit is inconsistent with the loc at the beginning of the next unit, determine it as a localization breakpoint; for each localization breakpoint, select a matching option from a pre-set transport mechanism library, and the transport mechanism includes the phosphoenolpyruvate-carbohydrate phosphotransferase system, the ATP-binding cassette transporter, and the channel protein. If there is a matching transport mechanism, insert a transport step at the localization breakpoint, mark Edge.type=transport, and supplement the transport direction, carrier protein, and other labels to form a complete pathway chain; if there is no matching transport mechanism, the connection relationship of the pathway segment is invalid, and the overall pathway is down-weighted or directly determined as an invalid pathway and removed.

[0041] Add the anaerobic attribute to all the pathway segments of the basic pathway units and the transport steps, and forcibly limit anaerobic=True, that is, only keep the reactions that can occur in an anaerobic environment; if a pathway segment can only occur in an aerobic environment, directly remove the pathway segment. At the same time, limit the pH value of the reaction environment of all the pathway segments, pH∈[4.5, 8.0]; by querying the enzyme activity database, such as BRENDA, verify the activity of the key enzyme in the pathway within the pH range is ≥50%; if the enzyme activity is lower than the threshold, down-weight the pathway segment.

[0042] Mark potential cooperative sites in the pathway after adding the anaerobic attribute, which are usually steps of intermediate transfer across species, and mark Edge.type=involved or shared; query the pre-set species-enzyme-positioning ability matrix to screen combinations that meet the following conditions: the species needs to have the enzyme required for the step, the enzyme action positioning needs to be consistent with the step loc label, and if cross-species transport is involved, there needs to be a compatible metabolite transfer mechanism between species. If there is a combination that meets the conditions, supplement the species label to the pathway node and the enzyme label to the edge to complete the cooperative pathway splicing; if there is no combination that meets the conditions, the cooperative branch cannot be continued and is directly truncated.

[0043] Integrate all the above constraint checks and assembly results to generate a structured pathway graph, which contains node labels and edge labels. Each node is annotated with: basic information, intermediate name, InChIKey; constraint information, loc label of generation step, anaerobic=True, pH adaptation range; synergy information, associated species. Each edge is annotated with: type label, Edge.type=reaction, transport, participation or sharing; enzyme label, CAZy enzyme family or EC number corresponding to the reaction; validity label, reserved, down-weighted or invalid.

[0044] Input the species x capability matrix, traverse each reaction step in the structured pathway graph, and filter out species containing the required CAZy / EC family of the step from the matrix; match the substrate preference of the enzyme with the step substrate; verify the consistency of the enzyme localization with the step loc label to obtain the enzyme, species candidate pair for the step. Calculate the step score C_step(i) for each candidate pair: ; Where step(i) represents the ith step in the pathway; RuleMatch(i) represents the rule hit confidence, with a value range of [0,1], based on template matching quality and number of supporting literature, for example, the higher the matching score of SMARTS, the more supporting literature, and the higher the final score; ECLevel(i) is the EC level matching degree, with a value range of [0,1], and the fine level scores higher than the coarse level; SubSim(i) represents the substrate similarity, with a value range of [0,1], the structural similarity between the candidate species and the rule representative species is calculated by Tanimoto coefficient; LocMatch(i) represents the localization matching degree, with a value range of [0,1], when the enzyme localization and step loc match completely, it gets 1 point, partially matches 0.5 points, and does not match 0 points; ΔG_proxy(i) represents the proxy value of Gibbs free energy change of the reaction, 1-ΔG_proxy(i) reflects the thermodynamic feasibility, the higher the value, the easier it is to be spontaneous; α1, α2, α3, α4, α5 are weight parameters, with a value range of [0,1], which can be normalized to ensure the sum is 1, and the default value is {0.25, 0.2, 0.2, 0.2, 0.15}. Output step-enzyme-species triplets, including candidate enzymes, species and corresponding C_step(i) scores associated with each step.

[0045] Aggregate step scores of step-enzyme-species triplets and incorporate global constraints to generate a pathway-level scoring model Score (P): Score(P)=w_rule +w_ec +w_sub + w_loc + w_thermo + w_len f(L)+w_abd g(Abundance); wherein, , , , , respectively represent the average value of the corresponding step indicator on the path; f(L) represents the length penalty function, the more steps the greater the penalty; g(Abundance) represents the individualized evidence aggregation value, the value range is [0, 1], which is the weighted average of bacterial abundance, enzyme expression and literature evidence; w_rule, w_ec, w_sub, w_loc, w_thermo, w_len, w_abd represent the weight vector, the value range is [0, 1] and the sum is 1, the default value is {0.22, 0.18, 0.14, 0.14, 0.12, 0.10, 0.10}: f(L)=exp(- (L-1)); wherein, L represents the number of steps of the path; the value range of λ is [0.1, 0.5], and the default is 0.25. Dirichlet sampling or evidence bootstrap resampling is used to generate 1000 times of simulation scores; the 95% confidence interval CI of the path score, i.e. the 2.5% and 97.5% quantile values, is calculated to evaluate the score stability; the paths are ranked in descending order of Score(P), and the Top-k paths are screened according to the confidence interval, i.e. the top k paths with the highest scores, the paths with high scores and narrow confidence intervals are preferentially retained, and the ranked path list is sorted.

[0046] The end product and intermediate product are extracted from the path graph, the set containing at least one of acetic acid, propionic acid, butyric acid, lactic acid and ethanol is screened, and the Top-k path of the product and its score and confidence interval are generated; based on the step-enzyme-species triplets, the key species and enzymes catalyzing the product generation are listed, and the network, pathway graph and heat map as shown in Figure 1 , Figure 2 and Figure 3 are generated, and then converted into a structured table CSV, a machine-readable JSON or an interactive report HTML format, a version number and parameter configuration are recorded to ensure the reproducibility of the results; the end / intermediate product set and associated information and multi-format visualization results and structured data model are output.

[0047] The above detailed description merely illustrates the technical solutions of the present application and is not limiting, and although the present application has been described in detail with reference to the examples, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or equivalently replaced without departing from the scope of the technical solutions of the present application, and all should be covered in the scope of the claims of the present application.

Claims

1. A method for establishing a visual model of the degradation pathway and products of carbohydrates, characterized in that, Includes the following steps: S1: Obtain the chemical structure data of the target carbohydrate, standardize the chemical structure data, and output a standardized molecular diagram; S2: Recursively cleave the standardized molecular graph to output a set of candidate intermediates and reaction sequences; S3: Assemble the candidate intermediate set according to the reaction sequence and output the path diagram; S4: Assign candidate enzymes and species to each step of the path graph and output the step-enzyme-substance triplet. S5: Perform step rating and polymerization decomposition on the step-enzyme-substance ternary set, and output confidence intervals; S6: Based on the confidence interval, output a visual model of the degradation pathway and products of the target sugar.

2. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 1, characterized in that, The steps for standardizing chemical structure data and outputting a standardized molecular diagram are as follows: Determine the molecular structure of the target carbohydrate substance based on the type of chemical structure data; The molecular structure of glycogen, its connection sites, and branching is analyzed and normalized into a directed graph of glycosyl-bond-site; The InChIKey directed graph of glycosyl-bond-site is generated to remove duplicates and output a normalized molecular graph.

3. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 1, characterized in that, The steps for recursively fragmenting the standardized molecular graph to output a set of candidate intermediates and reaction sequences are as follows: Subgraph matching is performed on the standardized molecular graph to label cleavable sites and candidate enzyme families, and Transform is executed to generate candidate intermediates. The candidate intermediates are expanded layer by layer by BFS until the preset number of steps are reached, and the set of candidate intermediates is output. Expand and sort the candidate intermediate set according to rule complexity, prune it with branch_prune probability, and output the reaction sequence.

4. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 2, characterized in that, The steps for assembling the candidate intermediate set and outputting the path diagram based on the reaction sequence are as follows: The reaction sequence incorporates transmembrane transport, extracellular-periplasmic-intracellular localization, anaerobic and pH constraints, and uses a multi-bacterial synergistic mode to perform reaction sequence assembly, outputting a pathway diagram.

5. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 4, characterized in that, The steps for assembling reaction sequences using a multi-strain synergistic model are as follows: A capability matrix is ​​constructed, which stores enzyme-species combinations. Each enzyme-species combination corresponds to localization feasibility data and transmembrane transport feasibility data. The localization feasibility data is used to verify whether the enzyme-species combination matches the extracellular, periplasmic, or intracellular localization requirements of the metabolic step. The transmembrane transport feasibility data is used to verify whether the enzyme-species combination matches the transmembrane transport direction and carrier occupancy requirements of the metabolic step. The enzyme-species combinations in the capability matrix were screened in the following order: Prioritize selecting combinations that match the location, then filter combinations that meet the requirements for cross-membrane transport, and associate the filtered combinations with the path edge to complete collaborative splicing; If an enzyme-species combination that does not meet the requirements for localization and transport is found during the screening process, the corresponding pathway branch will be removed or its weight will be reduced.

6. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 5, characterized in that, Assign candidate enzymes and species to each step of the path graph, and output the step-enzyme-substance triplet as follows: The species corresponding to candidate enzymes are screened from the capability matrix. The screening criteria include the matching between the CAZy / EC number of the step and the CAZy / E number of the enzyme in the capability matrix, the similarity between the substrate of the step and the substrate preference of the enzyme in the capability matrix, and the matching between the spatial positioning requirements of the step and the positioning signal of the enzyme in the capability matrix. Output the step-enzyme-species triplet, where the triplet contains the step identifier in the path diagram, the candidate enzyme assigned to the step, and the corresponding species.

7. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 1, characterized in that, The steps for which the step-enzyme-substance ternary set is evaluated and polymerized to determine confidence intervals are as follows: Establish a step-level scoring model C_step(i) consisting of a step-enzyme-substance ternary set. ; Where step(i) represents the i-th step in the path, RuleMatch(i) represents the rule hit confidence, ECLevel(i) represents the EC level matching degree, SubSim(i) represents the substrate similarity, LocMatch(i) represents the location matching degree; ΔG_proxy(i) represents the thermodynamic parameter surrogate value, and α1, α2, α3, α4, and α5 all represent weight parameters, and their values ​​range from [0,1]. Establish a path-level scoring model Score(P) for the step-enzyme-substance triplet. Score(P)=w_rule +w_ec +w_sub +w_loc +w_thermo + w_len f(L)+w_abd g(Abundance); in, , , , , Let f(L) represent the average value of the corresponding step index on the path, g(Abundance) represent the length penalty function, and w_rule, w_ec, w_sub, w_loc, w_thermo, w_len, and w_abd represent the weight vectors, all of which take values ​​in the range [0,1] and sum to 1. f (L)=exp(- (L-1)); Where L represents the number of path steps, The value range is [0.1, 0.5]; By performing Dirichlet sampling on the weight vector w, the confidence interval is set at 95% of the path-level score.

8. The method for establishing a visualization model of the degradation pathway and products of carbohydrates according to claim 1, characterized in that, Based on the confidence interval, the steps to output a visualization model of the degradation pathway and products of the target sugar are as follows: The metabolic pathways are sorted by confidence interval scores, and the set of end products and intermediate products contained therein are extracted and output to generate visualization results and structured data. The visualization results include at least one of network map, pathway map, and heat map, and the structured data includes at least one of CSV format file, JSON format file, and HTML table. The visualization results and structured data are used as a visualization model of the degradation pathways and products of the target sugars.