Glycomics agent construction method and system based on three-level coding, and storage medium

By using a three-level coding mapping network and Bayesian optimization, multi-source glycan data is integrated for cross-modal fusion and quantification, generating experimental parameters for verification. This solves the problem of multi-source data integration and autonomous decision-making in glycomics, and improves the accuracy and efficiency of glycan structure analysis.

CN122436008APending Publication Date: 2026-07-21NANJING SUPERYEARS GENE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610899078.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate multi-source heterogeneous glycomic data and cannot achieve autonomous decision-making verification experiments in complex and noisy environments, resulting in low accuracy and efficiency in glycan structure analysis.

Method used

A three-level coding mapping network maps glycan sequences, molecular graph topology, and three-dimensional spatial conformation to the same vector space, generating multimodal semantic representation vectors. This is combined with multi-source experimental spectral data for cross-modal fusion and uncertainty quantification. Bayesian optimization is used to generate experimental parameters, and the knowledge graph is verified and updated through an automated experimental platform.

Benefits of technology

It achieves high-precision glycan structure analysis in complex noise environments, provides autonomous experimental decision-making with quantifiable reliability, and improves the automation level and knowledge discovery rate of glycomics agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122436008A_ABST
    Figure CN122436008A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of multi-agent, and particularly relates to a glycomics agent construction method and system based on three-level coding and a storage medium, which maps sugar chain sequences, molecular graph topologies and three-dimensional space conformations to the same vector space through a three-level coding mapping network to generate multi-modal semantic representation vectors; inputs the representation vectors and multi-source spectral graph data into a multi-modal agent to output predicted sugar chain structures with structural unit confidence through cross-modal fusion and uncertainty quantification; generates research hypotheses automatically based on the prediction results and the confidence on a glycomics knowledge graph through k-hop random walk; generates recommended experimental parameters in an experimental parameter space through Bayesian optimization; issues the recommended experimental parameters to an automated experimental platform and receives results, evaluates deviation through KL divergence calculation, automatically updates knowledge graph weights and fine-tunes the multi-modal agent; and the application breaks the whole unmanned link from original spectral graphs to new knowledge discovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-agent technology, and particularly relates to a method, system and storage medium for constructing glycomics agents based on three-level coding. Background Technology

[0002] With the deep integration of artificial intelligence and life sciences, glycomics, as a cutting-edge field in the post-genomic era, is facing a profound transformation in its data-driven research paradigm. As the third life chain after nucleic acids and proteins, the structural analysis and functional study of glycans are crucial for understanding disease mechanisms and discovering carbohydrate-based drugs. However, unlike genomics and proteomics, which have widely benefited from deep learning technologies, the intelligentization process in glycomics has long lagged behind. The root cause lies in the fact that the structural complexity of glycans far exceeds that of nucleic acids and proteins. Specifically, the multidimensional heterogeneity of monosaccharide composition, linkage patterns, branching sites, and anodic conformations makes it impossible for single-dimensional sequence representations to fully capture the structural information of glycans. Simultaneously, the diverse and noisy data generated by various experimental methods such as mass spectrometry (MS), nuclear magnetic resonance (NMR), and liquid chromatography present a core bottleneck to the intelligentization of glycomics. These challenges have severely hindered the transformation of glycomics from a human experience-driven research paradigm to a data- and algorithm-driven one.

[0003] In existing technologies, various computational methods have been proposed for glycan structure analysis. For example, Chinese patent CN114166925B discloses a Denovo method and system for N-glycan structure identification based on mass spectrometry data. This method extracts the basic peaks and cross-peaks of fragment ions from mass spectrometry data and combines them with a generalized monosaccharide dictionary and pruning strategies to identify N-glycan structures. Essentially, this method is a traditional rule- and dictionary-driven identification technique that infers structure solely from single-dimensional mass spectrometry data, failing to integrate complementary information from multiple sources such as NMR and chromatography. Its rule-driven identification strategy exhibits significant limitations in generalization and robustness when facing noise interference in real experimental environments. Furthermore, this method only outputs a single identification result and lacks a quantitative evaluation mechanism for the reliability of the result. More importantly, this method serves only as an isolated structure identification tool, lacking the ability for autonomous experimental decision-making and knowledge evolution, and thus failing to constitute a closed-loop glycomics intelligent agent.

[0004] On the other hand, to improve the automation level of life science experiment design, existing technologies have also proposed experimental scheme planning methods based on knowledge graphs. For example, Chinese patent CN121278119B discloses a crop gene function research scheme planning method based on experimental reasoning chain. It constructs a knowledge graph by extracting research hypothesis-experimental operation-result interpretation triplets from historical literature, and generates multi-round experimental scheme sets through multi-context retrieval and semantic vector matching. This scheme is essentially a recommendation and ranking system based on existing literature knowledge. Its reasoning mechanism is limited to the retrieval and combination of known experimental schemes and does not have the ability to autonomously generate novel scientific hypotheses from data. The generation of its experimental schemes mainly relies on multi-dimensional scoring of historical experience and does not introduce mathematical optimization methods to quantitatively search and optimize the experimental parameter space. This method only stays at the level of paper scheme planning and does not include data interaction with physical experimental equipment or a closed-loop mechanism for automatic knowledge evolution from experimental results. Therefore, it also cannot provide a glycomics intelligent agent with autonomous closed-loop capabilities.

[0005] Therefore, how to construct a glycomics agent capable of autonomous decision-making verification experiments in complex noisy environments to improve the accuracy and efficiency of glycan structure analysis has become an urgent technical problem to be solved. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention proposes a method, system, and storage medium for constructing a glycomics agent based on three-level coding. This method maps glycan sequences, molecular graph topology, and three-dimensional spatial conformations to the same vector space through a three-level coding mapping network, generating multimodal semantic representation vectors. These representation vectors, along with multi-source spectral data, are input into the multimodal agent. Through cross-modal fusion and uncertainty quantification, a predicted glycan structure with structural unit confidence is output. Based on the prediction results and confidence levels, a k-hop random walk is performed on the glycomics knowledge graph to automatically generate research hypotheses. Bayesian optimization is used to generate recommended experimental parameters in the experimental parameter space. The results are then sent to an automated experimental platform, and the bias is evaluated by calculating KL divergence. The knowledge graph weights are automatically updated, and the multimodal agent is fine-tuned. This invention establishes a fully unmanned link from the original spectral graph to new knowledge discovery.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A method for constructing glycomics-based intelligent agents based on three-level coding includes:

[0009] Obtain a multimodal semantic representation vector of the target glycan, wherein the multimodal semantic representation vector integrates the sequence, graph topology and spatial conformation information of the target glycan;

[0010] Obtain multi-source experimental spectral data of the target glycan;

[0011] The trained multimodal agent is used to perform cross-modal fusion and uncertainty quantification on the multimodal semantic representation vector and the multi-source experimental spectrogram data to obtain the predicted structure of the target glycan and the confidence level of each structural unit in the predicted structure.

[0012] Based on the predicted structure and the confidence level, at least one research hypothesis is generated on the glycomics knowledge graph. The research hypothesis includes glycan functional associated entities to be experimentally verified and an optimizable experimental parameter space.

[0013] Within the optimizable experimental parameter space, a recommended set of experimental parameters is generated using Bayesian optimization with the experimental verification index corresponding to the predicted structure as the objective function.

[0014] The recommended set of experimental parameters is sent to the automated experimental platform, and the experimental results returned by the automated experimental platform after the experiment is performed are received.

[0015] Based on the deviation between the experimental results and the predicted structure, the relation weights in the glycomics knowledge graph are updated, and the deviation is used to fine-tune the multimodal agent.

[0016] Specifically, the multimodal semantic representation vector is generated by mapping sequence encoding, molecular graph topology and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network;

[0017] The generation of research hypotheses on the glycomics knowledge graph includes performing k-hop random walks; the glycomics knowledge graph is centered on glycan entities, and associates enzyme entities, protein entities, disease entities, and drug entities, and uses catalytic relationships, binding relationships, and upregulation relationships as the types of relationships between entities;

[0018] The method of generating a recommended set of experimental parameters using Bayesian optimization includes: within the optimizable experimental parameter space, using the experimental verification index as the objective function, performing Bayesian optimization iterative search using Gaussian process regression;

[0019] The method of fine-tuning the multimodal agent using the bias includes: parsing the experimental results into experimental verification index values, calculating the KL divergence between the experimental verification index values ​​and the predicted values ​​of the objective function, and when the KL divergence is greater than a preset threshold, updating the relation weights using the bias and fine-tuning the multimodal agent.

[0020] Specifically, a three-level coding mapping network maps sequence coding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space, including:

[0021] Obtain the sequence encoding data of the target glycan chain, wherein the sequence encoding data is the monosaccharide arrangement order and modification site information represented by a linear string in GlycoCT format;

[0022] Obtain the molecular graph topology data of the target glycan chain. The molecular graph topology data is an undirected connected graph with monosaccharide residues as nodes, glycosidic bonds as edges, and the connection sites and anodic configurations recorded in the edge attributes.

[0023] Obtain the three-dimensional spatial conformation data of the target glycan, wherein the three-dimensional spatial conformation data is the set of three-dimensional Cartesian coordinates of all heavy atoms in the target glycan.

[0024] Specifically, the three-level coding mapping network maps sequence coding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space, and also includes:

[0025] The GlycoCT format linear string is input into the sequence encoder, and a fixed-dimensional hidden state vector is output for each monosaccharide unit in the linear string. The hidden state vectors of all monosaccharide units are then averaged to obtain the sequence feature vector.

[0026] The undirected connected graph is input into a graph encoder. After multiple rounds of neighbor node message aggregation for each node in the undirected connected graph, the final hidden state vector of all nodes is obtained by average pooling to obtain the graph feature vector.

[0027] The three-dimensional Cartesian coordinate set is input into the spatial encoder. Under the constraint of ensuring rotational and translational equivariance, the coordinates of each atom and the relative positions of its neighboring atoms are encoded. The equivariant feature vectors of all atoms are then averaged and pooled to obtain the spatial configuration feature vector.

[0028] Specifically, the three-level coding mapping network maps sequence coding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space, and also includes:

[0029] Using three learnable projection matrices corresponding to sequence mode, graph topology mode and spatial configuration mode respectively, matrix multiplication is performed on the sequence feature vector, the graph feature vector and the spatial configuration feature vector respectively, mapping each mode feature vector to a shared semantic space of preset dimension d, and outputting three sets of intermediate vectors of dimension d, where dimension d is 512;

[0030] In a training batch containing multiple target glycan samples, any two of the three sets of intermediate vectors for the same target glycan are taken as positive sample pairs, and the corresponding intermediate vectors for different target glycans are taken as negative sample pairs. The InfoNCE extended loss function based on temperature scaling cosine similarity is used to calculate the contrast loss values ​​between the sequence intermediate vector and the graph intermediate vector, the graph intermediate vector and the spatial conformation intermediate vector, and the spatial conformation intermediate vector and the sequence intermediate vector, respectively. The average of the three sets of contrast loss values ​​is taken as the cross-modal contrast loss function value.

[0031] Specifically, the three-level coding mapping network maps sequence coding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space, and also includes:

[0032] With the goal of minimizing the cross-modal contrast loss function, the network weight parameters of the sequence encoder, the graph encoder, the spatial encoder, and the three projection matrices are updated synchronously through the backpropagation algorithm. The training continues until the cross-modal contrast loss function converges to a preset threshold range, so that the cosine similarity between any two pairs of intermediate vectors of the same target glycan in the vector space of dimension d reaches a preset high similarity condition, and the cosine similarity between intermediate vectors of different target glycans satisfies a preset low similarity condition.

[0033] Specifically, the three-level coding mapping network maps sequence coding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space, and also includes:

[0034] The three sets of intermediate vectors that meet the preset high similarity condition are input into the gated adaptive fusion layer. The gated adaptive fusion layer assigns an independent learnable scalar parameter to each set of input intermediate vectors. Softmax normalization is applied to the three learnable scalar parameters to generate a three-dimensional attention weight vector. After performing element-wise multiplication of each weight component in the three-dimensional attention weight vector with the corresponding intermediate vector, element-wise addition is performed on the three weighted vectors to generate a multimodal semantic representation vector of dimension d.

[0035] Specifically, the trained multimodal agents include cross-modal fusion sub-agents, structure prediction sub-agents, uncertainty quantification sub-agents, and structure judgment output sub-agents;

[0036] The cross-modal fusion sub-agent is configured to: receive the multimodal semantic representation vector and the multi-source experimental spectral data; map the mass spectrometry data, nuclear magnetic resonance spectral data and chromatographic parameter data in the multi-source experimental spectral data into mass spectrometry feature vectors, nuclear magnetic resonance feature vectors and chromatographic feature vectors respectively through their respective independent spectral feature extractors; use the multimodal semantic representation vector as a query vector; concatenate the mass spectrometry feature vector, the nuclear magnetic resonance feature vector and the chromatographic feature vector as a key vector and a value vector; perform cross-modal interaction through a multi-head cross-attention network; and output the glycan hidden state sequence after fusing spectral information, wherein the length of the glycan hidden state sequence is equal to the total number of monosaccharide residues of the target glycan.

[0037] Specifically, the structure prediction sub-agent is configured to: receive the glycan hidden state sequence output by the cross-modal fusion sub-agent; input the glycan hidden state sequence into an autoregressive sequence generation model based on a Transformer decoder; predict the monosaccharide type, glycosidic bond connection site, and anodic configuration of each monosaccharide residue of the target glycan step by step from the reduced end to the non-reduced end; and output a predicted glycan structure sequence of length L, where L is equal to the total number of monosaccharide residues of the target glycan. The output of the predicted glycan structure sequence at time step t is a three-dimensional probability vector, the first dimension of which is the probability distribution of the predicted t-th monosaccharide residue of the target glycan on a preset set of monosaccharide types; the second dimension of which is the probability distribution of the combination of glycosidic bond connection sites between the predicted t-th monosaccharide residue of the target glycan and the previous monosaccharide residue on a preset set of connection site combinations; and the third dimension of which is the probability distribution of the predicted anodic configuration of the predicted t-th monosaccharide residue of the target glycan on a preset set of anodic configurations.

[0038] Specifically, the uncertainty quantification sub-agent is configured as follows: during the autoregressive sequence generation process performed by the structure prediction sub-agent, Monte Carlo dropout is applied to each feedforward sublayer in the autoregressive sequence generation model based on the Transformer decoder. K random forward propagations are performed during the inference phase, and each random forward propagation independently outputs a glycan structure sequence sample. For each candidate monosaccharide type, each candidate glycosidic bond combination, and each candidate anodic configuration in the K glycan structure sequence samples at the t-th time step, the frequency of occurrence of each candidate monosaccharide type, each candidate glycosidic bond combination, and each candidate anodic configuration is counted in the K sampling. The frequency of occurrence is divided by K to obtain the confidence probability value of each candidate category of the t-th structural unit. A confidence probability distribution sequence of length L is output. The t-th element in the confidence probability distribution sequence is a triple containing the confidence distribution of monosaccharide type, the confidence distribution of combination site, and the confidence distribution of anodic configuration.

[0039] Specifically, the structure judgment output sub-agent is configured to: receive the predicted glycan structure sequence output by the structure prediction sub-agent and the confidence probability distribution sequence output by the uncertainty quantification sub-agent; take the minimum value of the triples of the t-th structural unit in the confidence probability distribution sequence as the comprehensive confidence score of the t-th structural unit; when the comprehensive confidence scores of all L structural units are greater than or equal to a preset confidence threshold, output the predicted glycan structure sequence as the predicted structure of the target glycan, and output the confidence probability distribution sequence as the confidence score corresponding to each structural unit in the predicted structure; when the comprehensive confidence score of at least one structural unit is less than the preset confidence threshold, mark the structural unit as a low-confidence unit, add a low-confidence mark to the low-confidence unit in the output predicted structure, and output the complete confidence probability distribution sequence.

[0040] A glycomics-based intelligent agent construction system based on three-level coding includes:

[0041] A multimodal coding module is used to obtain a multimodal semantic representation vector of the target glycan, wherein the multimodal semantic representation vector integrates the sequence, graph topology and spatial conformation information of the target glycan;

[0042] The spectral data acquisition module is used to acquire multi-source experimental spectral data of the target glycan;

[0043] The multimodal agent prediction module is used to perform cross-modal fusion and uncertainty quantification on the multimodal semantic representation vector and the multi-source experimental spectrogram data using a trained multimodal agent, so as to obtain the predicted structure of the target glycan and the confidence level of each structural unit in the predicted structure.

[0044] The hypothesis generation module is used to generate at least one research hypothesis on the glycomics knowledge graph based on the predicted structure and the confidence level. The research hypothesis includes glycan functional associated entities to be experimentally verified and an optimizable experimental parameter space.

[0045] The experimental design module is used to generate a recommended set of experimental parameters using Bayesian optimization within the optimizable experimental parameter space, with the experimental verification index corresponding to the predicted structure as the objective function.

[0046] The experiment execution module is used to send the recommended set of experimental parameters to the automated experiment platform and receive the experimental results returned by the automated experiment platform after the experiment is executed.

[0047] An iterative module is used to update the relation weights in the glycomics knowledge graph based on the deviation between the experimental results and the predicted structure, and to fine-tune the multimodal agent using the deviation.

[0048] A computer-readable storage medium storing computer instructions that, when executed, perform a method for constructing a glycomics-based intelligent agent based on three-level coding.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] The glycomics agent constructed in this invention unifies glycan sequences, graph topology, and spatial conformation into a multimodal semantic representation through a three-level coding mapping network. It then utilizes this representation to perform cross-modal fusion and uncertainty quantification on multi-source experimental spectra. While outputting high-precision predicted structures, it also provides confidence probability distributions for each structural unit, making the reliability of structural analysis quantifiable for the first time and providing an operable basis for autonomous experimental decision-making. Based on the predicted structures and confidence levels, the agent performs k-hop random walks on the glycomics knowledge graph to automatically generate research hypotheses containing functionally related entities and experimental parameter spaces, using experimental verification indicators as the objective function. By leveraging Bayesian optimization to efficiently search for optimal experimental conditions in a continuous parameter space, the agent autonomously generates a recommended set of experimental parameters. Through standard protocol integration with an automated experimental platform, the agent distributes experimental parameters, receives experimental results, and automatically evaluates experimental biases based on KL divergence, triggering knowledge graph relation weight updates and model fine-tuning. This enables the autonomous completion of unmanned closed-loop iterations of structure analysis, hypothesis generation, experimental optimization, and knowledge updating. The glycomics agent in this application significantly improves the accuracy of glycan structure analysis and the decision-making efficiency of verification experiments under complex noise environments, greatly enhancing the automation level and output rate of glycomics from spectrograms to new knowledge discovery. Attached Figure Description

[0051] Figure 1 This is a flowchart of the glycomics-based intelligent agent construction method according to Embodiment 1 of the present invention;

[0052] Figure 2 This is a flowchart of Embodiment 2 of the present invention, which maps the sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level encoding mapping network.

[0053] Figure 3 This is a block diagram of the glycomics intelligent agent construction system based on three-level coding in Embodiment 3 of the present invention. Detailed Implementation

[0054] Example 1

[0055] Please see Figure 1 One embodiment of the present invention provides a method for constructing a glycomics-based intelligent agent based on three-level coding, comprising:

[0056] S1. Obtain the sequence encoding, molecular graph topology, and three-dimensional spatial conformation data of the target glycan chain, and map the sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network to generate the multimodal semantic representation vector of the target glycan chain.

[0057] S2. Obtain multi-source experimental spectral data for the target glycan, wherein the multi-source experimental spectral data includes at least two of the following: mass spectrometry data, nuclear magnetic resonance spectral data, and chromatographic parameter data; input the multimodal semantic representation vector and the multi-source experimental spectral data into a pre-trained multimodal agent, and have the multimodal agent perform cross-modal fusion and uncertainty quantification to output the predicted glycan structure of the target glycan and the confidence probability distribution of each structural unit in the predicted glycan structure;

[0058] S3. Based on the predicted glycan structure and the confidence probability distribution, perform a k-hop random walk on the glycomics knowledge graph to generate at least one research hypothesis. The research hypothesis includes glycan functional entities to be experimentally verified and an optimizable experimental parameter space. The glycomics knowledge graph is centered on glycan entities and associates enzyme entities, protein entities, disease entities, and drug entities, with catalytic relationships, binding relationships, and upregulation relationships as the types of relationships between entities.

[0059] S4. Within the optimizable experimental parameter space, using the experimental verification index corresponding to the predicted glycan structure as the objective function, Bayesian optimization iterative search is performed using Gaussian process regression to generate a recommended set of experimental parameters.

[0060] S5. Send the recommended experimental parameter set to the automated experimental platform through a standardized communication protocol, and receive the original experimental result data returned by the automated experimental platform after the experiment is performed.

[0061] S6. Parse the original experimental results data into experimental verification index values, calculate the KL divergence between the experimental verification index values ​​and the predicted values ​​of the objective function; when the KL divergence is greater than a preset threshold, update the relation weights between entities in the glycomics knowledge graph using the deviation between the original experimental results data and the predicted glycan structure, and use the deviation as a supervision signal to fine-tune the multimodal agent; based on the updated glycomics knowledge graph and the fine-tuned multimodal agent, regenerate the research hypothesis to start the next round of closed loop.

[0062] Example 2

[0063] Further explanation is needed; please refer to [link / reference]. Figure 2 This embodiment maps the sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network, including:

[0064] S101. Obtain the sequence encoding data of the target glycan chain. The sequence encoding data is the monosaccharide arrangement order and modification site information represented by a linear string in GlycoCT format. In this embodiment, the GlycoCT format linear string strictly follows the GlycoCT naming convention established by the International Glycomics Union. Its structure consists of three core components: resource node RES, connection node LIN, and modification node REP. The RES node precisely defines each monosaccharide residue in the format of number + monosaccharide abbreviation + isomer type + ring type + absolute configuration. For example, 1b:x-dglc-HEX-1:5 represents the numbering. The residue with a value of 1 is at the reducing end, representing D-configuration glucose, existing as a pyranose ring, with carbon atom number 1 forming a ring with oxygen atom number 5. The LIN node defines glycosidic bond connections in the format of donor residue number:donor site / connection type / receptor residue number:receptor site. For example, LIN1:1d / 2+ / 2:1o indicates that carbon 1 of residue number 1 and carbon 1 of residue number 2 are connected by a 1,2 glycosidic bond, where the letter d represents the donor end and the letter o represents the recipient end. The REP node is used to label post-translational modification groups such as sulfation, phosphorylation, acetylation, and sialylation, and their specific connection sites on the corresponding monosaccharide residues. The specific path to obtain the sequence-encoded data includes: for known glycans already registered in the GlyTouCan international glycan structure database, the corresponding GlycoCT archive record is directly retrieved through the RESTful application interface provided by the database using the glycan registration number as the query key; for unknown glycans newly discovered in mass spectrometry experiments, the upstream glycan fragment ion automated analysis module generates a candidate structure set based on the fragmentation rules of tandem mass spectrometry. Each GlycoCT sequence in this candidate structure set is used as a candidate structure to be verified and input into this method for multimodal verification and confidence evaluation, rather than being used directly as a known truth value, thus avoiding circular dependencies. Before being input into the sequence encoder, the obtained GlycoCT format linear string needs to undergo lexical preprocessing. Specifically, the RES identifier, LIN identifier, REP identifier, monosaccharide abbreviation characters, numeric site values, and connection type symbols in the string are mapped to corresponding integer index numbers. The connection relationship information of the LIN node is fully encoded into a connection relationship term, which records the position index, carbon atom site, and anodic configuration of the two connected monosaccharide residues. Together with the residue term of the RES node, it forms a term sequence containing complete branch topology information, ultimately forming an integer sequence with a length not exceeding 256 terms, which serves as the standardized input for the sequence encoder. When the term sequence exceeds 256, it is truncated at the end and a truncation mark is added. Samples with incomplete branch information due to truncation are marked as abnormal samples and do not participate in training.In this embodiment, the upstream glycan fragment ion automated analysis module refers to an independent functional module that preprocesses the raw mass spectrometry data and generates candidate glycan structures before the structure analysis is performed in this method. In this embodiment, the workflow of the upstream glycan fragment ion automated analysis module is as follows: First, it receives the mass-to-charge ratio and relative abundance data of fragment ions generated by tandem mass spectrometry in collision-induced dissociation mode. Based on the inherent law of glycan mass spectrometry fragmentation—glycosidic bonds preferentially break during gas-phase fragmentation and generate characteristic B, Y, C, and Z ion series—it calculates the mass difference between adjacent fragment ions and matches them with a preset monosaccharide residue mass database to infer the type of each monosaccharide residue in the glycan sequence one by one. At the same time, it uses the diagnostic information of the linkage sites provided by the A and X ions generated by cross-loop fragmentation, combined with the specific constraint rules of glycosyltransferases in the glycan biosynthesis pathway, to generate five to ten candidate GlycoCT sequences containing complete monosaccharide sequences, glycosidic bond linkage sites, and branching structures. Each GlycoCT sequence in the candidate structure set output by this module is labeled with an initial score based on fragment ion matching degree. All of them are used as candidate structures to be verified and input into this method for subsequent multimodal verification and confidence evaluation, rather than being used directly as known truth values. Therefore, there is no circular dependency problem.

[0065] S102. Obtain the molecular graph topology data of the target sugar chain. The molecular graph topology data is an undirected connected graph with monosaccharide residues as nodes, glycosidic bonds as edges, and the connection sites and anodic configurations recorded in the edge attributes.

[0066] In this embodiment, the node's attribute information consists of a multidimensional one-heat encoded vector composed of three parts: monosaccharide type identifier, ring configuration identifier, and absolute configuration identifier. The monosaccharide type identifier covers 20 common monosaccharide types in glycomics research, including glucose, galactose, mannose, fucose, N-acetylglucosamine, N-acetylglucosamine, and sialic acid. The ring configuration identifier distinguishes between pyranose and furanose rings, and the absolute configuration identifier distinguishes between D and L configurations. The edge's attribute information consists of a multidimensional one-heat encoded vector composed of three parts: donor connection site number, acceptor connection site number, and anodic configuration identifier. The donor and acceptor connection site numbers are respectively assigned to the carbon atom numbers of the corresponding monosaccharide residues, and the anodic configuration identifier distinguishes between α and β configurations of glycosidic bonds. The specific construction process of this undirected connected graph is based on the parsing results of the previously generated GlycoCT format linear string: traversing each RES node in the GlycoCT format linear string one by one, creating a corresponding graph node for each monosaccharide residue and assigning it a corresponding node attribute vector; traversing each LIN node in the GlycoCT format linear string one by one, parsing out the donor residue number, donor connection site, acceptor residue number, acceptor connection site, and glycosidic bond type, establishing an undirected edge between the corresponding two graph nodes and assigning it a corresponding edge attribute vector; only when there is a clear glycosidic bond connection record between two monosaccharide residues in the GlycoCT format linear string is a corresponding edge created in the graph structure, ensuring that the graph topology faithfully reflects the true chemical connection relationship of the target sugar chain. The reason for using an undirected connected graph instead of a directed graph to represent the topology of sugar molecules is that glycosidic bonds, as covalent chemical bonds, inherently allow bidirectional information transfer between two connected residues. Message-passing neural networks naturally assume that the edges of the graph are bidirectional when performing neighbor message aggregation on nodes. Using an undirected graph is inherently consistent with the computation mechanism of message-passing neural networks, avoiding the loss of information flow paths caused by improper setting of directed edge directions.

[0067] S103. Obtain the three-dimensional spatial conformation data of the target sugar chain, wherein the three-dimensional spatial conformation data is the set of three-dimensional Cartesian coordinates of all heavy atoms in the target sugar chain;

[0068] In this embodiment, heavy atoms specifically refer to atoms with an atomic number greater than 1, that is, excluding all hydrogen atoms and retaining only non-hydrogen atoms such as carbon, nitrogen, oxygen, sulfur, and phosphorus atoms that constitute the sugar backbone and side chains. The acquisition path of this three-dimensional Cartesian coordinate set is based on the molecular diagram topology data generated in S102 and the GlycoCT format linear string generated in S101, and the specific process is as follows:

[0069] The first step involves using the monosaccharide arrangement and glycosidic bond connections recorded in the GlycoCT format linear string as structural constraints, and the monosaccharide types and ring configurations contained in the node attributes and the connection sites and anodic configurations contained in the edge attributes of the molecular graph topology data as stereochemical constraints. The distance geometry algorithm in the RDKit open-source cheminformatics toolkit is then used to generate the initial three-dimensional atomic coordinates of the target glycan chain. This distance geometry algorithm first generates lower and upper distance constraints for each pair of atoms based on the structural and stereochemical constraints. The lower distance constraint is the larger of the sum of the covalent bond length and the van der Waals radius of the corresponding atom pair, and the upper distance constraint is the covalent bond length plus a preset flexibility margin. Then, initial guesses of atomic coordinates satisfying all distance constraints are randomly generated in high-dimensional space. The high-dimensional coordinates are then iteratively reduced to three-dimensional space using a successive projection algorithm. After each dimensionality reduction, the coordinates of atom pairs violating the distance constraints are checked and corrected until the distances between all atom pairs satisfy the lower and upper distance constraints, thus constructing a chemically reasonable initial conformation. In this initial conformation, each heavy atom is assigned a three-dimensional Cartesian coordinate value (x, y, z), with the coordinate unit being angstroms.

[0070] The second step involves performing all-atom molecular dynamics simulations on the initial conformation under the combined influence of the AMBER14SB force field and the GB-Neck2 implicit generalized Born solvent model. The AMBER14SB force field is an optimized all-atom force field for protein and glycan systems, with its potential function including bond stretching, bond bending, dihedral torsion, van der Waals interaction, and electrostatic interaction terms. The force field parameters are taken from the AMBER14SB standard parameter set. The GB-Neck2 implicit generalized Born solvent model uses an iGB=8 parameter setting to implicitly simulate the average effect of the aqueous solvent environment on solute molecules, thus significantly reducing computational costs while maintaining simulation accuracy. The simulation was performed at a constant temperature of 300 Kelvin, controlled by a Langevin thermostat, with the collision frequency set to the reciprocal of 1.0 picosecond. The total simulation duration was 200 nanoseconds, with an integration time step of 2 femtoseconds. Covalent bonds involving hydrogen atoms were constrained using the SHAKE algorithm, freezing the vibrational degrees of freedom containing hydrogen bonds and allowing for a larger integration step. During the simulation, the instantaneous three-dimensional Cartesian coordinates of all heavy atoms in the system were recorded every 10 picoseconds, for a total of 20,000 conformational snapshots. Each conformational snapshot was stored in the trajectory file as an M-row by 3-column floating-point matrix, where M is the total number of heavy atoms in the target sugar chain, and the 3 columns correspond to the X-axis coordinate value, Y-axis coordinate value, and Z-axis coordinate value, respectively.

[0071] The third step involves extracting one frame every 500 frames from the equilibrium trajectory of the last 50 nanoseconds of the 20,000 conformational snapshots, for a total of 100 frames, to form an initial conformational ensemble. Using the first frame of this initial conformational ensemble as a reference conformation, the remaining 99 frames are then rotated, translated, and superimposed onto the reference conformation one by one. The specific operation of the overlay process is as follows: For the conformational snapshot to be overlaid, extract the three-dimensional Cartesian coordinates of all heavy atoms in the reference conformation. Calculate the optimal rotation matrix and optimal translation vector using the Kabsch algorithm. The optimal rotation matrix is ​​a 3x3 orthogonal matrix, and the optimal translation vector is a 3-dimensional vector. This optimal rotation matrix and optimal translation vector minimize the root mean square deviation between the coordinates of all heavy atoms in the snapshot to be overlaid after rotation and translation transformation and the coordinates of the corresponding heavy atoms in the reference conformation. Apply the calculated optimal rotation matrix and optimal translation vector to the coordinates of all heavy atoms in the snapshot to be overlaid to complete the overlay alignment. This overlay process eliminates the coordinate differences between each frame of conformational snapshots caused by the overall rotation and translation of the molecule, ensuring that the coordinate differences between each frame only reflect the changes in the degrees of freedom of the internal conformation of the sugar chain.

[0072] The fourth step involves taking 100 snapshots of the superimposed and aligned conformations, and extracting the probability of each conformation appearing in the trajectory of the last 50 nanoseconds of the molecular dynamics simulation. This probability is calculated based on the Boltzmann distribution of the conformation in the canonical ensemble, i.e., the weight w_i of the i-th conformation is equal to exp(-E_i / (k_B)). T)) divided by the total 100 frames of the image exp(-E_j / (k_B) The sum of T) is given, where E_i is the potential energy of the i-th frame conformation under the AMBER14SB force field, k_B is the Boltzmann constant, and T is the simulated temperature of 300 Kelvin. This yields the Boltzmann weights for each of the 100 frames of conformations, satisfying the normalization constraint that the sum of all weights equals 1.

[0073] The fifth step involves combining the 100 overlaid and aligned conformational snapshots and their corresponding Boltzmann weights to form a conformational ensemble. This conformational ensemble is stored in structured data format and consists of two parts: the first part is a 100 x M x 3 three-dimensional floating-point array, where the first dimension corresponds to the conformational frame number (1 to 100), the second dimension corresponds to the heavy atom number (1 to M), and the third dimension corresponds to the three-dimensional Cartesian coordinate components (X, Y, Z); the second part is a 100-dimensional floating-point array that stores the Boltzmann weights corresponding to each conformational frame. The complete data of the conformational ensemble is used as the input of the spatial encoder. The spatial encoder encodes each frame of the conformational ensemble using an isovariant graph neural network. Then, the Boltzmann weights of each frame of the conformation are used as weighting coefficients to perform weighted average aggregation on the isovariant feature vectors output by each frame of the encoding, resulting in a unified spatial conformational encoding vector. This method fully considers the conformational flexibility and conformational diversity of sugar chains under physiological conditions, and avoids the loss of structural information caused by representing the entire conformational ensemble with a single conformation.

[0074] S104. The GlycoCT format linear string is input into the sequence encoder, which is a Transformer-based encoder. The sequence encoder outputs a fixed-dimensional hidden state vector for each monosaccharide unit in the linear string. The hidden state vectors of all monosaccharide units are averaged and pooled to obtain the sequence feature vector. In this embodiment, the sequence encoder strictly inherits the GlycoCT format linear string and its lexicalization preprocessing result generated in the first step. The lexicalization preprocessing maps the monosaccharide identifier, connection symbol, and modification label corresponding to each RES node in the GlycoCT format linear string to independent integer indices. The connection relationship of the LIN node is fully encoded into connection relationship lexicals, forming an integer sequence containing branch topology information and with a length not exceeding 256. Each monosaccharide residue occupies a time step position in the sequence. This integer sequence is the direct input of the sequence encoder. The sequence encoder consists of six stacked Transformer encoder layers. Each layer contains an 8-head self-attention sublayer and a fully connected feedforward sublayer. The self-attention sublayer has 8 attention heads, and the feedforward sublayer has a hidden layer width of 2048. Each sublayer is followed by residual connections and layer normalization. The hidden state vector dimension of the entire encoder is fixed at 512. The input integer sequence is first mapped to a 512-dimensional word vector through a word embedding layer, and then added element-wise with a sinusoidal positional encoding to preserve the positional order information in the sequence. After feature extraction and context interaction through the six encoder layers, a 512-dimensional hidden state vector is output at each monosaccharide unit time step. Arithmetic average pooling is performed along the sequence length dimension on the hidden state vectors corresponding to all monosaccharide unit time steps to obtain a 512-dimensional sequence feature vector. This vector compresses the monosaccharide arrangement order and local contextual dependencies of the entire sugar chain sequence. During the pre-training phase, the sequence encoder employs a masked language modeling task for self-supervised training. 15% of the tokens in the input sequence are randomly masked. The encoder is then trained to reconstruct the original integer indices of the masked tokens based on the unmasked context tokens, enabling it to deeply capture the co-occurrence patterns and syntactic structures between monosaccharides in the sugar sequence. In this embodiment, for sugar sequences with a length exceeding 256 tokens after lexicalization, a tail-truncation strategy is used. The specific rules are as follows: the first 256 tokens are retained, and all tokens from the 257th token onwards are discarded. After truncation, a special truncation marker is appended to the end of the sequence, with its index value preset to the maximum index value of the current vocabulary + 1. For sugar sequences with fewer than 256 tokens, if their length is less than 256, padding is applied to the end of the sequence with a padding character whose index value is fixed at 0, until the total sequence length is exactly 256 tokens.When receiving the padded integer sequence, the sequence encoder simultaneously generates an attention mask vector of length 256. In this attention mask vector, the positions corresponding to the actual words and truncation marker words are set to 1, while the positions corresponding to the padding characters are set to 0. The self-attention mechanism ignores the positions with mask values ​​of 0 during calculation, ensuring that the padding characters do not participate in feature encoding. For sugar sequences where branch connection information is lost due to tail truncation, the system marks them as long-chain truncated samples. These samples do not participate in the construction of negative samples for cross-modal contrastive learning; they are only used for structure prediction and confidence evaluation through independent forward propagation. A long-chain truncated marker is appended to the output prediction result, indicating that the non-reduced end structural information of the sugar chain has potential uncertainties due to truncation, providing data quality guidance for subsequent hypothesis generation and experimental verification.

[0075] S105. Input the undirected connected graph into a graph encoder. The graph encoder is a message passing neural network. After performing multiple rounds of neighbor node message aggregation on each node in the undirected connected graph, the final hidden state vector of all nodes is averaged to obtain the graph feature vector.

[0076] In this embodiment, the graph encoder adopts a 5-layer graph isomorphic network (Edge-GIN) architecture with edge feature enhancement to encode a glycan molecule graph of arbitrary node size into a graph representation vector of fixed dimension, as shown in the following structure:

[0077] The initial feature vector of each monosaccharide residue node is a 32-dimensional one-hot encoded vector containing information on the monosaccharide type, ring configuration, and absolute configuration. This vector is mapped to a 64-dimensional dense embedding vector through a learnable linear transformation matrix (32 rows × 64 columns), serving as the initial hidden state vector for that node. Each element in this linear transformation matrix is ​​a trainable parameter, initialized using Xavier uniform initialization.

[0078] The graph encoder internally performs five rounds of neighbor node message aggregation. Specifically, in each round of message passing, for each node in the graph, it first determines the set of all its first-order neighbor nodes. For each first-order neighbor node, the hidden state vector of that neighbor node in the current round and the edge attribute vector of the edge pointing from that neighbor node to the current node are used as the input to the message function. The edge attribute vector contains three parts: donor connection site number, recipient connection site number, and anodic configuration identifier. Before being input into the message function, the edge attribute vector is first transformed into a 64-dimensional edge embedding vector through a learnable edge embedding layer. The weight matrix of this edge embedding layer has a size of 3×64. The message function is implemented as a multilayer perceptron with two hidden layers. The input of the first hidden layer is the vector obtained by concatenating the hidden state vector of the neighbor node in the current round and the 64-dimensional edge embedding vector along the dimensional direction. The width of the first hidden layer is 128, and the activation function is ReLU. The width of the second hidden layer is 128, and the activation function is ReLU. The output is a message vector with the same dimension as the hidden state vector of the current round. The aggregation function uses element-wise summation, adding the message vectors from all first-order neighbor nodes at their corresponding element positions to form a single aggregate vector. This aggregate vector is then concatenated with the node's own hidden state vector from the previous round along its dimension, resulting in a concatenated vector. This concatenated vector is fed into the update function, which is implemented as a multilayer perceptron with two hidden layers. The first hidden layer has a width twice the dimension of the current round's hidden state vector and uses ReLU activation. The second hidden layer has the same width as the current round's hidden state vector and also uses ReLU activation. The update function outputs the node's hidden state vector for the current round. The parameters of this edge embedding layer, the multilayer perceptron in the message function, and the multilayer perceptron in the update function are all learnable parameters throughout all five rounds of message passing in the graph encoder, and are optimized end-to-end through backpropagation.

[0079] The dimension change path of the node hidden state vector is as follows: In the first round of message passing, the output dimension of the message function is 128 dimensions, and the output dimension of the update function is 128 dimensions. Therefore, after the first round of message passing, the node hidden state vector expands to 128 dimensions. In the second round of message passing, the output dimension of the message function is 256 dimensions, and the output dimension of the update function is 256 dimensions. After the second round of message passing, the node hidden state vector expands to 256 dimensions. In the third to fifth rounds of message passing, the output dimensions of the message function and the update function remain unchanged at 256 dimensions, and the dimension of the node hidden state vector remains at 256 dimensions.

[0080] To eliminate the impact of variations in the number of nodes N in the input molecular graph on the output vector dimension, a global average pooling readout layer is added at the end of the graph encoder after 5 rounds of message passing. The specific operation of this readout layer is as follows: The final hidden state vectors of all N nodes after the 5th round of message passing are obtained, each vector having a dimension of 256. Arithmetic average pooling is then performed on these N 256-dimensional vectors along the node dimensions. Specifically, for each dimension index d (d=1,2,...,256), the arithmetic mean of the values ​​of all N nodes in that dimension is taken as the value of the d-th dimension of the output vector. This results in a fixed 256-dimensional graph feature vector that is strictly independent of the number of nodes N. This graph feature vector compresses the monosaccharide residue connection topology and branching structure information of the entire molecular graph, serving as a unified representation of the glycan graph modality, and can directly participate in subsequent cross-modal contrastive learning and multimodal fusion.

[0081] In the pre-training phase, the graph encoder employs a graph structure contrastive learning task for self-supervised training. For the molecular graph of the same glycan, 15% of the nodes are randomly selected, and their initial feature vectors are set to zero to achieve node feature masking. 10% of the edges are randomly selected and deleted to achieve edge perturbation, thus generating two not entirely identical augmented views. These two augmented views are then input into the graph encoder, and after five rounds of message passing and global average pooling, each generates a fixed 256-dimensional graph feature vector. The encoder is trained to maximize the cosine similarity between the graph feature vectors corresponding to the two augmented views of the same glycan, while simultaneously minimizing the cosine similarity with the graph feature vectors of other glycans within the batch, using them as negative samples. This self-supervised pre-training enables the graph encoder to capture the core topological features in the glycan molecular graph that are robust to local node and edge feature loss, providing a good parameter initialization for subsequent fine-tuning.

[0082] S106. The three-dimensional Cartesian coordinate set is input into a spatial encoder, which is an isovariant graph neural network. Under the constraint of ensuring rotational and translational isovariance, the coordinates of each atom and the relative positions of its neighboring atoms are encoded. The isovariant feature vectors of all atoms are averaged and pooled to obtain the spatial conformation feature vector. In this embodiment, the spatial encoder adopts an architecture based on SE3 isovariant Transformer with a network depth of 4 layers. It encodes the three-dimensional conformation of sugar chains with an arbitrary number of heavy atoms into a spatial conformation feature vector of fixed dimensions. The specific structure is as follows:

[0083] First, a unified atomic sorting rule is established to eliminate the influence of different sorting on the encoding results. Specifically, the spatial encoder receives the heavy atom coordinates of each frame in the conformational ensemble generated by molecular dynamics simulation. The atoms are sorted according to a unified chemical specification: first, they are arranged in ascending order by residue number, that is, according to the order of appearance of the RES node in the linear string of the target glycan GlycoCT format, from the first monosaccharide residue to the Lth monosaccharide residue; heavy atoms within the same residue are arranged in a fixed priority order of atom type, with priority from high to low as carbon, nitrogen, oxygen, sulfur, and phosphorus. For atomic positions that are present or missing due to differences in the type and number of monosaccharide residues in different glycans, a three-dimensional zero coordinate vector is used to fill the corresponding sorting position, and an atomic mask marker with a value of zero is generated for the filled position, while the atomic mask marker for the non-filled position has a value of one. After the above processing, the input data for each frame of conformation is a floating-point matrix of M rows multiplied by three columns and an atomic mask vector of length M, where the three columns correspond to the X-axis, Y-axis, and Z-axis coordinate values, respectively, with the coordinate unit being angstroms. After encoding each frame of the conformation ensemble separately, the spatial encoder uses the Boltzmann weights of each frame's conformation as weighting coefficients to perform weighted average aggregation on the equivariant feature vectors output by each frame's encoding, resulting in a unified spatial conformation encoding vector that fully considers the flexible features of the sugar chain and the diversity of conformations.

[0084] The core computational unit of each layer of the spatial encoder is the TensorField convolution. This TensorField convolution operation operates on all neighboring atoms within a spherical domain centered on each atom with a truncation radius of 5 angstroms. For the central atom, its neighboring atoms must simultaneously satisfy two conditions: the Euclidean distance between the central atom and its neighboring atom must be less than or equal to 5 angstroms, and the atom mask of the neighboring atom must be set to 1, meaning that the neighboring atom is a non-filled atom. The convolution kernel of the TensorField convolution is composed of a combination of learnable radial basis functions and spherical harmonic basis functions. The radial basis function maps the interatomic distance to a scalar weight. This function contains learnable weight parameters, center position parameters, and width parameters. It takes the interatomic distance as the input variable and outputs a scalar value that decays with distance. The spherical harmonic basis function expands the orientation angles (polar angle and azimuth angle) of the relative position vector into spherical harmonic basis function coefficients with a maximum order of three. The order can be zero, one, two, or three, and each order corresponds to a set of magnetic quantum number values, for a total of 16 spherical harmonic channels. This TensorField convolution can strictly follow the rotation and translation equivariance constraint during the calculation process. That is, when the input three-dimensional Cartesian coordinate set undergoes any spatial rotation or any spatial translation, the output equivariant feature vector will be transformed according to the same group representation rule, ensuring that the encoding result is completely independent of the initial orientation selection of the coordinate system.

[0085] The spatial encoder first maps the atom type of each heavy atom to a 16-dimensional initial scalar feature through an atom embedding layer. This atom embedding layer is a learnable embedding matrix with five rows corresponding to the five common glycan backbone elements: carbon, nitrogen, oxygen, sulfur, and phosphorus, and 16 columns. Each row stores a 16-dimensional embedding vector corresponding to the atom type. For filling positions where the atom mask is marked as zero, the initial scalar feature is set to a 16-dimensional all-zero vector.

[0086] A relative position vector is constructed by the difference between the 3D Cartesian coordinates of each atom and the 3D Cartesian coordinates of its neighboring atoms. This relative position vector is a 3D vector, with three components corresponding to the coordinate differences along the X, Y, and Z axes, respectively. The magnitude and orientation angle of this relative position vector are calculated. The magnitude is the square root of the sum of the squares of the three coordinate components. The orientation angle includes the polar angle and the azimuth angle. The polar angle is the angle between the relative position vector and the positive Z-axis, and the azimuth angle is the angle between the projection of the relative position vector onto the XY plane and the positive X-axis. The magnitude is input into the radial basis function to obtain scalar weights, and the polar angle and azimuth angle are input into the spherical harmonic basis function to obtain 16 spherical harmonic basis function coefficients. The scalar output of the radial basis function is multiplied element-wise with the 16 spherical harmonic basis function coefficients to obtain a 16-dimensional side feature vector, which is used as the side feature input for TensorField convolution.

[0087] After four layers of equivariant graph neural network encoding, specifically: In the k-th layer (k takes values ​​1, 2, 3, and 4), for each atom, its equivariant feature vector output from the layer below k (or the 16-dimensional initial scalar feature vector output from the atom embedding layer if k equals one) is input into a TensorField convolution along with the equivariant feature vectors output from all its neighboring atoms in the layer below k, and the 16-dimensional edge feature vectors between the atom and its neighbors. The output of the TensorField convolution is the equivariant feature vector of the atom in the k-th layer. This equivariant feature vector contains 16 spherical harmonic channels, each containing both scalar components and higher-order spherical harmonic tensor components. The scalar components capture information about the local chemical environment centered on the atom, including atom type, coordination number, and local geometry; the higher-order spherical harmonic tensor components capture spatial orientation information centered on the atom, including the spatial orientation of chemical bonds and the anisotropy of the spatial distribution of neighboring atoms. After four layers of encoding, each heavy atom outputs an equivariant feature vector containing 16 spherical harmonic channels.

[0088] To eliminate the impact of variations in the number of heavy atoms on the output dimension, a global average pooling readout layer is set at the end of the spatial encoder. Specifically, this readout layer first uses an atom mask to mask the equivariant feature vectors at the filling positions, that is, only the equivariant feature vectors of non-filled atoms marked as 1 in the atom mask are retained for pooling. Arithmetic average pooling is performed on the atom number dimension for all equivariant feature vectors of non-filled heavy atoms. During the pooling process, the L2 norm is taken for each spherical harmonic tensor component of each spherical harmonic channel. The L2 norm is calculated by squaring all components of the spherical harmonic tensor component, summing them, and then taking the square root, thereby transforming the higher-order spherical harmonic tensor components into rotation-invariant scalar values. Then, the scalar values ​​obtained after the L2 norm transformation of all 16 spherical harmonic channels are concatenated into a 512-dimensional spatial configuration feature vector. This spatial conformation feature vector compresses the three-dimensional folding morphology and spatial conformation information of the entire sugar chain. As a unified representation of the spatial modes of the sugar chain, its output dimension is strictly independent of the number of input heavy atoms M, and can directly participate in subsequent cross-modal comparative learning and multimodal fusion.

[0089] S107. Using three learnable projection matrices corresponding to the sequence mode, graph topology mode, and spatial configuration mode respectively, matrix multiplication is performed on the sequence feature vector, the graph feature vector, and the spatial configuration feature vector respectively, mapping each mode feature vector to a shared semantic space of a preset dimension d, and outputting three sets of intermediate vectors of dimension d. In this embodiment, the dimension d is 512, consistent with the dimension of the hidden state vectors output by the sequence encoder and spatial encoder, to ensure that information does not become a bottleneck due to dimension compression during forward propagation. The three learnable projection matrices are denoted as W_seq, W_graph, and W_3d, respectively, all of which are real number matrices of size 512 rows by 512 columns. Each element in the matrix is ​​a trainable parameter, initialized using Xavier uniform initialization, i.e., in... Uniform random sampling is performed within the specified interval to ensure that the variance of activation values ​​and gradient variance of each layer remain stable during forward and backward propagation. Continuing from the outputs of the three encoders mentioned above, the sequence feature vector has a dimension of 512, the graph feature vector has a dimension of 256, and the spatial configuration feature vector has a dimension of 512. Before being fed into their respective projection matrices, the graph feature vector first passes through a learnable linear dimension expansion layer, linearly mapping it from 256 dimensions to 512 dimensions. The dimension expansion matrix has a size of 256 rows multiplied by 512 columns. The dimension of the expanded graph feature vector is the same as that of the sequence and spatial configuration feature vectors, both being 512 dimensions. Subsequently, matrix multiplication operations are performed on the sequence feature vector, the expanded graph feature vector, and the spatial configuration feature vector with the projection matrices W_seq, W_graph, and W_3d, respectively, to obtain three intermediate vectors of dimension 512, denoted as z_seq, z_graph, and z_3d. L2 norm normalization was performed on the three sets of intermediate vectors to ensure that the L2 norm of each intermediate vector was equal to 1. This ensures that subsequent cross-modal semantic similarity calculations depend only on the consistency of vector directions and are not affected by differences in vector magnitudes. Taking the N-linked two-antenna complex glycan chain of the Fc region of human serum immunoglobulin G as a specific example, the first five elements of the L2 norm normalized intermediate vector z_seq obtained after projecting its sequence feature vector onto W_seq are 0.023, -0.015, 0.041, 0.008, and -0.033, respectively. The first five elements of the L2 norm normalized intermediate vector z_graph obtained after expanding the dimension and projecting the graph feature vector onto W_graph are 0.019, -0.022, 0.037, and -0.022, respectively. The first five elements of the L2 norm-normalized example values ​​of the intermediate vector z_3d obtained by projecting the spatial configuration feature vectors onto W_3d are 0.026, -0.011, 0.044, 0.005, and -0.036, respectively. Although the three sets of intermediate vectors come from three completely different modal encoders, they have been mapped to the same 512-dimensional vector space after linear spatial transformation of the projection matrix. This provides a vectorized representation with consistent dimensions and normalized modulus for subsequent cross-modal semantic alignment and gated adaptive fusion.

[0090] S108. In a training batch containing multiple target glycan samples, any two of the three sets of intermediate vectors for the same target glycan are taken as positive sample pairs, and the corresponding intermediate vectors for different target glycans are taken as negative sample pairs. The InfoNCE extended loss function based on temperature scaling cosine similarity is used to calculate the contrast loss values ​​between the sequence intermediate vector and the graph intermediate vector, the graph intermediate vector and the spatial conformation intermediate vector, and the spatial conformation intermediate vector and the sequence intermediate vector, respectively. The average of the three sets of contrast loss values ​​is taken as the cross-modal contrast loss function value. With minimizing the cross-modal contrast loss function value as the optimization objective, the network weight parameters of the sequence encoder, the graph encoder, the spatial encoder, and the three projection matrices are updated synchronously through the backpropagation algorithm. The training continues until the cross-modal contrast loss function value converges to a preset threshold range, so that the cosine similarity between any two pairs of the three sets of intermediate vectors for the same target glycan in the vector space of dimension d reaches a preset high similarity condition, and the cosine similarity between the intermediate vectors of different target glycans satisfies a preset low similarity condition.

[0091] In this embodiment, the cross-modal contrastive loss function described in S108 specifically adopts an extended form of the InfoNCE loss. During calculation, the InfoNCE loss values ​​applied to the three pairs of combinations—sequence intermediate vector z_seq and graph intermediate vector z_graph, graph intermediate vector z_graph and spatial intermediate vector z_3d, and spatial intermediate vector z_3d and sequence intermediate vector z_seq—are summed and their arithmetic mean is taken as the cross-modal contrastive loss function value. For any pair of intermediate vectors, a structural similarity-based screening strategy is introduced when constructing negative samples: the structural similarity measure between the anchor glycan and all candidate negative sample glycans within the batch is calculated, and only candidate glycans with a structural similarity below a preset threshold are included in the negative sample set; simultaneously, several negative samples with the highest structural similarity but still distinguishable are marked as hard negative samples, and a higher penalty weight is assigned to them in the loss function, forcing the model to learn to distinguish structurally similar glycan isomers. The temperature coefficient τ is set to 0.07, and the training batch size B is set to 256. These hyperparameters were determined through a grid search on a validation set containing 500 sugar chains. The evaluation metric is Recall@5 for the retrieval task. The search range for τ is {0.05, 0.07, 0.1}, the search range for the batch size is {128, 256, 512}, and the search range for the learning rate is {1e...}. -4 ,5e -5The optimal combination for validation performance was selected. In this embodiment, the temperature coefficient τ was set to 0.07, and its candidate search range was set to {0.05, 0.07, 0.1}. These two parameters were set based on the results of a grid search experiment performed on a validation set containing 500 glycan samples. The grid search used the recall rate (Recall@5) of the cross-modal retrieval task as the evaluation metric, and trained and evaluated the three candidate values ​​of the temperature coefficient separately. At the same time, it jointly searched for combinations of hyperparameters such as training batch size and learning rate, and finally selected the parameter combination that maximized the Recall@5 metric on the validation set as the optimal configuration. The experimental results showed that the model performance was optimal when the temperature coefficient was 0.07. This value achieved the best balance between the discrimination between positive and negative samples and the training stability within this range. Too low a temperature coefficient would cause the similarity distribution to be too steep, making it difficult for the model to learn subtle differences, while too high a temperature coefficient would cause the similarity distribution to be too smooth, weakening the ability of contrastive learning to capture semantic structure.

[0092] Taking the combination of z_seq and z_graph as an example, the cosine similarity between z_seq and z_graph of the same sugar chain is used as the numerator, and the sum of the cosine similarities between z_seq and z_graph of all sugar chains in the batch is used as the denominator. The negative logarithm with the natural constant e as the base is calculated to obtain the InfoNCE loss value of this combination; the other two combinations are calculated in the same way. In this embodiment, the backpropagation algorithm described in S108 uses the Adam optimizer with a learning rate of 0.0001 to synchronously update all learnable parameters of the sequence encoder, the graph encoder, the spatial encoder, and the three projection matrices. Training is terminated when the loss value falls into the preset threshold range [0.15, 0.25] and the fluctuation amplitude of 5 consecutive epochs is less than 0.02. At the same time, it is verified that the cosine similarity of the three sets of intermediate vectors of the same sugar chain is not less than 0.85. If both conditions are met, convergence is determined. This threshold is obtained by statistical analysis of 100 different initialization training. After training for a preset training period, the cosine similarity between the three sets of intermediate vectors of the same target sugar chain in the 512-dimensional vector space is constrained to no less than 0.85, while the cosine similarity between the intermediate vectors of different sugar chains is reduced to no more than 0.30, thereby achieving semantic alignment of sequence modalities, graph topology modalities and spatial configuration modalities in the 512-dimensional vector space.

[0093] S109. Input the three sets of intermediate vectors that meet the preset high similarity condition into the gated adaptive fusion layer. The gated adaptive fusion layer assigns an independent learnable scalar parameter to each set of intermediate vectors. Apply Softmax normalization to the three learnable scalar parameters to generate a three-dimensional attention weight vector. Perform element-wise multiplication of each weight component in the three-dimensional attention weight vector with the corresponding intermediate vector. Then perform element-wise addition on the three weighted vectors to generate a multimodal semantic representation vector of dimension d. The value of the i-th element in the multimodal semantic representation vector is the sum of the i-th element of the sequence intermediate vector, the i-th element of the graph intermediate vector, and the i-th element of the spatial configuration intermediate vector multiplied by their respective attention weight components.

[0094] In this embodiment, the gated adaptive fusion layer maintains three learnable scalar parameters, denoted as α, β, and γ, which are used to assign fusion weights to the sequence mode, graph topology mode, and spatial configuration mode, respectively. The initial values ​​of all three parameters are set to 1.0. During forward propagation, a Softmax normalization operation is applied to α, β, and γ to obtain the corresponding weight coefficients: w_seq equals e^(α) divided by e^(α) plus e^(β) plus e^(γ), w_graph equals e^(β) divided by e^(α) plus e^(β) plus e^(γ), and w_3d equals e^(γ) divided by e^(α) plus e^(β) plus e^(γ), satisfying the normalization constraint that w_seq plus w_graph plus w_3d equals 1. Subsequently, the intermediate vector z_seq is multiplied by the weight coefficient w_seq, z_graph by the weight coefficient w_graph, and z_3d by the weight coefficient w_3d. The three weighted vectors are then summed element-wise to generate a multimodal semantic representation vector with the same dimension of 512. This vector adaptively integrates information on the sequence composition, branching topology, and three-dimensional folding of the sugar chain within a unified semantic space. It should be further noted that in this embodiment, the gated adaptive fusion layer uses three global scalar parameters for intermodal weighted fusion, without introducing an intramodal attention mechanism for each dimension within the intermediate vector. This rationale is based on the semantic alignment achieved by the preceding cross-modal contrastive learning step. Specifically, after optimization using the cross-modal contrastive loss function, the cosine similarity between each pair of the three intermediate vectors z_seq, z_graph, and z_3d for the same target glycan in the 512-dimensional shared semantic space is no less than the corresponding cosine similarity threshold. This indicates that the sequence encoder, graph encoder, and spatial encoder have mapped the three heterogeneous modal data to a semantically consistent homogeneous vector space. Under this premise, each dimension of each intermediate vector encodes homogeneous semantic features of the glycan structure. Applying a uniform global scalar weight to the entire vector set is equivalent to scaling all dimensional features of the modality proportionally, without selectively suppressing the locally important features carried by a certain dimension within the modality. This design ensures the high efficiency of cross-modal fusion computation while avoiding the risk of overfitting due to the introduction of too many learnable parameters. The initial values ​​of the three scalar parameters α, β, and γ are all set to 1.0 instead of other values. The reason for this setting is that after cross-modal semantic alignment, there is no prior difference in importance between the sequence mode, graph topology mode, and spatial conformation mode for the representation of glycan structure. The three modes are at an equal starting point for fusion. After Softmax normalization, the equal initial values ​​make the initial values ​​of w_seq, w_graph, and w_3d all 1 / 3, ensuring that the fusion process starts in an unbiased state.In the subsequent task fine-tuning training of the multimodal agent, the three scalar parameters are automatically learned through the gradient backpropagation algorithm, with the cross-entropy loss of structure prediction as the optimization objective. This automatically learns the contribution of each modality to the final structure resolution accuracy, thereby achieving data-driven dynamic allocation of modal importance. For glycan samples where the molecular graph branch topology information plays a decisive role in the resolution result, the weight coefficient w_graph corresponding to the graph topology modality will automatically increase during training. For glycan samples with greater three-dimensional conformation flexibility and more conservative sequence features, the weight coefficient w_seq corresponding to the sequence modality will obtain a higher response value. This allows the multimodal semantic representation vector to adaptively adjust the contribution ratio of each modality according to the structural characteristics of different glycans.

[0095] This embodiment employs a three-level coding mapping network to extract targeted features from three heterogeneous data types—glycan sequence encoding, molecular graph topology, and 3D spatial conformation—using a Transformer-based sequence encoder, a message-passing neural network graph encoder, and an SE3-based variable graph neural network spatial encoder, respectively. This ensures that structural information from different physical entities is fully captured using neural network architectures best suited to their data formats. Three learnable projection matrices are used to uniformly map the sequence feature vectors, graph feature vectors, and spatial conformation feature vectors of varying dimensions to a 512-dimensional vector space and apply L2 norm normalization, eliminating the interference of differences in vector magnitude and dimension between different modal features on subsequent fusion. A cross-modal contrastive loss function is used to maximize the pairwise cosine similarity of the three sets of intermediate vectors for the same glycan while minimizing the differences between different glycans. Using vector similarity as the optimization objective, semantic alignment of sequence mode, graph topology mode, and spatial configuration mode is achieved in a 512-dimensional vector space. This ensures that the three modal representations of the same glycan are close to each other in the high-dimensional space, with a cosine similarity of no less than 0.85, overcoming the semantic gap between heterogeneous modes. Through a gated adaptive fusion layer, the three sets of intermediate vectors are summed element-wise with learnable Softmax normalized weights. This allows different glycans to adaptively adjust the contribution ratio of each mode according to their structural characteristics. For example, N-connected glycans with rigid branch topology are automatically assigned higher weights to the graph topology mode. This generates a multimodal semantic representation vector that integrates complete structural information from three aspects: sequence composition, branch topology, and three-dimensional folding. This provides a unified and lossless vectorized representation basis for cross-modal prediction of downstream glycan structures.

[0096] It should be further explained that the multimodal agent trained in this embodiment includes a cross-modal fusion sub-agent, a structure prediction sub-agent, an uncertainty quantification sub-agent, and a structure judgment output sub-agent. In this embodiment, the training process of the trained multimodal agent includes two parts: a pre-training stage and a task fine-tuning stage. The training data comes from glycomics public databases and structured datasets mined from literature, covering GlycoCT sequences of known glycans registered in the GlyTouCan international glycan structure database, corresponding tandem mass spectrometry (MS / MS) spectra, one-dimensional hydrogen nuclear magnetic resonance (1HNMR) spectra, and reversed-phase liquid chromatography (RPLC) retention time data, as well as manually labeled glycan structure sequence tags.

[0097] In this embodiment, the pre-training stage employs a self-supervised learning approach. The sequence encoder uses masked language modeling as its pre-training task, randomly masking 15% of the words in the input GlycoCT word sequence. The encoder is trained to restore the original integer indices of the masked words based on the unmasked context words, enabling the sequence encoder to capture the co-occurrence patterns and syntactic structures between monosaccharides in the glycan sequence. The graph encoder uses graph structure contrast learning as its pre-training task, applying random node feature masking and random edge perturbation to the molecular graph of the same glycan to generate two enhanced views. The encoder is trained to maximize the cosine similarity between the graph feature vectors of the two enhanced views of the same glycan, while minimizing the similarity with the graph feature vectors of other glycans in the same batch, enabling the graph encoder to capture the core topological features in the glycan molecular graph that are robust to local perturbations. The spatial encoder's pre-training employs a coordinate denoising task, applying Gaussian noise perturbation to the three-dimensional Cartesian coordinate set before inputting it into the spatial encoder. The encoder is trained to restore the original coordinate values, enabling the spatial encoder to learn the physicochemical constraints of the three-dimensional conformation of the glycan.

[0098] In this embodiment, the task fine-tuning stage employs supervised learning, using known glycan structure labels as supervisory signals. Multimodal semantic representation vectors and multi-source experimental spectrogram data are input into a multimodal agent composed of a cross-modal fusion sub-agent, a structure prediction sub-agent, an uncertainty quantification sub-agent, and a structure judgment output sub-agent. During forward propagation, the cross-modal fusion sub-agent fuses the intrinsic structural semantics of the glycan with the experimental spectrogram features through four rounds of sequence-to-sequence multi-head cross-attention interaction. The structure prediction sub-agent predicts the monosaccharide type, glycosidic bond connection site, and anodic configuration residue by residue from the non-reducing end to the reducing end, and adds an [EOS] marker as a generation termination condition. The uncertainty quantification sub-agent uses a deep ensemble method to perform prediction statistics on multiple sets of structure prediction sub-agents with different random initialization parameters to generate a confidence distribution. This confidence level reflects both accidental uncertainty and cognitive uncertainty. The structure judgment output sub-agent distinguishes between high-confidence and low-confidence units using a pre-set confidence threshold of 0.9. The cross-entropy loss between the predicted glycan structure sequence output by the structure prediction sub-agent and the known glycan structure label is used as the fine-tuning loss function. An Adam optimizer with a learning rate of 0.0001 is used, the training batch size is set to 256, and 100 rounds of iterative training are performed. After the loss function converges, the trained multimodal agent is obtained.

[0099] It should be further explained that, in this embodiment, the cross-modal fusion sub-agent is configured to: receive the multimodal semantic representation vector and the multi-source experimental spectral data; map the mass spectrometry data, nuclear magnetic resonance spectral data and chromatographic parameter data in the multi-source experimental spectral data into mass spectrometry feature vectors, nuclear magnetic resonance feature vectors and chromatographic feature vectors respectively through their respective independent spectral feature extractors; use the multimodal semantic representation vector as a query vector; and concatenate the mass spectrometry feature vector, the nuclear magnetic resonance feature vector and the chromatographic feature vector as a key vector and a value vector; perform cross-modal interaction through a multi-head cross-attention network; and output the glycan hidden state sequence after fusing spectral information. The length of the glycan hidden state sequence is equal to the total number of monosaccharide residues of the target glycan, and the dimension of the hidden state vector at each time step in the glycan hidden state sequence is 512.

[0100] In this embodiment, the multimodal semantic representation vector is a 512-dimensional real-valued vector output after processing by a three-level coding mapping network and a gated adaptive fusion layer. This vector integrates information on the sequence composition, branching topology, and three-dimensional spatial conformation of the target glycan. The multi-source experimental spectral data specifically includes data from three modalities: tandem mass spectrometry (MS / MS) spectra, one-dimensional hydrogen nuclear magnetic resonance (1HNMR) spectra, and reversed-phase liquid chromatography (RPLC) retention times. These three modalities provide structural constraint information of the target glycan from different dimensions: the MS / MS spectra record the mass-to-charge ratio and relative abundance of fragment ions generated during the gas-phase fragmentation of the glycan; the 1HNMR spectra record the chemical shifts and peak area integrals of each proton in the glycan under different chemical environments; and the RLC retention times record the elution time and peak shape parameters of the glycan under specific mobile phase gradient and stationary phase conditions.

[0101] In this embodiment, the spectral data is first encoded into feature sequences by their respective spectral feature extractors. Then, the spectral feature sequences are subjected to sequence-to-sequence multi-head cross-attention interaction with the hidden state sequences generated by the glycan structure decoder. This allows each generation position of the glycan to dynamically focus on information from different regions of the spectrum, achieving effective fusion. The mass spectrometry feature extractor uses a convolutional neural network containing two layers of one-dimensional convolution. Its input is a discretized representation of the tandem mass spectrometry (MS / MS) spectrum. Specifically, the mass-to-charge ratio range from 100 to 2000 is divided into 3800 discrete intervals at intervals of 0.5. The maximum relative abundance value of the fragment ions falling into each interval is taken as the feature value of that interval, forming a one-dimensional spectral vector of length 3800. The one-dimensional spectral vector first passes through a first one-dimensional convolutional layer. This layer has a kernel size of 5, a stride of 2, and 64 output channels. The kernel size of 5 indicates that this layer performs convolution operations on every 5 consecutive intervals of the spectral vector to capture the local fragmentation patterns between neighboring fragment ions. The stride of 2 indicates that the kernel slides 2 intervals each time to compress the sequence length to half of the original. The number of output channels of 64 indicates that this layer learns 64 different sets of convolutional kernel parameters to extract 64 different types of local fragment ion pattern features. After processing by this layer, the output feature matrix has a size of 1898 multiplied by 64. The feature matrix is ​​then passed through a second one-dimensional convolutional layer with a kernel size of 3, a stride of 2, and 128 output channels. The kernel size of 3 indicates that this layer performs convolution operations on every three consecutive feature vectors output from the previous layer to capture higher-level fragment ion combination patterns. The stride of 2 further compresses the sequence length to half its original value. The 128 output channels indicate that this layer learns 128 different sets of convolutional kernel parameters to extract 128 different fragment ion combination pattern features. After processing by this layer, the output is a feature matrix of size 948 x 128. Finally, global average pooling is performed on this feature matrix along the sequence length dimension, compressing the 128-dimensional feature vectors from 948 time steps into a single 128-dimensional mass spectrometry feature vector.

[0102] The NMR feature extractor employs a deep neural network with three fully connected layers. Its input is a discretized representation of the one-dimensional 1H NMR spectrum. Specifically, the chemical shift range from 0 ppm to 10 ppm is divided into 1000 discrete intervals at 0.01 ppm intervals. The peak intensity value at each chemical shift is taken as the feature value of that interval, forming a one-dimensional spectral vector of length 1000. This one-dimensional spectral vector is sequentially passed through the first, second, and third fully connected layers. The first fully connected layer has an input dimension of 1000 and an output dimension of 256; the second fully connected layer has an input dimension of 256 and an output dimension of 128; and the third fully connected layer has an input dimension of 128 and an output dimension of 128. A ReLU nonlinear activation function is applied after each fully connected layer to introduce nonlinear expressive power. After feature extraction and transformation layer by layer through the three fully connected layers, a 128-dimensional NMR feature vector is output.

[0103] The chromatographic feature extractor employs a deep neural network with two fully connected layers. Its input is a 3D vector consisting of three parameters related to the retention time in reversed-phase liquid chromatography: retention time, peak width, and peak symmetry factor. This 3D vector is sequentially passed through a first and a second fully connected layer. The first fully connected layer has an input dimension of 3 and an output dimension of 32, while the second fully connected layer has an input dimension of 32 and an output dimension of 128. A ReLU nonlinear activation function is applied after each fully connected layer. After layer-by-layer feature expansion and transformation through the two fully connected layers, a 128-dimensional chromatographic feature vector is output.

[0104] The 128-dimensional mass spectrometry feature vector, 128-dimensional nuclear magnetic resonance feature vector, and 128-dimensional chromatography feature vector output from the three spectral feature extractors are concatenated along their dimensional directions to form a 384-dimensional concatenated spectral vector. Subsequently, a learnable linear projection matrix (384 rows x 512 columns) is used to transform the 384-dimensional concatenated spectral vector to 512 dimensions, ensuring its dimensionality matches that of the multimodal semantic representation vector. The 512-dimensional multimodal semantic representation vector is used as the initial representation of the query matrix, and the linearly projected 512-dimensional concatenated spectral vector is used as the initial representation of both the key and value vectors. These are then fed into a multi-head cross-attention network with eight attention heads. In this multi-head cross-attention network, each attention head independently maps the query vector, key vector, and value vector to 64-dimensional query projection vectors, 64-dimensional key projection vectors, and 64-dimensional value projection vectors respectively through three learnable projection matrices. The size of each projection matrix is ​​the input dimension multiplied by 64. Within each attention head, the scaled dot product attention score between the query projection vector and the key projection vector is calculated, with a scaling factor of the square root of 64 (8). The attention score is then normalized using Softmax and used as a weighting coefficient to sum the value projection vectors, resulting in a 64-dimensional output vector for each attention head. The 64-dimensional output vectors from the eight attention heads are concatenated along the dimensional direction to form a 512-dimensional vector. This 512-dimensional vector is then mapped back to 512 dimensions through a learnable output projection matrix, completing one cross-attention interaction. The cross-attention interaction process of the aforementioned multi-head cross-attention network is repeated for four rounds. The 512-dimensional vector output in each round serves as the query input for the next round, while the key and value vectors remain unchanged in each round. This allows the multimodal semantic representation vector to extract the experimental spectrogram features most relevant to glycan structure analysis from the spliced ​​spectrogram vector round by round. After four rounds of cross-attention interaction, the output 512-dimensional vector is a single time-step representation of the glycan hidden state sequence that integrates spectrogram information and the intrinsic structural semantics of the glycan. The aforementioned cross-attention interaction process is applied in parallel to all L monosaccharide residue positions of the target glycan. Each monosaccharide residue position independently maintains a 512-dimensional query vector. The query vectors of all positions share the same set of spliced ​​spectrogram vectors as key and value vectors. After four rounds of cross-attention interaction, a glycan hidden state sequence of length L with 512 dimensions at each time step is output. Each time step of this sequence corresponds to a monosaccharide residue position of the target glycan, and the time step order is strictly arranged from the reduced end to the non-reduced end.

[0105] In this embodiment, the N-linked two-antenna complex glycan of the Fc region of human serum immunoglobulin G is used as a specific example. The total number of monosaccharide residues L in this glycan is equal to 14. Therefore, the generated hidden state sequence of the glycan contains 14 time steps. The first time step in the sequence corresponds to the N-acetylglucosamine residue at the reducing end, the second time step corresponds to the second N-acetylglucosamine residue linked to the reducing end by a β-1,4 bond, and so on until the 14th time step corresponds to the sialic acid residue at the very end of the non-reducing end. The hidden state vector of each time step is 512-dimensional. This hidden state sequence of the glycan chain fully integrates the three-level encoded multimodal semantic information of the glycan chain itself with information from the mass spectrometry fragmentation pattern, nuclear magnetic resonance chemical shift, and chromatographic retention behavior from experimental spectra. Among them, the mass spectrometry fragmentation pattern information provides the characteristics of glycosidic bond breakage sites for the time step vector of the N-acetylglucosamine residue at the reducing end; the nuclear magnetic resonance chemical shift information provides the identification characteristics of monosaccharide type and anodic configuration for the time step vector of each residue; and the chromatographic retention behavior information provides global constraint characteristics for the hydrophobicity and branching degree of the overall glycan chain. As the final output of the cross-modal fusion sub-agent, this sequence provides a complete information foundation for the subsequent structure prediction sub-agent.

[0106] It should be further explained that the structure prediction sub-agent in this embodiment is configured to: receive the hidden state sequence of the glycan chain output by the cross-modal fusion sub-agent; input the hidden state sequence of the glycan chain into an autoregressive sequence generation model based on a Transformer decoder; predict the monosaccharide type, glycosidic bond connection site, and anodic configuration of each monosaccharide residue of the target glycan chain step by step from the reduced end to the non-reduced end; and output a predicted glycan chain structure sequence of length L, where L is equal to the total number of monosaccharide residues of the target glycan chain. The output of the predicted glycan chain structure sequence at time step t is a three-dimensional probability vector, the first dimension of which is the probability distribution of the predicted t-th monosaccharide residue of the target glycan chain on a preset set of monosaccharide types; the second dimension of which is the probability distribution of the combination of glycosidic bond connection sites between the predicted t-th monosaccharide residue of the target glycan chain and the previous monosaccharide residue on a preset set of connection site combinations; and the third dimension of which is the probability distribution of the predicted anodic configuration of the predicted t-th monosaccharide residue of the target glycan chain on a preset set of anodic configurations.

[0107] In this embodiment, the predictive sub-agent strictly inherits the glycan hidden state sequence output by the cross-modal fusion sub-agent. This glycan hidden state sequence has a length of L and a hidden state vector dimension of 512 at each time step. The time step order strictly follows the monosaccharide residues of the target glycan from the reducing end to the non-reducing end. The preset monosaccharide type set includes 20 common monosaccharide types found in glycomics research, specifically glucose, galactose, mannose, fucose, N-acetylglucosamine, N-acetylgalactosamine, sialic acid, xylose, glucuronic acid, iduronic acid, N-acetylmannosamine, N-acetylfucosamine, galacturonic acid, mannuronic acid, guluronic acid, N-acetylgalactosamine-6-sulfate, N-acetylglucosamine-6-sulfate, N-acetylglucosamine-3-sulfate, N-acetylgalactosamine-4-sulfate, and N-acetylglucosamine-4-sulfate. The predefined set of glycosidic bond linkage combinations contains 48 possible combinations. These combinations are derived from the Cartesian product of six possible donor sites (C1, C2, C3, C4, C6, and C8) of the donor monosaccharide residue and five possible acceptor sites (C2, C3, C4, C6, and C8) of the acceptor monosaccharide residue, after removing combinations never seen in natural sugar chains. The predefined set of anodic configurations includes two glycosidic bond spatial orientations: α-configuration and β-configuration.

[0108] The autoregressive sequence generation model based on the Transformer decoder consists of six stacked Transformer decoder layers. Each layer contains, sequentially, an 8-head self-attention sublayer with causal masking, an 8-head cross-attention sublayer, and a fully connected feedforward sublayer. The causal masked self-attention sublayer has 8 attention heads. In the self-attention mechanism, the projection dimensions of the key vector, value vector, and query vector are all 64. The causal masking ensures that the prediction at time step t relies only on the generated information from the previous t-1 time steps and cannot access information from future time steps. The cross-attention sublayer has 8 attention heads. Its query vector comes from the output of the causal masked self-attention sublayer, and its key and value vectors come from the sugar chain hidden state sequence output by the cross-modal fusion sub-agent. The projection dimensions of both the key and value vectors are 64. The hidden layer width of the fully connected feedforward sublayer is 2048. Each sublayer is followed by residual connections and layer normalization, and the hidden state vector dimension of the entire decoder is fixed at 512.

[0109] The autoregressive sequence generation process begins at the reduction end. At the first time step, the model uses a special initial embedding vector as the output of the previous time step; this initial embedding vector is a 512-dimensional learnable parameter vector. At the subsequent t-th time step, the model element-wise adds the monosaccharide type embedding vector, the connector site combination embedding vector, and the anodic configuration embedding vector predicted at the (t-1)-th time step to form the comprehensive embedding vector for the (t-1)-th time step, which serves as the input for the t-th time step. The monosaccharide type embedding vector maps the predicted monosaccharide type index to a 512-dimensional vector using a learnable monosaccharide type embedding matrix; the connector site combination embedding vector maps the predicted connector site combination index to a 512-dimensional vector using a learnable connector configuration embedding matrix; and the anodic configuration embedding vector maps the predicted anodic configuration index to a 512-dimensional vector using a learnable anodic configuration embedding matrix. After processing layer by layer through six Transformer decoder layers, a 512-dimensional hidden state vector is output at the t-th time step. The hidden state vector is fed into three parallel linear classification heads: the monosaccharide type classification head is a 512x20 linear projection matrix that maps the 512-dimensional hidden state vector to a 20-dimensional score vector, and after Softmax normalization, the probability distribution of the t-th monosaccharide residue on the preset monosaccharide type set is obtained; the linkage site classification head is a 512x48 linear projection matrix that maps the 512-dimensional hidden state vector to a 48-dimensional score vector, and after Softmax normalization, the probability distribution of the glycosidic bond linkage site combination between the t-th monosaccharide residue and the previous monosaccharide residue on the preset linkage site combination set is obtained; and the anodic configuration classification head is a 512x2 linear projection matrix that maps the 512-dimensional hidden state vector to a 2-dimensional score vector, and after Softmax normalization, the probability distribution of the anodic configuration of the t-th monosaccharide residue on the preset anodic configuration set is obtained. The Softmax normalization operations of the three classification heads are independent of each other, and the three output probability distributions together constitute the three-dimensional probability vector at time step t. In this embodiment, the autoregressive sequence generation model based on the Transformer decoder adds a special sequence termination marker, namely the EOS marker, to the output layer vocabulary. The index value of this marker is preset to the maximum index value of the current vocabulary + 2, which is different from the index value of the truncated marker lexicon. The EOS marker participates in the forward propagation calculation of the model in both the training and inference phases, enabling the model to autonomously determine whether the current glycan structure sequence has been completely generated at any generation time step. In the inference phase, the autoregressive generation process strictly starts from the reduced end and predicts the type, connection site combination, and anodic configuration of the next monosaccharide residue towards the non-reduced end step by step. The probability distribution output by the monosaccharide type classification head at each time step includes the probability term corresponding to the EOS marker.When the predicted probability of the EOS marker reaches its maximum among all candidate categories at a certain time step, the generation process terminates immediately. All non-EOS monosaccharide residue sequences generated before this termination step are then concatenated in the order of generation from reduced to non-reduced ends, outputting the complete predicted glycan structure sequence of the target glycan. This EOS marker mechanism enables the model to adaptively determine the termination position of the generated sequence based on the input multimodal representation vector and multi-source experimental spectrogram data, producing prediction results with a length matching the actual number of monosaccharide residues in the target glycan. This fundamentally avoids the problems of infinite generation or fixed-length truncation caused by the lack of termination conditions, ensuring that the subsequent uncertainty quantification sub-agent can quantify the confidence of predicted structures of different lengths, and that the hypothesis generation stage can perform knowledge graph retrieval based on the determined complete predicted structure.

[0110] It should be further explained that the uncertainty quantification sub-agent in this embodiment is configured as follows: during the autoregressive sequence generation process of the structure prediction sub-agent, Monte Carlo dropout is applied to each feedforward sub-layer in the autoregressive sequence generation model based on the Transformer decoder, with a dropout rate set to 0.1. During the inference phase, K random forward propagations are performed, where K is 100. Each random forward propagation independently outputs a glycan structure sequence sample. For each candidate monosaccharide type, each candidate glycosidic bond combination, and each candidate anodic configuration in the K glycan structure sequence samples at the t-th time step, the frequency of occurrence of each candidate monosaccharide type, each candidate glycosidic bond combination, and each candidate anodic configuration is counted in the K sampling. The frequency of occurrence is divided by K to obtain the confidence probability value of each candidate category of the t-th structural unit. A confidence probability distribution sequence of length L is output. The t-th element in the confidence probability distribution sequence is a triple containing the confidence distribution of monosaccharide type, the confidence distribution of combination site, and the confidence distribution of anodic configuration. In this embodiment, the Monte Carlo dropout rate of 0.1, the number of random forward propagations of 100, and the preset confidence threshold of 0.9 are determined as follows: On a dataset containing 500 manually labeled glycan structures for verification, comparative experiments were conducted on all 16 parameter combinations with Monte Carlo dropout rates of 0.05, 0.1, 0.2, and 0.3 and random forward propagation numbers of 50, 100, 200, and 500, respectively. The area under the confidence calibration curve and the structure prediction accuracy were used as dual evaluation indicators. Experimental results show that when the Monte Carlo dropout rate is 0.1 and the number of random forward propagations is 100, the area under the confidence calibration curve reaches its maximum value of 0.94, while the structural prediction accuracy only decreases by 1.2 percentage points compared to the deterministic inference model, achieving the optimal balance between uncertainty quantification quality and computational efficiency. When the Monte Carlo dropout rate is below 0.1, the model's randomness is insufficient, leading to an underestimation of uncertainty; when it is above 0.1, the model's randomness is too large, resulting in a significant decrease in prediction accuracy. When the number of random forward propagations is below 100, the variance of the confidence estimate is too large, leading to statistical instability; when it is above 100, the confidence estimate has basically converged, and the marginal improvement is less than one-thousandth. Therefore, 100 times was selected as the optimal number of sampling times for cost-effectiveness. The determination of the preset confidence threshold of 0.9 is based on the following: On the above-mentioned validation dataset, each candidate threshold is traversed in the range of 0.50 to 0.99 with a step size of 0.05. The false positive rate (the proportion of actually correctly predicted units in low-confidence units) and the false negative rate (the proportion of actually incorrectly predicted units in high-confidence units) corresponding to each candidate threshold are calculated respectively. When the confidence threshold is set to 0.9, the false positive rate is 9.8% and the false negative rate is 2.1%. The sum of the two reaches the minimum value among all candidate thresholds, indicating that this threshold can effectively screen out structural units with insufficient prediction reliability while avoiding incorrectly labeling high-confidence units as low-confidence units as much as possible. This provides the optimal decision boundary for the reasonable allocation of experimental validation resources.

[0111] In this embodiment, because the termination position of each glycan structure sequence sample generated during the Monte Carlo discard random forward propagation process may be different, the lengths of the 100 glycan structure sequence samples generated from 100 samplings are not completely consistent. To solve the position alignment problem between sequences of different lengths, this embodiment adopts the following sequence alignment strategy: First, the sequence with the highest frequency among the 100 glycan structure sequence samples is selected as the reference sequence; then, the Needleman-Wunsch global sequence alignment algorithm is used to align the remaining 99 glycan structure sequence samples one by one with the reference sequence. In this alignment algorithm, the matching score is set to +1 point, the mismatch penalty is set to -1 point, the gap open penalty is set to -2 points, and the gap extension penalty is set to -1 point. Through dynamic programming, gap symbols are reasonably inserted into the sequence so that the same or similar structural units in each sequence are in the same alignment position after alignment. The gap symbol is used as an independent candidate category in subsequent frequency statistics. After the alignment of all 99 sequences is completed, the length of each sequence is extended to be consistent with the alignment length of the reference sequence. At each aligned position, based on the 100 samples (from 100 aligned sequences) corresponding to that position, the frequency of occurrence of each candidate monosaccharide type, each candidate glycosidic bond combination, and each candidate anomeric configuration in the 100 samples is calculated. The frequency is then divided by 100 to obtain the confidence probability value of each candidate category at that aligned position. This sequence alignment strategy ensures the comparability of confidence statistics for sampled sequences of different lengths at the same structural position, avoids the confidence statistical bias caused by direct truncation or padding, and guarantees the accuracy and reliability of the confidence probability distribution sequence output by the uncertainty quantification sub-agent at the structural unit level.

[0112] In this embodiment, the uncertainty quantification sub-agent strictly inherits the autoregressive sequence generation model based on the Transformer decoder within the structure prediction sub-agent. This autoregressive sequence generation model consists of six stacked Transformer decoder layers. Each decoder layer contains a causal masked self-attention sub-layer, a cross-attention sub-layer, and a fully connected feedforward sub-layer. The hidden layer width of the fully connected feedforward sub-layer is 2048. The Monte Carlo dropout operation is specifically applied to three residual connection positions in each decoder layer: after the self-attention sub-layer, after the cross-attention sub-layer, and after the feedforward sub-layer. For each position, each element in the 512-dimensional hidden state vector output is randomly set to zero with a probability of 0.1, and its original value is retained with a probability of 0.9, multiplied by 1, and divided by a scaling factor of 0.9 to maintain the expected activation value. During the training phase, the Monte Carlo dropout operation with a dropout rate of 0.1 participates in model training as a random regularization method, forcing the model to make correct predictions even when some hidden units are masked, thereby learning a more robust sugar chain structure representation. During the inference phase, the Monte Carlo dropout operation remains active, and 100 random forward propagations are performed. Because the combination of randomly dropped hidden units differs in each forward propagation, the network connection structure varies with each propagation, resulting in 100 distinct glycan structure sequence samples from each of the 100 forward propagations. At each step of autoregressive sequence generation, the monosaccharide type score vector, connection site combination score vector, and anodic configuration score vector output at each monosaccharide residue time step are all transformed into a probability distribution using Softmax normalization. Random sampling from this probability distribution is then used as the prediction output for that step, rather than a deterministic prediction based on the highest probability, ensuring the randomness of the sampling path in each forward propagation.

[0113] For 100 glycan structure sequence samples obtained from 100 random forward propagations, at the t-th time step, the frequency of occurrence of each candidate monosaccharide type in the 20 categories of the preset monosaccharide type set is counted. The frequency of occurrence is divided by 100 to obtain the monosaccharide type confidence distribution of the t-th structural unit. This distribution is a 20-dimensional vector, where each element takes a value between 0 and 1 and the sum of all elements is 1. Similarly, the frequency of occurrence of each candidate glycosidic bond linkage site combination in the 48 categories of the preset linkage site combination set is counted. The frequency of occurrence is divided by 100 to obtain the linkage site confidence distribution of the t-th structural unit. This distribution is a 48-dimensional vector. The frequency of occurrence of each candidate anodic configuration in the 2 categories of the preset anodic configuration set is counted. The frequency of occurrence is divided by 100 to obtain the anodic configuration confidence distribution of the t-th structural unit. This distribution is a 2-dimensional vector. The confidence distributions for monosaccharide type, linkage site, and anomeric configuration are combined into a triplet, which is then used as the t-th element in the confidence probability distribution sequence. This statistical process is repeated for all L time steps, outputting a confidence probability distribution sequence of length L.

[0114] It should be further explained that the structure judgment output sub-agent in this embodiment is configured to: receive the predicted glycan structure sequence output by the structure prediction sub-agent and the confidence probability distribution sequence output by the uncertainty quantification sub-agent; take the minimum value of the triples of the t-th structural unit in the confidence probability distribution sequence as the comprehensive confidence score of the t-th structural unit; when the comprehensive confidence scores of all L structural units are greater than or equal to a preset confidence threshold, output the predicted glycan structure sequence as the predicted structure of the target glycan, and output the confidence probability distribution sequence as the confidence score corresponding to each structural unit in the predicted structure; when the comprehensive confidence score of at least one structural unit is less than the preset confidence threshold, mark the structural unit as a low-confidence unit, add a low-confidence mark to the low-confidence unit in the output predicted structure, and output the complete confidence probability distribution sequence.

[0115] In this embodiment, the structure judgment output sub-agent strictly inherits the predicted glycan structure sequence output by the structure prediction sub-agent in a deterministic forward propagation mode and the confidence probability distribution sequence obtained by the uncertainty quantification sub-agent based on 100 Monte Carlo dropout random forward propagation statistics. The predicted glycan structure sequence is generated by the structure prediction sub-agent in inference mode with Monte Carlo dropout disabled, taking the maximum probability category from the three-dimensional probability vector at each time step. Specifically, at time step t, the monosaccharide type is taken from the monosaccharide type probability distribution with the maximum value, the glycosidic bond connection site combination is taken from the connection site combination probability distribution with the maximum value, and the anodic configuration is taken from the anodic configuration probability distribution with the maximum value. A preset confidence threshold is set to 0.9, determined according to industry standards for glycomic structure analysis. When the confidence level is below 0.9, the prediction of the structural unit is considered to have significant uncertainty, requiring further experimental verification. The operation of minimizing the triplet of the t-th structural unit specifically involves: extracting the maximum confidence value from the monosaccharide type confidence distribution, the maximum confidence value from the connection site confidence distribution, and the maximum confidence value from the anodic configuration confidence distribution of the t-th element in the confidence probability distribution sequence; and taking the minimum of these three maximum confidence values ​​as the comprehensive confidence score of the t-th structural unit. This operation ensures that the comprehensive confidence score of the structural unit only reaches the threshold when the predictions of the monosaccharide type, connection site, and anodic configuration all reach high confidence. Low confidence in any one dimension will lower the comprehensive confidence score. The low confidence marker is implemented by appending a low confidence marker bit with a value of 1 after the monosaccharide type field for structural units with a comprehensive confidence score lower than 0.9 in the output predicted glycan structure sequence; for structural units with a comprehensive confidence score greater than or equal to 0.9, the marker bit is set to 0.Taking the N-linked two-antenna complex glycan of the Fc region of human serum immunoglobulin G as a specific example, the total number of monosaccharide residues L in this glycan is equal to 14. In the confidence probability distribution sequence output by the uncertainty quantification sub-agent, the 6th structural unit is a mannose residue linked to the core mannose via an α-1,6 bond. In its monosaccharide type confidence distribution, the confidence of mannose is 0.96; in the linkage site confidence distribution, the confidence of the combination of carbon 1 and carbon 6 is 0.85; and in the anodic configuration confidence distribution, the confidence of the α configuration is 0.92. Therefore, the overall confidence of this structural unit is... The score of 0.85, the minimum of the three maximum confidence values ​​of 0.96, 0.85, and 0.92, is less than the preset confidence threshold of 0.9. Therefore, this structural unit is marked as a low-confidence unit, and a low-confidence marker of 1 is added to the position of the mannose residue in the output predicted structure. The overall confidence scores of the other 13 structural units are all not lower than 0.9 and are marked as 0. The system also uses the complete 14 triplet confidence probability distribution sequences as the confidence output of each structural unit, providing a precise uncertainty quantification basis for subsequent research hypothesis generation and experimental parameter optimization.

[0116] This embodiment utilizes a cross-modal fusion sub-agent and a multi-head cross-attention network to sequentially fuse the glycan three-level encoded multimodal semantic representation vector with features from mass spectrometry, nuclear magnetic resonance, and chromatography. This enables structure prediction to simultaneously utilize the intrinsic structural semantics of the glycan and external experimental observation constraints, overcoming the structural ambiguity problem caused by a single modal information source. The structure prediction sub-agent uses an autoregressive sequence generation model based on a Transformer decoder to predict monosaccharide type, glycosidic bond connection site, and anodic configuration residue by residue from the reducing end to the non-reducing end, achieving end-to-end probabilistic analysis of the complete chemical structure of the glycan. Uncertainty quantification sub-agent maintains Monte Carlo dropout activation and executes the inference process 100 times. Random forward propagation assigns a confidence probability value to each candidate category of each structural unit by statistical sampling frequency, making the reliability of structure prediction measurable for the first time at the monosaccharide residue level and chemical feature dimension. The structure judgment output sub-agent calculates the comprehensive confidence score by taking the minimum value of the maximum confidence score in the three dimensions, and automatically distinguishes high-confidence units from low-confidence units with a threshold of 0.9. Even after adding a label to low-confidence units, the complete confidence distribution sequence is still output. This allows the downstream hypothesis generation and experimental optimization stages to accurately avoid uncertain regions and focus experimental verification resources on the reliable structural parts, thus providing an operable confidence decision basis for the fully unmanned closed loop from spectrum analysis to knowledge discovery.

[0117] It should be further explained that, based on the predicted structure and the confidence level, this embodiment generates at least one research hypothesis on the glycomics knowledge graph, including:

[0118] S301. Obtain the confidence score of each structural unit in the predicted structure of the target glycan. Compare the overall confidence score of all L structural units with a preset confidence threshold one by one. Mark structural units with an overall confidence score lower than the preset confidence threshold as uncertain nodes. Extract the monosaccharide type, glycosidic bond connection site, and anodic configuration information corresponding to the uncertain nodes to form an uncertain structural feature set. In this embodiment, the preset confidence threshold is set to 0.9. L is the total number of monosaccharide residues in the target glycan. The overall confidence score is taken from the overall confidence score of each structural unit output by the structural judgment output sub-agent. The comparison operation starts from the first structural unit and traverses sequentially to the Lth structural unit. For each structural unit, the following judgment logic is executed: if the overall confidence score of the structural unit is greater than or equal to 0.9, the structural unit is marked as a deterministic node; if the overall confidence score of the structural unit is less than 0.9, the structural unit is marked as an uncertain node. For structural units marked as uncertain nodes, the monosaccharide type is extracted from the predicted glycan structure sequence output by the structural prediction sub-agent; the glycosidic bond combination between the structural unit and the preceding monosaccharide residue is extracted from the predicted glycan structure sequence output by the structural prediction sub-agent; and the anodic configuration of the structural unit is extracted from the predicted glycan structure sequence output by the structural prediction sub-agent. These three pieces of information—monosaccharide type, glycosidic bond combination, and anodic configuration—are treated as a single uncertain feature entry. The uncertain structural feature set consists of all uncertain feature entries corresponding to uncertain nodes. This set is stored in list form, where each element is a triplet, recording the position number of the uncertain node in the glycan sequence, the monosaccharide type of the node, the combination of the node's bond sites with the preceding monosaccharide residue, and the anodic configuration.

[0119] S302. Using the monosaccharide type and glycosidic bond connection site combination of the deterministic structural units in the predicted structure as query conditions, retrieve glycan entity nodes containing the same monosaccharide arrangement pattern and the same connection topology in the glycomics knowledge graph. Use the retrieved glycan entity nodes as graph anchors, where the graph anchors are the set of starting nodes for a k-hop random walk. In this embodiment, deterministic nodes are structural units with a comprehensive confidence score greater than or equal to 0.9 among all L structural units. The extraction method for the monosaccharide type and glycosidic bond connection site combination of the deterministic structural units is as follows: from the predicted glycan structure sequence output by the structural prediction agent, extract the monosaccharide type field and connection site combination field of each deterministic structural unit in ascending order of its position number in the glycan sequence, forming a deterministic structural feature sequence. The monosaccharide arrangement pattern is defined as the linear arrangement order of the monosaccharide type field in the deterministic structural feature sequence, and the connection topology is defined as the linear arrangement order of the connection site combination field in the deterministic structural feature sequence. The specific operation for retrieving glycan entity nodes containing the same monosaccharide arrangement pattern and the same connection topology in the glycomics knowledge graph is as follows: Using the monosaccharide arrangement pattern as the first search condition and the connection topology as the second search condition, an exact match search is performed on the set of glycan entity nodes in the glycomics knowledge graph. Glycan entity nodes that simultaneously meet both conditions of identical monosaccharide arrangement pattern and identical connection topology are output as search results. When no glycan entity node simultaneously meets both conditions is found in the search results, the search conditions are gradually relaxed. First, the connection topology condition is relaxed to allow donor or acceptor sites in the connection site combination to differ by one carbon atom. Second, the monosaccharide arrangement pattern condition is relaxed to allow monosaccharide types to be substituted within similarity groups in a preset monosaccharide type set. Similarity grouping refers to dividing the 20 monosaccharide types in the preset monosaccharide type set into four similarity groups based on structural similarity: hexose group, N-acetylglucosamine group, sialic acid group, and uronic acid group. The knowledge graph anchor is a set of retrieved glycan entity nodes, which is stored in the form of a list. Each element in the list is a unique identifier for a glycan entity node and its internal index number in the knowledge graph.

[0120] In this embodiment, the construction process of the glycomics knowledge graph is as follows: Glycan entities and their structural information are extracted from the GlyTouCan international glycan structure database (version number and release date are fixed in specific implementations). The screening criteria are glycan entries in the database that have been manually verified and whose registration status is "confirmed"; enzyme entities and their catalytic activity information are extracted from the UniProt database. The screening criteria are entries with an annotation score greater than or equal to 3 points and experimentally verified enzyme activity annotations; protein entity information is extracted from the UniProt database. The screening criteria are experimentally verified glycan-binding proteins, and their glycan binding specificity information is extracted from the GlycanBindingDB database; disease entities and their association with glycans are extracted through literature mining. Literature abstracts obtained by searching the PubMed database using a combination of specific disease names and glycan structure keywords are used. After manual verification, associated entries supported by statistically significant differential expression evidence are retained; drug entity information is extracted from the DrugBank database. The screening criteria are carbohydrate drugs that have been approved for marketing or have entered clinical trials, and glycan-targeting drugs. The types of relationships between entities were determined as follows: catalytic relationships were determined based on enzymatic experimental data collected in the UniProt database, with enzymatic reaction types such as glycosyltransferases and glycoside hydrolases as specific subtypes; binding relationships were determined based on affinity assay experimental data collected in the GlycanBindingDB database, with a dissociation constant of less than 10 μmol / L as the screening threshold for high-confidence binding relationships; upregulation relationships were determined based on differential expression analysis results from literature mining, with a statistically significant p-value of less than 0.05 and a fold change in expression level greater than 2 as the criteria for identifying upregulation relationships, and a p-value of less than 0.05 and a fold change in expression level less than 0.5 as the criteria for identifying downregulation relationships. The initial relation weights are assigned based on the success rate of experimental verification and the credibility of the evidence for the corresponding relation type: the basic evidence for catalytic relations is in vitro enzymatic reaction experiments, which have a short verification cycle and mature technology, so the initial weight is assigned 0.5; the basic evidence for binding relations is protein interaction experiments such as surface plasmon resonance or isothermal titration calorimetry, which are technically more difficult and have a longer verification cycle than enzymatic reaction experiments, so the initial weight is assigned 0.3; the basic evidence for upregulation relations is cell or animal model experiments, which have the highest verification cost and technical complexity, and the coefficient of variation of experimental results is relatively large, so the initial weight is assigned 0.2. These initial weights are dynamically updated in subsequent closed-loop iterations based on actual experimental verification results. The database version number, retrieval strategy, and screening criteria used in this knowledge graph construction process are fixed during implementation to ensure that those skilled in the art can reconstruct a glycomics knowledge graph matching the technical solution of this invention based on the above steps.

[0121] S303. Obtain all outgoing edges pointing to external entities from the current entity node in the glycomics knowledge graph. Select relationship edges whose edge type belongs to a preset relationship type set as candidate transfer edges. The preset relationship type set includes at least catalytic, binding, and upregulation relationships. Determine the initial transfer weight corresponding to each candidate transfer edge according to a preset edge type weight mapping table. Normalize the initial transfer weights of all candidate transfer edges of the current node, and use the normalized weight values ​​as the sampling probability corresponding to each candidate transfer edge. Based on the sampling probability, randomly select one of the candidate transfer edges as the transfer edge for this walk. Walk along the selected transfer edge from the current entity node to the adjacent entity node, and record the transfer edge type and the identifier of the reached adjacent entity node sequentially in the current path sequence. The adjacent entity node reached is taken as the new current entity node, and the above walking steps are repeated until the preset walking termination condition is met. The walking termination condition includes: the walking is terminated in advance when the current entity node is a terminal entity type node, and the terminal entity type node includes disease entities and drug entities; or the walking is terminated and the current path sequence is discarded when the cumulative walking steps reach the preset maximum walking steps k and the current entity node does not belong to the terminal entity type node. The path sequence that is not discarded is retained as a valid random walking path. Each valid random walking path consists of the identifier of the starting sugar chain entity node, the intermediate entity nodes passed in sequence and the transition edge type, and the identifier of the final terminal entity node reached, wherein the intermediate entity nodes are enzyme entities or protein entities, and the terminal entity node is a disease entity or drug entity.

[0122] In this embodiment, the random walk operation described in S303 is implemented as follows:

[0123] The graph anchor set is a list of sugar chain entity nodes obtained after filtering by both monosaccharide arrangement pattern and connection topology. The preset maximum number of walk steps k is set to 3, meaning that each random walk path undergoes a maximum of 3 transitions. In the preset edge type weight mapping table, the initial transition weight of catalytic relation edges is set to 0.5, the initial transition weight of combination relation edges is set to 0.3, and the initial transition weight of up-adjustment relation edges is set to 0.2, with the sum of the three weight values ​​being 1.0. In this embodiment, the weighting scheme is determined based on the priority of glycomics experimental verification: Catalytic relationship edges directly link to enzymatic reaction experiments. The catalytic relationship between glycans and enzymes is the most common and easiest type of relationship to verify through in vitro enzymatic reaction experiments in glycan function research, therefore it is assigned the highest weight of 0.5; Interaction experiments involving proteins and glycans linked by relationship edges have slightly higher experimental technical difficulty and verification cycle than enzymatic reaction experiments, therefore they are assigned the second highest weight of 0.3; The regulatory effect of glycans linked by relationship edges on protein expression or disease biomarker levels usually requires cell or animal model support for experimental verification, resulting in the highest verification cost and technical complexity, therefore it is assigned the lowest weight of 0.2. In this embodiment, the preset maximum number of steps k is set to 3 because the biological function linkage in the glycomics knowledge graph, starting from glycan entities, passing through intermediate nodes such as enzyme or protein entities, and finally reaching disease or drug entities, typically completes the transfer of key nodes in no more than three steps in known classical regulatory mechanisms. For example, the typical pathway of glycan synthesis catalyzed by glycosyltransferases and then regulating disease biomarker expression through protein binding can be achieved within two steps. Limiting the transition to a maximum of three steps ensures that the hypothesis path generated by the random walk fully covers both direct and indirect functional associations with clear biological significance, while effectively suppressing redundant intermediate entities and noisy associations introduced by excessively long paths, thus achieving the optimal balance between the simplicity and experimental verifiability of the hypothesis mechanism.

[0124] Using each sugar chain entity node in the aforementioned graph anchor set as the starting node, 50 independent random walks are performed. The specific walk process is as follows:

[0125] The current entity node is initialized as the starting glycan entity node. All outgoing edges pointing to external entities from the current entity node in the glycomics knowledge graph are obtained. Edges belonging to the catalytic, binding, and upregulation relationships are selected as candidate transfer edges. The composition of the candidate transfer edges varies depending on the type of the current entity node: when the current entity node is a glycan entity, its outgoing edges only include catalytic relationship edges pointing to enzyme entities and binding relationship edges pointing to protein entities. After normalizing the initial transfer weights of these two candidate transfer edges, the sampling probability corresponding to the catalytic relationship edge is 0.5 divided by 0.5 plus 0.3, which equals 0.625; the sampling probability corresponding to the binding relationship edge is 0.3 divided by 0.5 plus 0.3, which equals 0.375. When the current entity node is an enzyme entity, its outgoing edges include upregulation edges pointing to protein entities. For relational edges and associative relational edges, the normalized sampling probability of the upward relational edge is 0.2 divided by 0.2 plus 0.3, which equals 0.4, and the sampling probability of the associative relational edge is 0.3 divided by 0.2 plus 0.3, which equals 0.6. When the current entity node is a protein entity, its outgoing edges include upward relational edges pointing to disease entities and associative relational edges pointing to drug entities. After normalization, the sampling probability of the upward relational edge is 0.2 divided by 0.2 plus 0.3, which equals 0.4, and the sampling probability of the associative relational edge is 0.3 divided by 0.2 plus 0.3, which equals 0.6.

[0126] After determining the sampling probability corresponding to each candidate transition edge, a random number following a uniform distribution is generated within the interval of 0 to 1. The sampling probabilities of each candidate transition edge are sequentially accumulated to form a cumulative probability interval. Based on the cumulative probability interval into which the random number falls, a candidate transition edge is selected as the transition edge for this walk. The walk proceeds along the selected transition edge from the current entity node to the adjacent entity node, and the type of the transition edge and the identifier of the adjacent entity node reached are sequentially recorded in the current path sequence. The adjacent entity node reached during the walk is then used as the new current entity node, and the above walking steps are repeated.

[0127] When a walk reaches an entity node that is a disease entity or a drug entity, the early termination condition in the walk termination conditions is met, and the walk path is terminated but retained. When the cumulative number of walk steps reaches the preset maximum number of walk steps k equal to 3, and the current entity node is still an enzyme entity or a protein entity, the invalid termination condition in the walk termination conditions is met, the walk is terminated, and the current path sequence is discarded.

[0128] Each successfully retained valid random walk path is recorded in a structured sequence format. The record format is as follows: the identifier of the starting sugar chain entity node, followed by alternating records of the relation edge type along each step and the identifier of the intermediate entity node reached. The last element of the sequence is the identifier of the final entity node. After performing 50 independent random walks on each of the starting nodes, all valid random walk paths that have not been discarded are summarized into a candidate path set. After removing completely duplicate paths, these are used as input for the hypothetical path filtering and sorting operation described in S304.

[0129] S304. Filter all generated random walk paths, remove paths containing monosaccharide type or connection site information of any uncertain node in the uncertain structural feature set, sort the remaining paths in ascending order of the number of hops between the end entity node and the starting sugar chain entity node, sort paths with the same number of hops in descending order of the product of the weights of each edge type in the path, and select the top N paths as candidate hypothesis paths, where N is the preset number of hypotheses generated.

[0130] In this embodiment, the path filtering operation described in S304 specifically includes:

[0131] Obtain the candidate path set generated by S303 and the uncertain structural feature set generated by S301; wherein, the uncertain structural feature set stores the monosaccharide type, glycosidic bond connection site combination and anodic configuration information of all structural units marked as uncertain nodes in the form of a list, and each element in the list is an uncertain feature entry in the form of a triplet.

[0132] For each random walk path in the candidate path set, perform the following filtering judgment:

[0133] Extract the starting glycan entity node of the random walk path, and query the complete glycan structure information corresponding to the starting glycan entity node in the glycomics knowledge graph. The complete glycan structure information includes the monosaccharide type sequence and the combination sequence of all monosaccharide residues of the glycan represented by the glycan entity node.

[0134] Each uncertain feature entry in the uncertain structural feature set is traversed. The monosaccharide type in the current uncertain feature entry is compared with the monosaccharide type sequence, and the glycosidic bond linkage site combination in the current uncertain feature entry is compared with the linkage site combination sequence. If, in the complete glycan structure information, there exists a structural segment that simultaneously satisfies that its monosaccharide type is consistent with the monosaccharide type of the current uncertain feature entry and that its linkage site combination is consistent with the glycosidic bond linkage site combination of the current uncertain feature entry, then the random walk path is determined to contain uncertain structural features, and the path is removed from the candidate path set. If, after traversing all uncertain feature entries in the uncertain structural feature set, no structural segment that simultaneously satisfies the above two consistency conditions is found in the complete glycan structure information, then the random walk path is determined not to contain uncertain structural features and is retained.

[0135] In this embodiment, uncertain nodes correspond to structural units in the glycan structure with a predicted confidence level lower than a preset threshold, indicating significant uncertainty in the structural identification results of this region. If a candidate hypothesis path relies on a graph anchor glycan with the same monosaccharide type and connection mode as the uncertain node at the corresponding structural position, then the structural information of the anchor glycan in that region is also unreliable. Functional association hypotheses derived from unreliable structural information lack a credible experimental verification basis and should therefore be eliminated. In this embodiment, the technical reasoning logic based on the filtering operation is as follows: the biological function of a glycan is determined by its precise chemical structure. When the predicted confidence level of a certain structural region of the target glycan is lower than a preset threshold, it indicates significant uncertainty in the identification results of the monosaccharide type, glycosidic bond connection site, and anomeric configuration in that region. If an anchor glycan in the knowledge graph has the same monosaccharide type and connection mode as the aforementioned uncertain node at the corresponding structural position, then the structural information of the anchor glycan in that region corresponds to the structural uncertainty of the target glycan. Since the true configuration of the structure in this region cannot be determined, the functional association hypothesis of the sugar chain derived from this unreliable structural region lacks a credible chemical structural basis, and its experimental verification results cannot be accurately attributed to the functional contribution of specific structural units. Therefore, pre-eliminating graph paths containing uncertain structural features during the hypothesis generation stage ensures that the final output research hypotheses are all based on high-confidence structural units with sufficiently reliable structural identification. This allows subsequent Bayesian optimization and automated experimental verification resources to focus on predicting clear structure-function associations, avoiding invalid experimental verification and erroneous knowledge graph updates caused by the transmission of structural identification uncertainty to the hypothesis generation stage.

[0136] The path sorting operation is divided into two levels: The first level sorting is based on the number of hops between the terminal entity node and the starting sugar entity node. The number of hops is defined as the number of edges in the path. Sorting by the number of hops from smallest to largest means prioritizing the hypothesis that connects the sugar entity to the disease entity or drug entity with the shortest path. Short path hypotheses have more direct biological mechanisms and fewer variables for experimental verification. The second level sorting is performed between paths with the same number of hops. The sorting is based on the product of the edge type weights in the path. The edge type weights are 0.5 for catalytic relations, 0.3 for binding relations, and 0.2 for upregulation relations. The product of the edge type weights is calculated by multiplying the type weights of the edges along each step of the path. For example, if a path contains two steps, the first step is a step along the catalytic relation edge with a weight of 0.5, and the second step is a step along the binding relation edge with a weight of 0.3. Then the product of the edge type weights of this path is 0.5 multiplied by 0.3, which equals 0.15. Sort the edge type weights from largest to smallest, which means prioritizing the hypothesis path with higher experimental verification priority and more mature technical route. The larger the weight product, the higher the proportion of high-weight edge types in the path, and the easier the corresponding experimental verification scheme is to implement.

[0137] N represents the preset number of hypotheses to generate, set to 5. This value balances the efficiency of subsequent Bayesian optimization in searching the experimental parameter space with the diversity of hypotheses. Too small an N value limits the breadth of hypothesis exploration, while too large an N value increases the burden on downstream experimental design and validation. When selecting the top N paths, priority is given to the first-level hop count, followed by the second-level weight product, until N paths are selected or no more paths are available. The selected N paths are output as candidate hypothesis paths and passed to subsequent research hypothesis structure generation.

[0138] S305. For each candidate hypothesis path, the glycan structure corresponding to the starting glycan entity node is taken as the known premise of the hypothesis. The disease entity or drug entity at the end of the path is taken as the glycan functional association entity to be experimentally verified. The enzyme reaction condition parameter space corresponding to the enzyme entities and protein entities along the path is taken as the optimizable experimental parameter space. The enzyme reaction type corresponding to the catalytic relationship edge, the protein binding target corresponding to the binding relationship edge, and the regulatory direction corresponding to the upregulation relationship edge in the path are combined to form the relationship type to be verified of the hypothesis, generating a research hypothesis. The research hypothesis includes the glycan functional association entity to be experimentally verified and the optimizable experimental parameter space. In this embodiment, the research hypothesis strictly follows the scientific discovery logic of "statistical association first, causal verification second," and does not directly equate the association relationship in the knowledge graph with the causal relationship. In the knowledge graph, the upregulation relationship edge between glycan entities and disease entities represents the following meaning: In existing literature or databases, there is statistical evidence that changes in the expression level or content of this glycan are significantly positively correlated with certain biomarkers or pathological indicators of the disease. The statistical significance p-value of this association is less than 0.05 and the fold change is greater than 2. However, whether the biological effect is due to changes in the glycan leading to disease progression or disease state leading to changes in glycan expression, and the specific molecular regulatory mechanism, have not yet been experimentally verified. Therefore, the structured format of the research hypothesis is: "Based on the complete structure of glycan X, existing statistical association evidence indicates a significant positive correlation between this glycan and disease Y. Its nature and causal direction need to be determined by subsequent experiments. This hypothesis aims to verify, under recommended experimental parameters, whether this glycan has a regulatory effect on relevant biomarkers of disease Y through enzyme-catalyzed reaction experiments, protein binding experiments, or cell function experiments." The "Relationship Type to be Verified" field in the structured dictionary of the research hypothesis is clearly marked with two tags: "Relationship Nature to be Verified" and "Causal Direction to be Verified," to distinguish between known statistical associations and biological causal relationships to be verified. This hypothesis formulation shifts the goal of subsequent Bayesian optimization and automated experimental verification from "confirming a pre-set causal assertion" to "testing whether there is a real biological causal relationship behind a statistical association." This fundamentally avoids a large number of false positive verification results and waste of experimental resources caused by misjudging correlation as causation, ensuring that the new knowledge ultimately incorporated into the glycomics knowledge graph has a causal evidence basis supported by experimental verification, rather than merely an accumulation of statistical artifacts.

[0139] In this embodiment, the research hypothesis structuring generation operation described in S305 specifically includes:

[0140] Obtain the N candidate hypothesis paths output by S304, where N is 5. For each of the N candidate hypothesis paths, perform the following operations to generate a corresponding research hypothesis:

[0141] Using the first node identifier in the path string of the current candidate hypothesis path as the query key, the corresponding glycan entity node is retrieved in the glycomics knowledge graph; the complete GlycoCT format glycan structure string, total number of monosaccharide residues, monosaccharide type sequence, and connection site combination sequence are extracted from the attributes of the retrieved glycan entity node; the extracted information is encapsulated into the known premise field of the current research hypothesis, which is used to record the glycan molecular structure on which the research hypothesis is based, providing a precise chemical structural basis for the experimental verification of the research hypothesis.

[0142] The extraction method for disease or drug entities at the end of the path is as follows: using the identifier of the last node in the path string as the query key, the terminal entity node is retrieved in the glycomics knowledge graph. If the node is a disease entity, the disease name, disease ontology identifier, and list of known related biomarkers stored in the node's attributes are extracted. If the node is a drug entity, the drug name, drug chemical structure, list of known drug targets and indications stored in the node's attributes are extracted. The above information is encapsulated into the hypothetical glycan functional association entity field to be experimentally verified.

[0143] The construction method for the enzyme reaction condition parameter space corresponding to the enzyme and protein entities traversed by the path is as follows: traverse all intermediate entity nodes in the path string except for the starting and ending nodes. For each intermediate entity node, use the node identifier as the query key to search in the glycomics knowledge graph. If the node is an enzyme entity, extract the enzyme classification number, optimal pH range, optimal temperature range, cofactor requirement list, and known substrate specificity constants stored in the enzyme entity's attributes. Use the optimal pH range as the search boundary for the pH parameter space, the optimal temperature range as the search boundary for the temperature parameter space, the cofactor concentration range in the cofactor requirement list as the search boundary for the cofactor concentration parameter space, and the Michaelis constant range in the substrate specificity constants as the search boundary for the substrate concentration parameter space. Combine these search boundaries to form the enzyme reaction condition parameter subspace for the enzyme entity. If the node is a protein entity, extract the protein binding buffer pH range, binding reaction temperature range, and protein to ligand molar ratio range stored in the protein entity's attributes. Combine these range parameters to form the protein binding condition parameter subspace for the protein entity. The union of the parameter subspaces of all intermediate entity nodes traversed along the path is used as the optimizable experimental parameter space for this hypothesis. This optimizable experimental parameter space is stored in the form of a structured dictionary. The keys of the dictionary are parameter names, including at least one of pH value, temperature, cofactor concentration, substrate concentration, and protein-ligand molar ratio. The values ​​of the dictionary are triples containing a minimum value, a maximum value, and a search step size. The minimum value is taken from the lower boundary of the corresponding parameter subspace, the maximum value is taken from the upper boundary of the corresponding parameter subspace, and the search step size is set to the maximum value minus the minimum value divided by 20.

[0144] The specific method for combining the enzyme reaction type corresponding to the catalytic relationship edge, the protein binding target corresponding to the binding relationship edge, and the regulatory direction corresponding to the upregulation relationship edge in the path into the hypothesis's unverified relationship type is as follows: Parse the relationship edge type along each step of the path string. If the step moves along the catalytic relationship edge, extract the enzyme reaction type name stored in the catalytic relationship edge attribute, including one of glycosyltransferase reaction, glycoside hydrolysis reaction, glycosyl oxidation reaction, and glycosyl phosphorylation reaction. If the step moves along the binding relationship edge, extract the protein binding target name and dissociation constant range stored in the binding relationship edge attribute. If the step moves along the upregulation relationship edge, extract the regulatory direction stored in the upregulation relationship edge attribute, including one of positive upregulation and negative inhibition, as well as the effect size range of the regulatory effect. Arrange the type information of all relationship edges in the path according to the path transfer order into a list of unverified relationship types. Each element in the list is a triple containing the relationship type, participating entity, and quantitative parameter range.

[0145] A complete research hypothesis consists of four fields: a known premise field, a glycan functional association entity field to be experimentally verified, an optimizable experimental parameter space field, and a relation type field to be verified. These fields are stored and output in a structured dictionary format. The generated research hypothesis is then passed to the Bayesian optimization operation described in S4. The glycan functional association entity field and the relation type field are used to determine the objective function definition in the Bayesian optimization, and the optimizable experimental parameter space field is used to determine the experimental parameter search domain in the Bayesian optimization.

[0146] This embodiment transforms the confidence distribution of structural predictions into a set of uncertain structural features through confidence graph mapping. This allows the hypothesis generation process to accurately avoid glycan structural regions with low prediction reliability, thus avoiding the risk of deriving functional association hypotheses based on unreliable structural information. By using knowledge graph anchor point positioning, high-confidence structural units are used as query conditions to retrieve structurally similar known glycan entities in the glycomics knowledge graph, ensuring that hypothesis generation is based on experimentally traceable structures. Furthermore, a weighted random walk along catalytic, binding, and upregulation edges, with a probability sampling mechanism of 0.5 for catalytic edges, 0.3 for binding edges, and 0.2 for upregulation edges, guides the hypothesis exploration direction to prioritize experimental verification. The most feasible enzymatic reaction path is identified, thus overcoming the limitation of existing technologies that can only retrieve and combine historical documents, and achieving cross-entity type association discovery beyond the boundaries of known knowledge. Hypothesis path filtering and sorting, using the product of hop count and edge type weights as a dual sorting criterion, ensures that the output candidate hypotheses possess both the simplicity of biological mechanisms and the quantifiable priority of experimental verification. Through structured generation of research hypotheses, random walk paths are automatically transformed into complete research hypotheses containing known premises, glycan functional association entities to be experimentally verified, an optimizable experimental parameter space, and the relationship type to be verified. This provides a precise domain of objective function and boundary of parameter search space for downstream Bayesian optimization, enabling seamless integration between hypothesis generation and experimental design. In this embodiment, the preset weights for catalytic relationship edges (0.5), binding relationship edges (0.3), and upsetting relationship edges (0.2) are based on the fact that the three relationship types have significant differences in technical maturity, implementation cost, and success rate in glycomics experimental verification, requiring differentiated assignment based on verification priority. Among the various experimental approaches, the catalytic relationship edge corresponds to the enzyme-catalyzed reaction experiment. This experimental system is the most mature under in vitro conditions, with a short verification cycle and high reproducibility, therefore it is assigned the highest weight of 0.5. The protein-glycan interaction experiment corresponding to the relationship edge has a slightly higher technical difficulty and verification cycle than the enzyme-catalyzed reaction experiment, therefore it is assigned the second highest weight of 0.3. The verification of the regulatory effect of upregulating glycans on protein expression or disease biomarker levels usually requires cell or animal models, has the highest verification cost and the greatest technical complexity, therefore it is assigned the lowest weight of 0.2. This weighting scheme ensures that the random walk prioritizes sampling along the catalytic relationship direction with the highest experimental feasibility when generating hypotheses, ensuring that the final research hypothesis can be verified in a closed loop using the most convenient experimental methods, thereby improving the overall efficiency of the glycomic agent from hypothesis generation to experimental confirmation.

[0147] It should be further explained that this embodiment uses the experimental verification index corresponding to the predicted structure as the objective function, and uses Bayesian optimization to generate a recommended set of experimental parameters, including:

[0148] S401. Each parameter name in the optimizable experimental parameter space is used as a decision variable. The continuous multidimensional space spanned by all decision variables is used as the search domain for Bayesian optimization. The quantitative measurement index of the correlation strength verification experiment between the predicted glycan structure and the glycan functional association entity to be verified is used as the objective function. The quantitative measurement index is one of the following: enzyme-catalyzed reaction product generation rate, protein binding dissociation constant, or fold change in the expression level of disease markers. In this embodiment, the research hypothesis is stored in the form of a structured dictionary, containing four fields: known premise field, glycan functional association entity to be verified field, optimizable experimental parameter space field, and relationship type to be verified field. The optimizable experimental parameter space field is extracted as follows: the optimizable experimental parameter space field in the research hypothesis structured dictionary is read. This field is a dictionary with parameter name as key and triples containing minimum value, maximum value, and search step size as value. Each key in this dictionary is used as a decision variable. The name of the decision variable is completely consistent with the parameter name. The parameter name includes at least one of pH value, temperature, cofactor concentration, substrate concentration, and protein ligand molar ratio. By using the minimum value of each decision variable in its corresponding triplet as the lower bound and the maximum value as the upper bound, a continuous multidimensional space with a dimension equal to the number of decision variables is formed. This continuous multidimensional space is the search domain of Bayesian optimization. The number of dimensions of the search domain is equal to the number of decision variables. For example, when the optimizable experimental parameter space contains four parameters—pH, temperature, cofactor concentration, and substrate concentration—the search domain is a 4-dimensional continuous space. Its first dimension is pH, ranging from the minimum to the maximum value; the second dimension is temperature, ranging from the minimum to the maximum value; the third dimension is cofactor concentration, ranging from the minimum to the maximum value; and the fourth dimension is substrate concentration, ranging from the minimum to the maximum value.

[0149] In this embodiment, for discrete decision variables in the optimizable experimental parameter space, such as enzyme type and buffer type, which cannot take arbitrary values ​​within a continuous numerical range, the following discrete variable continuousization processing strategy is adopted: First, a unique integer index is assigned to each candidate value of the discrete decision variable, with the index numbered continuously from 0 to the total number of candidate values ​​minus 1; then, the integer index is converted into a one-hot vector using one-hot encoding, the dimension of which is equal to the total number of candidate values, and only the element corresponding to the index position in the vector is 1, while the rest are 0; next, the one-hot vector is input into a trainable embedding matrix, the number of rows of which is the total number of candidate values, and the number of columns is the embedding dimension required for the unified representation of continuous variables in the search domain. After matrix multiplication, a continuous embedding vector is output, so that different values ​​of the discrete variable are mapped to different positions in the continuous vector space, and the semantic similarity between different values ​​is reflected by the cosine similarity between the embedding vectors. In the search domain construction process of Bayesian optimization, the original discrete variable values ​​are replaced by this continuous embedding vector, which together with the other continuous decision variables spans a unified continuous multidimensional search space. During the optimization iteration process, the Gaussian process surrogate model directly searches and recommends candidate experimental parameter combinations in this continuous embedding space. When it is necessary to restore the recommended experimental parameter combinations to discrete parameter values ​​that can be executed by an automated experimental platform, the cosine similarity between the recommended embedding vector and the corresponding candidate value embedding vector in each row of the embedding matrix is ​​calculated one by one, and the candidate value with the largest cosine similarity is selected as the final recommended value for the discrete decision variable. This approach enables Bayesian optimization to handle both continuous and discrete experimental parameters within a unified continuous space framework, without the need to introduce an additional classification Bayesian optimization algorithm, ensuring complete coverage of the experimental parameter space and optimization efficiency.

[0150] The extraction method for glycan functional association entities to be experimentally verified is as follows: The glycan functional association entity field to be experimentally verified is read from the structured dictionary of the research hypothesis. This field contains the type identifier and attribute information of the terminal entity node; the type identifier is either a disease entity or a drug entity. The extraction method for predicted glycan structures is as follows: The known premise field is read from the structured dictionary of the research hypothesis. This field contains the predicted glycan structure sequence and confidence probability distribution sequence output by the structure judgment output sub-agent.

[0151] The type of the objective function is determined jointly by the type identifier of the glycan functional association entity to be verified and the type of relationship to be verified. When the glycan functional association entity to be verified is a drug entity and the type of relationship to be verified includes the enzyme-catalyzed reaction type corresponding to the catalytic relationship edge, the quantitative indicator is selected as the enzyme-catalyzed reaction product generation rate, in micromoles per liter per minute. When the glycan functional association entity to be verified is a disease entity and the type of relationship to be verified includes the protein binding target corresponding to the binding relationship edge, the quantitative indicator is selected as the protein binding dissociation constant, in nanomoles per liter. When the glycan functional association entity to be verified is a disease entity and the type of relationship to be verified includes the regulatory direction corresponding to the upregulation relationship edge, the quantitative indicator is selected as the fold change in the expression level of the disease biomarker, in dimensionless units. The input of the objective function is a vector of experimental parameter combinations in the search domain, with the vector dimension equal to the number of decision variables, and the output of the objective function is the scalar value of the quantitative indicator.

[0152] S402. Randomly sample M initial experimental parameter combinations within the search domain, using these M initial experimental parameter combinations as the initial sampling point set. Calculate the objective function observation value corresponding to each initial sampling point. Use the M initial sampling points and their corresponding objective function observation values ​​to form an initial training dataset. Use the initial training dataset to fit a Gaussian process surrogate model. The mean function of the Gaussian process surrogate model is a zero function, and the covariance function of the Gaussian process surrogate model is the sum of the radial basis function kernel and the constant kernel. The initial value of the length scale parameter of the radial basis function kernel is set to 1 / 10 of the range of values ​​for each dimension of the search domain, and the initial value of the variance parameter of the constant kernel is set to 1.0.

[0153] In this embodiment, M is set to 10. This value balances the dimensionality of the search domain with the sufficiency of the initial exploration: when the search domain has 4 dimensions, 10 initial sampling points can provide an average coverage density of approximately 2.5 sampling points per dimension, providing sufficient information for the Gaussian process surrogate model to initially infer the overall trend and local variation characteristics of the objective function within the search domain. Random sampling employs the Latin hypercube sampling method, specifically dividing each dimension of the search domain into M equal-probability intervals. A value is randomly selected from each equal-probability interval, and the sampled values ​​from each dimension are randomly arranged and combined to form M initial experimental parameter combinations. This method ensures that the initial sampling points are unique and have uniform coverage across all dimensions of the search domain. Each initial experimental parameter combination is represented by a real-number vector with the same number of dimensions as the number of decision variables. The first element of the vector corresponds to the sampled value of the first decision variable, the second element corresponds to the sampled value of the second decision variable, and so on.

[0154] The specific method for calculating the objective function observation value corresponding to each initial sampling point is as follows: Input each initial experimental parameter combination vector into the objective function defined in S401. When the quantitative indicator is the rate of enzyme-catalyzed reaction product generation, the objective function observation value is obtained by performing an enzyme-catalyzed reaction experiment on an automated experimental platform according to the pH, temperature, cofactor concentration, and substrate concentration conditions set by the initial experimental parameter combination, and measuring the increase in product concentration per unit time, in micromoles per liter per minute. When the quantitative indicator is the protein binding dissociation constant, the objective function observation value is obtained by performing a surface plasmon resonance experiment on an automated experimental platform according to the pH, temperature, and protein-ligand molar ratio conditions set by the initial experimental parameter combination, and measuring the ratio of binding to dissociation rate constants, in nanomoles per liter. When the quantitative indicator is the fold change in disease biomarker expression, the objective function observation value is obtained by performing an enzyme-linked immunosorbent assay (ELISA) or a quantitative real-time polymerase chain reaction (qPCR) on an automated experimental platform after treating the cell model according to the conditions set by the initial experimental parameter combination, and measuring the ratio of biomarker expression levels between the treatment group and the control group, in dimensionless units.

[0155] The initial training dataset is stored in a two-element tuple format. The first element is an input matrix with M rows and D columns, where M equals 10 (the number of sampling points) and D is the dimension of the search domain, i.e., the number of decision variables. Each row stores an initial experimental parameter combination vector. The second element is a target function observation vector of length M, where the i-th element corresponds to the target function observation of the initial experimental parameter combination vector in the i-th row of the input matrix. The fitting process of the Gaussian process surrogate model is as follows: the input matrix and target function observation vector from the initial training dataset are input into the Gaussian process regression algorithm. Under the optimization objective of maximizing the log marginal likelihood, the L-BFGS-B optimizer is used to iteratively update the covariance function hyperparameters, including the length scaling parameter of the radial basis function kernel and the variance parameter of the constant kernel. The maximum number of iterations of the L-BFGS-B optimizer is set to 1000, and the convergence tolerance is set to 0.00001.

[0156] In this embodiment, the radial basis function kernel K_RBF is mathematically expressed as follows: K_RBF equals the variance parameter multiplied by an exponential function, where the exponent term of the exponential function is -0.5 multiplied by the square of the Euclidean distance between the two input points scaled by a length scale parameter. The length scale parameter is a positive real vector with a length equal to the number of dimensions D of the search domain. The initial value of the j-th element in this length scale parameter vector is set to the difference between the upper and lower bounds of the range of values ​​in the j-th dimension of the search domain, divided by 10. The constant kernel K_constant is mathematically expressed as follows: K_constant equals the variance parameter, which is a positive real scalar with an initial value of 1.0. The covariance function K_total is the sum of the radial basis function kernel K_RBF and the constant kernel K_constant, i.e., K_total equals K_RBF plus K_constant. This summation allows the Gaussian process surrogate model to simultaneously capture the nonlinear variation of the objective function and the overall mean shift. The mean function is a zero function, meaning that, without any prior observations of the objective function, the prior expected value of the objective function at any input point in the search domain is 0.

[0157] S403. Obtain the predicted mean function and predicted standard deviation function of the current Gaussian process surrogate model in the search domain; determine the current optimal objective function observation value from all experimental parameter combinations that have completed objective function evaluation; generate a set of unevaluated candidate experimental parameter combinations in the search domain, and perform the following operations for each candidate experimental parameter combination in the set of unevaluated candidate experimental parameter combinations:

[0158] S404. Input the candidate experimental parameter combination into the prediction mean function to obtain the prediction mean corresponding to the candidate experimental parameter combination; input the candidate experimental parameter combination into the prediction standard deviation function to obtain the prediction standard deviation corresponding to the candidate experimental parameter combination.

[0159] S405. Calculate the difference between the predicted mean corresponding to the candidate experimental parameter combination and the observed value of the current optimal objective function; based on the difference and the predicted standard deviation corresponding to the candidate experimental parameter combination, substitute them into the analytical formula of the expected improvement function to calculate the expected improvement value of the candidate experimental parameter combination on the observed value of the current optimal objective function.

[0160] S406. After traversing and calculating the expected improvement value corresponding to each candidate experimental parameter combination in the set of unevaluated candidate experimental parameter combinations, with the goal of maximizing the expected improvement value, select the candidate experimental parameter combination with the largest expected improvement value from the set of unevaluated candidate experimental parameter combinations as the next recommended experimental parameter.

[0161] In this embodiment, the method for determining the current optimal objective function observation value in S403 is as follows: extract all M objective function observation values ​​from the objective function observation value vector of the initial training dataset. If the optimization direction of the objective function is maximization, then the maximum value among the M objective function observation values ​​is taken as the current optimal objective function observation value; if the optimization direction of the objective function is minimization, then the minimum value among the M objective function observation values ​​is taken as the current optimal objective function observation value. Taking the N-linked two-antenna complex glycan of the Fc region of human serum immunoglobulin G as a specific example, the objective function is the fold change in the expression level of the disease marker, and the optimization direction is maximization. The 10 objective function observation values ​​obtained in S402 are 3.2, 1.8, 4.5, 2.1, 3.9, 2.7, 1.3, 4.1, 2.5, and 3.6 respectively. Therefore, the current optimal objective function observation value is taken as the maximum value of 4.5.

[0162] The predicted mean function and predicted standard deviation function of the Gaussian process surrogate model are determined based on the standard derivation formula of Gaussian process regression. For any unevaluated candidate experimental parameter combination within the search domain, the predicted mean function gives a mean equal to the covariance vector between the candidate experimental parameter combination and the M sampled points multiplied by the inverse of the sampled point covariance matrix, and then multiplied by the objective function observation vector; the predicted standard deviation function gives a standard deviation equal to the square root of the predicted variance, which is equal to the prior variance of the candidate experimental parameter combination itself minus the covariance vector between the candidate experimental parameter combination and the M sampled points multiplied by the inverse of the sampled point covariance matrix, and then multiplied by the transpose of the covariance vector.

[0163] In this embodiment, the expected improvement function of S405 is calculated as follows: For each candidate experimental parameter combination, the difference between the predicted mean of the candidate experimental parameter combination and the observed value of the current optimal objective function is calculated. If the optimization direction is maximization, the difference is equal to the predicted mean minus the observed value of the current optimal objective function; if the optimization direction is minimization, the difference is equal to the observed value of the current optimal objective function minus the predicted mean. When the predicted standard deviation of the candidate experimental parameter combination is greater than 0, the difference is calculated by dividing the predicted standard deviation to obtain the standardized improvement amount. The expected improvement value of the candidate experimental parameter combination is equal to the difference multiplied by the cumulative distribution function of the standard normal distribution at the standardized improvement amount plus the predicted standard deviation multiplied by the probability density function of the standard normal distribution at the standardized improvement amount. When the predicted standard deviation of the candidate experimental parameter combination is equal to 0, if the difference is greater than 0, the expected improvement value is equal to the difference; if the difference is less than or equal to 0, the expected improvement value is equal to 0. The cumulative distribution function of the standard normal distribution is numerically approximated using the rational function approximation formula of Abramowitz and Stegun, and the probability density function of the standard normal distribution is calculated using the analytical formula of the standard normal distribution.

[0164] The search method for maximizing the expected improvement value described in S406 is as follows: 10,000 candidate experimental parameter combinations are randomly generated within the search domain in a uniform distribution. The expected improvement value is calculated for each of the 10,000 candidate experimental parameter combinations. A complete traversal search strategy is used to select the candidate experimental parameter combination with the largest expected improvement value as the next recommended experimental parameter. The output format of the recommended experimental parameter is a real number vector with the same number of dimensions as the number of decision variables. The first element of this vector corresponds to the recommended value of the first decision variable, the second element corresponds to the recommended value of the second decision variable, and so on.

[0165] S407. The recommended experimental parameters are used as new sampling points. The observed values ​​of the objective function corresponding to the new sampling points are obtained. The new sampling points and their observed values ​​of the objective function are added to the initial training dataset to form an updated training dataset. The updated training dataset is used to refit the Gaussian process surrogate model and update the covariance function hyperparameters of the Gaussian process surrogate model. The covariance function hyperparameters include the length scaling parameter of the radial basis function kernel and the variance parameter of the constant kernel. In this embodiment, the specific form of the new sampling points is the recommended experimental parameter vector obtained after the expected improvement function maximization search in S403. The dimension of this vector is equal to the number of decision variables, and the value of each element in the vector is between the lower and upper bounds of the corresponding search domain dimension. The specific method of obtaining the observed values ​​of the objective function corresponding to the new sampling points is completely consistent with the method of calculating the observed values ​​of the objective function of the initial sampling points in S402: the new sampling point vector is input into the objective function defined in S401, and the corresponding correlation strength verification experiment is performed on the automated experimental platform according to all the parameter conditions set for the new sampling points to determine the observed values ​​of the objective function. When the quantitative indicator is the rate of product formation in an enzyme-catalyzed reaction, the objective function observation value is obtained by measuring the product formation rate in an enzyme-catalyzed reaction experiment performed on an automated experimental platform according to the pH, temperature, cofactor concentration, and substrate concentration conditions set at the newly added sampling points. The unit is micromoles per liter per minute. When the quantitative indicator is the protein binding dissociation constant, the objective function observation value is obtained by measuring the ratio of the binding to dissociation rate constants in a surface plasmon resonance experiment performed on an automated experimental platform according to the conditions set at the newly added sampling points. The unit is nanomoles per liter. When the quantitative indicator is the fold change in the expression level of a disease biomarker, the objective function observation value is obtained by measuring the ratio of the biomarker expression levels in the treated group to the control group in an enzyme-linked immunosorbent assay or a quantitative real-time polymerase chain reaction (qPCR) experiment performed on an automated experimental platform after treating the cell model according to the conditions set at the newly added sampling points. The unit is dimensionless.

[0166] The updated training dataset is formed as follows: The input matrix in the initial training dataset is expanded from M rows to M+1 rows, where M is 10. The vector of newly added sampling points is stored in the M+1th row of the expanded input matrix, with the first to D elements corresponding to the values ​​of the D decision variables. The objective function observation vector in the initial training dataset is expanded from length M to M+1, and the objective function observation value corresponding to the newly added sampling point is stored in the M+1th element position of the expanded objective function observation vector. The updated training dataset is stored in the same two-element tuple format as the initial training dataset, with the first element being the input matrix of length M+1 multiplied by D columns, and the second element being the objective function observation vector of length M+1.

[0167] The specific method for refitting the Gaussian process surrogate model using the updated training dataset is exactly the same as the method for fitting the Gaussian process surrogate model in S402: The input matrix and objective function observation vector from the updated training dataset are input into the Gaussian process regression algorithm. Under the optimization objective of maximizing the logarithmic marginal likelihood, the L-BFGS-B optimizer is used to iteratively update the covariance function hyperparameters. The maximum number of iterations for the L-BFGS-B optimizer is set to 1000, and the convergence tolerance is set to 0.00001. During the refitting process, the initial values ​​of the covariance function hyperparameters are no longer the initial settings from S402, but rather the covariance function hyperparameter values ​​obtained from the previous fitting are used as the starting point for this fitting. That is, the length scaling parameter vector of the radial basis function kernel and the variance parameter of the constant kernel are iteratively updated starting from the convergence value of the previous fitting. After refitting, the covariance function hyperparameters of the Gaussian process surrogate model are updated. The values ​​of each dimension in the length scale parameter vector of the radial basis function kernel reflect the change in the smoothness of the objective function along each dimension after incorporating the information of the new sampling points. The variance parameter of the constant kernel reflects the change in the overall noise level of the objective function after incorporating the information of the new sampling points.

[0168] S408. Repeat S403 and S407 until a preset convergence condition is met. The preset convergence condition is that the maximum value of the expected improvement function in T consecutive iterations is lower than a predetermined expected improvement threshold, where T is 5. The predetermined expected improvement threshold is set to 1 / 100 of the current optimal objective function observation value. When the preset convergence condition is met, the iteration stops, and the experimental parameter combination corresponding to the current optimal objective function observation value is output as a recommended experimental parameter set. In this embodiment, the iteration count is initialized to 0 before entering S403 for the first time, and incremented by 1 after each complete execution of S403 and S404. The term "T consecutive iterations" means backtracking T iterations from the most recent iteration. In each of these T iterations, the maximum expected improvement value of the objective function is lower than a predetermined expected improvement threshold. T is set to 5. This value is determined because if the expected improvement value of any unsampled point in the search domain is lower than one-hundredth of the threshold in five consecutive iterations, it indicates that there are no unexplored regions with significant improvement potential within the search domain under the posterior distribution of the current Gaussian process surrogate model. Continuing iterations will most likely only yield a small improvement in the objective function value. The predetermined expected improvement threshold is set to 1 / 100 of the current optimal objective function observation value. This relative threshold design allows the convergence criterion to adapt to the dimensions and numerical range of the objective function. When the objective function value is large, it tolerates a larger absolute improvement; when the objective function value is small, it requires a more refined improvement. The current optimal objective function observation value is re-determined after each iteration (S404) based on all objective function observation values ​​in the updated training dataset. If the optimization direction is maximization, the maximum value is taken; if the optimization direction is minimization, the minimum value is taken.

[0169] The maximum value of the expected improvement function is obtained as follows: During each execution of S403, the expected improvement value is calculated for each of the 10,000 randomly generated candidate experimental parameter combination vectors within the search domain, and the maximum expected improvement value obtained in that search is recorded. The maximum expected improvement value of each iteration is stored in a sliding window queue of length T in iteration order. The queue adopts a first-in, first-out strategy; when the maximum expected improvement value of a new iteration enters the queue, the earliest maximum expected improvement value in the queue is removed. After each completion of S404, it is checked whether the sliding window queue is full of T maximum expected improvement values. If it is not full, the next iteration continues. If it is full, the T maximum expected improvement values ​​in the queue are compared one by one to see if they are all lower than a predetermined expected improvement threshold. If all are lower, the preset convergence condition is met; if at least one is greater than or equal to the predetermined expected improvement threshold, the preset convergence condition is not met, and the next iteration continues.

[0170] The iteration stops when the preset convergence condition is met. The specific method for outputting the experimental parameter set as the combination of experimental parameters corresponding to the current optimal objective function observations is as follows: Extract the objective function observation vector from the updated training dataset. If the optimization direction is maximization, find the index position of the maximum value in the objective function observation vector; if the optimization direction is minimization, find the index position of the minimum value. Use the row vector corresponding to this index position in the input matrix as the recommended experimental parameter set. The output format of the recommended experimental parameter set is a structured dictionary containing parameter names and parameter values. The keys of the dictionary are the names of the decision variables defined in S401, and the values ​​are the corresponding values ​​of these decision variables in the optimal experimental parameter combination vector. This recommended experimental parameter set is the output data sent to the automated experimental platform.

[0171] In this embodiment, it is recommended that the experimental parameter set be stored in the form of a structured dictionary. The keys of the dictionary are the names of decision variables, including at least one of pH value, temperature, cofactor concentration, substrate concentration, and protein ligand molar ratio. The values ​​of the dictionary are the numerical values ​​of the corresponding decision variables. The standardized communication protocol adopts a hypertext transfer protocol based on a representational state transition architecture. The data exchange format adopts JavaScript object representation. The root path of the communication interface is uniformly the experimental task submission interface under port 8080 of the control server's Internet Protocol address of the automated experimental platform, and the request method is the POST method. The specific sending operation is as follows: serialize the structured dictionary of the recommended experimental parameter set into a JavaScript object representation string, construct a request body dictionary containing a task type field, an experimental parameter set field, and a timestamp field. The value of the task type field is set according to the quantitative measurement index type determined in S401, which is one of enzyme-catalyzed reaction experiment, surface plasmon resonance experiment, enzyme-linked immunosorbent assay, or fluorescence quantitative polymerase chain reaction experiment. The experimental parameter set field is directly taken from the structured dictionary of the recommended experimental parameter set, and the timestamp field is taken from the current Coordinated Universal Time (UTC) time string. After serializing the request body dictionary into a JavaScript object representation string, it is used as the body of the HTTP POST request and sent to the control server of the automated experimental platform.

[0172] After receiving the recommended experimental parameter set, the automated experimental platform's task scheduling module parses the request body dictionary, extracts the task type field and experimental parameter set field, and assigns the experimental task to the corresponding experimental execution module based on the value of the task type field. When the task type is an enzyme-catalyzed reaction experiment, the experimental execution module is a system combining a liquid handling workstation and an ELISA reader. The liquid handling workstation automatically prepares the reaction system and controls the reaction temperature and time according to the pH, temperature, cofactor concentration, and substrate concentration parameters in the experimental parameter set. After the reaction, the ELISA reader measures the absorbance value of the product and converts it into a product concentration value, outputting the raw data of the enzyme-catalyzed reaction product generation rate. When the task type is a surface plasmon resonance (SPR) experiment, the experimental execution module is a SPR instrument, which automatically executes the sample introduction, binding, dissociation, and regeneration process according to the pH, temperature, and protein-ligand molar ratio parameters in the experimental parameter set, outputting the raw sensor data. When the task type is an enzyme-linked immunosorbent assay (ELISA) experiment, the experimental execution module is a system combining a liquid handling workstation and an ELISA reader. After processing the cell model according to all the parameters in the experimental parameter set, it performs antibody incubation and colorimetric reactions, outputting the raw absorbance data. When the task type is a real-time polymerase chain reaction (RTP) experiment, the experimental execution module is a real-time polymerase chain reaction (RTP) instrument. It performs nucleic acid extraction, reverse transcription, and amplification reactions according to all parameters in the experimental parameter set, and outputs the raw data of the amplification cycle number corresponding to the fluorescence signal of each reaction tube reaching the preset fluorescence threshold.

[0173] After the automated experimental platform completes the experiment, it encapsulates the raw experimental results into a response body dictionary containing a task identifier field, a raw data field, an execution status field, and a completion timestamp field. The response body dictionary is then serialized into a JavaScript object representation string and returned as an HTTP response. The system in this method receives this HTTP response, parses the response body dictionary, extracts the values ​​of the raw data fields, and passes them as the raw experimental results data to subsequent operations. The timeout for receiving operations is set to 3600 seconds. If no response is received within 3600 seconds, communication is considered a failure, and a retry mechanism is triggered. The maximum number of retries is set to 3, with a 60-second interval between each retry. After all 3 retries fail, the experimental loop is terminated, and an exception log is recorded.

[0174] It should be further explained that, in this embodiment, the relation weights in the glycomics knowledge graph are updated based on the deviation between the experimental results and the predicted structure, and the deviation is used to fine-tune the multimodal agent, specifically including:

[0175] S601. Receive the raw experimental result data returned by the automated experimental platform. Based on the type identifier of the glycan functional association entity to be verified in the research hypothesis and the type of the relationship to be verified, determine the experimental verification indicator type. The experimental verification indicator type is one of the following: enzyme-catalyzed reaction product generation rate, protein binding dissociation constant, or fold change in disease marker expression level. Extract experimental verification indicator values ​​from the raw experimental result data according to the data parsing rules corresponding to the experimental verification indicator type. The data parsing rules include: when the experimental verification indicator type is enzyme-catalyzed reaction product generation rate, extract the increase in product concentration per unit time from the raw experimental result data as the experimental verification indicator value, in micromoles per liter per minute; when the experimental verification indicator type is protein binding dissociation constant, extract the ratio of the binding rate constant to the dissociation rate constant from the raw experimental result data as the experimental verification indicator value, in nanomoles per liter; when the experimental verification indicator type is fold change in disease marker expression level, extract the ratio of the average expression level of the marker in the treatment group to the average expression level of the marker in the control group from the raw experimental result data as the experimental verification indicator value, dimensionless unit.

[0176] S602. Obtain the predicted value of the objective function, which is the posterior mean of the Gaussian process surrogate model at the recommended experimental parameter set. Treat both the experimental verification index value and the predicted value of the objective function as continuous variables, and use the KL divergence calculation method based on the Gaussian distribution assumption to measure the degree of deviation between them. Specifically, construct a predicted Gaussian distribution using the predicted value of the objective function as the mean and the variance estimate of the historical experimental error as the variance; construct an observed Gaussian distribution using the experimental verification index value as the mean and the same variance estimate of the historical experimental error as the variance; calculate the KL divergence between the two Gaussian distributions. The KL divergence is calculated as follows: KL divergence equals the square of the difference between the observed mean and the predicted mean divided by twice the variance. The variance estimate of the historical experimental error is pre-calibrated based on historical repeated experimental data of the same experimental type and stored in the system configuration parameters; in this embodiment, the variance estimate is 1.0. When the experimental verification index value is exactly equal to the predicted value of the objective function, the KL divergence is zero. When the experimental verification index value or the target function prediction value is less than or equal to zero, an offset constant is first added to both to convert them into positive values ​​before calculating according to the above formula. The offset constant value is the larger of twice the absolute value of the target function prediction value and twice the absolute value of the experimental verification index value.

[0177] S603. Compare the calculated KL divergence with a preset KL divergence threshold, which is set to 0.5. This threshold is determined by collecting 20 known accurate sets of glycan verification data from historical experiments, calculating the KL divergence value for each set using the Gaussian distribution KL divergence calculation method described above, and taking the 95th percentile as the threshold, which is calculated to be 0.47. This threshold is then rounded down to 0.5, ensuring a false positive trigger rate of approximately 5%. When the KL divergence is less than or equal to the preset KL divergence threshold of 0.5, the experimental result is determined to be consistent with the predicted structure, indicating that the deviation between the actual experimental result and the model prediction result is within an acceptable statistical fluctuation range. This does not trigger knowledge graph weight updates and model fine-tuning, ending the current closed loop. When the KL divergence is greater than the preset KL divergence threshold of 0.5, a significant deviation exists between the experimental result and the predicted structure, triggering knowledge graph weight updates and model fine-tuning operations.

[0178] S604. Update the relation weights between entities in the glycomics knowledge graph using the deviation between the original experimental results data and the predicted glycan structure. Specifically, extract the relation edges along each step of the random walk path corresponding to the research hypothesis, multiply the current relation weight of each step of the relation edge in the glycomics knowledge graph by a decay coefficient. The decay coefficient is equal to the ratio of the experimental verification index value to the predicted value of the objective function. The decay coefficient is truncated to the range of 0.1 to 2.0 to obtain the updated relation weights and write them into the weight attribute field of the corresponding relation edge in the glycomics knowledge graph. In this embodiment, to avoid the random error of a single experimental result causing the knowledge graph relation weights to be updated incorrectly, the update of relation weights does not depend on a single experiment. Instead, a dual robust update strategy based on multiple independent repeated verifications and Bayesian fusion is adopted. The specific implementation method is as follows: First, multiple independent repeated verification mechanism. For each hypothetical path that triggers a weight update operation in each closed-loop iteration, the system requires at least three independent repeated experiments to be performed for the experimental verification corresponding to that hypothetical path. The experimental parameter set for each repeated experiment remains unchanged. The automated experimental platform executes these experiments independently under the same conditions and returns the original experimental results data. For each repeated experiment, the corresponding experimental verification index value is parsed according to step S601. The arithmetic mean and standard deviation of the three experimental verification index values ​​are calculated. The arithmetic mean is used as the comprehensive experimental verification index value for that hypothetical path to participate in the calculation of the attenuation coefficient or enhancement coefficient. At the same time, if the coefficient of variation among the three repeated experiments, i.e., the standard deviation divided by the arithmetic mean, exceeds the preset coefficient of variation threshold of 0.15, it is determined that the random error of the experimental results is too large. The weight update operation for this round is suspended, and the hypothetical path is marked as "to be further verified" and updated again after more experimental evidence is accumulated in subsequent closed-loop iterations. Second, Bayesian fusion update strategy. When experimental conditions are limited and three or more independent repeated experiments cannot be performed, a Bayesian fusion method is adopted to update the relation weights. The current weight value of the relation edge in the knowledge graph is considered as the prior distribution. A likelihood function is constructed using the deviation between the experimental verification index value of the current single experiment and the predicted value of the objective function. Bayes' theorem is used to fuse the prior distribution and the likelihood function to obtain the posterior distribution. The mean of this posterior distribution is taken as the updated relation weight and written into the knowledge graph. Specifically, the noise variance parameter in the likelihood function is taken as the variance estimate of similar historical experiments. This estimate is statistically calibrated based on public datasets of similar experiments during system initialization and stored in the system configuration parameters. This Bayesian fusion update strategy can effectively integrate historical prior knowledge with current experimental evidence, reducing the impact of random fluctuations in a single experiment on weight updates.The aforementioned dual robust update strategy ensures that the updates of relation weights in the glycomics knowledge graph are always based on sufficient and reliable experimental evidence, avoiding a vicious cycle of continuously declining credibility of knowledge graph relations due to accidental errors in a single experiment, distortion of subsequent random walk sampling probabilities, and deviation of the closed-loop knowledge evolution direction from reality.

[0179] S605. Using the aforementioned bias as a supervisory signal to fine-tune the multimodal agent, specifically: using the monosaccharide type, glycosidic bond connection site, and anodic configuration of low-confidence units in the predicted glycan structure whose overall confidence score is lower than a preset confidence threshold of 0.9 as fine-tuning labels, and using the multi-source experimental spectral data and the multimodal semantic representation vector as fine-tuning inputs, supervised fine-tuning is performed on the spectral feature extractor parameters of the cross-modal fusion sub-agent and the projection matrix parameters of the multi-head cross-attention network, as well as the linear classifier head projection matrix parameters and embedding matrix parameters in the autoregressive sequence generation model of the structure prediction sub-agent, with a learning rate of 0 during fine-tuning. The Adam optimizer in step 00001 is fine-tuned with 50 iterations and a batch size of 8 per iteration. It's important to further explain that the technical logic of using functional validation experimental bias to drive fine-tuning of the structural prediction model is based on the following inference chain: In each closed loop, the generation of the research hypothesis strictly depends on the high-confidence units in the predicted glycan structure; the experimental validation target is directly related to the function of a specific structural domain in the predicted structure. When the KL divergence between the experimental results and the predicted values ​​exceeds a preset threshold, it indicates a significant deviation between the predicted structure and the experimental results. At this point, the deviation is first attributed to the possibility of errors in the structural unit with the lowest confidence in the predicted structure, rather than directly rejecting the entire functional hypothesis. The rationale for this attribution strategy is that the low-confidence units in the structural prediction are precisely the areas where the model itself lacks confidence in its prediction accuracy. These areas are most likely to contain erroneous structural identification results, leading to deviations in the experimental validation of the functional hypothesis derived from this structure. Therefore, using the true structure of low-confidence units as fine-tuning labels can guide the model to make targeted corrections at the weakest points in structure prediction, continuously improving the accuracy of structure prediction in closed-loop iterations, while avoiding interference with the correctly learned structural knowledge of high-confidence units. This fine-tuning mechanism transforms the deviation signals from functional verification experiments into diagnostic information for weaknesses in the structure prediction model, realizing a feedback loop from macro-functional verification to micro-structural correction.

[0180] In this embodiment, to avoid the catastrophic forgetting problem caused by using only low-confidence unit labels for fine-tuning—that is, the model loses previously learned high-confidence knowledge while updating the predictive ability of low-confidence units—the fine-tuning process in this embodiment adopts a dual anti-forgetting mechanism based on elastic weight consolidation regularization and mixed replay of historical high-confidence samples. The specific implementation is as follows: First, elastic weight consolidation regularization. Before fine-tuning begins, a batch of samples is randomly selected from the high-confidence sample pool accumulated in the historical closed-loop iterations. Using the parameter values ​​of the current multimodal agent as historical reference parameters, the sensitivity of each network weight parameter to the prediction loss function of historical high-confidence samples is calculated using the Fisher information matrix estimation method. The diagonal element approximation method of this Fisher information matrix is ​​as follows: for each sample in the historical high-confidence sample pool, the gradient of the loss function with respect to each weight parameter is calculated, the square of the gradient is taken, and then the arithmetic mean of the squared gradient values ​​of all samples is taken to obtain the Fisher information value corresponding to each weight parameter. During fine-tuning, a regularization penalty term is added to the original loss function. This penalty term is calculated as follows: for each network weight parameter, the square of the difference between its current value and the historical reference parameter value is multiplied by the corresponding Fisher information value, and then multiplied by a regularization strength coefficient. The sum of these calculations for all parameters is used as the regularization penalty term. In this embodiment, the regularization strength coefficient is set to 0.1. The purpose of this regularization penalty term is to impose greater update resistance on key parameters that contribute significantly to historical high-confidence knowledge, while allowing larger updates for parameters less correlated with historical knowledge. This preserves parameter adjustment space for learning new knowledge while protecting existing knowledge. Second, historical high-confidence samples are mixed and replayed. In each round of fine-tuning iteration, an equal number of complete glycan samples are randomly selected from the historical high-confidence sample pool as well as the current low-confidence fine-tuning samples. The multi-source experimental spectrogram data and structural labels of both samples are concatenated and merged to form a mixed training batch. The mixed training batch consists of half low-confidence samples and half high-confidence samples, with samples within the batch randomly arranged. This hybrid replay strategy allows the model to continuously review and reinforce previously correctly learned high-confidence structural knowledge while updating the predictive ability of low-confidence units, preventing the model parameters from excessively shifting towards low-confidence samples. In this embodiment, the historical high-confidence sample pool is maintained as follows: in each closed-loop iteration, the complete glycan multimodal semantic representation vector, multi-source experimental spectrogram data, and experimentally verified structural labels corresponding to the structural units judged as high-confidence by the structural judgment output sub-agent are stored in the sample pool; the sample pool is managed using a first-in-first-out queue structure, with a maximum capacity of 1000 glycan samples. When the sample pool is full, the earliest stored sample is automatically removed.The aforementioned dual anti-forgetting mechanism ensures that while the multimodal agent continuously learns new knowledge through closed-loop iteration, it does not lose the previously acquired high-confidence glycan structure analysis ability. The overall prediction accuracy of the model remains stable and improves rather than degrades as the number of iterations increases.

[0181] S606. After completing the update of the relation weights of the glycomics knowledge graph and the fine-tuning of the multimodal agent, based on the updated glycomics knowledge graph and the fine-tuned multimodal agent, return to execute the operation described in S3 of generating at least one research hypothesis on the glycomics knowledge graph based on the predicted glycan structure and the confidence probability distribution, and start the next round of the closed loop from graph analysis to experimental verification.

[0182] As a concrete example, this embodiment uses the complex N-linked diantenna glycan of the Fc region of human serum immunoglobulin G as the target glycan, fully demonstrating the unmanned closed-loop process from the original spectrum to the discovery of new glycobiological knowledge. This glycan contains 14 monosaccharide residues and is a typical complex N-glycan that plays a key role in antibody-dependent cytotoxicity.

[0183] First, the system uses the RESTful application interface of the GlyTouCan database to retrieve the corresponding linear string in GlycoCT format using the registration number G00001 of the sugar chain as the query key. The RES node of this string defines all 14 monosaccharide residues from the reducing N-acetylglucosamine to the non-reducing sialic acid. The LIN node defines 13 glycosidic bonds, including the core fucose connected to the reducing N-acetylglucosamine via α-1,6 bonds, and the two antenna branches connected to the core mannose via α-1,3 and α-1,6 bonds. The REP node marks the α-2,6-sialic acidification modification on the non-reducing galactose residues. The GlycoCT string is preprocessed by lexicalization. The monosaccharide identifier of the RES node, the connection relationship information of the LIN node (including the position index, carbon atom site and anodic configuration of the two connected monosaccharide residues), and the modification label of the REP node are mapped to the corresponding integer indices to form an integer sequence of length 14, which does not exceed 256 words. The sequence is padded to 256 words with padding characters and the corresponding attention mask vector is generated. Simultaneously, based on the GlycoCT string, molecular graph topology data was constructed, generating an undirected connected graph with 14 monosaccharide residues as nodes and 13 glycosidic bonds as edges. The one-hot encoded vector of each node contains information on the type of monosaccharide, ring configuration, and absolute configuration. The one-hot encoded vector of each edge contains information on the donor connection site, acceptor connection site, and anodic configuration. Specifically, the edge attribute record between the core fucose and the reducing end N-acetylglucosamine shows that the donor connection site is at carbon 1, the acceptor connection site is at carbon 6, and the anodic configuration is α. The edge attribute record between the branch end galactose and sialic acid shows that the donor connection site is at carbon 2, the acceptor connection site is at carbon 3, and the anodic configuration is α. Furthermore, based on the molecular graph topology data and GlycoCT string of the sugar chain, initial three-dimensional atomic coordinates were generated using the RDKit distance geometry algorithm. A 200-nanosecond all-atom molecular dynamics simulation was performed under the combined action of the AMBER14SB force field and the GB-Neck2 implicit generalized Born solvent model. From the last 50 nanosecond equilibrium trajectory, 100 conformational snapshots were uniformly extracted. After rotation, translation, and alignment using the Kabsch algorithm, the Boltzmann weights of each frame were calculated, forming a conformational ensemble containing 441 heavy atoms. The coordinates of the carbon-6 atom of the core fucose in the first frame of the conformational ensemble are X-axis 12.345, Y-axis -3.678, and Z-axis 25.901. In this conformational ensemble, atoms in each frame are arranged in ascending order of residue number, and within the same residue, priority is given to carbon, nitrogen, oxygen, sulfur, and phosphorus. Missing atoms are filled with zero coordinates and the atomic mask is set to 0.

[0184] Secondly, the three modal data are input into a three-level encoding mapping network. The sequence encoder performs 6 layers of Transformer encoding on the lexicalized integer sequence. Each layer contains an 8-head self-attention network and a feedforward network with a hidden layer width of 2048. It outputs a 512-dimensional hidden state vector for each monosaccharide residue position, and obtains a 512-dimensional sequence feature vector through average pooling. The graph encoder employs a 5-layer graph isomorphic network with edge feature enhancement. The initial 32-dimensional one-hot encoding of 14 nodes is linearly transformed into a 64-dimensional embedding vector. In each of the 5 rounds of message passing, the edge attribute vector is converted into a 64-dimensional edge embedding vector and then input together with the features of neighboring nodes into a multilayer perceptron containing 2 hidden layers and a width of 128 to generate a message vector. After element-wise summation and aggregation, the vector is concatenated with the node's state from the previous round and fed into the update function. The dimension of the node's hidden state is expanded from 64 dimensions to 128 dimensions in the first round, to 256 dimensions in the second round, and remains at 256 dimensions in rounds 3 to 5. Finally, the global average pooling readout layer takes the arithmetic mean of the 256-dimensional hidden state vectors of the 14 nodes along the node dimensions to obtain a fixed 256-dimensional graph feature vector. The spatial encoder employs a 4-layer SE3 equivariant Transformer. For each of the 441 atoms in the ensemble of the conformation, a spatial connectivity graph with a radius of 5 Å is constructed, consisting of approximately 8500 spatial adjacency edges centered on each atom. The atom type is mapped to a 16-dimensional initial scalar feature through an embedding layer. The magnitude of the relative position vector is input into the radial basis function to obtain scalar weights. The orientation angle is expanded to the coefficients of the highest-order 3 spherical harmonic basis function to obtain a 16-dimensional edge feature vector. After being encoded layer by layer by 4 layers of TensorField convolution, each atom outputs an equivariant feature vector containing 16 spherical harmonic channels. Then, after global average pooling readout layer, the spherical harmonic tensor components of each spherical harmonic channel are converted into rotation-invariant scalar values ​​by taking the L2 norm, and then concatenated. Finally, the Boltzmann weights of each frame of the conformation are used for weighted averaging to obtain a fixed 512-dimensional spatial conformation feature vector. The sequence feature vector is mapped by the projection matrix W_seq, the graph feature vector is expanded to 512 dimensions and then mapped by the projection matrix W_graph, and the spatial configuration feature vector is mapped by the projection matrix W_3d, generating three sets of 512-dimensional L2 norm normalized intermediate vectors z_seq, z_graph, and z_3d. After cross-modal contrastive learning based on structural similarity to select negative samples, the cosine similarity between z_seq and z_graph for this glycan is 0.89, the cosine similarity between z_graph and z_3d is 0.91, and the cosine similarity between z_3d and z_seq is 0.86, all exceeding the semantic alignment threshold of 0.85. The gated adaptive fusion layer uses the Softmax normalized weights w_seq (0.31), w_graph (0.40), and w_3d (0.29) to weight and sum the three sets of intermediate vectors, generating a 512-dimensional multimodal semantic representation vector.

[0185] Third, the system acquires multi-source experimental spectral data of the target glycan and inputs it into a trained multimodal agent for structure prediction and confidence assessment. The multi-source experimental spectral data includes tandem mass spectrometry (MS / MS) spectra, one-dimensional hydrogen nuclear magnetic resonance (1H NMR) spectra, and reversed-phase liquid chromatography (RP-LC) retention time data of the glycan. The mass spectra are discretized into relative abundance vectors of 3800 mass-to-charge ratio intervals, and a 128-dimensional mass spectrometry feature vector is extracted through two layers of one-dimensional convolution and global average pooling. The nuclear magnetic resonance (NMR) spectra are discretized into peak intensity vectors of 1000 chemical shift intervals, and a 128-dimensional NMR feature vector is extracted through a three-layer fully connected network. The three parameters of the chromatographic retention time are used to extract a 128-dimensional chromatographic feature vector through a two-layer fully connected network. Three 128-dimensional feature vectors are concatenated into a 384-dimensional spectral vector, which is then linearly projected to 512 dimensions. This vector is then used in four rounds of sequence-to-sequence multi-head cross-attention interaction with the glycan structure query sequence formed by expanding and replicating the multimodal semantic representation vector 14 times. In each round, eight attention heads learn the attention weights between the query and the key and sum the values ​​in a weighted manner. The output is a glycan hidden state sequence with a length of 14 and 512 dimensions at each time step. The structure prediction agent uses an autoregressive sequence generation model based on a Transformer decoder to predict residues one by one from the non-reduced end to the reduced end. During the generation process, the model predicts N-acetylglucosamine from the first time step with the initial embedding vector as input. From the second to the fifth step, it predicts the second N-acetylglucosamine and the core mannose branch region in sequence. In the sixth step, when predicting the mannose residue connected to the core mannose, it outputs a mannose probability of 0.96, a carbon 1-to-carbon 6 combination probability of 0.85, and an α configuration probability of 0.92. It continues to generate subsequent residues until the 14th step, when it predicts sialic acid with a probability of 0.93, a carbon 2-to-carbon 3 combination probability of 0.88, and an α configuration probability of 0.90. The model automatically terminates generation when the probability of outputting the EOS tag in the monosaccharide type classification header reaches the maximum value, and outputs a predicted glycan structure sequence of length 14. The uncertainty quantification sub-agent maintains Monte Carlo dropout activation during the inference phase with a dropout rate of 0.1, performing 100 random forward propagations. Since the hidden units randomly dropped in each propagation are different, 100 prediction sequences of potentially different lengths are generated. Using the sequence with the highest frequency as a reference, the Needleman-Wunsch algorithm is used to align the remaining sequences one by one. A matching score of 1 point, a mismatch penalty of -1 point, an open gap penalty of -2 points, and a gap extension penalty of -1 point are applied. After alignment, the frequency of each candidate monosaccharide type, connection site combination, and anodic configuration in the 100 samples is counted at each position, and the confidence score is obtained by dividing by 100. Specifically, for the 6th structural unit, the confidence score distribution for monosaccharide type is 0.96 for mannose, for connection site confidence distribution is 0.85 for the carbon-1 to carbon-6 combination, and for anodic configuration confidence distribution is 0.92 for the α configuration.The structural judgment output sub-agent takes the minimum value of the maximum confidence of the three dimensions as the comprehensive confidence score. The 6th structural unit has a score of 0.85, which is lower than the preset confidence threshold of 0.9. It is marked as a low confidence unit and marked with a 1. The comprehensive confidence scores of the remaining 13 structural units are all not lower than 0.9.

[0186] Fourth, research hypotheses are generated on the glycomics knowledge graph based on predicted structure and confidence level. Combinations of monosaccharide types and connection sites of 12 deterministic structural units are used as query conditions, and a dual-match search is performed on the glycomics knowledge graph. Three graph anchors—Glycan-001, Glycan-003, and Glycan-007—that satisfy the same monosaccharide arrangement pattern and branching topological fingerprint are found. Using probabilities of 0.5 for catalytic relationship edges, 0.3 for binding relationship edges, and 0.2 for upregulation relationship edges, 50 independent random walks are performed starting from each of the three anchor points. During the walk, candidate transition edges for glycan entities are catalytic relationship edges pointing to enzyme entities (sampling probability 0.625) and binding relationship edges pointing to protein entities (sampling probability 0.375). Candidate transition edges for enzyme entities are upregulation relationship edges pointing to protein entities (sampling probability 0.4) and binding relationship edges (sampling probability 0.6). Candidate transition edges for protein entities are upregulation relationship edges pointing to disease entities (sampling probability 0.4) and binding relationship edges pointing to drug entities (sampling probability 0.6). The walk terminates early when it reaches a disease or drug entity. After filtering by uncertainty structural features, 35 candidate paths are retained. These paths are then sorted by the number of hops from smallest to largest, and for those with the same number of hops, by the product of the edge type weights from largest to smallest. The top 5 candidate hypothesis paths are selected. The top-ranked pathway is Glycan-001 transferred via a catalytic relationship to Enzyme-A, then upregulated to Disease-X, with 2 jumps and a weighted product of 0.10. Based on this, the system generates the following hypothesis: Given the complete structure of the glycan chain, existing statistical association evidence indicates a significant positive correlation between this glycan chain and chronic inflammatory disease X; the nature of its action and the direction of its causality need to be determined by subsequent experiments. The optimizable experimental parameter space includes pH 6.0–7.5, temperature 30–40 degrees Celsius, manganese ion concentration 1–5 mmol / L, and substrate concentration 10–100 μmol / L.

[0187] Fifth, using the fold change in disease biomarker expression levels as the objective function, ten initial experimental parameter combinations were generated within a 4-dimensional search domain using Latin hypercube sampling. The observed objective function values ​​were measured to be 3.2, 1.8, 4.5, 2.1, 3.9, 2.7, 1.3, 4.1, 2.5, and 3.6 fold changes, respectively. After fitting a Gaussian process surrogate model, an iterative search was performed with the objective of maximizing the expected improvement value. In each round, 10,000 candidate points were generated within the search domain to calculate the expected improvement value. The candidate point corresponding to the maximum expected improvement value was selected as the recommended experimental parameter for real-world experimental verification. After eight rounds of iteration, the maximum expected improvement values ​​for five consecutive rounds were 0.038, 0.029, 0.021, 0.015, and 0.011, respectively, all lower than one-hundredth of the current optimal observed value of 5.9, i.e., 0.059, thus satisfying the convergence condition. The optimal combination of experimental parameters is: pH 6.88, temperature 35.7 degrees Celsius, manganese ion concentration 2.6 mmol / L, and substrate concentration 58.2 μmol / L. This recommended set of experimental parameters is sent to port 8080 of the automated experimental platform control server via an HTTP POST request in JSON format.

[0188] Sixth, after receiving the task, the liquid handling workstation of the automated experimental platform, coupled with the microplate reader system, automatically prepares a buffer solution with a pH of 6.88. The cell model is then treated under conditions of 35.7 degrees Celsius, a manganese ion concentration of 2.6 mmol / L, and a substrate concentration of 58.2 μmol / L. Antibody incubation and colorimetric reaction are performed in the enzyme-linked immunosorbent assay (ELISA), and the absorbance is measured at 450 nm using the microplate reader. The returned raw experimental data includes absorbance values ​​of 0.852, 0.791, and 0.834 for the three replicates in the treatment group, and 0.156, 0.148, and 0.162 for the three replicates in the control group.

[0189] Seventh, the system's pre-set raw data analysis module converts absorbance values ​​into biomarker concentration values ​​using a standard curve. It calculates the average absorbance of the treatment group (0.826) divided by the average absorbance of the control group (0.155), yielding an experimental verification index value of 5.33 times change. Using this value as the observed mean, the predicted value of the objective function (5.9) as the predicted mean, and the historical experimental error variance estimate (1.0) as the variance, the system calculates the Gaussian distribution KL divergence. The KL divergence is calculated as the square of (5.33 minus 5.9) divided by 2.0, which equals 0.162. This value is less than the preset KL divergence threshold of 0.5, indicating that the experimental results are consistent with the predicted structure. Since the validation was successful, the system upweighted the catalytic relationship edge from Glycan-001 to Enzyme-A and the upregulation relationship edge from Enzyme-A to Disease-X in the hypothesis pathway with enhancement coefficients. This experimentally validated biological functional association of the glycan chain positively upregulating chronic inflammatory disease X through glycosyltransferase GT-A was incorporated into the glycomics knowledge graph as new knowledge. The complete experimental data from this successful validation was stored in the historical high-confidence sample pool for subsequent continuous learning and knowledge accumulation, thus concluding this round of closed-loop processing. If the number of newly added confidence knowledge graph edges in subsequent closed-loop iterations is less than 2 per round for three consecutive rounds, or if the total number of iterations reaches 20 rounds, the system automatically stops and outputs a complete functional association report for the glycan chain, completing the fully automated discovery process from the original spectrum to new glycobiological knowledge.

[0190] Example 3

[0191] Please see Figure 3 Another embodiment of the present invention provides: a glycomics intelligent agent construction system based on three-level coding, comprising:

[0192] The multimodal coding module is used to acquire the sequence code, molecular graph topology and three-dimensional spatial conformation data of the target glycan, and to map the sequence code, molecular graph topology and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network to generate the multimodal semantic representation vector of the target glycan;

[0193] The spectral data acquisition module is used to acquire multi-source experimental spectral data for the target glycan, wherein the multi-source experimental spectral data includes at least two of mass spectrometry data, nuclear magnetic resonance spectral data, and chromatographic parameter data.

[0194] The multimodal agent prediction module is used to input the multimodal semantic representation vector and the multi-source experimental spectrogram data into a pre-trained multimodal agent, which performs cross-modal fusion and uncertainty quantification, and outputs the predicted glycan structure of the target glycan and the confidence probability distribution of each structural unit in the predicted glycan structure.

[0195] The hypothesis generation module is used to perform a k-hop random walk on the glycomics knowledge graph based on the predicted glycan structure and the confidence probability distribution to generate at least one research hypothesis. The research hypothesis includes glycan functional entities to be experimentally verified and an optimizable experimental parameter space. The glycomics knowledge graph is centered on glycan entities and associates enzyme entities, protein entities, disease entities, and drug entities, and uses catalytic relationships, binding relationships, and upregulation relationships as the types of relationships between entities.

[0196] The experimental design module is used to generate a recommended set of experimental parameters by performing Bayesian optimization iterative search using Gaussian process regression within the optimizable experimental parameter space, with the experimental verification index corresponding to the predicted glycan structure as the objective function.

[0197] The experiment execution module is used to send the recommended set of experimental parameters to the automated experiment platform through a standardized communication protocol, and to receive the raw experimental result data returned by the automated experiment platform after the experiment is executed.

[0198] An iterative module is used to parse the original experimental result data into experimental verification index values, calculate the KL divergence between the experimental verification index values ​​and the predicted values ​​of the objective function; when the KL divergence is greater than a preset threshold, the relation weights between entities in the glycomics knowledge graph are updated using the deviation between the original experimental result data and the predicted glycan structure, and the deviation is used as a supervision signal to fine-tune the multimodal agent. Based on the updated knowledge graph and the fine-tuned multimodal agent, a research hypothesis is regenerated to start the next round of closed loop.

[0199] Example 4

[0200] A smart terminal includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a glycomics-based intelligent agent construction method.

[0201] A computer-readable storage medium storing computer instructions that, when executed, perform a method for constructing a glycomics-based intelligent agent based on three-level coding.

[0202] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments under the guidance of the present invention without departing from the spirit and scope of the present invention. All of these variations are within the protection scope of the present invention.

[0203] If the technical solution disclosed herein involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution disclosed herein involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of explicit consent. For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A method for constructing glycomics-based intelligent agents based on three-level coding, characterized in that, include: Obtain a multimodal semantic representation vector of the target glycan, wherein the multimodal semantic representation vector integrates the sequence, graph topology and spatial conformation information of the target glycan; Obtain multi-source experimental spectral data of the target glycan; The trained multimodal agent is used to perform cross-modal fusion and uncertainty quantification on the multimodal semantic representation vector and the multi-source experimental spectrogram data to obtain the predicted structure of the target glycan and the confidence level of each structural unit in the predicted structure. Based on the predicted structure and the confidence level, at least one research hypothesis is generated on the glycomics knowledge graph. The research hypothesis includes glycan functional associated entities to be experimentally verified and an optimizable experimental parameter space. Within the optimizable experimental parameter space, a recommended set of experimental parameters is generated using Bayesian optimization with the experimental verification index corresponding to the predicted structure as the objective function. The recommended set of experimental parameters is sent to the automated experimental platform, and the experimental results returned by the automated experimental platform after the experiment is performed are received. Based on the deviation between the experimental results and the predicted structure, the relation weights in the glycomics knowledge graph are updated, and the deviation is used to fine-tune the multimodal agent.

2. The method for constructing a glycomics-based intelligent agent based on three-level coding as described in claim 1, characterized in that, The multimodal semantic representation vector is generated by mapping sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network; The generation of research hypotheses on the glycomics knowledge graph includes performing k-hop random walks; the glycomics knowledge graph is centered on glycan entities, and associates enzyme entities, protein entities, disease entities, and drug entities, and uses catalytic relationships, binding relationships, and upregulation relationships as the types of relationships between entities; The method of generating a recommended set of experimental parameters using Bayesian optimization includes: within the optimizable experimental parameter space, using the experimental verification index as the objective function, performing Bayesian optimization iterative search using Gaussian process regression; The method of fine-tuning the multimodal agent using the bias includes: parsing the experimental results into experimental verification index values, calculating the KL divergence between the experimental verification index values ​​and the predicted values ​​of the objective function, and when the KL divergence is greater than a preset threshold, updating the relation weights using the bias and fine-tuning the multimodal agent.

3. The method for constructing a glycomics-based intelligent agent based on three-level coding as described in claim 2, characterized in that, The process of mapping sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network includes: Obtain the sequence encoding data of the target glycan chain, wherein the sequence encoding data is the monosaccharide arrangement order and modification site information represented by a linear string in GlycoCT format; Obtain the molecular graph topology data of the target glycan chain. The molecular graph topology data is an undirected connected graph with monosaccharide residues as nodes, glycosidic bonds as edges, and the connection sites and anodic configurations recorded in the edge attributes. Obtain the three-dimensional spatial conformation data of the target glycan, wherein the three-dimensional spatial conformation data is the set of three-dimensional Cartesian coordinates of all heavy atoms in the target glycan.

4. The method for constructing a glycomics-based intelligent agent based on three-level coding as described in claim 3, characterized in that, The method of mapping sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network also includes: The GlycoCT format linear string is input into the sequence encoder, and a fixed-dimensional hidden state vector is output for each monosaccharide unit in the linear string. The hidden state vectors of all monosaccharide units are then averaged to obtain the sequence feature vector. The undirected connected graph is input into a graph encoder. After multiple rounds of neighbor node message aggregation for each node in the undirected connected graph, the final hidden state vector of all nodes is obtained by average pooling to obtain the graph feature vector. The three-dimensional Cartesian coordinate set is input into the spatial encoder. Under the constraint of ensuring rotational and translational equivariance, the coordinates of each atom and the relative positions of its neighboring atoms are encoded. The equivariant feature vectors of all atoms are then averaged and pooled to obtain the spatial configuration feature vector.

5. The method for constructing a glycomics-based intelligent agent as described in claim 4, characterized in that, The method of mapping sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network also includes: Using three learnable projection matrices corresponding to sequence mode, graph topology mode and spatial configuration mode respectively, matrix multiplication is performed on the sequence feature vector, the graph feature vector and the spatial configuration feature vector respectively, mapping each mode feature vector to a shared semantic space of preset dimension d, and outputting three sets of intermediate vectors of dimension d, where dimension d is 512; In a training batch containing multiple target glycan samples, any two of the three sets of intermediate vectors for the same target glycan are taken as positive sample pairs, and the corresponding intermediate vectors for different target glycans are taken as negative sample pairs. The InfoNCE extended loss function based on temperature scaling cosine similarity is used to calculate the contrast loss values ​​between the sequence intermediate vector and the graph intermediate vector, the graph intermediate vector and the spatial conformation intermediate vector, and the spatial conformation intermediate vector and the sequence intermediate vector, respectively. The average of the three sets of contrast loss values ​​is taken as the cross-modal contrast loss function value.

6. The method for constructing a glycomics-based intelligent agent as described in claim 5, characterized in that, The method of mapping sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network also includes: With the goal of minimizing the cross-modal contrastive loss function, the network weight parameters of the sequence encoder, the graph encoder, the spatial encoder, and the three projection matrices are updated synchronously through the backpropagation algorithm. The training continues until the cross-modal contrastive loss function converges to a preset threshold range, so that the cosine similarity between any two pairs of intermediate vectors of the same target glycan in the vector space of dimension d reaches a preset high similarity condition, and the cosine similarity between intermediate vectors of different target glycans satisfies a preset low similarity condition.

7. The method for constructing a glycomics-based intelligent agent as described in claim 6, characterized in that, The method of mapping sequence encoding, molecular graph topology, and three-dimensional spatial conformation data to the same vector space through a three-level coding mapping network also includes: The three sets of intermediate vectors that meet the preset high similarity condition are input into the gated adaptive fusion layer. The gated adaptive fusion layer assigns an independent learnable scalar parameter to each set of input intermediate vectors. Softmax normalization is applied to the three learnable scalar parameters to generate a three-dimensional attention weight vector. After performing element-wise multiplication of each weight component in the three-dimensional attention weight vector with the corresponding intermediate vector, element-wise addition is performed on the three weighted vectors to generate a multimodal semantic representation vector of dimension d.

8. The method for constructing a glycomics-based intelligent agent as described in claim 7, characterized in that, The trained multimodal agent includes a cross-modal fusion sub-agent, a structure prediction sub-agent, an uncertainty quantification sub-agent, and a structure judgment output sub-agent; The cross-modal fusion sub-agent is configured to: receive the multimodal semantic representation vector and the multi-source experimental spectral data; map the mass spectrometry data, nuclear magnetic resonance spectral data and chromatographic parameter data in the multi-source experimental spectral data into mass spectrometry feature vectors, nuclear magnetic resonance feature vectors and chromatographic feature vectors respectively through their respective independent spectral feature extractors; use the multimodal semantic representation vector as a query vector; concatenate the mass spectrometry feature vector, the nuclear magnetic resonance feature vector and the chromatographic feature vector as a key vector and a value vector; perform cross-modal interaction through a multi-head cross-attention network; and output the glycan hidden state sequence after fusing spectral information, wherein the length of the glycan hidden state sequence is equal to the total number of monosaccharide residues of the target glycan.

9. The method for constructing a glycomics-based intelligent agent as described in claim 8, characterized in that, The structure prediction sub-agent is configured to: receive the glycan hidden state sequence output by the cross-modal fusion sub-agent; input the glycan hidden state sequence into an autoregressive sequence generation model based on a Transformer decoder; predict the monosaccharide type, glycosidic bond connection site, and anodic configuration of each monosaccharide residue of the target glycan step by step from the reduced end to the non-reduced end; and output a predicted glycan structure sequence of length L, where L is equal to the total number of monosaccharide residues of the target glycan. The output of the predicted glycan structure sequence at time step t is a three-dimensional probability vector, the first dimension of which is the probability distribution of the predicted t-th monosaccharide residue of the target glycan on a preset set of monosaccharide types; the second dimension of which is the probability distribution of the combination of glycosidic bond connection sites between the predicted t-th monosaccharide residue of the target glycan and the previous monosaccharide residue on a preset set of connection site combinations; and the third dimension of which is the probability distribution of the predicted anodic configuration of the predicted t-th monosaccharide residue of the target glycan on a preset set of anodic configurations.

10. The method for constructing a glycomics-based intelligent agent as described in claim 9, characterized in that, The uncertainty quantification sub-agent is configured to: apply Monte Carlo dropout to each feedforward sublayer in the Transformer-based autoregressive sequence generation model during the autoregressive sequence generation process of the structure prediction sub-agent; perform K random forward propagations during the inference phase, with each random forward propagation independently outputting a glycan structure sequence sample; and calculate the occurrence frequency of each candidate monosaccharide type, each candidate glycosidic bond combination, and each candidate anodic configuration at the t-th time step in the K glycan structure sequence samples in the K samplings; divide the occurrence frequency by K to obtain the confidence probability value of each candidate category of the t-th structural unit; and output a confidence probability distribution sequence of length L, where the t-th element in the confidence probability distribution sequence is a triple containing the confidence distribution of monosaccharide type, the confidence distribution of combination site, and the confidence distribution of anodic configuration.

11. The method for constructing a glycomics-based intelligent agent as described in claim 10, characterized in that, The structure judgment output sub-agent is configured to: receive the predicted glycan structure sequence output by the structure prediction sub-agent and the confidence probability distribution sequence output by the uncertainty quantification sub-agent; take the minimum value of the triples of the t-th structural unit in the confidence probability distribution sequence as the comprehensive confidence score of the t-th structural unit; when the comprehensive confidence scores of all L structural units are greater than or equal to a preset confidence threshold, output the predicted glycan structure sequence as the predicted structure of the target glycan, and output the confidence probability distribution sequence as the confidence score corresponding to each structural unit in the predicted structure; when the comprehensive confidence score of at least one structural unit is less than the preset confidence threshold, mark the structural unit as a low-confidence unit, add a low-confidence mark to the low-confidence unit in the output predicted structure, and output the complete confidence probability distribution sequence.

12. A glycomics intelligent agent construction system based on three-level coding, used to implement the glycomics intelligent agent construction method based on three-level coding as described in any one of claims 1-11, characterized in that, include: A multimodal coding module is used to obtain a multimodal semantic representation vector of the target glycan, wherein the multimodal semantic representation vector integrates the sequence, graph topology and spatial conformation information of the target glycan; The spectral data acquisition module is used to acquire multi-source experimental spectral data of the target glycan; The multimodal agent prediction module is used to perform cross-modal fusion and uncertainty quantification on the multimodal semantic representation vector and the multi-source experimental spectrogram data using a trained multimodal agent, so as to obtain the predicted structure of the target glycan and the confidence level of each structural unit in the predicted structure. The hypothesis generation module is used to generate at least one research hypothesis on the glycomics knowledge graph based on the predicted structure and the confidence level. The research hypothesis includes glycan functional associated entities to be experimentally verified and an optimizable experimental parameter space. The experimental design module is used to generate a recommended set of experimental parameters using Bayesian optimization within the optimizable experimental parameter space, with the experimental verification index corresponding to the predicted structure as the objective function. The experiment execution module is used to send the recommended set of experimental parameters to the automated experiment platform and receive the experimental results returned by the automated experiment platform after the experiment is executed. An iterative module is used to update the relation weights in the glycomics knowledge graph based on the deviation between the experimental results and the predicted structure, and to fine-tune the multimodal agent using the deviation.

13. A computer-readable storage medium, characterized in that, It stores computer instructions, which, when executed, perform the glycomics-based intelligent agent construction method according to any one of claims 1-12.

Citation Information

Patent Citations

  • A Denovo method and system for N-glycan structure identification based on mass spectrometry data

    CN114166925B

  • A crop gene function research scheme planning method based on experimental reasoning chain

    CN121278119B