Multi-source mass spectrum metadata fusion method and system

By employing a hierarchical processing and adaptive fusion strategy, combined with Markov random field spatiotemporal consistency verification and mass transfer entropy assessment, the problem of insufficient accuracy and reliability in multi-source mass spectrometry metadata fusion is solved, achieving efficient and precise development of mass spectrometry detection.

CN121935847AActive Publication Date: 2026-04-28NATIONAL INSTITUTE OF METROLOGY CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NATIONAL INSTITUTE OF METROLOGY CHINA
Filing Date
2026-01-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the fusion methods for multi-source mass spectrometry metadata fail to effectively consider the inherent correlation and spatiotemporal consistency between metadata, resulting in insufficient accuracy and reliability of spectral analysis and detection results. Furthermore, the lack of an adaptive adjustment mechanism makes it difficult to cope with the complex needs of different detection scenarios.

Method used

A fusion strategy is adopted, which employs hierarchical processing, dual-channel adaptive fusion, Markov random field spatiotemporal consistency verification, and quality transfer entropy to guide the fusion. Spatiotemporal consistency verification is performed through the belief propagation algorithm to dynamically evaluate the quality of metadata, and spatial alignment and dimensionality compression are performed through the tensor ring decomposition operator.

Benefits of technology

It significantly improves the credibility and availability of multi-source mass spectrometry metadata, enhances the accuracy and stability of mass spectrometry detection, and meets the requirements of high-quality scientific research and compliance auditing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935847A_ABST
    Figure CN121935847A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source mass spectrum metadata fusion method and system, and the method comprises the steps: extracting data, carrying out the hierarchical processing, loading a same timestamp and an instrument unique identifier, packaging, partitioning and storing the data into a metadata cache pool, inputting the mass spectrum metadata of each layer into a dual-channel adaptive fusion module, and obtaining preliminary fusion metadata; constructing a Markov random field undirected graph, iteratively calculating a global consistency score through a belief propagation algorithm, carrying out space-time consistency verification on preliminary fusion metadata, calculating a mass transfer entropy and guiding a fusion strategy, and outputting credible metadata; and carrying out space alignment and dimension compression on the credible metadata and the corresponding original mass spectrum data through a tensor ring decomposition operator, binding to generate a unified data packet, and storing the unified data packet in a fusion database. The method not only can improve the efficiency and accuracy of multi-source mass spectrum metadata fusion, but also has good interpretability, and can be directly applied to a multi-source mass spectrum metadata fusion system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mass spectrometry data analysis technology, and in particular to a method and system for multi-source mass spectrometry metadata fusion. Background Technology

[0002] With the widespread application of mass spectrometry technology in key fields such as life sciences, environmental monitoring, food safety, and pharmaceutical research and development, unprecedented demands have been placed on the accuracy, completeness, and reliability of mass spectrometry data. During mass spectrometry detection, the synergistic effect of multi-source metadata, including instrument parameters, sample information, and environmental conditions, directly impacts the accuracy of spectral analysis and the reliability of detection results. Therefore, the efficient fusion of multi-source mass spectrometry metadata is of great significance for promoting the upgrading of detection capabilities in various application fields.

[0003] Multi-source mass spectrometry (MS / MS) metadata is complex and diverse in dimensions, and the metadata from different sources suffers from problems such as format differences, spatiotemporal asynchrony, and inconsistent reliability. Traditional MS / MS metadata processing methods have the following drawbacks in practical applications: First, traditional methods often employ strategies of directly superimposing single-dimensional data or simple weighted fusion, failing to consider the inherent correlation and spatiotemporal consistency between metadata, severely impacting the accuracy of spectral analysis and detection results. Second, traditional methods lack dynamic evaluation and adaptive adjustment mechanisms for metadata quality, making it difficult to cope with the complex needs of different detection scenarios, and unable to effectively filter unreliable data, thus reducing the stability and robustness of the MS / MS detection system. Therefore, this invention proposes a method and system for multi-source MS / MS metadata fusion. Through a series of innovative technologies such as hierarchical processing, dual-channel adaptive fusion, Markov random field spatiotemporal consistency verification, and mass transfer entropy-guided fusion strategy, it achieves efficient integration and optimization of multi-source metadata, breaking through the limitations of traditional fusion methods, significantly improving the reliability and usability of fused data, providing a new solution for the intelligent and precise development of MS / MS detection technology, and having significant practical implications for promoting the deep application of MS / MS technology in various industries. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for multi-source mass spectrometry metadata fusion.

[0005] To achieve the above objectives, the present invention is implemented according to the following technical solution: This invention includes the following steps: The raw mass spectrometry spectrum data and corresponding multi-source mass spectrometry metadata are extracted through the mass spectrometer interface module. The multi-source mass spectrometry metadata is processed in layers, and each layer of mass spectrometry metadata and the corresponding spectrum data are loaded with the same timestamp and instrument unique identifier, packaged and partitioned and stored in the metadata cache pool. The mass spectrometry metadata of each layer is input into the dual-channel adaptive fusion module to obtain adaptive fusion weights, and the layered mass spectrometry metadata is weighted and fused to obtain preliminary fusion metadata. A Markov random field undirected graph is constructed based on the initial fused metadata node set. The global consistency score is iteratively calculated using the belief propagation algorithm to perform spatiotemporal consistency verification on the initial fused metadata. Calculate the quality transfer entropy of the preliminary fusion metadata that has passed the spatiotemporal consistency check, and guide the fusion strategy based on the quality transfer entropy to re-layer the corresponding mass spectrometry metadata and output reliable metadata; The trusted metadata and the corresponding original mass spectrometry data are spatially aligned and dimensionally compressed using the tensor ring decomposition operator, then bound to generate a unified data package and stored in the fusion database. The multi-source mass spectrometry metadata includes instrument parameters, sample information, and environmental conditions; The layered processing specifically involves preprocessing the corresponding multi-source mass spectrometry metadata according to the L1-instrument parameter layer, L2-sample information layer, and L3-environmental condition layer. The dual-channel adaptive fusion module includes a DS evidence channel and a dynamic Bayesian channel.

[0006] Furthermore, the method for obtaining preliminary fused metadata includes: The specific method for inputting the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain the adaptive fusion weights; Discretize the continuous hidden state space of the dynamic Bayesian channel into credibility assumptions to obtain the hypothesis space. , , , These correspond to credible assumptions, uncredible assumptions, and uncertain assumptions, respectively. Inputting each mass spectrometry metadata into the DS evidence channel yields the basic probability assignment, while inputting multi-source mass spectrometry metadata into the dynamic Bayesian channel yields the hypothesis probability, expressed as: ; ; in For mass spectrometry metadata The basic probability assignment represents the mass spectrometry metadata. For the hypothesis The level of support For the dynamic Bayesian channel pair hypothesis The assumed probability, Prior weights for mass spectrometry metadata are dynamically updated based on historical accuracy. The deviation between the current observation and the historical baseline value. , For mass spectrometry metadata Historical mean and standard deviation, For the probability of observation, for The observed variable at time, for The hidden state vector at time t. for Hidden state transition probability at time step. for, This is the state transition matrix from the previous time step. The process noise covariance matrix is... For the set of hidden state vectors, for Observed at all times The posterior distribution of the hidden states; The final fusion probability is calculated based on the basic probability assignment and the assumed probability, and adaptive fusion weights are generated, expressed as follows: ; ; in For mass spectrometry metadata In the assumption The final fusion probability, KL divergence measures the difference in probability distributions between the DS channel and the Bayesian channel. It is the numerical stability constant. For the Dempster-Shafer combinatorial operator, For mass spectrometry metadata parameter The generated adaptive fusion probability weights, This is the time-related decay coefficient. This is the time interval since the last calibration. For parameters The sensitivity coefficient, This is the prior probability; Based on the adaptive fusion probability weight, the mass spectrometry metadata from different instrument sources with the same parameters in the same layer is weighted and averaged to obtain preliminary fused metadata, and then spliced ​​in the same layer to obtain a preliminary fused metadata vector.

[0007] Furthermore, the method for constructing an undirected graph of a Markov random field includes: Each parameter of the initially integrated metadata is mapped to a graph node; the graph node numbering follows the "layer-category-serial number" rule; Edge relationships are determined using two methods: correlation analysis based on physical laws and correlation mining of historical data; the edge relationships include intra-layer edges and inter-layer edges.

[0008] 5. Furthermore, the method for performing spatiotemporal consistency verification includes: The expression for calculating the node potential energy and edge potential energy is as follows: ; ; in To initially integrate metadata The nodal potential energy corresponding to the nodes in the graph. To initially integrate metadata , The edge potential energy corresponding to the nodes and edges in the graph. , For graph nodes Predicted values ​​and tolerance bandwidth parameters, , The parameter is the historical mean. , The standard deviation parameter, The correlation coefficient of the graph nodes; The node information is iteratively updated using the belief propagation algorithm. After the node information converges, the marginal probability is calculated, expressed as: ; ; in To initially integrate metadata Graph nodes are Message update after the next iteration For nodes The set of neighboring nodes Exclude target nodes The subsequent set of neighboring nodes, For metadata Marginal probability of graph nodes This represents the number of iterations required for the node information to converge. The marginal probability of a graph node is taken as the corresponding global consistency score. Graph nodes with global consistency scores lower than the preset warning score threshold are removed, and the preliminary fusion metadata of the remaining graph nodes is output.

[0009] Furthermore, the method for calculating the quality transfer entropy of the preliminary fused metadata through spatiotemporal consistency verification includes: The marginal information entropy and conditional information entropy of each layer of data are calculated using the following expression: ; ; in for Layer parameters Mass spectrometry metadata Marginal information entropy, To initially integrate metadata Spectral metadata corresponding to the fusion mode Conditional information entropy, For discretized interval indexing, For the number of discretized intervals, For mass spectrometry metadata Falling into the range The empirical probability is obtained through batch sample statistics. To determine the number of discretized dimensions of the fused data, For mass spectrometry metadata Corresponding to the initial integration of metadata, To initially integrate metadata fusion models, For mass spectrometry metadata Located in the interval and initial integration of metadata lie in The joint probability of the fusion mode To initially integrate metadata lie in Mass spectrometry metadata in fusion mode Located in the interval The probability of; The difference between marginal information entropy and conditional information entropy is taken as the initial fused metadata. and corresponding mass spectrometry metadata The mass transfer entropy.

[0010] Furthermore, the method for outputting trusted metadata includes: By comparing the quality transfer entropy of each layer parameter with the preset entropy threshold, the preliminary fused metadata with a quality transfer entropy greater than the preset entropy threshold is directly output as trusted metadata. Conversely, if the fusion strategy is triggered, the adjusted preliminary fusion metadata will be associated with the fusion strategy operation features and output as trusted metadata; the fusion strategy operation features include parameter data layer adjustment records, parameter prior weight correction records, retest results, and data tags; The steps of the fusion strategy are as follows: The parameter-corresponding preliminary fusion metadata and mass spectrometry metadata are divided into the L3-environmental condition layer, the corresponding storage path is adjusted, Kalman filtering is used to correct the prior weights of the parameter-corresponding mass spectrometry metadata, and the source parameter-corresponding mass spectrometry metadata is retrieved from the automatically retained sample for secondary verification and data labeling.

[0011] Secondly, a multi-source mass spectrometry metadata fusion system includes: Cache module: used to extract raw mass spectrometry spectrum data and corresponding multi-source mass spectrometry metadata, perform layered processing on the multi-source mass spectrometry metadata, load the same timestamp and instrument unique identifier for each layer of mass spectrometry metadata and the corresponding spectrum data, and package and partition them to be stored in the metadata cache pool. Preliminary fusion module: This module is used to input the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain adaptive fusion weights, and then weight and fuse the mass spectrometry metadata of each layer to obtain preliminary fusion metadata. Spatiotemporal verification module: used to construct a Markov random field undirected graph based on the initial fused metadata node set, iteratively calculate the global consistency score through the belief propagation algorithm, and perform spatiotemporal consistency verification on the initial fused metadata; Quality assessment module: used to calculate the quality transfer entropy of the preliminary fusion metadata that has passed the spatiotemporal consistency check, and guide the fusion strategy according to the quality transfer entropy, re-layer the corresponding mass spectrometry metadata, and output reliable metadata; Storage module: Used to spatially align and compress the trusted metadata and the corresponding original mass spectrometry data using the tensor ring decomposition operator, bind them to generate a unified data package and store it in the fusion database.

[0012] The beneficial effects of this invention are: This invention provides a method and system for multi-source mass spectrometry metadata fusion. Compared with existing technologies, this invention has the following technical advantages: This invention, by processing multi-source mass spectrometry metadata in layers and using a dual-channel adaptive fusion module to calculate fusion weights, can fully explore the intrinsic value of metadata at each layer, effectively distinguish the credibility of metadata, and avoid the one-sidedness of a single fusion method. At the same time, by using a belief propagation algorithm for spatiotemporal consistency verification, inconsistent and unreliable metadata can be accurately eliminated, significantly improving the reliability of the initial fused metadata and laying a solid foundation for subsequent data processing. This invention introduces quality transfer entropy as a metadata quality assessment index. By comparing the difference between marginal information entropy and conditional information entropy, it can dynamically determine the quality level of the initially fused metadata and trigger targeted adjustments to the fusion strategy (adjusting data storage paths, correcting prior weights, and secondary verification), forming a closed loop of "fusion-assessment-feedback-optimization" and realizing adaptive optimization of the fusion strategy. This invention uses the tensor ring decomposition operator to spatially align and compress the dimensions of trusted metadata and the original mass spectrometry data. While ensuring data integrity, it effectively reduces data dimensionality, storage resource consumption, and subsequent computational complexity, thereby improving data processing efficiency. The key steps in the entire fusion process of this invention are all recorded and bound to the final output "trusted metadata". This provides a complete trust profile for each data point, enhances the interpretability of the data, and facilitates users to review results, trace the source of problems, and conduct in-depth analysis, thus meeting the core requirements of high-quality scientific research and compliance audit for data traceability. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating the steps of a multi-source mass spectrometry metadata fusion method according to the present invention. Detailed Implementation

[0014] The present invention will be further described below through specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.

[0015] The present invention provides a method and system for multi-source mass spectrometry metadata fusion, comprising the following steps: like Figure 1 As shown, this embodiment includes the following steps: The raw mass spectrometry spectrum data and corresponding multi-source mass spectrometry metadata are extracted through the mass spectrometer interface module. The multi-source mass spectrometry metadata is processed in layers, and each layer of mass spectrometry metadata and the corresponding spectrum data are loaded with the same timestamp and instrument unique identifier, packaged and partitioned and stored in the metadata cache pool. The mass spectrometry metadata of each layer is input into the dual-channel adaptive fusion module to obtain adaptive fusion weights, and the layered mass spectrometry metadata is weighted and fused to obtain preliminary fusion metadata. A Markov random field undirected graph is constructed based on the initial fused metadata node set. The global consistency score is iteratively calculated using the belief propagation algorithm to perform spatiotemporal consistency verification on the initial fused metadata. Calculate the quality transfer entropy of the preliminary fusion metadata that has passed the spatiotemporal consistency check, and guide the fusion strategy based on the quality transfer entropy to re-layer the corresponding mass spectrometry metadata and output reliable metadata; The trusted metadata and the corresponding original mass spectrometry data are spatially aligned and dimensionally compressed using the tensor ring decomposition operator, then bound to generate a unified data package and stored in the fusion database. The multi-source mass spectrometry metadata includes instrument parameters, sample information, and environmental conditions; The layered processing specifically involves preprocessing the corresponding multi-source mass spectrometry metadata according to the L1-instrument parameter layer, L2-sample information layer, and L3-environmental condition layer. The dual-channel adaptive fusion module includes a DS evidence channel and a dynamic Bayesian channel.

[0016] In this embodiment, the method for obtaining preliminary fused metadata includes: The specific method for inputting the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain the adaptive fusion weights; Discretize the continuous hidden state space of the dynamic Bayesian channel into credibility assumptions to obtain the hypothesis space. , , , These correspond to credible assumptions, uncredible assumptions, and uncertain assumptions, respectively. Inputting each mass spectrometry metadata into the DS evidence channel yields the basic probability assignment, while inputting multi-source mass spectrometry metadata into the dynamic Bayesian channel yields the hypothesis probability, expressed as: ; ; in For mass spectrometry metadata The basic probability assignment represents the mass spectrometry metadata. For the hypothesis The level of support For the dynamic Bayesian channel pair hypothesis The assumed probability, Prior weights for mass spectrometry metadata are dynamically updated based on historical accuracy. The deviation between the current observation and the historical baseline value. , For mass spectrometry metadata Historical mean and standard deviation, For the probability of observation, for The observed variable at time, for The hidden state vector at time t. for Hidden state transition probability at time step. for, This is the state transition matrix from the previous time step. The process noise covariance matrix is... For the set of hidden state vectors, for Observed at all times The posterior distribution of the hidden states; The final fusion probability is calculated based on the basic probability assignment and the assumed probability, and adaptive fusion weights are generated, expressed as follows: ; ; in For mass spectrometry metadata In the assumption The final fusion probability, KL divergence measures the difference in probability distributions between the DS channel and the Bayesian channel. It is the numerical stability constant. For the Dempster-Shafer combinatorial operator, For mass spectrometry metadata parameter The generated adaptive fusion probability weights, This is the time-related decay coefficient. This is the time interval since the last calibration. For parameters The sensitivity coefficient, This is the prior probability; Based on the adaptive fusion probability weight, the mass spectrometry metadata of different instrument sources with the same parameters in the same layer is weighted and averaged to obtain the preliminary fused metadata, and then spliced ​​in the same layer to obtain the preliminary fused metadata vector. In practical evaluations, taking the fusion of multi-source mass spectrometry metadata from the Thermo Q Exactive and Agilent 6495C dual-source system as an example, the raw mass spectrometry data extracted through the mass spectrometer interface module refers to the unprocessed ion current signal data directly generated by the mass spectrometer analyzer, specifically including but not limited to: Mass spectrum 1 (MS1): Full scan spectrum with a mass-to-charge ratio (m / z) range of 50-1500 Da and a resolution of 70,000 (at m / z 200), stored as an mzML file in Centroid or Profile mode; Mass spectrum 2 (MS2): Fragment ion spectrum generated based on data-dependent acquisition (DDA) mode, with step-wise collision energies of 15 / 30 / 45 eV, stored in MGF format; Ion chromatogram: Continuous signal curves of total ion current intensity (TIC) and extracted ion current (XIC) over time; The "multi-source" in multi-source mass spectrometry metadata specifically refers to different independently operating liquid chromatography-tandem mass spectrometry systems (different instrument sources), namely the Thermo Scientific Q Exactive HF-X high-resolution mass spectrometer (serial number ESI-20240815-001) and the Agilent 6495C triple quadrupole mass spectrometer (serial number G6495C-20231108). The specific contents of the layered pretreatment include: deviation calculation and standardization of L1-instrument parameter layer (capillary voltage, ion source temperature, collision energy, mass calibration parameters), data encoding of L2-sample information layer (pretreatment method, solvent type, concentration, biological matrix type), and deviation calculation and standardization of L3-environmental conditions layer (temperature, humidity, air pressure, operator, maintenance records). The ISO 8601 standard timestamp format (e.g., 2024-08-15T14:30:25.123Z) is adopted, accurate to the millisecond level. The instrument's unique identifier uses the first 16 characters of the mass spectrometer's serial number hash value (e.g., the MD5 hash of ESI-20240815-001, 3a7f9c2d8e1b5f6a) as the primary key for the metadata record. The naming rule for the spectral data file is: [hash identifier]_[timestamp]_[sample number].mzML, for example: 3a7f9c2d8e1b5f6a_20240815T143025_S12345.mzML; A three-level cache directory structure (instrument parameter layer, sample information layer, and environmental condition layer) is established on the local server. Each metadata file adopts the Apache Parquet columnar storage format (containing millisecond timestamp, instrument identifier, parameter name, observation value, and reference baseline value). Discretize the continuous hidden state space of the dynamic Bayesian channel into credibility assumptions to obtain the hypothesis space. The expression is: ; in for The mapping function, for The hidden state vector at time t. , For historical steady-state parameters, The steady-state coefficient, The hidden state standard deviation, This is the uncertainty tolerance threshold. , , These correspond to credible assumptions, uncredible assumptions, and uncertain assumptions, respectively. When obtaining preliminary fusion metadata, taking the Thermo Q Exactive source as an example, the prior weights of the mass spectrometry metadata are used in the DS evidence channel. Calculation under the assumed conditions The basic probability distributions are 0.818, 0.144, and 0.038. The hidden state vector includes three spatial model parameters: capillary aging factor (range 0-1, 0 for brand new, 1 for severely aged), ion source contamination index (calculated based on abnormal fluctuations in atomization current), and vacuum cavity leakage rate (based on the speed deviation of molecular turbopump). Based on this, the hypothesis probabilities of the dynamic Bayesian channel under each hypothesis are calculated to be 0.792, 0.168, and 0.040. The conflict factor is taken as 0.124 (indicating a 12.4% conflict between the two sources under the unreliable hypothesis). After merging and normalizing the two data sources, the basic probability combination is obtained as [0.756, 0.189, 0.055]. The final fusion probability of the Thermo Q Exactive source under each hypothesis is calculated as [0.792, 0.168, 0.04]. Taking the adaptive fusion weight calculation of the L1-instrument parameter "collision energy" as an example, the reciprocal of the instrument calibration half-life is taken as the aging decay coefficient. Time since last calibration Substituting the corresponding sensitivity coefficients, the adaptive fusion weights of the Thermo Q Exactive source and the Agilent 6495C source are calculated to be 0.765 and 0.312, respectively. The collision energies of 30.2 eV and 29.8 eV from the two data sources in the L1-instrument parameter layer are extracted and averaged to obtain the preliminary fusion metadata (collision energy), i.e., (0.756*30.2+0.312*29.8) / (0.756+0.312)=30.07 eV. The weighted average metadata of all parameters in the L1-instrument parameter layer is calculated in the same way to obtain the preliminary fusion metadata vector [3.49kV,30.07eV,321.5℃,0.998]. Similarly, the preliminary fusion metadata vectors of the L2-sample information layer and the L3-environmental condition layer are calculated.

[0017] In this embodiment, the method for constructing an undirected graph of a Markov random field includes: Each parameter of the initially integrated metadata is mapped to a graph node; the graph node numbering follows the "layer-category-serial number" rule; Edge relationships are determined using two methods: correlation analysis based on physical laws and correlation mining of historical data; the edge relationships include intra-layer edges and inter-layer edges. In actual evaluation, in the L1-instrument parameter layer, all parameters are continuous variables; in the L2-sample information layer, the pretreatment method code and solvent type code are discrete variables, while the logarithmic value of sample concentration and the inhibitory factor of biological matrix are continuous variables; in the L3-environmental conditions layer, temperature, humidity, air pressure, and maintenance records (number of days since the last maintenance) are continuous variables, while the operator (operator experience level) is a discrete variable. The intralayer edges of the L1-instrument parameter layer include (instrument parameter coupling relationships): capillary voltage-ion source temperature synergistic edge (strong positive correlation), collision energy-mass calibration error compensation edge (moderate negative correlation), capillary voltage-collision energy influence edge (weak negative correlation), etc. The intralayer edges of the L2 sample information layer include (sample information interaction relationships): pretreatment method - concentration correction edge (moderate negative correlation), solvent type - matrix inhibition edge (strong positive correlation), etc. The intra-layer edges of the L3-environmental conditions layer include (environmental condition synergy relationships): temperature-humidity coupling edge (strong negative correlation), personnel experience-maintenance cycle edge (moderate negative correlation), etc. Different layer-edge relationships include: capillary voltage-matrix inhibition compensation edge (strong positive correlation), collision energy-concentration adaptation edge (moderate negative correlation), ion source temperature-ambient temperature compensation edge (strong positive correlation), quality calibration error-maintenance cycle warning edge (moderate positive correlation), sample concentration-personnel experience adaptation edge (moderate negative correlation), etc.

[0018] 6. In this embodiment, the method for performing spatiotemporal consistency verification includes: The expression for calculating the node potential energy and edge potential energy is as follows: ; ; in To initially integrate metadata The nodal potential energy corresponding to the nodes in the graph. To initially integrate metadata , The edge potential energy corresponding to the nodes and edges in the graph. , For graph nodes Predicted values ​​and tolerance bandwidth parameters, , The parameter is the historical mean. , The standard deviation parameter, The correlation coefficient of the graph nodes; The node information is iteratively updated using the belief propagation algorithm. After the node information converges, the marginal probability is calculated, expressed as: ; ; in To initially integrate metadata Graph nodes are Message update after the next iteration For nodes The set of neighboring nodes Exclude target nodes The subsequent set of neighboring nodes, For metadata Marginal probability of graph nodes This represents the number of iterations required for the node information to converge. The marginal probability of a graph node is taken as the corresponding global consistency score. Graph nodes with global consistency scores lower than the preset warning score threshold are removed, and the preliminary fusion metadata of the remaining graph nodes is output. In a practical evaluation, 150 plasma samples collected over three days using a dual-source mass spectrometry system (Thermo Q Exactive HF-X and Agilent 6495C) were used to fully demonstrate the marginal probability calculation, iterative convergence process, and global consistency evaluation of the capillary voltage node. Sample numbers: Plasma-20240815-001 to Plasma-20240815-150; Time span: 2024-08-15-08:00 to 2024-08-17-18:00; The average voltage deviation after weighted fusion of this batch is 0.042kV, the historical mean deviation is 0.006kV, the tolerance bandwidth is 0.05kV, and the standard deviation of the deviation is 0.024kV. The neighbor graph nodes of the capillary voltage node include the ion source temperature node (ion source temperature fusion deviation), the collision energy node (collision energy shift), and the biological matrix type node (biological matrix inhibition factor). The node potential energies of the four nodes are calculated to be 0.772, 0.527, 0.835, and 0.98, and the edge potential energies are 0.053, 1.339, and 0.681. The node information is updated iteratively using the belief propagation algorithm. First, each node sends an initial message to its neighboring nodes, where... The process is iterated and updated multiple times until the message updates of each node are less than the convergence threshold of 10. -3 The iteration stops when convergence is reached (after 11 iterations), and the marginal probability of the capillary voltage node after convergence is calculated to be 0.0085. At this point, the marginal probability of 0.0085 is an unnormalized confidence value and needs to be converted into a globally consistent score, expressed as follows: ; in For global consistency scoring, The effective interval; Based on historical mean parameters 0.006, standard deviation parameter The global consistency score corresponding to a marginal probability of 0.0085 is calculated to be 0.682 (i.e., the observed value). The data falls near the right boundary of the interval, with a confidence ratio of approximately 68.2% within the valid interval. This is greater than the preset warning score threshold of 0.5, but lower than the preset qualified score threshold of 0.7. The initial fused metadata of the capillary voltage node passes the spatiotemporal consistency check, but the data consistency is low, triggering a secondary verification (automatically retrieving the original voltage feedback value of this batch to confirm whether it is caused by instantaneous grid fluctuations) and record tracing (marking the mass spectrometry metadata of this node as a yellow abnormal warning and writing it into the quality control log). Repeat the above steps to calculate the global consistency score for all nodes, and remove the preliminary fusion metadata of maintenance record nodes (corresponding data: number of days since the last maintenance) and humidity nodes (corresponding data: relative humidity deviation) from this batch of data.

[0019] In this embodiment, the method for calculating the quality transfer entropy of the preliminary fused metadata after spatiotemporal consistency verification includes: The marginal information entropy and conditional information entropy of each layer of data are calculated using the following expression: ; ; in for Layer parameters Mass spectrometry metadata Marginal information entropy, To initially integrate metadata Spectral metadata corresponding to the fusion mode Conditional information entropy, For discretized interval indexing, For the number of discretized intervals, For mass spectrometry metadata Falling into the range The empirical probability is obtained through batch sample statistics. To determine the number of discretized dimensions of the fused data, For mass spectrometry metadata Corresponding to the initial integration of metadata, To initially integrate metadata fusion models, For mass spectrometry metadata Located in the interval and initial integration of metadata lie in The joint probability of the fusion mode To initially integrate metadata lie in Mass spectrometry metadata in fusion mode Located in the interval The probability of; The difference between marginal information entropy and conditional information entropy is taken as the initial fused metadata. and corresponding mass spectrometry metadata The mass transfer entropy; In practical evaluation, taking the calculation of mass transfer entropy of initial fusion metadata from capillary voltage as an example, the data was divided into 10 equal intervals according to the historical value range. Fusion metadata from 150 plasma samples was statistically analyzed (a total of 300 mass spectrometry metadata from the two instrument sources; the same initial fusion metadata corresponds to two mass spectrometry metadata, and half the value was used when analyzing the mass spectrometry metadata). The sample frequency within each interval was calculated to determine the empirical probability that the mass spectrometry metadata falls into the corresponding interval (for example, if the sample frequency in the second interval is 22, the empirical probability that the mass spectrometry metadata falls into the corresponding interval is 0.147). K-means clustering was used to cluster the initial fusion metadata into 5 patterns (normal and stable, slight drift, matrix interference, environmental fluctuation, and abnormal pattern). The frequency of each pattern was counted to calculate the mass transfer entropy. The joint probability of the spectral metadata being located in the corresponding interval and the preliminary fusion metadata being located in the corresponding fusion mode (e.g., the normal stable sample frequency in the third interval is 41, so the corresponding joint probability is 0.28) is calculated. The probability of the mass spectrometry metadata being located in the corresponding interval when the preliminary fusion metadata is located in the corresponding fusion mode is calculated (e.g., the normal stable sample frequency in the third interval is 42, and the normal stable sample frequency is 89, so the corresponding conditional probability is 0.472). From this, the marginal information entropy of the L1 layer parameter capillary voltage is calculated to be 2.391 bits, and the conditional information entropy is calculated to be 2.041 bits. The mass transfer entropy of the L1 layer parameter capillary voltage is obtained by subtracting the values ​​to be 0.35 bits. Similarly, the mass transfer entropy of the remaining preliminary fusion metadata that have passed the spatiotemporal consistency check is calculated.

[0020] In this embodiment, the method for outputting trusted metadata includes: By comparing the quality transfer entropy of each layer parameter with the preset entropy threshold, the preliminary fused metadata with a quality transfer entropy greater than the preset entropy threshold is directly output as trusted metadata. Conversely, if the fusion strategy is triggered, the adjusted preliminary fusion metadata will be associated with the fusion strategy operation features and output as trusted metadata; the fusion strategy operation features include parameter data layer adjustment records, parameter prior weight correction records, retest results, and data tags; The steps of the fusion strategy are as follows: The parameter-corresponding preliminary fusion metadata and mass spectrometry metadata are divided into the L3-environmental condition layer, the corresponding storage path is adjusted, Kalman filtering is used to correct the prior weight of the parameter-corresponding mass spectrometry metadata, and the source-corresponding mass spectrometry metadata is retrieved from the automatically retained sample for secondary verification and data labeling. In actual assessments, when implementing the fusion strategy: The preliminary fusion metadata and mass spectrometry metadata corresponding to the L1-instrument parameter layer and L2-sample information layer parameters whose mass transfer entropy is less than or equal to the preset entropy threshold are assigned to the L3-environmental condition layer, and the storage path of the metadata cache pool is updated. The prior weights of the mass spectrometry metadata corresponding to the parameters are corrected using Kalman filtering. The specific steps are as follows (taking the correction of the prior weights of the temperature parameter mass spectrometry metadata as an example): (1) Establish the state-space model of the parameters (state equation and observation equation), and calculate the process noise Q and observation noise Z as 0.01. 2 0.05 2 Determine initial conditions: prior weights before correction. The error covariance is 0.72. It is 0.05 2 ; (2) Calculate the prediction step: Calculate the prediction prior weights based on the initial conditions. Prediction error covariance ; (3) Calculate the update step: First, calculate the Kalman gain. The expected mass transfer entropy is determined to be 0.15 bits. Based on the mass transfer entropy of -0.214 bits, the mass transfer entropy difference is calculated to be -0.364 bits, and the prior weights are updated accordingly. Update error covariance ; The retained samples of the 150 plasma samples in this batch were automatically located according to the sample number index table. 5% of the samples (8 samples, numbered 15 intervals) from the trigger batch were randomly selected for retesting. The same instrument (Thermo Q Exactive HF-X) was used, but the chromatographic column was changed and recalibrated to isolate the influence of environmental parameters. During the retest, the environmental temperature and humidity monitoring (controlled variable) was turned off, and only the data of the internal sensor of the instrument was recorded. The preliminary fusion metadata (capillary voltage deviation) was calculated to be 0.008kV (0.042kV before the retest). The absolute value of the deviation between the two measurements was calculated to be 0.034kV, which is greater than the instrument precision threshold of 0.02kV. Thus, the retest results showed that the environmental parameters of the L3 layer (laboratory temperature) interfered with the voltage stability, and this batch of data was marked as low-quality data. For the initial fused metadata with a preset entropy threshold, it is directly output as trusted metadata; otherwise, the fusion strategy is triggered, and the adjusted initial fused metadata is associated with the fusion strategy operation features and output as trusted metadata. Taking the fusion adjustment results of 150 plasma samples as an example, a total of 8 parameters of reliable metadata are output. A reliable metadata vector for each sample is constructed (with a dimension of 9, including 8 parameters and 1 global consistency score). The metadata of the 150 samples is stacked into a third-order tensor. The original mass spectrum was compressed into a low-rank approximation and then reconstructed into a third-order tensor. ; Define the tensor ring decomposition operator heterogeneous tensors and Mapping to a unified latent space Tensor ring rank Cross-validation was used to determine (balancing compression ratio and information retention, with a compression ratio target of 15%), thus constructing the kernel tensor sequence: , , ; Cross-modal similarity loss is used to force the metadata and spectral graph to be spatially aligned, as expressed by: ; in For cross-modal similarity loss, This is a modulo-1 product, i.e., the product of a tensor and a matrix. The metadata projection matrix is ​​initialized using a random Gaussian matrix. The spectral projection matrix is ​​initialized using a pre-trained autoencoder. The rank penalty coefficient, It is a random number; The kernel tensor is updated using alternating least squares (ALS) until the change in cross-modal similarity loss is less than 0.01, thus obtaining the unified latent space tensor. The dimension of the latent vector corresponding to each sample is compressed from the original 9+256=265 dimensions to 8 dimensions; From the unified hidden space tensor Extract each sample eigenvectors (latent representations) The 64-dimensional latent vectors are mapped to an 8-dimensional semantic space through sparse linear projection, i.e. ,in The projection matrix is ​​a block diagonal structure, with dimensions 1-4 corresponding to compressed instrument parameter semantics, dimensions 5-6 corresponding to compressed sample information semantics, dimension 7 corresponding to environmental residual semantics, and dimension 8 corresponding to quality control scoring semantics. Generating a fusion fingerprint ensures metadata-spectrum Figure 1 One-to-one correspondence, the expression is: ; in For the sample Fingerprint fusion For hash256 function, For the sample Trusted metadata vectors, The content hash of the original spectrum file; A single data packet file is generated for each batch. The spatially aligned and dimensionally compressed trusted metadata is bound to the corresponding original mass spectrometry data to generate a unified data packet and stored in the fusion database.

[0021] Secondly, a multi-source mass spectrometry metadata fusion system includes: Cache module: used to extract raw mass spectrometry spectrum data and corresponding multi-source mass spectrometry metadata, perform layered processing on the multi-source mass spectrometry metadata, load the same timestamp and instrument unique identifier for each layer of mass spectrometry metadata and the corresponding spectrum data, and package and partition them to be stored in the metadata cache pool. Preliminary fusion module: This module is used to input the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain adaptive fusion weights, and then weight and fuse the mass spectrometry metadata of each layer to obtain preliminary fusion metadata. Spatiotemporal verification module: used to construct a Markov random field undirected graph based on the initial fused metadata node set, iteratively calculate the global consistency score through the belief propagation algorithm, and perform spatiotemporal consistency verification on the initial fused metadata; Quality assessment module: used to calculate the quality transfer entropy of the preliminary fusion metadata that has passed the spatiotemporal consistency check, and guide the fusion strategy according to the quality transfer entropy, re-layer the corresponding mass spectrometry metadata, and output reliable metadata; Storage module: Used to spatially align and compress the trusted metadata and the corresponding original mass spectrometry data using the tensor ring decomposition operator, bind them to generate a unified data package and store it in the fusion database.

[0022] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for multi-source mass spectrometry metadata fusion, characterized in that, Includes the following steps: S1. Extract the raw mass spectrometry spectrum data and the corresponding multi-source mass spectrometry metadata through the mass spectrometer interface module. Perform layered processing on the multi-source mass spectrometry metadata. Load the same timestamp and instrument unique identifier onto each layer of mass spectrometry metadata and the corresponding spectrum data. Package and partition the metadata cache pool. S2. Input the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain the adaptive fusion weight, and then weight and fuse the mass spectrometry metadata of each layer to obtain the preliminary fusion metadata. S3. Construct a Markov random field undirected graph based on the preliminary fused metadata node set, and iteratively calculate the global consistency score through the belief propagation algorithm to perform spatiotemporal consistency verification on the preliminary fused metadata. S4. Calculate the quality transfer entropy of the preliminary fusion metadata that has passed the spatiotemporal consistency check, and guide the fusion strategy according to the quality transfer entropy, re-layer the corresponding mass spectrometry metadata, and output the reliable metadata. S5. Spatial alignment and dimensional compression of trusted metadata and corresponding original mass spectrometry data are performed using tensor ring decomposition operators, and then a unified data package is generated and stored in the fusion database. The multi-source mass spectrometry metadata includes instrument parameters, sample information, and environmental conditions; The layered processing specifically involves preprocessing the corresponding multi-source mass spectrometry metadata according to the L1-instrument parameter layer, L2-sample information layer, and L3-environmental condition layer. The dual-channel adaptive fusion module includes a DS evidence channel and a dynamic Bayesian channel.

2. The method for multi-source mass spectrometry metadata fusion according to claim 1, characterized in that, The method for obtaining preliminary fused metadata includes: The specific method for inputting the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain the adaptive fusion weights; Discretize the continuous hidden state space of the dynamic Bayesian channel into credibility assumptions to obtain the hypothesis space. , , , These correspond to credible assumptions, uncredible assumptions, and uncertain assumptions, respectively. Inputting each mass spectrometry metadata into the DS evidence channel yields the basic probability assignment, while inputting multi-source mass spectrometry metadata into the dynamic Bayesian channel yields the hypothesis probability, expressed as: ; ; in For mass spectrometry metadata The basic probability assignment represents the mass spectrometry metadata. For the hypothesis The level of support For the dynamic Bayesian channel pair hypothesis The assumed probability, Prior weights for mass spectrometry metadata are dynamically updated based on historical accuracy. The deviation between the current observation and the historical baseline value. , For mass spectrometry metadata Historical mean and standard deviation, For the probability of observation, for The observed variable at time, for The hidden state vector at time t. for Hidden state transition probability at time step. for, This is the state transition matrix from the previous time step. The process noise covariance matrix is... For the set of hidden state vectors, for Observed at all times The posterior distribution of the hidden states; The final fusion probability is calculated based on the basic probability assignment and the assumed probability, and adaptive fusion weights are generated, expressed as follows: ; ; in For mass spectrometry metadata In the assumption The final fusion probability, KL divergence measures the difference in probability distributions between the DS channel and the Bayesian channel. It is the numerical stability constant. For the Dempster-Shafer combinatorial operator, For mass spectrometry metadata parameter The generated adaptive fusion probability weights, This is the time-related decay coefficient. This is the time interval since the last calibration. For parameters The sensitivity coefficient, This is the prior probability; Based on the adaptive fusion probability weight, the mass spectrometry metadata of different instrument sources with the same parameters in the same layer is weighted and averaged to obtain the preliminary fused metadata, and then spliced ​​in the same layer to obtain the preliminary fused metadata vector.

3. The method for multi-source mass spectrometry metadata fusion according to claim 1, characterized in that, The method for constructing an undirected graph of a Markov random field includes: Each parameter of the initially integrated metadata is mapped to a graph node; the graph node numbering follows the "layer-category-serial number" rule; Edge relationships are determined using two methods: correlation analysis based on physical laws and correlation mining of historical data; the edge relationships include intra-layer edges and inter-layer edges.

4. The method for multi-source mass spectrometry metadata fusion according to claim 1, characterized in that, The method for performing spatiotemporal consistency verification includes: The expression for calculating the node potential energy and edge potential energy is as follows: ; ; in To initially integrate metadata The nodal potential energy corresponding to the nodes in the graph. To initially integrate metadata , The edge potential energy corresponding to the nodes and edges in the graph. , For graph nodes Predicted values ​​and tolerance bandwidth parameters, , The historical mean parameter, , The standard deviation parameter, The correlation coefficient of the graph nodes; The node information is iteratively updated using the belief propagation algorithm. After the node information converges, the marginal probability is calculated, expressed as: ; ; in To initially integrate metadata Graph nodes are Message update after the next iteration For nodes The set of neighboring nodes Exclude target nodes The subsequent set of neighboring nodes, For metadata Marginal probability of graph nodes This represents the number of iterations required for the node information to converge. The marginal probability of a graph node is taken as the corresponding global consistency score. Graph nodes with global consistency scores lower than the preset warning score threshold are removed, and the preliminary fusion metadata of the remaining graph nodes is output.

5. The method for multi-source mass spectrometry metadata fusion according to claim 1, characterized in that, The method for calculating the quality transfer entropy of the preliminary fused metadata through spatiotemporal consistency verification includes: The marginal information entropy and conditional information entropy of each layer of data are calculated using the following expression: ; ; in for Layer parameters Mass spectrometry metadata Marginal information entropy, To initially integrate metadata Spectral metadata corresponding to the fusion mode Conditional information entropy, For discretized interval indexing, For the number of discretized intervals, For mass spectrometry metadata Falling into the range The empirical probability is obtained through batch sample statistics. To determine the number of discretized dimensions of the fused data, For mass spectrometry metadata Corresponding to the initial integration of metadata, To initially integrate metadata fusion models, For mass spectrometry metadata Located in the interval and initial integration of metadata lie in The joint probability of the fusion mode To initially integrate metadata lie in Mass spectrometry metadata in fusion mode Located in the interval The probability of; The difference between marginal information entropy and conditional information entropy is taken as the initial fused metadata. and corresponding mass spectrometry metadata The mass transfer entropy.

6. The method for multi-source mass spectrometry metadata fusion according to claim 1, characterized in that, The method for outputting trusted metadata includes: By comparing the quality transfer entropy of each layer parameter with the preset entropy threshold, the preliminary fused metadata with a quality transfer entropy greater than the preset entropy threshold is directly output as trusted metadata. Conversely, if the fusion strategy is triggered, the adjusted preliminary fusion metadata will be associated with the fusion strategy operation features and output as trusted metadata; the fusion strategy operation features include parameter data layer adjustment records, parameter prior weight correction records, retest results, and data tags; The steps of the fusion strategy are as follows: The parameter-corresponding preliminary fusion metadata and mass spectrometry metadata are divided into the L3-environmental condition layer, the corresponding storage path is adjusted, Kalman filtering is used to correct the prior weights of the parameter-corresponding mass spectrometry metadata, and the source parameter-corresponding mass spectrometry metadata is retrieved from the automatically retained sample for secondary verification and data labeling.

7. A multi-source mass spectrometry metadata fusion system for performing the method according to any one of claims 1-6, characterized in that, include: Cache module: used to extract raw mass spectrometry spectrum data and corresponding multi-source mass spectrometry metadata, perform layered processing on the multi-source mass spectrometry metadata, load the same timestamp and instrument unique identifier for each layer of mass spectrometry metadata and the corresponding spectrum data, and package and partition them to be stored in the metadata cache pool. Preliminary fusion module: This module is used to input the mass spectrometry metadata of each layer into the dual-channel adaptive fusion module to obtain adaptive fusion weights, and then weight and fuse the mass spectrometry metadata of each layer to obtain preliminary fusion metadata. Spatiotemporal verification module: used to construct a Markov random field undirected graph based on the initial fused metadata node set, iteratively calculate the global consistency score through the belief propagation algorithm, and perform spatiotemporal consistency verification on the initial fused metadata; Quality assessment module: used to calculate the quality transfer entropy of the preliminary fusion metadata that has passed the spatiotemporal consistency check, and guide the fusion strategy based on the quality transfer entropy to re-layer the corresponding mass spectrometry metadata and output reliable metadata; Storage module: Used to spatially align and compress the trusted metadata and the corresponding original mass spectrometry data using the tensor ring decomposition operator, bind them to generate a unified data package and store it in the fusion database.

Citation Information

Patent Citations

  • Multi-source information fusion data space analysis method and system based on knowledge graph

    CN120217290A

  • Original mass spectrum data classification method based on multi-channel embedded representation

    CN120524271A

  • Card type DSL generation method for mass spectrum metadata standardization

    CN121212128A

  • Methods and systems for the industrial internet of things

    WO2019094729A1