Mass spectrum data solving method and solving model training method, device and equipment, medium

By extracting mass spectrometry parameters to construct spectral interpretation instructions and using pre-trained models for inference, the problem of poor accuracy in spectral interpretation results for small molecule compounds was solved. This resulted in an automated and reproducible spectral interpretation process, improved the accuracy and consistency of spectral interpretation results, reduced human bias, and saved time.

CN122091002APending Publication Date: 2026-05-26EXPERIMENTAL RES CENT CHINA ACAD OF CHINESE MEDICAL SCI +2
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EXPERIMENTAL RES CENT CHINA ACAD OF CHINESE MEDICAL SCI
Filing Date
2026-04-27
Publication Date
2026-05-26

Smart Images

  • Figure CN122091002A_ABST
    Figure CN122091002A_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, device, and medium for mass spectrometry data interpretation and interpretation model training. This disclosure achieves an automated and reproducible interpretation process from raw mass spectrometry data to compound structures and related information by extracting standardized mass spectrometry parameters, constructing interpretation instructions, and using interpretation models for inference. This reduces the interpretation threshold and human bias for laboratories at different levels, enhances the consistency of cross-platform and cross-batch data results, improves the accuracy of interpretation results, and saves time and increases interpretation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, equipment, and medium for mass spectrometry data interpretation and interpretation model training. Background Technology

[0002] In the annotation of small molecule compounds, the core of spectral deduction is to deduce the chemical structure of the compound from the mass spectrometry signal obtained from the detection instrument by analyzing the signal characteristics. Ultimately, the discrete mass spectrometry signal is transformed into clear structural information, providing a direct basis for the qualitative identification of the compound.

[0003] During the spectral interpretation process, it is usually necessary to compare the characteristics of the mass spectrometry signal with the characteristics of known compound signals in the standard spectral library, and combine multi-stage mass spectrometry to verify the fragmentation path, ion mobility to assist in structure differentiation and other techniques, and make subjective judgments based on expert experience.

[0004] However, expert experience relies heavily on case sharing, and spectral interpretation is highly subjective. The interpretation results for the same complex mass spectrum can vary greatly, making it impossible to guarantee the accuracy of the interpretation results. Summary of the Invention

[0005] To address the aforementioned technical issues, this disclosure provides a method, apparatus, device, and medium for mass spectrometry data interpretation and interpretation model training, thereby improving the accuracy of interpretation results.

[0006] In a first aspect, embodiments of this disclosure provide a method for interpreting mass spectrometry data, including: Mass spectrometry parameters are extracted from the raw mass spectrometry data. These parameters include the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of the fragment ions, their relative intensities, spectral retention time, and ion mode. Decoding commands are constructed based on preset command templates and mass spectrometry parameters; The spectral interpretation command is input into the pre-trained spectral interpretation model to obtain the spectral interpretation result output by the spectral interpretation model. The spectral interpretation command is used to guide the spectral interpretation model to infer the spectral interpretation result based on the mass spectrometry parameters. The spectral interpretation result includes the compound molecular formula corresponding to the original mass spectrometry data.

[0007] Secondly, embodiments of this disclosure provide a method for training a spectral analysis model, including: The initial model was self-supervised training was performed based on the mass spectrometry interpretation database to obtain a pre-trained model. A sample dataset was constructed based on the mass spectrometry signal features and annotation results of the original files. The annotation results include the target compounds corresponding to the mass spectrometry signal features and the fragmentation patterns of the target compounds. The pre-trained model is trained based on the sample dataset, a pre-set instruction set, and pre-set structured association knowledge to obtain an intermediate model; The parameters of the intermediate model are updated based on the output score of the intermediate model until the training termination condition is met, resulting in a well-trained spectral model. The output score represents the similarity between the output of the intermediate model and the labeled result.

[0008] Thirdly, embodiments of this disclosure provide a mass spectrometry data interpretation apparatus, comprising: The extraction module is used to extract mass spectrometry parameters from the raw mass spectrometry data. These parameters include the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of fragment ions, their relative intensities, spectral retention time, and ion mode. The first construction module is used to construct spectral interpretation instructions based on preset instruction templates and mass spectrometry parameters; The spectral decomposition module is used to input spectral decomposition commands into a pre-trained spectral decomposition model to obtain the spectral decomposition results output by the model. The spectral decomposition commands are used to guide the model to infer the spectral decomposition results based on the mass spectrometry parameters. The spectral decomposition results include the compound molecular formula corresponding to the original mass spectrometry data.

[0009] Fourthly, embodiments of this disclosure provide a spectral analysis model training apparatus, comprising: The first training module is used to perform self-supervised training on the initial model based on the mass spectrometry decomposition database to obtain a pre-trained model. The second construction module is used to construct a sample dataset based on the mass spectrometry signal features and annotation results of the original files. The annotation results include the target compounds corresponding to the mass spectrometry signal features and the fragmentation rules of the target compounds. The second training module is used to train the pre-trained model based on the sample dataset, the preset instruction set, and the preset structured association knowledge to obtain the intermediate model. The update module is used to update the parameters of the intermediate model based on the output score of the intermediate model until the training termination condition is met, and the trained spectral model is obtained. The output score represents the similarity between the output of the intermediate model and the labeled result.

[0010] Fifthly, embodiments of this disclosure provide an electronic device, including: Memory; Processor; and Computer programs; The computer program is stored in memory and configured to be executed by a processor to implement the methods described in the first and / or second aspects.

[0011] In a sixth aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the methods described in the first aspect and / or the second aspect.

[0012] The mass spectrometry data interpretation and interpretation model training methods, apparatus, equipment, and media provided in this disclosure, by extracting standardized mass spectrometry parameters, constructing interpretation instructions, and using interpretation models for inference, realize an automated and reproducible interpretation process from raw mass spectrometry data to compound structures and related information. This reduces the interpretation threshold and human bias for laboratories at different levels, enhances the consistency of cross-platform and cross-batch data results, improves the accuracy of interpretation results, and saves time for manual interpretation, thereby improving interpretation efficiency. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0014] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart of the mass spectrometry data interpretation method provided in the embodiments of this disclosure; Figure 2 A flowchart of a spectral analysis model training method provided in another embodiment of this disclosure; Figure 3 This is a schematic diagram of a spectral model training method provided in another embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of the mass spectrometry data interpretation device provided in the embodiments of this disclosure; Figure 5 This is a schematic diagram of the structure of the spectral model training device provided in the embodiments of this disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0016] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0017] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0018] The interpretation of spectra in the annotation of small molecule compounds specifically includes calculating possible molecular formulas based on accurate mass, assigning corresponding structural fragments by the relative abundance and fragmentation patterns of fragment ions, such as determining whether a fragment comes from the parent nucleus or a side chain, and distinguishing isomers by combining retention time. Ultimately, the discrete mass spectrometry signal is transformed into clear structural information, providing a direct basis for the qualitative analysis of the compound.

[0019] Spectral analysis relies primarily on three aspects: hardware-wise, it depends on high-resolution mass analyzers to achieve accurate mass measurement and multi-stage fragment acquisition, ensuring the accuracy of signal interpretation; data-wise, it relies on standard spectral libraries such as MassBank's "known compound-mass spectrum" matching and the accumulated "compound category-fragmentation rules" knowledge base; methodologically, it relies on techniques such as multi-stage mass spectrometry to verify fragmentation pathways and ion mobility-assisted structure differentiation to improve the reliability of structure derivation. Current experience accumulated in spectral analysis includes: mastering the characteristic fragmentation rules of common compounds such as flavonoids (glucose desaccharides), esters (dealkoxylation), and alkaloids (dealkylation), which can quickly narrow down the structural range; and familiarity with the effects of different ionization methods—electrospray ionization (ESI), atmospheric pressure chemical ionization (APCI), and atmospheric pressure photoionization ionization (APPI)—on fragment generation, for example, ESI readily generates adduct ions. Based on the above experience, co-emission peak interpretation techniques are developed, such as combining the slight differences in retention time with the uniqueness of fragment ions to separate mixed mass spectrometry signals and improve the interpretation accuracy of single compounds in complex matrices.

[0020] However, its implementation suffers from problems such as the tacit nature of knowledge, difficulties in its transmission, and the subjectivity of spectral interpretation criteria. Spectral interpretation relies on experts' subjective judgment of complex fragments, such as the attribution of fragments from novel heterocyclic or bridged ring skeletal compounds, and the rationale for weak signal fragments with an intensity less than 5% of the base peak. The core logic is difficult to express in a structured way; novices need long-term accumulation to master the "interpretation of abnormal fragmentation pathways," and expert experience is largely based on case sharing. There is no unified spectral interpretation training system, leading to significant differences in interpretation results for the same complex mass spectrum. Furthermore, there is no unified standard for spectral library matching thresholds, making the interpretation of low-match fragment spectra easily influenced by personal experience. Mass spectrometric fragments of isomers have identical compositions, with only slight differences in abundance. The subjective judgment of abundance thresholds during spectral interpretation can easily lead to structural misjudgments.

[0021] To address the aforementioned issues, this disclosure provides a method for interpreting mass spectrometry data. The method will be described below with reference to specific embodiments.

[0022] Figure 1A flowchart illustrating the mass spectrometry data interpretation method provided in this embodiment of the disclosure. Figure 1 As shown, the specific steps included in this method are as follows: S101. Extract mass spectrometry parameters from raw mass spectrometry data.

[0023] The mass spectrometry parameters include the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of the fragment ions, their relative intensities, the spectral retention time, and the ion mode.

[0024] Raw mass spectrometry data are raw signal files directly acquired by detection instruments such as liquid chromatography-mass spectrometry (LC-MS), containing complete mass spectrometry signal information. Optionally, the raw mass spectrometry data are in mzML format.

[0025] The parent ion is a complete ion of a compound molecule that has not broken down after ionization in an ion source. The mass-to-charge ratio of the parent ion refers to the ratio of the mass of the parent ion to its charge.

[0026] Fragment ions are secondary ions formed after the parent ion undergoes chemical bond breakage in the collision chamber. The relative intensity of fragment ions is the percentage of the signal intensity of each fragment ion relative to the total signal intensity of all fragment ions, used to reflect the abundance distribution of fragment ions.

[0027] The retention time of a chromatogram is the time it takes for a compound to reach the detector after separation in the chromatographic column. It is related to the compound's physicochemical properties such as polarity and molecular weight, and can be used to assist in screening compound categories.

[0028] Ion mode refers to the ionization mode of compounds in an ion source, which is divided into positive ion mode and negative ion mode. The ionization behavior of compounds differs under different modes, affecting the direction of deduction of the fragmentation path.

[0029] The raw mass spectrometry data is analyzed using analytical tools to extract the corresponding mass spectrometry parameters. Optionally, the pymzML library from the Python ecosystem is used as the core analytical tool, combined with NumPy for efficient numerical computation and processing. Specifically, pymzML.run.Reader is used to read the mzML file and parse the first-level mass spectrometry scan event (denoted as MS1) and the second-level mass spectrometry scan event (denoted as MS2).

[0030] MS1 is an event in which all ions entering the mass spectrometer during the chromatographic elution process are scanned across the entire range, and MS2 is an event in which the precursor ions selected by the MS1-level scan are subjected to targeted fragmentation scanning.

[0031] Obtain the mass-to-charge ratio of the parent ion from the MS2 scan, extract fragment ion peaks with relative intensities greater than a preset intensity threshold, and record the mass-to-charge ratio and relative intensity of the corresponding fragment ions.

[0032] The retention times corresponding to the precursor ions were matched and extracted from the total ion chromatogram (TIC) of the MS1 scan.

[0033] Optionally, parameters such as ion mode and collision energy can also be read from the metadata of the mzML file. Simultaneously, threshold filtering and smoothing algorithms are used to reduce noise signals such as low-intensity fragment peaks in the spectrum.

[0034] Finally, a structured JSON file containing mass spectrometry parameters is generated for later use.

[0035] S102. Construct spectral interpretation instructions based on preset instruction templates and mass spectrometry parameters.

[0036] The pre-defined instruction template refers to a structured language framework based on the professional logic of mass spectrometry analysis. Based on this template, numerical and categorical features are uniformly organized into natural language instructions, which serve as input to the spectral analysis model. This instruction template integrates multi-dimensional features with structured language descriptions, maintaining semantic coherence and domain logic, without requiring complex feature embedding or cross-modal alignment computation.

[0037] Numerical features include the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of fragment ions, their relative intensities, and spectral retention time; categorical features include ion patterns.

[0038] S103. Input the spectral decomposition command into the pre-trained spectral decomposition model to obtain the spectral decomposition result output by the spectral decomposition model.

[0039] The spectral interpretation command is used to guide the spectral interpretation model to infer the spectral interpretation results based on the mass spectrometry parameters. The spectral interpretation results include the compound molecular formula corresponding to the original mass spectrometry data.

[0040] Optionally, the spectral results may also include the target fragmentation pathway and fragmentation mechanism analysis for each compound's molecular formula.

[0041] The target fragmentation pathway refers to the sequence of chemical bond breaking of compounds and the process of fragment ion generation that the model infers, which are consistent with the original mass spectrometry data. It is used to explain how the parent ion breaks to form various fragment ions.

[0042] Analysis of cleavage mechanisms refers to the explanation of the chemical principles underlying the target cleavage pathway, such as the thermodynamic basis of chemical bond breaking and the reaction characteristics of functional groups, making the decision-making process transparent and facilitating understanding and verification by chemical experts.

[0043] The spectral decomposition command in the preceding steps is used as the input to the spectral decomposition model, and the spectral decomposition command is used to guide the spectral decomposition model to reason based on the mass spectrometry parameters to obtain the most likely molecular formula and its supporting evidence.

[0044] The process of reasoning based on mass spectrometry parameters in the spectral analysis model includes: calculating the corresponding molecular weight according to the ion mode and the mass-to-charge ratio of the parent ion; analyzing fragment ion peaks to infer possible fragmentation paths and fragmentation patterns; combining retention time and molecular weight information to exhaustively list possible molecular formulas under constraints, and finally outputting the most likely compound molecular formula and its supporting evidence.

[0045] This disclosure, through the extraction of standardized mass spectrometry parameters, the construction of interpretation instructions, and the use of interpretation models for inference, achieves an automated and reproducible interpretation process from raw mass spectrometry data to compound structures and related information. This reduces the interpretation threshold and human bias for laboratories at different levels, enhances the consistency of cross-platform and cross-batch data results, improves the accuracy of interpretation results, and saves time for manual interpretation, thereby increasing interpretation efficiency.

[0046] Based on the above embodiments, the spectral decomposition model is used to implement the following steps: S1031. Determine at least one ion peak correlation based on the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of the fragment ions, and their relative intensities.

[0047] Specifically, the mass difference between the fragment ions and the parent ion is calculated based on the mass-to-charge ratio of the parent ion and the mass-to-charge ratio of the fragment ions; at least one ion peak correlation is determined based on the mass difference and relative intensity.

[0048] Among them, ion peak correlation characterizes the correlation between fragment ions with relative intensities greater than a preset intensity threshold and the parent ion, indicating that the fragment ion originates from the breakage of the parent ion.

[0049] Mass difference refers to the numerical difference between the mass-to-charge ratio of a single fragment ion and that of the parent ion. The spectral analysis model calculates the mass difference between the fragment and the parent ion through a self-attention mechanism, reflecting the amount of structural fragment missing by the fragment ion relative to the parent ion.

[0050] A preset intensity threshold is used to filter low-intensity noise signals, such as instrument noise or background interference. For example, the preset intensity threshold is set to 5%.

[0051] If the mass difference corresponds to the molecular weight of a known functional group or structural fragment, then a correlation is determined between the fragment ion and the parent ion.

[0052] S1032. Determine the reference chemical bond breaking site and reference cleavage path based on the ion mode and ion peak correlation.

[0053] Reference bond breakage site refers to the specific location of the parent ion's bond breakage, inferred from the correlation between ion patterns and ion peaks. Reference fragmentation pathway refers to the breakage process in which the parent ion breaks to form fragment ions.

[0054] Because compounds exhibit different ionization characteristics under different ion modes, at least one type of cleavage reaction can be determined based on ion models. By analyzing the mass difference in the correlation of ion peaks, missing structural fragments can be deduced, leading to the identification of the corresponding broken chemical bonds. Based on the cleavage reaction type and the broken chemical bonds, and combined with prior chemical knowledge, reference bond cleavage sites and reference fragmentation pathways can be obtained. For example, it can be determined whether a fragment originates from a Maxwell rearrangement or a reverse Diels-Alder reaction.

[0055] S1033. Based on the spectral retention time, ion mode, and mass of the parent ion, determine at least one candidate molecular formula.

[0056] Specifically, based on the spectral retention time, at least one target compound category is determined from multiple candidate compound categories; and at least one candidate molecular formula is calculated based on the ion mode and the mass of the parent ion.

[0057] Among them, the compounds corresponding to the candidate molecular formulas belong to at least one target compound category.

[0058] Candidate compound categories include ketones, alkaloids, and lipids.

[0059] The target compound category refers to the category of compounds determined after screening based on spectral retention time.

[0060] Specifically, using a pre-defined compound category-retention time correlation model, categories of candidate compounds that do not match the retention time of the current spectrum are removed, resulting in at least one target compound category.

[0061] The mass of the parent ion is obtained after correcting for its mass-to-charge ratio based on the ion model. A candidate molecular formula refers to a chemical formula that conforms to the rules regarding the mass of the parent ion and the elemental composition of the target compound category.

[0062] Specifically, based on the elemental composition rules of the target compound category and combined with prior chemical knowledge, at least one candidate molecular formula that meets the mass of the parent ion is listed to obtain a set of molecular formulas.

[0063] S1034. Based on the reference fragmentation path and the spectrum retention time, determine the compound molecular formula corresponding to the original mass spectrometry data from at least one candidate molecular formula.

[0064] In addition, the target fragmentation pathway and fragmentation mechanism corresponding to each compound molecular formula were analyzed.

[0065] The target fragmentation pathway refers to the fragmentation pathway that matches the mass spectrometry parameters after verification by candidate molecular formula matching. Fragmentation mechanism analysis refers to the explanation of the chemical principles of the target fragmentation pathway, including the thermodynamic basis of chemical bond breaking and the reaction characteristics of functional groups.

[0066] The theoretical fragmentation pathways of each candidate molecular formula are verified to be consistent with the reference fragmentation pathways. At the same time, the deviation between the retention time of the candidate molecular formula standard and the spectral retention time is verified to be less than the preset deviation threshold. Finally, the molecular formulas of the compounds that meet the requirements and the corresponding compound names are determined.

[0067] This disclosure utilizes the reasoning capabilities of spectral analysis models, based on standardized computational logic and chemical principles, and deeply integrates multi-dimensional data such as retention time to achieve accurate simulation of the fracture behavior of multifunctional compounds and reliable traceability of ion sources. This ensures both the reproducibility of the operation and the accuracy of the results.

[0068] Furthermore, addressing the technical bottlenecks of complex structure analysis difficulties and the strong subjectivity in isomer differentiation, this disclosure leverages the powerful reasoning capabilities of a large language model, deeply integrating multi-dimensional data such as retention time to achieve accurate simulation of the fracture behavior of multifunctional compound compounds and reliable tracing of ion origins. The system significantly improves the accuracy of complex structure analysis and isomer differentiation precision by quantitatively evaluating fragment abundance differences and multi-dimensional feature correlations.

[0069] Figure 2 A flowchart of a spectral analysis model training method provided in another embodiment of this disclosure is shown below. Figure 2 As shown, the method includes the following steps: S201. Self-supervised training of the initial model is performed based on the mass spectrometry interpretation database to obtain the pre-trained model.

[0070] The mass spectrometry interpretation database is a comprehensive corpus that integrates mass spectrometry interpretation-related literature, standard mass spectrometry data, public spectral library information, instrument characteristic data, etc., including but not limited to: mass spectrometry data interpretation-related corpus, literature on compound fragmentation patterns, textual descriptions of standard mass spectrometry data, expert experimental data, data on the correlation between instrument parameters and signal characteristics, and textual data from public spectral libraries.

[0071] The initial model refers to a basic large language model with fundamental natural language understanding and logical reasoning capabilities. Self-supervised training refers to a training method that does not require manual labeling, allowing the model to automatically learn the potential patterns and relationships in the data through its own analysis and mining.

[0072] Through self-supervised training, the initial model acquires a basic understanding of the correlation between mass spectrometry signals and compounds, establishes basic language representations and chemical prior knowledge for mass spectrometry analysis, and obtains a pre-trained model with basic understanding of mass spectrometry interpretation and preliminary reasoning ability.

[0073] S202. Construct a sample dataset based on the mass spectrometry signal features and annotation results of the original files.

[0074] The annotation results include the target compounds corresponding to the mass spectrometry signal features and the fragmentation patterns of the target compounds.

[0075] Mass spectrometry signal characteristics refer to normalized mass spectrometry parameters extracted from the original file, such as the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of fragment ions, their relative intensities, spectral retention times, and ion modes. The annotation results are the actual spectral interpretation results determined by experts based on the characteristics of the mass spectrometry signal, including the molecular formula and chemical name of the target compound, the fragmentation path of the target compound, the breakup site, etc.

[0076] The mass spectrometry signal features are correlated with the annotation results to form a sample dataset for further model fine-tuning.

[0077] Optionally, the sample dataset includes an adversarial dataset. First, adversarial data is generated, which includes synthetic mass spectrometry data and / or noisy sample data. The adversarial data is then further mixed with mass spectrometry signal features to obtain the adversarial dataset.

[0078] Adversarial data refers to mass spectrometry data containing invalid signals or interference information generated to improve the model's anti-interference ability and generalization performance, such as chemically meaningless fragment peaks, interfering ion signals, baseline drift, peak overlap, weak signal interference, etc.

[0079] Adversarial data is mixed with labeled mass spectrometry signal features to obtain an adversarial dataset. Then, noise is uniformly processed on both the adversarial dataset and the clean dataset, such as by applying threshold filtering and smoothing algorithms to simulate the data processing flow of real instruments. This forces the pre-trained model to learn to distinguish between effective signals and interference, thereby significantly improving the robustness and generalization ability of the final spectral resolution model in complex real-world scenarios. It also enhances the model's ability to resolve weak signal fragments and its anti-interference performance in complex samples, ensuring the reliability and stability of the method in practical applications.

[0080] The clean dataset consists of labeled mass spectrometry signal features and does not contain adversarial data.

[0081] S203. Train the pre-trained model based on the sample dataset, the preset instruction set, and the preset structured association knowledge to obtain the intermediate model.

[0082] The pre-set instruction set refers to a set of structured instruction templates designed based on the tasks corresponding to the inference nodes in the mass spectrometry interpretation process, which guide the model to complete specific inference behaviors. It is a Chain-of-Thought (COT) inference mechanism instruction, used to guide the pre-trained model to complete the inference task during the training process, so that the pre-trained model can complete the spectral interpretation inference process according to standardized logic and obtain the spectral interpretation result based on the input sample data.

[0083] The reasoning tasks include one or more of the following: judging the correctness of ion peak parameter extraction, matching fragment peaks with parent ion peaks, inferring fragmentation paths, matching retention time with compound category, determining ion mode with parent ion peak type, calculating molecular formulas, matching compound names, and generating spectral interpretation results.

[0084] Optionally, the preset instruction set also includes classification instructions, which are used to require the model to perform normal inference on the sample data in the pure dataset and to identify unanalyzable noise or point out its irrationality on the sample data in the adversarial dataset when the sample dataset includes both adversarial dataset and pure dataset.

[0085] Pre-defined structured association knowledge refers to a set of knowledge that integrates key rules in the field of mass spectrometry analysis, including the correspondence between functional groups and fragment peak features, and the association rules between retention time intervals and compound polarity. This pre-defined structured association knowledge is injected into the parameter space of the pre-trained model to enhance its ability to map signal features to structural properties.

[0086] Specifically, the pre-defined structured knowledge is transformed into a vector representation that the model can understand. During training, the vector representation is fused with the model's word embedding vectors to ensure that the model can automatically invoke relevant knowledge rules when processing mass spectrometry parameters.

[0087] For example, Quantized LLMs with Low-Rank Adapters (QLoRA) can be used for efficient parameter fine-tuning, and pre-defined structured relational knowledge can be injected into the model through knowledge distillation.

[0088] Optionally, queries for cleavage patterns can be encapsulated as Model Control Protocol (MCP) functions for the model to call during inference, thereby improving the model's stability and accuracy.

[0089] S204. Update the parameters of the intermediate model based on the output score of the intermediate model until the training termination condition is met, and obtain the trained spectral model.

[0090] Among them, the output score characterizes the similarity between the output of the intermediate model and the labeled results. For example, the intermediate model's output is scored by experts, and the scoring indicators include the degree of matching of the cleavage path and the consistency of compounds, or the model performs a self-evaluation based on the standard data corresponding to the labeled results.

[0091] For example, a reinforcement learning from human feedback (RLHF) mechanism can be used to construct a reward model to train and optimize intermediate models. The reward criteria include the degree of matching between the inferred fragmentation path and the actual fragmentation pattern of the standard, the consistency between the compound inference results and the real information of the standard, the completeness and logic of the explanation of the spectral interpretation criteria, and the stability of the spectral interpretation results of the same sample in different batches.

[0092] Optionally, a binary classification model based on the Bidirectional Encoder Representations from Transformers (BERT) architecture can be used to output a reward value, which can then be used to update the parameters of the intermediate model.

[0093] The parameters of the intermediate model are adjusted based on the error reflected in the output score. For example, the model strategy is iteratively optimized using the reinforcement learning algorithm (Proximal Policy Optimization, PPO) to improve the output score of the intermediate model, reduce the confidence bias of the spectral results, enhance the model's robustness to spectral analysis of weak signal fragments and complex matrix samples, and improve its generalization ability in practical applications.

[0094] Training termination conditions refer to the preset criteria for judging whether the model training is complete, such as the output score of the validation set in the sample dataset reaching a threshold, the score fluctuation stabilizing, or the number of iterations reaching a preset number.

[0095] The iteration stops after the training termination condition is met, and the trained spectral model is obtained.

[0096] This disclosure embodiment uses a unified model to model massive amounts of professional literature, standard spectral libraries, and expert interpretation logic, covering multi-dimensional spectral knowledge such as ionization mode, fragmentation path, fragment characteristics, and retention time. It transforms traditionally difficult-to-express expert experience into calculable and reusable structured knowledge, significantly reducing the dependence of mass spectrometry analysis on senior experts and effectively solving the problem of poor consistency of spectral results between different laboratories and different operators.

[0097] In addition, this embodiment of the present disclosure introduces a COT inference mechanism and an MCP method embedded in a large model to guide the model to perform step-by-step inference, self-verification and reflective optimization, thereby ensuring the logical rigor and output accuracy of the spectral interpretation process.

[0098] In some embodiments, this disclosure provides a system for deep analysis and reasoning of mass spectrometry data using a hierarchical architecture, including an mzML structured analysis layer, a spectroscopic pre-trained large language model layer, a feature embedding and semantic fusion layer, and a reasoning and MCP control layer.

[0099] The mzML structured parsing layer is used to extract key parameters from the original file, including the mass-to-charge ratio of the parent ion and fragment ions, their relative intensities, retention times, ion patterns, etc., and convert them into normalized structured data that the model can process.

[0100] The pre-trained large language model layer for spectral analysis is based on the Transformer architecture and relies on a massive amount of professional mass spectrometry corpus, covering fragmentation rules, instrument response characteristics, standard spectral library data, etc., for pre-training, so that the intermediate model can deeply grasp the mapping relationship between compound fracture rules and mass spectrometry signals.

[0101] The feature embedding and semantic fusion layer, based on a pre-defined instruction template, organizes numerical features, including mass-to-charge ratio, relative intensity, and retention time, as well as categorical features, i.e., ion patterns, into natural language instructions, which serve as input to the spectral analysis model. This instruction template integrates multi-dimensional features with structured language descriptions, maintaining semantic coherence and domain logic, without requiring complex feature embedding or cross-modal alignment calculations.

[0102] The reasoning and MCP control layer introduces the COT reasoning mechanism and the MCP method embedded in the spectral interpretation model to guide the spectral interpretation model to perform step-by-step reasoning, self-verification, and reflective optimization, ensuring the logical rigor and output accuracy of the spectral interpretation process.

[0103] Figure 3 This is a schematic diagram of a spectral analysis model training method provided in another embodiment of this disclosure. Figure 3 As shown, data parsing and annotation were performed based on the original mzML file and literature data to extract mass spectrometry parameters such as the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of fragment ions, relative intensity, spectral retention time, and ion mode. Noise processing was performed on the extracted data to eliminate the influence of low-peak intensity data.

[0104] Further, inference mechanism instructions (COT instructions) are generated. A structured instruction set is constructed based on chemical rules, including but not limited to the relationship between ion modes and parent ions, the relationship between fragment ion peaks and fragmentation patterns, and the relationship between retention time and compounds. Synthesized adversarial noise data is mixed into the sample data. During model training, COT instructions are used to guide the model to perform step-by-step inference based on the sample data, and finally a trained spectral model is obtained.

[0105] Figure 4 This is a schematic diagram of the mass spectrometry data interpretation device provided in this embodiment of the disclosure. The mass spectrometry data interpretation device provided in this embodiment of the disclosure can execute the processing flow provided in the method embodiment of the mass spectrometry data interpretation device, such as... Figure 4As shown, the mass spectrometry data interpretation device 40 includes: an extraction module 41, a first construction module 42, and an interpretation module 43. The extraction module 41 extracts mass spectrometry parameters from the raw mass spectrometry data, including the mass-to-charge ratio of the parent ion, the mass-to-charge ratio and relative intensity of fragment ions, spectral retention time, and ion mode. The first construction module 42 constructs interpretation instructions based on a preset instruction template and the mass spectrometry parameters. The interpretation module 43 inputs the interpretation instructions into a pre-trained interpretation model to obtain the interpretation results output by the model. The interpretation instructions guide the model to infer the interpretation results based on the mass spectrometry parameters. The interpretation results include the compound molecular formula corresponding to the raw mass spectrometry data.

[0106] Optionally, the spectral interpretation module 43 is specifically used to determine at least one ion peak correlation based on the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of the fragment ions, and their relative intensities; to determine the reference chemical bond breaking site and the reference fragmentation path based on the ion mode and the ion peak correlation; to determine at least one candidate molecular formula based on the spectrum retention time, ion mode, and the mass of the parent ion; and to determine the compound molecular formula corresponding to the original mass spectrometry data from at least one candidate molecular formula based on the reference fragmentation path and the spectrum retention time.

[0107] Optionally, the spectrum interpretation module 43 is specifically used to calculate the mass difference between the fragment ion and the mother ion based on the mass-to-charge ratio of the mother ion and the mass-to-charge ratio of the fragment ion; and to determine at least one ion peak correlation based on the mass difference and relative intensity. The ion peak correlation characterizes the correlation between the fragment ion and the mother ion whose relative intensity is greater than a preset intensity threshold.

[0108] Optionally, the spectral interpretation module 43 is specifically used to determine at least one target compound category from multiple candidate compound categories based on the spectral retention time; and to calculate at least one candidate molecular formula based on the ion mode and the mass of the parent ion, wherein the compound corresponding to the candidate molecular formula belongs to at least one target compound category.

[0109] Optionally, the spectrum interpretation module 43 is also used to generate the target fragmentation path and fragmentation mechanism analysis corresponding to each compound molecular formula.

[0110] Figure 4 The mass spectrometry data interpretation device shown in the embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.

[0111] Figure 5 This is a schematic diagram of the structure of the spectral model training device provided in the embodiments of this disclosure. The spectral model training device provided in the embodiments of this disclosure can execute the processing flow provided in the embodiments of the spectral model training device method, such as... Figure 5As shown, the spectral analysis model training device 50 includes: a first training module 51, a second construction module 52, a second training module 53, and an update module 54. The first training module 51 is used to perform self-supervised training on the initial model based on the mass spectrometry spectral analysis database to obtain a pre-trained model. The second construction module 52 is used to construct a sample dataset based on the mass spectrometry signal features and annotation results of the original files. The annotation results include the target compounds corresponding to the mass spectrometry signal features and the fragmentation rules of the target compounds. The second training module 53 is used to train the pre-trained model according to the sample dataset, a preset instruction set, and preset structured association knowledge to obtain an intermediate model. The update module 54 is used to update the parameters of the intermediate model based on the output result score of the intermediate model until the training termination condition is met to obtain a trained spectral analysis model. The output result score represents the similarity between the output result of the intermediate model and the annotation result.

[0112] Optionally, the sample dataset includes an adversarial dataset, and the second building module 52 is also used to generate adversarial data, which includes synthetic mass spectrometry data and / or noisy sample data; the adversarial data is mixed with mass spectrometry signal features to obtain the adversarial dataset.

[0113] Optionally, a preset instruction set is used to guide the pre-trained model to complete inference tasks, which include one or more of the following: judging the correctness of ion peak parameter extraction, matching fragment peaks with parent ion peaks, inferring fragmentation paths, matching retention time with compound category, determining ion mode with parent ion peak type, calculating molecular formulas, matching compound names, and generating spectral interpretation results.

[0114] Optionally, the second training module 53 is also used to inject preset structured association knowledge into the parameter space of the pre-trained model. The preset structured association knowledge includes the correspondence between functional groups and fragment peak features, and the association rules between retention time intervals and compound polarity.

[0115] Figure 5 The spectral model training device shown in the embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.

[0116] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The electronic device provided in this embodiment of the mass spectrometry data interpretation method can execute the processing flow provided in the embodiment of the mass spectrometry data interpretation method, such as... Figure 6 As shown, the electronic device 60 includes: a memory 61, a processor 62, a computer program, and a communication interface 63; wherein the computer program is stored in the memory 61 and is configured to be executed by the processor 62 as described above, using the mass spectrometry data descaling method and / or the descaling model training method.

[0117] In addition, this disclosure also provides a computer-readable storage medium storing a computer program thereon, which is executed by a processor to implement the mass spectrometry data despectroscopy method and / or despectroscopy model training method described in the above embodiments.

[0118] Furthermore, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the mass spectrometry data despectroscopy method and / or despectroscopy model training method as described above.

[0119] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0120] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of mass spectral data deconvolution, characterized by, The method comprises: extracting mass spectrum parameters from original mass spectrum data, the mass spectrum parameters comprising mass-to-charge ratios of parent ions, mass-to-charge ratios of fragment ions and relative intensities, spectrum retention time, ion mode; constructing a spectrum resolving instruction based on a preset instruction template and the mass spectrum parameters; inputting the spectrum resolving instruction into a pre-trained spectrum resolving model to obtain a spectrum resolving result output by the spectrum resolving model, the spectrum resolving instruction being used to guide the spectrum resolving model to infer the spectrum resolving result based on the mass spectrum parameters, the spectrum resolving result comprising a compound molecular formula corresponding to the original mass spectrum data.

2. The method of claim 1, wherein, The spectrum resolving model is used to: determine at least one ion peak association according to the mass-to-charge ratios of the parent ions, the mass-to-charge ratios of the fragment ions and the relative intensities; determine a reference chemical bond cleavage site and a reference fragmentation path according to the ion mode and the ion peak association; determine at least one candidate molecular formula based on the spectrum retention time, the ion mode and the mass of the parent ions; determine a compound molecular formula corresponding to the original mass spectrum data from the at least one candidate molecular formula based on the reference fragmentation path and the spectrum retention time.

3. The method of claim 2, wherein, The determination of the at least one ion peak association according to the mass-to-charge ratios of the parent ions, the mass-to-charge ratios of the fragment ions and the relative intensities comprises: calculating a mass difference between the fragment ions and the parent ions according to the mass-to-charge ratios of the parent ions and the mass-to-charge ratios of the fragment ions; determining at least one ion peak association based on the mass difference and the relative intensities, the ion peak association representing an association relationship between the fragment ions with the relative intensities greater than a preset intensity threshold and the parent ions.

4. The method of claim 2, wherein, The determination of the at least one candidate molecular formula based on the spectrum retention time, the ion mode and the mass of the parent ions comprises: determining at least one target compound category from a plurality of candidate compound categories based on the spectrum retention time; calculating at least one candidate molecular formula according to the ion mode and the mass of the parent ions, the compound corresponding to the candidate molecular formula belonging to the at least one target compound category.

5. The method of claim 2, wherein, The spectrum resolving model is further used to: generate a target fragmentation path and fragmentation mechanism analysis corresponding to each compound molecular formula.

6. A method for training a deconvolutional model, comprising: The method comprises: performing self-supervised training on an initial model based on a mass spectrum resolving database to obtain a pre-training model; constructing a sample data set based on mass spectrum signal features of an original file and annotation results, the annotation results comprising a target compound corresponding to the mass spectrum signal features and fragmentation rules of the target compound; training the pre-training model based on the sample data set and a preset instruction set and a preset structured association knowledge to obtain an intermediate model; updating parameters of the intermediate model based on an output result score of the intermediate model until a training end condition is met, to obtain a trained spectrum resolving model, the output result score representing a similarity degree between an output result of the intermediate model and the annotation results.

7. A mass spectral data deconvolution apparatus, characterized by, The method comprises: An extraction module is configured to extract mass spectrum parameters from original mass spectrum data, the mass spectrum parameters including mass-to-charge ratios of parent ions, mass-to-charge ratios and relative intensities of fragment ions, spectrum retention time, and ion mode; A first construction module is configured to construct a deconvolution instruction based on a preset instruction template and the mass spectrum parameters; A deconvolution module is configured to input the deconvolution instruction into a pre-trained deconvolution model to obtain a deconvolution result output by the deconvolution model, the deconvolution instruction being used to guide the deconvolution model to infer the deconvolution result based on the mass spectrum parameters, the deconvolution result including a compound molecular formula corresponding to the original mass spectrum data.

8. A deconvolution model training apparatus characterized by comprising: The method comprises: A first training module is configured to perform self-supervised training on an initial model based on a mass spectrum deconvolution database to obtain a pre-training model; A second construction module is configured to construct a sample data set based on mass spectrum signal features of an original file and a labeled result, the labeled result including a target compound corresponding to the mass spectrum signal features and a fragmentation rule of the target compound; A second training module is configured to train the pre-training model based on the sample data set, a preset instruction set, and a preset structured correlation knowledge to obtain an intermediate model; An updating module is configured to update parameters of the intermediate model based on an output result score of the intermediate model until a training end condition is met to obtain a trained deconvolution model, the output result score representing a similarity degree between an output result of the intermediate model and the labeled result.

9. An electronic device, comprising: The method comprises: a memory; a processor; and a computer program; The computer program is stored in the memory and is configured to be executed by the processor to implement the method of any one of claims 1-6. The computer program is executed by the processor to implement the method of any one of claims 1-6.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, ​