Secondary mass spectrum auxiliary analysis method based on artificial intelligence model SPHERE-MS

By constructing a virtual reference mass spectrometry library using the SPHERE-MS model, the problems of dependence on experimental reference spectra and low resolution efficiency in existing mass spectrometry analysis methods are solved. This enables high-throughput, automated, and high-accuracy analysis of unknown compounds, and is applicable to fields such as drug quality supervision, drug development, and life sciences.

CN122050601APending Publication Date: 2026-05-15JIANGSU INST OF FOOD & DRUG SUPERVISION & INSPECTION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU INST OF FOOD & DRUG SUPERVISION & INSPECTION
Filing Date
2026-01-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing mass spectrometry analysis methods rely on human experience, and the analysis process is complex and inefficient, making it difficult to meet the needs of rapid screening of large-scale samples. In particular, when dealing with complex samples, accuracy and repeatability are difficult to guarantee. Existing models suffer from high computational complexity and insufficient coverage in high-resolution mass spectrometry prediction.

Method used

The SPHERE-MS model based on graph neural networks is adopted. Through standardized data acquisition and cleaning, potential ion fragments are generated by a hybrid algorithm of single bond breakage-secondary fragment library-secondary loss library. The graph neural network is used to predict fragment abundance and construct a virtual reference mass spectrometry library to achieve automated high-precision analysis.

Benefits of technology

It achieves high-throughput, automated, and high-accuracy analysis of unknown compounds, reduces reliance on human experience, and significantly improves analysis coverage and computational efficiency, making it suitable for fields such as drug quality supervision, drug development, and life sciences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050601A_ABST
    Figure CN122050601A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of mass spectrum analysis, and provides a secondary mass spectrum auxiliary analysis method based on an artificial intelligence model SPHERE-MS. According to the method, the auxiliary analysis of a high-resolution secondary mass spectrum is realized through standardized sample pretreatment, unified collection of secondary mass spectrum data, standardized data cleaning and enhancement and combination with a virtual reference mass spectrum library established by an SPHERE-MS map neural network model; therefore, unknown impurities and illegal additives in drugs and unknown compounds in complex matrix samples can be quickly determined. The most probable molecular structure can be returned only by providing the experimental spectrogram, and the method has the advantages of convenience in use, high flux, high accuracy and high system adaptability, and provides a promotable auxiliary secondary mass spectrometry analysis method for drug supervision, drug research and development and life science.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mass spectrometry analysis technology, specifically relating to a secondary mass spectrometry-assisted analysis method based on the artificial intelligence model SPHERE-MS. Background Technology

[0002] Mass spectrometry, as a crucial technique for elucidating the molecular structure of substances and characterizing the chemical composition of complex systems, has been widely applied in life sciences, natural product chemistry, drug development, and drug regulation. Through precise measurement of molecular ions and their fragment ions, mass spectrometry provides abundant structural information, serving as a vital basis for identifying unknown compounds in complex samples. However, in practical applications, the mass spectrometry analysis process still heavily relies on human experience, resulting in complex procedures and low efficiency. This is particularly true when dealing with sample systems containing complex compositions and numerous interfering substances, where accuracy and reproducibility are difficult to guarantee. While the rapid development of high-resolution mass spectrometry and high-throughput experimental technologies has significantly improved the ability to acquire experimental data, the efficiency of mass spectrometry data analysis has not kept pace, highlighting the growing contradiction between massive amounts of mass spectrometry data and limited analytical capabilities. This problem is especially evident in the field of drug quality and safety regulation. Unknown impurities, degradation products, or illegal additives that may exist in drugs may possess potential genotoxicity, carcinogenicity, or other safety risks; for example, nitrosamines have been proven to be highly carcinogenic. Furthermore, criminals modify known drug skeletons to create novel illegal additives (such as vardenafil derivatives), whose mass spectrometric characteristics are often difficult to identify using conventional comparison methods, posing a significant challenge to regulatory detection. Therefore, establishing a technical means for rapid and reliable analysis of unknown structures is of great practical significance. Currently, the secondary mass spectrometry analysis of unknown compounds mainly relies on two methods: manual spectral interpretation and library matching prediction. Manual spectral interpretation requires highly skilled personnel, is time-consuming, and highly subjective, making it difficult to meet the needs of rapid screening of large-scale samples. Library matching methods rely on the coverage of existing standard mass spectrometry libraries, and their analytical capabilities are limited by the number and diversity of molecular structures already included in experimental libraries. Since the theoretically possible chemical structure space is much larger than the scale of existing experimental libraries, relying solely on measured libraries is insufficient for practical applications, necessitating computational methods to predict virtual mass spectra to expand the reference library. Regarding the prediction problem from molecular structure to mass spectra, various computational models have been proposed. One type of method discretizes the continuous m / z range into fixed intervals and uses neural networks to predict the peak intensity distribution of each interval. This type of method achieves good overall performance when the fixed interval is set wide, but its resolution is limited, making it difficult to accurately reflect the fine fragment peak structure in high-resolution mass spectrometry; when the interval is further refined, the model performance drops significantly. Another type of method predicts the structure or chemical formula of ion fragments that may be generated during mass spectrometry fragmentation and further estimates their relative abundance. Since the m / z values ​​of fragments can be calculated accurately, this strategy has a strong advantage in high-resolution mass spectrometry prediction and has a certain degree of physical and chemical interpretability. However, existing fragment generation-based prediction methods still have certain limitations.On the one hand, simulating molecular fragmentation solely through the breaking of a single chemical bond is insufficient to cover ions generated by multi-stage fragmentation and rearrangement reactions in actual mass spectrometry experiments. On the other hand, directly considering the combined breaking of multiple chemical bonds would lead to an exponential increase in the number of potential fragments, significantly increasing computational complexity and making it unsuitable for large-scale molecular databases. Furthermore, rearrangement phenomena such as hydrogen migration, which are prevalent in mass spectrometry, are difficult to describe effectively using simple bond-breaking models, resulting in the omission of potential fragments.

[0003] Therefore, how to construct a fragment ion space with high coverage and strong interpretability while ensuring computational efficiency, and further realize automated and high-precision analysis of secondary mass spectrometry to reduce reliance on human experience, has become a key technical problem that urgently needs to be solved in the field of mass spectrometry analysis. Solving this problem has significant application value for improving the analysis efficiency and accuracy of unknown impurities, illegal additives, and unknown compounds in complex samples. Summary of the Invention

[0004] The purpose of this invention is to provide a secondary mass spectrometry-assisted analysis method based on the artificial intelligence model SPHERE-MS. This method standardizes the sample pretreatment process, uniformly collects secondary mass spectrometry data, and combines standardized data cleaning and data enhancement strategies. It utilizes a SPHERE-MS model built based on a graph neural network to generate a virtual reference mass spectrum library, thereby achieving assisted analysis of high-resolution secondary mass spectrometry spectra. This solves the problems of existing mass spectrometry analysis methods, such as dependence on experimental reference spectra for unknown compounds, low analysis efficiency, and limited applicability.

[0005] Another objective of this invention is to generate potential ion fragments using a hybrid algorithm of single bond breakage, secondary fragmentation library, and secondary loss library introduced in the SPHERE-MS model, and to predict the relative abundance of each ion fragment. This enables the construction of a large-scale, high-coverage virtual secondary mass spectrometry reference library while balancing computational efficiency and the interpretability of fragmentation mechanisms, thus providing reliable theoretical support for spectral comparison of unknown compounds.

[0006] Another objective of this invention is to provide, with only the secondary mass spectra of the experiment to be analyzed, the experimental spectra as input, and to compare and rank them with the established virtual reference mass spectrometry library, outputting the candidate molecular structures with the highest similarity ranking as the analysis results, thereby achieving rapid qualitative analysis of unknown impurities, illegal additives, and unknown compounds in complex matrix samples in pharmaceuticals.

[0007] Through the above technical solution, the present invention can achieve high-throughput, automated and high-accuracy analysis of unknown compounds without the need for additional experimental control spectra. It has the advantages of simple operation, strong adaptability, fast analysis speed and strong scalability, and can be widely used in fields such as drug quality supervision, drug research and development and life sciences.

[0008] To achieve the above-mentioned objectives, this invention proposes a second-order mass spectrometry-assisted analysis method based on the artificial intelligence model SPHERE-MS, characterized in that: (1) Construction of the training set for the SPHERE-MS model A dataset for training the SPHERE-MS model is constructed based on known compounds and their corresponding experimental secondary mass spectrometry data, which are obtained from public mass spectrometry databases. Potential ion fragments are generated by a hybrid algorithm of single bond breaking, secondary fragmentation library, and secondary loss library to form the ion fragment space of the parent ion. (2) Training of SPHERE-MS model and fragment abundance prediction The two-dimensional structure of the parent ion is represented as a molecular graph, and the experimental conditions of mass spectrometry are used as experimental parameter variables and input into the SPHERE-MS model. Through feature embedding, graph structure message passing and multi-scale feature fusion, the mapping relationship between molecular structure, experimental conditions and fragmentation behavior is learned. The intensity of the second-order mass spectrometry peak corresponding to the ion fragment space obtained in step (1) is used as the training set to train the graph neural network model so that the model can predict the relative abundance of each potential ion fragment in the ion fragment space. (3) Construction of virtual reference mass spectrometry library Using the trained SPHERE-MS model, a large number of candidate molecular structures are predicted in batches to generate corresponding virtual secondary mass spectra. Molecular structure information is then associated with and stored with the virtual mass spectra to construct a searchable virtual reference mass spectrum library. (4) Preparation of unknown samples and acquisition of secondary mass spectrometry data The unknown sample to be analyzed is subjected to standardized preprocessing, and its secondary mass spectrometry data is obtained by high-resolution mass spectrometry. The secondary mass spectrometry data is used as the input of the spectrum to be analyzed after a unified preprocessing process. (5) Reference mass spectrometry library matching and auxiliary analysis The secondary mass spectrum of the unknown sample obtained in step (4) is compared with the virtual reference mass spectrum library constructed in step (3) for similarity. The candidate molecular structures are sorted according to the similarity score, and the molecular structures with the highest similarity ranking are output as auxiliary analysis results of the unknown sample.

[0009] The method is characterized in that, The construction of the training set for the SPHERE-MS model described in step (1) is implemented as follows: A. Select the NIST-20 database as the training set; B. Construction of Ion Fragment Space: Potential ion fragments are generated through a hybrid algorithm of single bond breaking, secondary fragment library, and secondary loss library, and a secondary mass spectrometry prediction model is used to predict the abundance of each ion fragment. Specifically, firstly, a set of potential primary ion fragments is obtained by single-break of chemical bonds or destruction of ring structures in the molecular structure, and after adjusting the number of restrictive hydrogen atoms; then, high-frequency secondary fragments are subtracted from the ion fragments in the primary ion fragment set to generate all possible ion fragments.

[0010] The method is characterized in that the construction of the training set for the SPHERE-MS model in step (1) is implemented according to the following steps: A. NIST-20 undergoes the following filtering steps: only [M+H]+ and [MH]- type precursor ions are considered; molecules containing atoms other than {C, H, N, O, P, S, F, Cl, Br, I} are excluded; the filtered mass spectrometry data are used to statistically construct a secondary loss library; B. Primary fragments are generated through single bond and ring cleavage: ① The parent ion is represented as a molecular graph structure, where atoms are nodes and chemical bonds are edges. Edge removal is performed on the molecular graph to simulate the fragmentation process in mass spectrometry. ② A single chemical bond breaking algorithm is used to generate primary fragments, i.e., by removing single chemical bonds in non-cyclic structures, the parent ion is fragmented into charged fragment ions and neutral loss molecules. ③ For cyclic structures present in the molecule, a preset number of chemical bonds in each ring structure are set as breakable bonds. Breaking the chemical bonds within the ring generates primary fragment ions produced by ring fragmentation. ④ During the single chemical bond breaking and ring structure breaking processes, a limiting hydrogen atom number adjustment mechanism is introduced to adjust the generated... Fragment ions are defined by three forms of hydrogen atom change: the number of hydrogen atoms remains unchanged, one hydrogen atom is added, or one hydrogen atom is removed, to characterize the hydrogen migration phenomenon that may occur during actual fragmentation; ⑤ All fragment ions generated through the breaking of the single chemical bond, the breaking of the ring structure, and the rearrangement of hydrogen atoms, together with the parent ion itself, constitute the primary fragment set; ⑥ Based on a fixed-scale secondary neutral loss fragment library pre-constructed from an experimental mass spectrometry database, the secondary neutral loss is applied sequentially to the primary fragment set, and by subtracting the corresponding neutral loss molecular formula from the primary fragments, all potential secondary fragment ion sets are generated, constituting the ion fragment space of the parent ion.

[0011] The method is characterized by comprising the following steps: Step (2) Training and fragment abundance prediction of the SPHERE-MS model, implemented as follows: ① The two-dimensional structure of the parent ion is represented as an undirected molecular graph, where nodes represent heavy atoms and edges represent interatomic chemical bonds. Node features include at least the atomic element type, hybridization state, formal charge, and number of bonded hydrogen atoms. Edge features include bond type and whether it belongs to a ring structure. Mass spectrometry experimental parameters are supernode variables of the undirected molecular graph, including instrument type, parent ion type, and normalized collision energy. Based on this, the first few feature vectors and eigenvalues ​​of the Laplacian matrix of the molecular graph are further extracted to characterize the overall topological structure of the molecule. ② Fusion of graph neural network feature extraction and multi-scale characterization: The node features, edge features, and mass spectrometry experimental parameter variables are mapped to a unified-dimensional latent feature space through an embedded network. In the process, a graph neural network is used to perform multi-layer message passing and feature updating on the molecular graph, so as to simultaneously integrate the node's own attributes, neighborhood structure information, and chemical bond features. After the node features are extracted, the features of nodes belonging to the same first-level fragment are aggregated to obtain fragment-level feature representations. Then, attention-weighted pooling is applied to all nodes in the molecular graph to obtain molecular graph-level feature representations. The molecular graph-level structural features are fused with the mass spectrometry experimental parameter variable features to form global supernode features of the molecular graph. ③ Prediction of the generation probability of first-level fragments: Based on the similarity relationship between the fragment-level features and the supernode features, the probability distribution of each first-level fragment as a real fragmentation product is predicted. The calculation formula is as follows: ⑤ Conditional probability prediction of hydrogen atom adjustment and second-order neutral loss: For each type of primary fragment, different output layers are used to predict the conditional probability distributions of different hydrogen atom adjustment states and the conditional probability distributions of different second-order neutral losses; the conditional probabilities are calculated by the output layer composed of a multilayer sensing mechanism, and their expressions are as follows: ⑥ Joint probability modeling and output of fragment ion abundance: The first-order fragment selection probability, the hydrogen atom adjustment conditional probability, and the second-order neutral loss conditional probability are treated as independent events, and their joint probability is used as the predicted abundance of the corresponding ion fragment. The expression is as follows: ⑦ Training and Inference of SPHERE-MS: During the training phase, based on the mapping relationship between fragments and true mass spectrum peaks, the predicted abundances of multiple fragments belonging to the same element are merged to obtain the predicted abundance. The cross-entropy loss function is used to calculate the difference between the predicted abundance and the true abundance, which is used as the optimization objective to optimize the model parameters. The loss function is calculated as follows: ⑧ During the inference stage, the predicted abundance of ion fragments with the same chemical formula are merged, and the intensity of the final output mass spectrum peaks is normalized to obtain the predicted secondary mass spectrum.

[0012] The method is characterized by comprising the following steps: Step (2) Training of the SPHERE-MS model is implemented according to the following steps: A. Mapping of possible ion fragments with NIST-20 mass spectrum peaks: For each parent ion structure and its corresponding mass spectrum, possible fragment ions are obtained using a hybrid algorithm combining single-batch bond breaking, secondary fragment library, and secondary loss library. ,in Indicates possible fragment ions. and They represent Chemical formula and precise mass; peak signals in mass spectrometry , where m p The charge-to-mass ratio of the peak, Indicates the labeled chemical formula, Indicates peak intensity; Ion fragments are determined based on the following two criteria. and experimental mass spectrometry peak signals The correspondence between them: (1) Or (2) or The predicted set of potential ion fragments is matched with experimental mass spectrometry peaks; successfully matched peaks are used as the true values, while unmatched peaks are ignored. Then, the intensities of the successfully matched peaks are normalized so that their sum equals 1. Finally, a virtual peak with zero intensity is created. To match ion fragments that do not correspond to any peaks: .

[0013] B. Encoding of molecular structure and tandem mass spectrometry experimental parameter variables: The two-dimensional molecular structure of the precursor ion is described using a graph data structure. Node features are characterized using one-hot encoding to represent atom type, hybridization state, charge information, and the number of bonding hydrogen atoms. Edge features are characterized using one-hot encoding to represent chemical bond type and whether it belongs to a ring structure. Simultaneously, information such as instrument type, precursor ion type, and collision energy during mass spectrometry acquisition are also input into the model as experimental parameters using one-hot encoding to enhance the model's adaptability to changes in experimental conditions. The two-dimensional molecular structure is represented as an undirected graph, where each node corresponds to a heavy atom, and each edge represents a chemical bond between two atoms. The first eight eigenvectors and eigenvalues ​​of the graph Laplacian operator are used as features to capture the structural characteristics of the molecule. Mass spectrometry experimental parameters are treated as graph-level features. C. Feature Extraction Module: The initial encoding of molecular graph nodes, edges, and mass spectrometry experimental parameter variables is converted through independent embedding layers. Each embedding layer consists of an MLP block with GraphNorm: Where h represents the dimension of the latent features (h = 512), and the dropout rate is set to 0.1. Unless otherwise specified, these settings are the default settings for other modules in SPHERE-MS; The first 8 eigenvectors of the Graph Laplacian operator and eigenvalues Convert using SignNet The embedding layer in SignNet consists of an MLP block without a GraphNorm layer: The obtained embedded features are then used for feature extraction via a graph attention network module: The extracted node features are represented as follows The following abbreviation is x. First-level fragmented features are obtained through summation and aggregation. This preserves information about fragment size. Then, the AttentionalAggregation operator is used to pool the graph nodes to obtain graph-level supernode features. Subsequently, through layer combination with residual connections... and To obtain graph-level features : Each Level 1 Fragment The probability is calculated based on the similarity between the extracted graph-level features and the features of each first-level fragment:

[0014] D. Prediction module for ion fragment probability: Output layer with residual connections to predict hydrogen atom number adjustment And the loss of different secondary fragments Conditional probability distribution: Ion fragments The probability (i.e., peak intensity) can be viewed as the joint probability of three independent events: primary fragmentation. Adjustment of the number of hydrogen atoms and lost secondary fragments : According to ion fragments and experimental mass spectrometry peaks The matching relationships between them are used to predict peak intensity. :

[0015] The cross-entropy loss function is used as the loss function for model training:

[0016] During the reasoning phase, due to ion fragments The matching relationship between the peaks and the experimental mass spectrometry peaks is unknown, and those with the same chemical formula will be... Ion fragments The peak intensity is obtained by summing the intensities. Then, the obtained peak intensities are normalized so that their sum equals 1.0.

[0017] Mass Spectrometry Cosine Similarity: Mass spectrometry cosine similarity (using the CosineHungarian algorithm provided in the matchms library) is used to evaluate the accuracy of the predicted results against the actual mass spectra. Tolerance Window size set to 0.1 Da:

[0018] E. Model Implementation and Training: SPHERE-MS is primarily developed using the PyTorch deep learning framework. PyTorchGeometric is used to process graph data and graph neural network models, while PyTorchLightning is used to simplify the model training process. F. Model Evaluation: Benchmark models: GRAFF-MS, FIORA, and CFM-ID were used as benchmark models for comparison. Library matching: For each mass spectrum in the NIST-20 test set, 49 “decoy” isomers with the highest Tanimoto similarity to the real molecular structure were collected from the PubChem database; this formed a library of 50 molecules for each test mass spectrum; the similarity between the predicted mass spectra of the molecular structures in the library and the real mass spectra was sorted to calculate the frequency of the top k real molecules at different k values.

[0019] The method is characterized in that, Step (3) Preparation of unknown samples and acquisition of secondary mass spectrometry data: Standardized sample pretreatment for unknown samples to be analyzed: Test solution: Take about 5 mg of the sample to be tested, add 250 ml of 0.1% formic acid-water:acetonitrile 40:60 to dissolve, centrifuge at 13000g for 10 min, filter, discard the initial filtrate, and take the subsequent filtrate as the test solution; Acquisition of secondary mass spectrometry data: The obtained secondary mass spectrometry data were preprocessed uniformly, including isotope peak merging, noise peak filtering, and peak intensity normalization, to obtain the experimental secondary mass spectra to be analyzed.

[0020] The method is characterized in that, Reference mass spectrometry library matching and auxiliary analysis Step (4) Reference Mass Spectrometry Library Matching and Auxiliary Analysis: The virtual reference spectrum library obtained in step (3) is matched with the measured spectrum for similarity. The similarity is calculated using cosine similarity and then sorted from high to low similarity. The candidate molecular structures and confidence levels are output, thereby achieving auxiliary analysis of the secondary mass spectrometry and helping to quickly identify unknown compounds. The formula for calculating cosine similarity is: Error window Set to 0.01 Da, where m and I are the atomic weight and normalized abundance of the peak, respectively. Beneficial effects

[0021] Addressing the limitations of traditional bond-breaking strategies, and inspired by the vocabulary-based model GRAFF-MS, we propose SPHERE-MS (Secondary-loss Powered Hybrid Edge Removal Engine for predicting Mass Spectra). As the name suggests, SPHERE-MS is a graph neural network model that employs a hybrid approach combining bond-breaking and vocabulary-based strategies to generate potential fragments and predict peak intensities.

[0022] For each input precursor ion, a first-level fragment set is generated by breaking individual chemical bonds or ring structures. Then, a set of predefined small neutral losses, called "secondary losses," are subtracted from each first-level fragment. All potential fragments are eventually generated by subtracting every possible secondary fragment from every primary fragment. This hybrid strategy allows SPHERE-MS to simulate multi-level fragmentation processes without enumerating all combinations of multiple chemical bonds, thus significantly reducing computational complexity. Evaluations on test sets (including the NIST-20 tandem mass spectrometry library, MSnLib, and CASMI2022) show that SPHERE-MS predicts mass spectra with higher similarity to real mass spectra compared to the benchmark models FIORA (which relies solely on chemical bond breaking) and GRAFF-MS (which relies entirely on the vocabulary) (Murphy, M.; Jegelka, S.; Fraenkel, E.; Kind, T.; Healey, D.; Butler, T., Efficiently predicting high resolution mass spectra with graph neural networks. Int.Conf. Mach. Learn. 2023, 25549 - 25562.). Furthermore, analysis of the relationship between predicted mass spectra accuracy and similarity to molecular structure and the training set indicates that SPHERE-MS has stronger generalization ability than both FIORA and GRAFF-MS. These advantages in accuracy, generalization, and efficiency make SPHERE-MS a valuable tool for mass spectrometry prediction. By simulating the multi-stage fragmentation process in mass spectra through a secondary loss library rather than recursively exhaustively enumerating bond breaking, SPHERE-MS significantly improves efficiency, avoiding the exponential computational costs associated with traditional fragmentation-based methods. Notably, tests showed that SPHERE-MS predicted 1.5 billion mass spectra in approximately two hours on a desktop workstation, highlighting its predictive efficiency for large-scale mass library expansion.

[0023] The advantages are as follows: (1) Eliminate dependence on experimental reference spectral library and significantly improve resolution coverage: This invention constructs a virtual reference mass spectrometry library through SPHERE-MS deep learning model, which can realize the spectral analysis of unknown compounds. It is particularly suitable for application scenarios where experimental spectra are missing or reference spectra are difficult to obtain.

[0024] (2) Balancing the interpretability of fragmentation mechanism with computational efficiency: The potential ionic fragments are modeled by a hybrid algorithm of single bond breaking, secondary fragmentation library, and secondary loss library. While effectively covering the main fragmentation path, the exponential computational complexity caused by multi-level recursive fracture is avoided, thus achieving a balance between computational efficiency and physicochemical interpretability.

[0025] (3) Improved matching accuracy by using graph neural network for refined abundance prediction: This invention uses graph neural network to jointly model molecular structure and experimental conditions, which can more accurately predict the relative abundance of ion fragments, thereby improving the similarity matching accuracy between virtual spectrum and experimental spectrum.

[0026] (4) The analysis can be completed by relying solely on experimental spectra, with low usage threshold and high throughput: In practical applications, the present invention only requires inputting the secondary mass spectrum to be analyzed, and can automatically output the candidate molecular structure results, which is suitable for high-throughput analysis needs and significantly reduces the reliance on human experience and expert knowledge.

[0027] (5) Wide applicability and strong system adaptability: The method of this invention is applicable to secondary mass spectrometry data acquired under different instrument types, different ionization modes, and different collision energies, and has good cross-platform adaptability. This invention can be widely applied in fields such as drug quality supervision, drug development, screening for illegal additives, analysis of complex matrix samples, and life science research, and has significant engineering application value and promotion prospects. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating the process of the present invention. Figure 2 Peak coverage achieved on the NIST-20 dataset for different fragmentation generation strategies; Figure 3 Two representative prediction examples are shown; Figure 4 (a) Distribution of ECFP4 Tanimoto structural similarity between molecules in the NIST-20, MSnLib, and CASMI2022 test sets and training sets. (b) Cosine similarity between ground-truth spectra and SPHERE-MS predicted spectra for subsets of molecules with low structural similarity. N represents the number of test spectra in each subset.

[0029] Figure 5 The dependence of spectral cosine similarity on the structural similarity between the molecule and the training set.

[0030] Figure 6. Top-k retrieval accuracy of SPHERE-MS, GRAFF-MS, and FIORA-OS in the NIST-20 library matching task. Detailed Implementation

[0031] The present invention will now be described in further detail with reference to the embodiments. The embodiments are only for illustrating the present invention and are not intended to limit the present invention.

[0032] Example 1: Development and Training of SPHERE-MS I. Acquisition and Construction of Training Data This invention uses three datasets: (1) NIST-20, used as the training set, validation set and internal test set for SPHERE-MS; (2) MSnLib; and (3) CASMI 2022, both of which are used as external test sets.

[0033] A. Select the NIST-20 database as the training set. Publicly available secondary mass spectrometry databases were selected as training data sources, with the NIST-20 database (https: / / www.nist.gov / programs-projects / nist23-updates-nist-tandem-and-electron-ionization-spectral-libraries) being the preferred choice. This database contains a large amount of molecular structure information for known compounds and their corresponding experimental secondary mass spectrometry data. For each mass spectrometry record in the database, the following information was extracted for model training: molecular structure information of the precursor ion, precursor ion form, mass-to-charge ratio and peak intensity data from the experimental secondary mass spectrometry, and the instrument type and collision energy parameters used during mass spectrometry acquisition. This information, after being processed, constituted the original training dataset for the SPHERE-MS model.

[0034] B. Select MSnLib and CASMI 2022 as the test set. MSnLib and CASMI 2022: MSnLib is a large-scale MS nThe library was developed using a high-throughput data acquisition and processing workflow. The MSnLib test set was selected based on the following criteria: (1) MSLEVEL was set to 2; (2) SPECTYPE was limited to “SINGLE_BEST_SCAN” and “SAME_ENERGY”; (3) precursor ion type was limited to [M+H]+ and [MH]-; (4) spectra different from any molecule in NIST-20 were selected based on the “SMILES” field. CASMI 2022 provides raw liquid chromatography-precise secondary mass spectrometry data in positive and negative electrospray ionization modes, including 500 compounds and their retention times and peaks. The raw CASMI 2022 .mzml file was read using the OpenMS library. Multiple measurements with a precursor ion m / z tolerance of 0.01 Da and a retention time window of 0.2 min at the same collision energy were combined into a single mass spectrum. If the parent ion type is not [M+H]+ or [MH]-, or if the corresponding molecular structure appears in the NIST-20 dataset, the spectrum is discarded. After these filtering steps, the MSnLib test set consists of 21,723 mass spectra corresponding to 4,470 molecular structures, while the CASMI 2022 test set contains 718 mass spectra corresponding to 245 molecular structures.

[0035] II. Construction of Potential Ion Fragment Space For each parent ion molecular structure in the training dataset, a hybrid algorithm in SPHERE-MS, consisting of single bond breaking, secondary fragmentation library, and secondary loss library, is used to construct a potential ion fragment space. Specifically, this involves: firstly, obtaining a set of potential primary ion fragments by single-break of chemical bonds or destruction of ring structures in the molecular structure, and adjusting the number of restrictive hydrogen atoms; and then subtracting frequently occurring secondary fragments from the ion fragments in the primary ion fragment set to generate all possible ion fragments.

[0036] The specific processing steps are as follows: A. Select the NIST-20 database as the training set. The NIST-20 database is filtered as follows: (1) Only the parent ions of the [M+H]+ and [MH]- types are considered; (2) Molecules containing atoms other than {C, H, N, O, P, S, F, Cl, Br, I} are excluded; The filtered mass spectrometry data are used to statistically establish a secondary loss library. B. Primary fragments are generated through single bond and ring cleavage: ① The parent ion is represented as a molecular graph structure, where atoms are nodes and chemical bonds are edges. Edge removal is performed on the molecular graph to simulate the fragmentation process in mass spectrometry. ② A single chemical bond breaking algorithm is used to generate primary fragments, i.e., by removing single chemical bonds in non-cyclic structures, the parent ion is fragmented into charged fragment ions and neutral loss molecules. ③ For cyclic structures present in the molecule, a preset number of chemical bonds in each ring structure are set as breakable bonds. Primary fragment ions generated by ring fragmentation are generated by breaking the chemical bonds within the rings. ④ A limiting hydrogen atom number adjustment mechanism is introduced during the single chemical bond breaking and ring structure breaking processes. The generated fragment ions are limited to three forms of hydrogen atom changes: the number of hydrogen atoms remains unchanged, one hydrogen atom is added, or one hydrogen atom is removed, to characterize the hydrogen migration phenomenon that may occur during the actual fragmentation process; ⑤ All fragment ions generated through the breaking of the single chemical bond, the breaking of the ring structure, and the rearrangement of hydrogen atoms are combined with the parent ion to form a primary fragment set; ⑥ Based on a fixed-scale secondary neutral loss fragment library pre-constructed from an experimental mass spectrometry database, the secondary neutral loss is applied sequentially to the primary fragment set, and all potential secondary fragment ion sets are generated by subtracting the corresponding neutral loss molecular formula from the primary fragments.

[0037] In SPHERE-MS, mass spectrometry prediction is defined as the task of predicting the probability distribution of possible fragment ions from the parent ion. The first key step in SPHERE-MS is to generate possible fragment ions based on the parent ion structure, simulating the fragmentation process of a molecule breaking down into charged products and one or more neutral loss molecules in mass spectrometry. The feasibility of algorithms for obtaining ion fragments through single bond breaking has been demonstrated by the FIORA model [Nowatzky, Y.; Russo, FF; Lisec, J.; Kier, A.; Reinert, K.; Muth, T.; Benner, P., FIORA: Local neighborhood-based prediction of compound mass spectra from single fragmentation events. Nat. Commun 2025, 16, 2298.]. However, this algorithm neglects ion fragments generated by ring structure breaking. To address this issue, we simultaneously break two chemical bonds belonging to the same ring structure to generate fragment ions. On the other hand, hydrogen atom rearrangement may occur during fragmentation. The single bond breaking process in acyclic systems proceeds through two typical pathways, as shown in Equation 1a. Compared to molecular fragments, ionic fragments may gain or lose a hydrogen atom. In the process of ring structure breaking to generate ionic fragments, the ionic fragment may gain or lose a hydrogen atom, or seemingly no hydrogen transfer occurs (Equation 1b). Furthermore, the precursor ion may undergo multiple fragmentation pathways. For example, if the parent ion undergoes "+1H" fragmentation followed by "-1H" fragmentation, the net change in the number of hydrogen atoms will be zero. Considering all possible cases would make the algorithm overly complex and difficult to deploy. Therefore, we adjust the number of hydrogen atoms in all generated fragments to three types: "no change in the number of hydrogen atoms," "+1H," and "-1H." The first-level fragment set is obtained by breaking the chemical bonds and rings of the parent ion to generate ionic fragments, while simultaneously adjusting the number of hydrogen atoms. In addition, the parent ion itself is also considered an element in the first-level fragment set.

[0038] C. A fixed list of neutral molecular formulas, called the “secondary loss library,” is defined based on NIST-20. For each mass spectrometer in NIST-20, the following steps are performed: (1) Generate primary fragments as described above. (2) Match primary fragments and retrogradely according to the chemical formulas labeled on the mass spectrometer peaks in NIST-20. Peaks without chemical formulas are matched with the mass of the primary fragments based on the peak mass-to-charge ratio, with a threshold of 20 ppm or 0.01 Da. (3) Peaks that are successfully matched are taken as “true primary fragments,” while peaks that fail to match but have chemical formulas are assumed to be generated by primary fragments by subtracting “secondary losses,” the weight of which is set to the normalized peak intensity. Peaks that fail to match and have no label are ignored. (4) Merge the secondary loss weights in the repeated mass spectrometers and normalize their weights by the number of repetitions. (5) Filter secondary losses according to the “N rule”: if the number of nitrogen or phosphorus atoms is odd, then the number of hydrogen and halogen atoms should also be odd. This rule filters out non-neutral stimulus losses. After the above statistical process, a list of 144 secondary losses was defined, including: the top 100 weighted secondary losses containing only C, H, O, and N; the top 11 weighted secondary losses containing F or Cl; the top 8 weighted secondary losses containing Br or F; the top 3 weighted secondary losses containing P; the top 2 weighted secondary losses containing I; and an empty entry indicating no stimulus loss. For each input molecular structure, after generating primary fragments using single-bond and ring-fragmentation algorithms, each secondary loss is subtracted from the primary fragments to obtain a set of possible ionic fragments for that molecule.

[0039] III. Construction and Training of the SPHERE-MS Graph Neural Network Model SPHERE-MS is an artificial intelligence model that uses a hybrid fragmentation generation strategy to produce ion fragments and extracts latent features from molecular maps and mass spectrometry experimental parameters based on a graph attention network framework. It then predicts the probability of each possible ion fragment appearing, thus completing the positive secondary mass spectrometry prediction. The specific steps are as follows: A. Mapping of possible ion fragments with NIST-20 mass spectrum peaks: For each parent ion structure and its corresponding mass spectrum, possible fragment ions are obtained using a hybrid algorithm combining single-batch bond breaking, secondary fragment library, and secondary loss library. ,in and They represent Chemical formula and precise mass; peak signals in mass spectrometry ,in The charge-to-mass ratio of the peak, Indicates the labeled chemical formula, Indicates peak intensity.

[0040] Ion fragments are determined based on the following two criteria. and experimental mass spectrometry peak signals The correspondence between them: (1) Or (2) or The predicted set of potential ion fragments is matched with experimental mass spectrometry peaks. Successfully matched peaks are used as the true values, while unmatched peaks are ignored. Then, the intensities of the successfully matched peaks are normalized so that their sum equals 1. Finally, a virtual peak with zero intensity is created. To match ion fragments that do not correspond to any peaks: .

[0041] B. Encoding of molecular structure and tandem mass spectrometry experimental parameter variables: The two-dimensional molecular structure of the precursor ion is described using a graph data structure. Node features are characterized using one-hot encoding to represent atom type, hybridization state, charge information, and the number of bonding hydrogen atoms. Edge features are characterized using one-hot encoding to represent chemical bond type and whether it belongs to a ring structure. Simultaneously, information such as instrument type, precursor ion type, and collision energy during mass spectrometry acquisition are also input into the model as experimental parameters using one-hot encoding to enhance the model's adaptability to changes in experimental conditions.

[0042] Specifically: Two-dimensional molecular structures are represented as undirected graphs, where each node corresponds to a heavy atom and each edge represents a chemical bond between two atoms. The first eight eigenvectors and eigenvalues ​​of the graph Laplacian operator are used as features to capture the structural characteristics of the molecule. Mass spectrometry experimental parameters are treated as graph-level features. The initial encodings of nodes, edges, and mass spectrometry experimental parameters are summarized in Table 1.

[0043] Table 1 Initial features used for atoms and bonds in the two-dimensional molecular diagram and mass spectrometry covariates

[0044] C. Feature Extraction Module: The initial encoding of molecular graph nodes, edges, and mass spectrometry experimental parameter variables is converted through independent embedding layers. Each embedding layer consists of an MLP block with GraphNorm: Where h represents the dimension of the latent features (h = 512), and the dropout rate is set to 0.1. Unless otherwise specified, these settings are the default settings for other modules in SPHERE-MS.

[0045] The first 8 eigenvectors of the Graph Laplacian operator and eigenvalues Convert using SignNet The embedding layer in SignNet consists of an MLP block without a GraphNorm layer: The obtained embedded features are then used for feature extraction via a graph attention network module: The extracted node features are represented as follows (Hereafter abbreviated as x), first-level fragmented features are obtained through summation and aggregation. This preserves information about fragment size. Then, the AttentionalAggregation operator is used to pool the graph nodes to obtain graph-level supernode features. Subsequently, through layer combination with residual connections... and To obtain graph-level features :

[0046] Each Level 1 Fragment The probability is calculated based on the similarity between the extracted graph-level features and the features of each fragment:

[0047] D. Prediction module for ion fragment probability: Output layer with residual connections to predict hydrogen atom number adjustment And the loss of different secondary fragments Conditional probability distribution:

[0048] Ion fragments The probability (i.e., peak intensity) can be viewed as the joint probability of three independent events: primary fragmentation. Adjustment of the number of hydrogen atoms and lost secondary fragments : According to ion fragments and experimental mass spectrometry peaks The matching relationships between them are used to predict peak intensity. : The cross-entropy loss function is used as the loss function for model training: During the reasoning phase, due to ion fragments The matching relationship between the peaks and the experimental mass spectrometry peaks is unknown, and those with the same chemical formula will be... Ion fragments The peak intensity is obtained by summing the intensities. Then, the obtained peak intensities are normalized so that their sum equals 1.0.

[0049] Mass Spectrometry Cosine Similarity: Mass spectrometry cosine similarity (using the CosineHungarian algorithm provided in the matchms library) is used to evaluate the accuracy of the predicted results against the actual mass spectra. Tolerance Window size set to 0.1 Da:

[0050] E. Model Implementation and Training: SPHERE-MS is primarily developed using the PyTorch deep learning framework. PyTorchGeometric is used to handle graph data and graph neural network models, while PyTorchLightning simplifies the model training process. During training, a batch size of 32 is used, and an early stopping strategy is employed (training terminates if the validation set metric does not improve after 10 batches). The metric is the median mass spectrometric cosine similarity on the validation set. Training is performed on a workstation equipped with an NVIDIA RTX 4070 Ti SUPERGPU and an Intel i7-13700KF CPU.

[0051] F. Model Evaluation: Benchmark Models: In this invention, GRAFF-MS, FIORA, and CFM-ID are used as benchmark models for comparison. GRAFF-MS is a vocabulary-based model, while FIORA is a recently reported model based on single-bond fragmentation. SPHERE-MS shares similarities with both GRAFF-MS and FIORA in terms of fragmentation generation algorithms. Following the instructions in the benchmark model's GitHub documentation, we retrained GRAFF-MS and FIORA (hereinafter referred to as FIORA-R) using the NIST-20 training and validation sets. We also used the open-source version of FIORA (hereinafter referred to as FIORA-OS) for comparison, with the training set of FIORA-OS being MSnLib. CFM-ID is an earlier machine learning model for mass spectrometry prediction, trained on the METLIN dataset, and is currently commonly used as a benchmark model for comparison. Since CFM-ID predicts mass spectra only at the qualitative collision energy level, the normalized collision energies of the mass spectra used to test CFM-ID are limited to three ranges: (0.18, 0.22), (0.33, 0.37), and (0.48, 0.52).

[0052] Library matching: We followed the method described in the reference [Goldman, S.; Li, J.; Coley, CW, Generating Molecular Fragmentation Graphs with Autoregressive Neural Networks. Anal. Chem. 2024, 96, 3419-3428.]: For each mass spectrum in the NIST-20 test set, 49 “decoy” isomers with the highest Tanimoto similarity to the real molecular structures were collected from the PubChem database. This formed a library of 50 molecules for each test mass spectrum. The similarity between the predicted mass spectra and the real mass spectra of the molecular structures in the library was ranked to calculate the frequency of the top k real molecules at different k values.

[0053] IV. Results: 1. Peak coverage between actual fragment ions and predicted fragment ions Since SPHERE-MS predicts tandem mass spectra by forecasting the probability distribution of possible fragment ions derived from the parent ion, the extent to which the chemical formula space of its virtually generated fragments adequately covers the chemical formula space of the experimentally observed peaks is fundamental to its performance. In addition to fragmentation through single bond breaking and hydrogen atom number adjustment, SPHERE-MS also expands the theoretical fragment chemical formula space by incorporating ring breaking and secondary loss libraries. Testing has shown that the chemical formula space of the ion fragments generated by SPHERE-MS covers nearly 100% of the actual fragment chemical formula space in NIST-20.

[0054] As shown in Figure 2, more complex fragmentation strategies can achieve higher peak coverage, but at the cost of increased computation time. The simplest method, S (considering only single-bond breaking), has an average peak coverage of 0.286, missing most peak signals in the experimental mass spectra. Although the weighted average peak coverage of S is 0.825, this is mainly due to the high peak intensity of the parent ion itself in tandem mass spectrometry. SH, based on S, allows for the adjustment of individual hydrogen atoms. The fragments produced by this strategy are far superior to those of S in terms of coverage, indicating that introducing a hydrogen atom adjustment mechanism can better simulate the hydrogen rearrangement phenomenon (Equations 1a and 1b) that occurs in the ion fragmentation pathway, and the resulting fragments are more consistent with experimental observations. Although SH_H2 and FIORA are implemented differently, both models actually use the same fragmentation strategy. Introducing the secondary loss "H2" allows for greater adjustment of the number of hydrogen atoms in the generated fragments, but SH_H2 only slightly improves the peak coverage compared to SH. Building upon SH_H2, SHR_H2 further introduces a ring-fragmentation algorithm, allowing ring structures to break into two fragments, significantly improving the average peak coverage from 0.615 to 0.695. This result demonstrates that ring fragmentation can effectively capture ring fragmentation processes occurring in mass spectrometry. In summary, SPHERE-MS employs a strategy combining bond breaking, hydrogen atom number adjustment, and ring fragmentation to generate a first-level fragment set.

[0055] However, the primary fragments generated by SHR_H2 still miss 30% of the peak signals observed in the experimental spectra. Theoretically, the above fragmentation algorithm can be recursively applied to previously generated fragments to simulate a multi-stage fragmentation process, but the computational cost will increase exponentially with the number of chemical bonds in the molecule and the depth of recursion. Inspired by the GRAFF-MS algorithm, we define a fixed list of neutral molecular formulas, called "secondary losses". By subtracting these small neutral losses from the primary fragments, a variety of possible ionic fragments can be generated efficiently without explicitly modeling multi-stage fragmentation.

[0056] We first validated the effectiveness of our method by observing whether the neutral molecules in the secondary loss library reflected common and chemically significant neutral losses observed in real tandem mass spectrometry. The five most frequently occurring secondary losses by weighting were "H2O, H2, CO, NH3, and CO2," all of which are common neutral loss fragments during mass spectrometry fragmentation. This result reflects the interpretability of the strategy of constructing the secondary loss table. When comparing the performance of applying secondary loss libraries of different sizes, we found that when using a secondary loss library of size 100, the median coverage of both unweighted and weighted peaks reached close to 100%. It is worth noting that only a small number of molecules contain elements such as bromine or phosphorus, and simply counting the frequency of secondary losses would introduce bias. Therefore, in addition to considering the frequency of various secondary losses, we supplemented the top 100 secondary losses consisting only of C, H, O, and N by including secondary losses containing F, Cl, Br, I, S, and P. The final secondary loss library size was 144. Using the secondary loss library, the fragments generated by SPHERE-MS almost completely cover the experimental fragment space, with the 90th percentile of weighted and unweighted peak coverage approaching 100%. Therefore, we ultimately chose this secondary loss library of size 144 for application to SPHERE-MS.

[0057] 2. Performance of forward tandem mass spectrometry prediction On the NIST-20 internal test set, the mean cosine similarity of SPHERE-MS was 0.688 (Table 2), indicating that SPHERE-MS successfully captured most of the fragment ions generated during tandem mass spectrometry and accurately predicted their relative peak intensities.

[0058] Table 2. Cosine similarity between predicted and actual mass spectra on the NIST-20 retained test set, external test sets MSnLib, and CASMI2022.

[0059] a: Training and evaluation were performed using the same NIST-20 training and validation sets; b: Training was conducted on the MSnLib dataset; c: Training was conducted on the METLIN dataset. As shown in Table 2, on almost all test sets, SPHERE-MS and GRAFF-MS predicted mass spectra with higher cosine similarity to the true mass spectra than other benchmark models. The mean and median cosine similarity values ​​of SPHERE-MS predicted mass spectra were 0.687 and 0.789, respectively, slightly higher than GRAFF-MS (mean and median were 0.658 and 0.773, respectively). A similar situation was observed in CASMI2022, where SPHERE-MS slightly outperformed GRAFF-MS. On MSnLib, SPHERE-MS's performance metrics were significantly higher than GRAFF-MS. Notably, SPHERE-MS and GRAFF-MS were trained using the exact same NIST-20 training set, highlighting SPHERE-MS's superior generalization ability. The open-source version of FIORA achieves higher performance metrics on the MSnLib dataset than the other two datasets, likely because FIORA was developed using MSnLib as its training set. Nevertheless, SPHERE-MS still outperforms FIORA on MSnLib. This result further demonstrates SPHERE-MS's performance advantage in forward tandem mass spectrometry prediction. A more comprehensive modeling of the fragmentation process enhances SPHERE-MS performance. In addition to single-bond breakage, SPHERE-MS introduces ring breakage and secondary losses, effectively simulating the multi-stage fragmentation process of molecules in mass spectrometry. CFM-ID relies on heuristic bond-breaking rules to enumerate possible fragments and estimate fragmentation probabilities, making it one of the earliest fragmentation-based mass spectrometry prediction models. However, due to the complexity of real-world fragmentation pathways, manually defined rules struggle to capture the diverse fragmentation behaviors, resulting in CFM-ID's predictive performance lagging behind data-driven deep learning models like SPHERE-MS and GRAFF-MS.

[0060] On the CASMI2022 dataset, the performance metrics of all tested models were significantly lower than those on NIST-20. Literature reports that CASMI2022 presents a high degree of structural diversity, making it extremely challenging. However, MSnLib showed the lowest overall structural similarity to the training set among the three test datasets. Yet, the performance degradation of various models on MSnLib was significantly less than that on CASMI2022. To explain this phenomenon, we constructed a low-similarity subset by selecting molecules with a similarity of less than 0.4 to the training set molecule ECFP4 Tanimoto. On the low-similarity subsets of NIST-20 and MSnLib, SPHERE-MS performed comparably, while the test metrics significantly decreased on the low-similarity subset of CASMI2022 (Figure 4b). Therefore, we hypothesize that the quality of the test spectra is another factor. Specifically, the mass spectra in NIST-20 were obtained using standards and manually processed and labeled, while the mass spectra in MSnLib underwent automated processing and manual calibration. In contrast, the CASMI2022 test mass spectrometry is generated by directly merging multiple tandem mass spectrometry scans, a process that may introduce more noise. Differences in data quality can cause anomalies in the spectrum similarity calculation, reducing the prediction accuracy of all models, including SPHERE-MS.

[0061] 3. Dependence on molecular similarity To further evaluate the performance of SPHERE-MS in predicting molecules that differ significantly from the training data, we compared the mean cosine similarity between molecules with varying similarities to the training set and their predicted mass spectra (Figure 5). Since both SPHERE-MS and GRAFF-MS are data-driven deep learning models, they exhibit similar overall trends: prediction accuracy improves with increasing structural similarity to the training data. Notably, the performance gap between SPHERE-MS and GRAFF-MS gradually narrows with increasing molecular similarity. For a subset with very low structural similarity (ECFP4 Tanimoto similarity < 0.3), SPHERE-MS predicted a mean cosine similarity of 0.525, higher than GRAFF-MS (mean cosine similarity 0.363). In contrast, for molecules with high structural similarity to the training set (similarity > 0.9), GRAFF-MS even slightly outperformed SPHERE-MS in terms of mean cosine similarity of its predicted mass spectra. Since SPHERE-MS and GRAFF-MS use the same training data and employ similar graph neural network architectures for feature extraction, their performance differences stem from their different fragmentation generation strategies. GRAFF-MS is a purely statistically dependent modeling approach: on NIST-20, it constructs a fixed-size vocabulary containing 10,000 ion fragments and neutral losses based solely on frequency of occurrence. This approach excels at predicting structures highly similar to molecules in the training set, but the limited number of compound structures in NIST-20 means that a purely statistically dependent approach may introduce bias, reducing its applicability to structurally dissimilar compounds. In contrast, SPHERE-MS incorporates prior knowledge from the field of mass spectrometry into a data-driven framework. Specifically, (i) it generates primary fragments by adjusting single bond breaking, ring breaking, and the number of restricted hydrogen atoms; and (ii) it builds a library of 144 chemically interpretable neutral losses. The selection of these secondary losses is based not only on frequency statistics but also on chemical rules (including nitrogen rules) and additional constraints on rare elements such as sulfur, phosphorus, and halogens. These designs help improve the robustness and generalization ability of SPHERE-MS, especially for molecules with low structural similarity.

[0062] 4. Library matching To evaluate the ability of mass spectrometry prediction models to resolve the spectra of unknown compounds using a reference library, we conducted a library matching test. For each mass spectra tested in the NIST-20 test set, the model needed to identify real molecules from a candidate pool containing 50 structures, including the correct molecules and 49 structurally most similar "decoy" molecules retrieved from PubChem. Predicted mass spectra were ranked based on the cosine similarity between the predicted and experimental mass spectra, and performance was evaluated by recording the frequency of the top k correct molecular structures at different k values. The top-1 accuracies for SPHERE-MS, GRAFF-MS, and FIORA-OS were 39.4%, 34.4%, and 32.2%, respectively. Compared to the other two models, SPHERE-MS improved the top-1 accuracy by 14.5% and 22.4%, respectively, indicating a significant advantage in distinguishing structurally similar compounds. When considering the top 10 candidate compounds, SPHERE-MS achieved a success rate of approximately 80%. Besides retrieval accuracy, building large reference libraries requires models with high computational efficiency. In this task, SPHERE-MS used a single RTX 4070 Ti SUPER GPU to predict approximately 1.5 billion images in less than two hours, demonstrating its computational efficiency suitable for large-scale mass spectrometry library construction. In summary, these results demonstrate the effectiveness of SPHERE-MS in distinguishing structurally similar compounds by mass spectra, highlighting its practical value in mass spectrometry structure resolution.

[0063] Example 2: Establishing a virtual reference mass spectrum library based on SPHERE-MS to assist in secondary mass spectrometry analysis (1) Construction of the virtual reference mass spectrometry library: A large number of small molecule compounds were obtained from the publicly available compound structure database PubChem and their molecular structures were input one by one into the SPHERE-MS model trained in Example 1. For each candidate molecule, the following steps were performed: ① Construct a set of potential ion fragments based on the molecular structure; ② Predict the relative abundance of each potential ion fragment using the SPHERE-MS graph neural network model; ③ Generate the corresponding virtual secondary mass spectrometry library based on the prediction results. Through the above steps, a virtual reference mass spectrometry library based on the PubChem dataset was constructed.

[0064] (2) Preparation of the test sample and acquisition of secondary mass spectrometry data: After standardized sample pretreatment, the secondary mass spectrometry data of the test sample were acquired using a high-resolution mass spectrometer. The experimental procedure is as follows: Take about 5 mg of the test sample and dissolve it in 250 ml of 0.1% formic acid-water:acetonitrile (40:60). Centrifuge at 13000 g for 10 min, filter, discard the initial filtrate, and take the subsequent filtrate as the test solution. The test sample was measured by mass spectrometry using an AB Sciex TripleTOF 5600 instrument. Instrument parameters: Scan range: 100~2000 Da; Mode: positive ion; Spray gas: 50 psi; Auxiliary heating gas: 50 psi; Curtain gas: 30 psi; Ion source temperature: 350℃; Ionization voltage: 5500 V; Declustering voltage: 80 V; Collision voltage: 15 eV. Chromatographic column: ACQUITY UPLC BEH C18 (2.1 mm × 100 mm, 1.7 μm); mobile phase: 0.1% formic acid-water:acetonitrile (40:60); flow rate: 0.3 ml / min; column temperature: 30 ℃; injection plate temperature: 5 ℃; injection volume: 50 μl. The obtained secondary mass spectrometry data underwent unified preprocessing, including isotope peak merging, noise peak filtering, and peak intensity normalization, to obtain the experimental secondary mass spectra to be analyzed.

[0065] (3) Matching analysis between experimental spectra and virtual reference mass spectrometry library Reference Mass Spectrometry Library Matching and Auxiliary Analysis: The obtained virtual reference spectrum library is matched with the measured spectra using cosine similarity calculation. The spectra are then sorted from highest to lowest similarity, and candidate molecular structures and their confidence levels are output. This provides auxiliary analysis for secondary mass spectrometry, aiding in the rapid identification of unknown compounds. The formula for calculating cosine similarity is as follows: Error window Set to 0.01 Da, where m and I are the atomic weight and normalized abundance of the peak, respectively.

[0066] Based on the similarity scores between spectra, candidate molecular structures are automatically ranked, and the top-ranked candidate structures are output as the analysis results. The spectra matching results show that the virtual secondary mass spectra corresponding to aspirin rank first in the similarity ranking, with a significantly higher similarity score than other candidate molecular structures. Based on this, the system automatically identifies the sample as aspirin, achieving correct identification of the sample's molecular structure.

Claims

1. A secondary mass spectrometry-assisted analysis method based on the artificial intelligence model SPHERE-MS, characterized in that, Includes the following steps: (1) Constructing the training set for the SPHERE-MS model Based on known compounds and their corresponding experimental secondary mass spectrometry data, a dataset for training the SPHERE-MS model was constructed. The experimental secondary mass spectrometry data was obtained from a public mass spectrometry database. Potential ion fragments were generated using a hybrid algorithm of single bond breaking, secondary fragment library, and secondary loss library, and an ion fragment space corresponding to the parent ion was constructed. (2) Training of SPHERE-MS model and fragment abundance prediction The two-dimensional structure of the parent ion is transformed into a molecular graph, and the experimental conditions of mass spectrometry are used as experimental parameter variables and input into the SPHERE-MS model. Through feature embedding, graph structure message passing and multi-scale feature fusion, the mapping relationship between molecular structure, experimental conditions and molecular fragmentation behavior is learned. The intensity of the second-order mass spectrometry peak corresponding to the ion fragment space obtained in step (1) is used as the training label to train the SPHERE-MS model, so that the model has the ability to predict the relative abundance of each potential ion fragment in the ion fragment space. (3) Constructing a virtual reference mass spectrometry library Using the trained SPHERE-MS model, batch secondary mass spectrometry predictions are performed on a large number of candidate molecular structures to generate virtual secondary mass spectra for each candidate molecular structure. The molecular structure information is associated with the corresponding virtual secondary mass spectra and stored to construct a searchable virtual reference mass spectrometry library. (4) Preparation of unknown samples and acquisition of secondary mass spectrometry data Standardized pretreatment was performed on the unknown samples to be analyzed, and secondary mass spectrometry data were acquired using a high-resolution mass spectrometer. The secondary mass spectrometry data are processed through a unified preprocessing procedure and then used as the spectrum to be analyzed. (5) Reference mass spectrometry library matching and auxiliary analysis The secondary mass spectrum of the unknown sample obtained in step (4) is compared with the virtual reference mass spectrum library constructed in step (3) for similarity. The candidate molecular structures are sorted according to the similarity score, and the molecular structures with the highest similarity ranking are output as auxiliary analysis results of the unknown sample.

2. The method according to claim 1, characterized in that, The construction of the training set for the SPHERE-MS model in step (1) is implemented as follows: A. Select the NIST-20 database as the data source for the training set; B. Construction of Ion Fragment Space: A hybrid algorithm of single bond breaking, secondary fragment library, and secondary loss library is used to generate potential ion fragments, and a secondary mass spectrometry prediction model is constructed to predict the abundance of each ion fragment. The specific implementation process is as follows: First, by single-time breaking of chemical bonds in the molecular structure and breaking of ring structures in the molecular structure, combined with adjustment of the number of restrictive hydrogen atoms, a set of potential primary ion fragments is obtained; then, high-frequency secondary fragments are subtracted from the ion fragments in the set of primary ion fragments to obtain all possible ion fragments, thereby completing the construction of the ion fragment space.

3. The method according to claim 2, characterized in that, The construction of the training set for the SPHERE-MS model is carried out according to the following steps: A. When selecting the NIST-20 database as the training set data source, NIST-20 undergoes the following filtering steps: only the parent ions of the [M+H]+ and [MH]- types are considered; molecules containing atoms other than {C, H, N, O, P, S, F, Cl, Br, I} are excluded; the filtered mass spectrometry data is used to statistically establish a secondary loss library; B. Primary fragments are generated through single bond and ring cleavage: ① The parent ion is represented as a molecular graph structure, where atoms are nodes and chemical bonds are edges. Edge removal is performed on the molecular graph to simulate the fragmentation process in mass spectrometry. ② A single chemical bond breaking algorithm is used to generate primary fragments, i.e., by removing single chemical bonds in non-cyclic structures, the parent ion is fragmented into charged fragment ions and neutral loss molecules. ③ For cyclic structures present in the molecule, a preset number of chemical bonds in each ring structure are set as breakable bonds. Breaking the chemical bonds within the ring generates primary fragment ions produced by ring fragmentation. ④ During the single chemical bond breaking and ring structure breaking processes, a limiting hydrogen atom number adjustment mechanism is introduced to adjust the generated... Fragment ions are defined by three forms of hydrogen atom change: the number of hydrogen atoms remains unchanged, one hydrogen atom is added, or one hydrogen atom is removed, to characterize the hydrogen migration phenomenon that may occur during actual fragmentation; ⑤ All fragment ions generated through the breaking of the single chemical bond, the breaking of the ring structure, and the rearrangement of hydrogen atoms, together with the parent ion itself, constitute the primary fragment set; ⑥ Based on a fixed-scale secondary neutral loss fragment library pre-constructed from an experimental mass spectrometry database, the secondary neutral loss is applied sequentially to the primary fragment set, and by subtracting the corresponding neutral loss molecular formula from the primary fragments, all potential secondary fragment ion sets are generated, constituting the ion fragment space of the parent ion.

4. The method according to claim 1, characterized in that, Includes the following steps: Step (2) Training the SPHERE-MS model and predicting fragment abundance is carried out as follows: ① The two-dimensional structure of the precursor ion is represented as an undirected molecular graph, where nodes represent heavy atoms and edges represent interatomic chemical bonds. Node features include at least the atomic element type, hybridization state, formal charge, and number of bonded hydrogen atoms. Edge features include bond type and whether it belongs to a ring structure. Mass spectrometry experimental parameters are undirected molecular graph supernode variables, including instrument type, precursor ion type, and normalized collision energy. Based on this, the first few feature vectors and eigenvalues ​​of the molecular graph Laplacian matrix are further extracted to characterize the overall topological structure of the molecule. ② Graph neural network feature extraction and multi-scale representation fusion: Node features, edge features, and mass spectrometry experimental parameter variables are mapped to a unified-dimensional latent feature space through an embedded network. Graph neural networks are used to perform multi-layer message passing and feature updates on the molecular graph to simultaneously fuse node attributes, neighborhood structure information, and chemical bond features. After node feature extraction, node features belonging to the same first-level fragment are aggregated to obtain fragment-level feature representations. Attention-weighted pooling is applied to the entire molecular graph nodes to obtain molecular graph-level feature representations. The molecular graph-level structural features are fused with the mass spectrometry experimental parameter variable features to form global supernode features of the molecular graph. ③ Prediction of the generation probability of first-level fragments: Based on the similarity relationship between the fragment-level features and the supernode features, the probability distribution of each first-level fragment as a real fragmentation product is predicted, and the calculation formula is as follows: ⑤ Conditional probability prediction of hydrogen atom adjustment and second-order neutral loss: For each type of primary fragment, different output layers are used to predict the conditional probability distributions of different hydrogen atom adjustment states and the conditional probability distributions of different second-order neutral losses; the conditional probabilities are calculated by the output layer composed of a multilayer sensing mechanism, and their expressions are as follows: ⑥ Joint probability modeling and output of fragment ion abundance: The first-order fragment selection probability, the hydrogen atom adjustment conditional probability, and the second-order neutral loss conditional probability are treated as independent events, and their joint probability is used as the predicted abundance of the corresponding ion fragment. The expression is as follows: ⑦ Training and Inference of SPHERE-MS: During the training phase, based on the mapping relationship between fragments and true mass spectrum peaks, the predicted abundances of multiple fragments belonging to the same element are merged to obtain the predicted abundance. The cross-entropy loss function is used to calculate the difference between the predicted abundance and the true abundance, which is used as the optimization objective to optimize the model parameters. The loss function is calculated as follows: ⑧ During the inference stage, the predicted abundance of ion fragments with the same chemical formula are merged, and the intensity of the final output mass spectrum peaks is normalized to obtain the predicted secondary mass spectrum.

5. The method according to claim 1, characterized in that, Includes the following steps: Step (2) Training the SPHERE-MS model is carried out as follows: A. Mapping of possible ion fragments with NIST-20 mass spectrum peaks: For each parent ion structure and its corresponding mass spectrum, possible fragment ions are obtained using a hybrid algorithm combining single-batch bond breaking, secondary fragment library, and secondary loss library. ,in Indicates possible fragment ions. and They represent Chemical formula and precise mass; peak signals in mass spectrometry , where m p The charge-to-mass ratio of the peak, Indicates the labeled chemical formula, Indicates peak intensity; Ion fragments are determined based on the following two criteria. and experimental mass spectrometry peak signals The correspondence between them: (1) Or (2) or The predicted set of potential ion fragments is matched with experimental mass spectrometry peaks; successfully matched peaks are used as the true values, while unmatched peaks are ignored. Then, the intensities of the successfully matched peaks are normalized so that their sum equals 1. Finally, a virtual peak with zero intensity is created. To match ion fragments that do not correspond to any peaks: ; B. Encoding of molecular structure and tandem mass spectrometry experimental parameter variables: The two-dimensional molecular structure of the precursor ion is described using a graph data structure. The node features are characterized by one-hot encoding to represent the atom type, hybridization state, charge information, and number of bonding hydrogen atoms, while the edge features are characterized by one-hot encoding to represent the chemical bond type and whether it belongs to a ring structure. At the same time, information such as the instrument type, precursor ion type, and collision energy during the mass spectrometry acquisition process are also input into the model as experimental parameters using one-hot encoding to enhance the model's adaptability to changes in experimental conditions. The two-dimensional molecular structure is represented as an undirected graph, where each node corresponds to a heavy atom and each edge represents a chemical bond between two atoms; the first eight eigenvectors and eigenvalues ​​of the graph Laplacian operator are used as features to capture the structural features of the molecule; Mass spectrometry experimental parameters are treated as graphical features; C. Feature Extraction Module: The initial encoding of molecular graph nodes, edges, and mass spectrometry experimental parameter variables is converted through independent embedding layers. Each embedding layer consists of an MLP block with GraphNorm: Where h represents the dimension of the latent feature (h = 512), and the dropout rate is set to 0.1; unless otherwise specified, these settings are the default settings for other modules in SPHERE-MS; The first 8 eigenvectors of the Graph Laplacian operator and eigenvalues Convert using SignNet The embedding layer in SignNet consists of an MLP block without a GraphNorm layer: The obtained embedded features are then used for feature extraction via a graph attention network module: The extracted node features are represented as follows The following abbreviation is x. First-level fragmented features are obtained through summation and aggregation. This preserves information about fragment size; then, the AttentionalAggregation operator is used to pool the graph nodes to obtain graph-level supernode features. Subsequently, through layer combination with residual connections and To obtain graph-level features : Each Level 1 Fragment The probability is calculated based on the similarity between the extracted graph-level features and the features of each first-level fragment: D. Prediction module for ion fragment probability: Output layer with residual connections to predict hydrogen atom number adjustment And the loss of different secondary fragments Conditional probability distribution: Ion fragments The probability (i.e., peak intensity) can be viewed as the joint probability of three independent events: primary fragmentation. Adjustment of the number of hydrogen atoms and lost secondary fragments : According to ion fragments and experimental mass spectrometry peaks The matching relationships between them are used to predict peak intensity. : The cross-entropy loss function is used as the loss function for model training: During the reasoning phase, due to ion fragments The matching relationship between the peaks and the experimental mass spectrometry peaks is unknown, and those with the same chemical formula will be... Ion fragments The peak intensity is obtained by summing the intensities; then the obtained peak intensities are normalized so that the sum equals 1.

0. Mass spectrometry cosine similarity: Mass spectrometry cosine similarity (using the CosineHungarian algorithm provided in the matchms library) is used to evaluate the accuracy of the predicted results with the actual mass spectra; tolerance Window size set to 0.1 Da: E. Model Implementation and Training: SPHERE-MS is primarily developed using the PyTorch deep learning framework. PyTorchGeometric is used to process graph data and graph neural network models, while PyTorchLightning is used to simplify the model training process. F. Model Evaluation: Benchmark models: GRAFF-MS, FIORA, and CFM-ID were used as benchmark models for comparison. Library matching: For each mass spectrum in the NIST-20 test set, 49 "decoy" isomers with the highest Tanimoto similarity to the real molecular structure were collected from the PubChem database; this formed a library of 50 molecules for each test mass spectrum; the similarity between the predicted mass spectra of the molecular structures in the library and the real mass spectra was sorted to calculate the frequency of the real molecules in the top k at different k values.

6. The method according to claim 1, characterized in that, Step (4) Preparation of unknown samples and acquisition of secondary mass spectrometry data: Standardized sample pretreatment for unknown samples to be analyzed: Test solution: Take about 5 mg of the sample to be tested, add 250 ml of 0.1% formic acid-water:acetonitrile 40:60 to dissolve, centrifuge at 13000g for 10 min, filter, discard the initial filtrate, and take the subsequent filtrate as the test solution; Acquisition of secondary mass spectrometry data: The obtained secondary mass spectrometry data were preprocessed uniformly, including isotope peak merging, noise peak filtering, and peak intensity normalization, to obtain the experimental secondary mass spectra to be analyzed.

7. The method according to claim 1, characterized in that, Step (5) Reference Mass Spectrometry Library Matching and Auxiliary Analysis: The virtual reference spectrum library obtained in step (3) is matched with the measured spectrum for similarity. The similarity is calculated using cosine similarity and then sorted from high to low similarity. The candidate molecular structures and confidence levels are output, thereby achieving auxiliary analysis of the secondary mass spectrometry and helping to quickly identify unknown compounds. The formula for calculating cosine similarity is: Error window Set to 0.01 Da, where m and I are the atomic weight and normalized abundance of the peak, respectively.