A method and system for visual classification of compounds

By using a compound visualization classification method and generating molecular networks using algorithms such as SIRIUS, the challenge of compound identification in non-targeted mass spectrometry data has been solved, enabling a switch from non-targeted to targeted analysis and improving the accuracy and efficiency of compound identification.

CN115691702BActive Publication Date: 2025-12-02ZHEJIANG CHINESE MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211428657.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-12-02
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and focus on compounds or chemicals of interest to researchers in non-targeted mass spectrometry data analysis. Manual identification is time-consuming and results are influenced by subjective factors. Small spectral libraries lead to low accuracy, and computer prediction methods lack effective integration.

Method used

A compound visualization classification method is adopted. Raw mass spectrometry data is acquired, preprocessed and screened, and advanced algorithms such as SIRIUS are used to generate molecular networks to achieve the visualization classification and focusing of compounds.

Benefits of technology

It enables a switch from non-targeted to targeted analysis, intuitively focusing on compounds or chemical categories of interest to researchers, improving the accuracy and efficiency of compound identification, and reducing the time and subjective influence of manual identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691702B_ABST
    Figure CN115691702B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for visualizing and classifying compounds, specifically in the field of compound classification. The method includes: preprocessing raw mass spectrometry data to obtain compound information; selecting the compound with the highest molecular formula score as the optimal molecular formula; selecting the compound with the highest probability of structure belonging to the optimal molecular formula as the structure dataset; filtering the categories of compounds belonging to the optimal molecular formula based on priority parameters, posterior probability, a set threshold, and the probability of the compound belonging to a category, resulting in a category dataset; filtering the category dataset according to set conditions, resulting in a filtered category dataset; clustering the filtered category dataset to obtain multiple clusters; generating a molecular network based on the multiple clusters; and mapping the structure dataset to the molecular network. This invention enables a switch from non-targeted analysis to targeted analysis, intuitively focusing on compounds or chemical classes of interest to researchers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of compound classification, and in particular to a method and system for visually classifying compounds. Background Technology

[0002] Non-targeted mass spectrometry (LC-MS / MS) is a powerful tool in biological research, but researchers often spend considerable time analyzing the datasets. The analysis of LC-MS / MS data is particularly complex due to the large volume of data, the complexity of the spectra, and the diversity of compound structures. Identification of fragment spectra is a major challenge. Several strategies exist for analyzing fragment spectra, including: 1) Spectral library matching. This method remains mainstream and offers high accuracy. However, compared to structural databases (PubChem has over 100 million records), spectral libraries are too small, limiting the application of mass spectrometry. 2) Matching fragment spectra with computer-simulated fragment spectra. 3) Using machine learning to predict fragment spectra. Computer prediction methods are rapidly developing. The cutting-edge technology SIRIUS 4, combined with many advanced artificial intelligence algorithms, achieves 70% accuracy when searching within structural libraries. This method helps identify metabolites outside the scope of spectral libraries. While computer prediction techniques facilitate chemical identification, a method that integrates and utilizes the latest technologies in biological research—specifically, the discovery of biomarkers in non-targeted mass spectrometry datasets—remains lacking. Manual identification and screening of biomarkers is very time-consuming, and the results are susceptible to subjective factors. In terms of identification, molecular networks are increasingly popular due to their visualization and data transparency.

[0003] The history of chemical classification can be traced back to at least the mid-20th century, with the Derwent World Patent Index (DWPI) developing a chemical fragment coding system in 1963. In recent years, more systematic approaches to chemical classification have emerged, such as Gene Ontology (GO), combining it with taxonomy and ontology. ClassyFire, due to its computational availability and systematic nature, is increasingly used for composite annotation in both massive and non-massive datasets. Chemical taxonomy and ontology are beneficial. For example, a hierarchical classification-based approach called Qemistree has been proposed to handle chemical relationships within a dataset. However, chemical taxonomy or ontology is not a one-size-fits-all solution for pharmacological or biological research. Many key metabolites or drugs within chemical classes are distributed across different levels, such as “bile acids, alcohols and derivatives” (subclass), “indole and derivatives” (class), and “acylcarnitines” (level 5). These categories represent families of compounds with similar biological functions or activities; however, functionally or activity-independent compounds are scattered across different branches at different taxonomic levels.

[0004] Therefore, there is a need for a method that allows switching from non-targeted to targeted analysis, enabling researchers to intuitively focus on compounds or chemicals of interest. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for visual classification of compounds, so as to switch from non-targeted analysis to targeted analysis and intuitively focus on compounds or chemical categories of interest to researchers.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] A method for visually classifying compounds, comprising:

[0008] Obtain raw mass spectrometry data of the compound;

[0009] The raw mass spectrometry data is preprocessed to obtain compound information; the compound information includes the compound molecular formula, molecular formula score, probability of the compound belonging to the category, and probability of the compound structure.

[0010] The molecular formula with the highest score among the stated molecular formulas is selected as the optimal molecular formula.

[0011] The compound structure with the highest probability of selecting the optimal molecular formula is used as the structure dataset.

[0012] The categories of compounds with the optimal molecular formula are filtered based on priority parameters, posterior probabilities, set thresholds, and the probability of the category to which the compound belongs, to obtain a category dataset;

[0013] The category dataset is filtered according to set conditions to obtain a filtered category dataset; the set conditions include the position of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores;

[0014] The filtered category dataset is clustered to obtain multiple clusters;

[0015] A molecular network is generated based on multiple said clusters; the structure dataset is mapped to the molecular network; the molecular network is used to visualize the categories and structures of the compounds; the points of the molecular network include the categories and structures of the compounds; the edges of the molecular network are determined by the secondary fragment similarity of the original mass spectrometry data of different compounds.

[0016] Optionally, the preprocessing of the raw mass spectrometry data to obtain compound information specifically includes:

[0017] The raw mass spectrometry data is then converted to a format to obtain Extensible Markup Language (EXPLAIN).

[0018] The extensible markup language was used for feature detection using MZmine2 and for analysis using SIRIUS to obtain compound information.

[0019] Optionally, the step of filtering the categories of compounds with the optimal molecular formula based on priority parameters, posterior probabilities, set thresholds, and the probability of the category to which the compound belongs, to obtain a category dataset, specifically includes:

[0020] Based on the priority parameters, the categories of compounds with the optimal molecular formula are initially screened to obtain preliminary screening results;

[0021] The preliminary screening results are then subjected to a secondary screening based on the posterior probability, the set threshold, and the probability of the compound belonging to a certain category, to obtain a category dataset.

[0022] Optionally, the step of filtering the category dataset according to the set conditions to obtain the filtered category dataset specifically includes:

[0023] Delete the categories representing the positions of the chemical functional groups from the category dataset to obtain the first-level filtering result;

[0024] Delete the categories with the largest number of category features and the categories with the smallest number of category features from the first-level filtering results to obtain the second-level filtering results;

[0025] The third-level filtering result is obtained by deleting categories that contain all the same features in the second-level filtering result;

[0026] Calculate the similarity score between each pair of categories;

[0027] Remove categories from the three-level filtering results whose similarity scores are less than the minimum reach rate to obtain the filtered category dataset.

[0028] A compound visualization classification system, comprising:

[0029] The acquisition module is used to acquire the raw mass spectrometry data of the compound;

[0030] The preprocessing module is used to preprocess the raw mass spectrometry data to obtain compound information; the compound information includes the compound molecular formula, molecular formula score, probability of the compound belonging to the category, and probability of the compound structure.

[0031] The optimal molecular formula determination module is used to select the molecular formula with the highest molecular formula score among the molecular formulas of the compound as the optimal molecular formula.

[0032] The structure dataset determination module is used to select the compound structure with the highest probability of the optimal molecular formula as the structure dataset.

[0033] The filtering module is used to filter the category of the compound with the best molecular formula according to the priority parameter, posterior probability, set threshold and the probability of the category of the compound, so as to obtain the category dataset;

[0034] The filtering module is used to filter the category dataset according to set conditions to obtain a filtered category dataset; the set conditions include the position of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores;

[0035] The clustering module is used to cluster the filtered category dataset to obtain multiple clusters;

[0036] A generation module is used to generate a molecular network based on multiple said clusters; the structure dataset is mapped to the molecular network; the molecular network is used to visualize the categories and structures of the compounds; the points of the molecular network include the categories and structures of the compounds; the edges of the molecular network are determined by the secondary fragment similarity of the original mass spectrometry data of different compounds.

[0037] Optionally, the preprocessing module specifically includes:

[0038] The format conversion unit is used to convert the original mass spectrometry data into a format extensible markup language.

[0039] The feature detection and analysis unit is used to perform feature detection on the extensible markup language using MZmine2 and analysis using SIRIUS to obtain compound information.

[0040] Optionally, the filtering module specifically includes:

[0041] The preliminary screening unit is used to perform preliminary screening of the category of the compound with the optimal molecular formula according to the priority parameter, and to obtain preliminary screening results;

[0042] A secondary screening unit is used to perform secondary screening on the preliminary screening results based on the posterior probability, the set threshold, and the probability of the compound belonging to the category, to obtain a category dataset.

[0043] Optionally, the filtering module specifically includes:

[0044] A first-level filtering unit is used to delete categories representing the positions of the chemical functional groups in the category dataset to obtain the first-level filtering result;

[0045] The secondary filtering unit is used to delete the categories with the largest number of category features and the categories with the smallest number of category features from the primary filtering result, thereby obtaining the secondary filtering result;

[0046] The third-level filtering unit is used to delete categories that contain all the same features in the second-level filtering results, thus obtaining the third-level filtering results;

[0047] The calculation unit is used to calculate the similarity score between each pair of categories;

[0048] The filter category dataset determination unit is used to delete categories with similarity scores less than the minimum reach rate from the three-level filtering results, thereby obtaining the filter category dataset.

[0049] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0050] This invention acquires raw mass spectrometry data of compounds; preprocesses the raw mass spectrometry data to obtain compound information; the compound information includes the compound molecular formula, molecular formula score, probability of the compound belonging to a category, and probability of the compound structure; selects the compound molecular formula with the highest molecular formula score as the optimal molecular formula; selects the compound structure with the highest probability of the optimal molecular formula as the structure dataset; filters the compound categories of the optimal molecular formula according to priority parameters, posterior probability, a set threshold, and the probability of the compound belonging to a category to obtain a category dataset; filters the category dataset according to set conditions to obtain a filtered category dataset. The set conditions include the location of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores. The filtered category dataset is clustered to obtain multiple clusters. A molecular network is generated based on these clusters. The structure dataset is mapped to the molecular network. The molecular network is used to visualize the categories and structures of the compounds. The points of the molecular network include the categories and structures of the compounds. The edges of the molecular network are determined by the secondary fragment similarity of the original mass spectrometry data of different compounds, thereby enabling a switch from non-targeted analysis to targeted analysis, intuitively focusing on compounds or chemical classes of interest to the researcher. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 Flowchart of the compound visualization classification method provided by the present invention;

[0053] Figure 2 This is a schematic diagram illustrating the practical application of the compound visualization classification method provided by the present invention;

[0054] Figure 3 A schematic diagram of the data hierarchy of the compound visualization classification method provided by this invention;

[0055] Figure 4 A visualization diagram illustrating a specific example of the compound visualization classification method provided by this invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] The purpose of this invention is to provide a method and system for visual classification of compounds, so as to switch from non-targeted analysis to targeted analysis and intuitively focus on compounds or chemical categories of interest to researchers.

[0058] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] like Figures 1 to 3 As shown, the present invention provides a method for visual classification of compounds, comprising:

[0060] Step 101: Obtain the raw mass spectrometry data of the compound.

[0061] Step 102: Preprocess the raw mass spectrometry data to obtain compound information; the compound information includes the compound molecular formula, molecular formula score, probability of the compound's category, and probability of the compound's structure.

[0062] Step 102 specifically includes:

[0063] The raw mass spectrometry data was converted to Extensible Markup Language (MZML) using MSconvertProteowizard.

[0064] The extensible markup language was used for feature detection using MZmine2 and for analysis using SIRIUS to obtain compound information.

[0065] Feature detection is performed using MZmine2 (version 2.53). A SIRIUS analysis workflow is executed, involving SIRIUS, ZODIAX, CSI:fingerID, and CANOPUS.

[0066] Step 103: Select the molecular formula with the highest molecular formula score among the compound molecular formulas as the optimal molecular formula.

[0067] The MCnebula (Multi-chemical nebula) processing workflow is implemented within an R package. In the R console or Studio, by loading the MCnebula package and using its various functions, data preparation, integration, and visualization can be achieved.

[0068] For each feature, multiple molecular formula candidates may exist as a result of the calculation. MCnebula considers both ZODIAC and CSI:fingerID scores to obtain the optimal molecular formula. If any structural candidate is retrieved by CSI:fingerID, MCnebula prioritizes the structural formula with the highest score by default. If no structural candidate is found, MCnebula selects the molecular formula with the highest ZODIAC score. The priority of selecting the molecular formula with the highest ZODIAC or CSI:fingerID score can be manually reversed. The selection of the optimal molecular formula will determine the selection of structural formulas and PPCP data in the following algorithms. Specifically, since multiple molecular formula candidates are generated during the prediction process of each compound, and each molecular formula has its own PPCP candidate dataset and structural candidates, the determination of the optimal molecular formula is particularly important, as it is at a crossroads in upstream analysis. Subsequently, the optimal molecular formulas for all features are collected into the MCnebula molecular formula set (.MCn.formula_set).

[0069] Step 104: Select the compound structure with the highest probability of the optimal molecular formula as the structure dataset.

[0070] Based on the .MCn.formula_set, for each feature, only the best molecular formula is considered, and MCnebula selects the best structure (i.e., the structure with the highest score) from the candidate CSI:fingerID chemical structures. The selected structures are then collected into the MCnebula structure set (.MCn.structure_set).

[0071] Step 105: Filter the categories of compounds with the optimal molecular formula according to priority parameters, posterior probabilities, set thresholds, and the probability of the category to which the compound belongs, to obtain a category dataset.

[0072] Step 105 specifically includes:

[0073] The compounds with the optimal molecular formula are initially screened according to the priority parameters to obtain preliminary screening results.

[0074] The preliminary screening results are then subjected to a secondary screening based on the posterior probability, the set threshold, and the probability of the compound belonging to a certain category, to obtain a category dataset.

[0075] The posterior probability of classification prediction (PPCP) for each compound category is compiled. Similarly, based on the .MCn.formula_set, for each feature, considering only the optimal molecular formula, MCnebula extracts PPCP data for all categories (this dataset is a text dataset, read in R and merged with the corresponding directory). This data is collected as the MCnebulaPPCP dataset (.MCn.ppcp_dataset).

[0076] Summarize the category dataset. In the .MCn.ppcp_dataset, for each feature, there are thousands of posterior probabilities for class predictions. A threshold is set (default is T). ppcp The data is filtered using a parameter of 0.5. Additionally, a priority parameter for chemical classification is set (default is P). hierarchy.priority =c(6,5,4,3), equivalent to ClassyFire's level 5, subclass, class, superclass, etc., are used to filter and sort these categories. The original .MCn.ppcp_dataset contains a large amount of substructure or dominant structure class prediction data. This step aims to obtain those classes that are beneficial for identification. After filtering, the dataset is collected into a category dataset (.MCn.nebula_class).

[0077] Step 106: Filter the category dataset according to the set conditions to obtain the filtered category dataset; the set conditions include the position of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores.

[0078] Step 106 specifically includes:

[0079] The categories representing the positions of the chemical functional groups in the dataset are deleted to obtain the first-level filtering result.

[0080] The categories with the highest number of category features and the categories with the lowest number of category features in the first-level filtering results are deleted to obtain the second-level filtering results.

[0081] The third-level filtering result is obtained by deleting categories that contain all the same features in the second-level filtering result.

[0082] Calculate the similarity score for each pair of categories.

[0083] Remove categories from the three-level filtering results whose similarity scores are less than the minimum reach rate to obtain the filtered category dataset.

[0084] To summarize nebula-index, although the original .MCn.ppcp_dataset was filtered in the previous step, all these categories are still too redundant to provide an overall visualization of the classification of non-targeted LC-MS datasets. In this step, we will implement automatic filtering through the following steps.

[0085] Remove classes representing the positions of chemical functional groups. In fact, MS / MS spectroscopy is not good at distinguishing positional isomers. Due to the nature of the International Union of Applied Chemistry (IUPAC) rules, this measure is implemented by removing class names that involve Arabic numerals in pattern matching.

[0086] Filtering is performed based on features set by the maximum and minimum ownership values ​​of a class. The previously filtered classes are then iterated through the `.MCn.ppcp_dataset`. For any given class, when the PPCP of a feature reaches T... ppcp At this point, the feature will be organized into an index for this class. Then, the feature numbers in the indices of all classes are listed, and it is determined whether the category will be filtered out. Minimum occupancy T min.absence It is determined by absolute numbers, while the maximum possession T mmax.absence It is determined by a relative number (e.g., 20% of all features). The former aims to filter out categories with sparse features, while the latter aims to filter out compound categories with excessive coverage.

[0087] Classes containing nearly identical characteristics are eliminated. The highest chemical classification level is determined (T by default). iden.top.hierarchy =4, which is the level in ClassyFire) and the same factor (T by default) iden.factor =0.7) standard. All below T iden.top.hierarchy The classes are compared in a binary manner. When they have more than T... iden.factor When classes share the same characteristics, those with fewer characteristics will be filtered out.

[0088] Feature classes with low structural identification levels are filtered out. In most cases, incorrect molecular formulas will cause fingerprint predictions from the corresponding fragmentation tree to fail. Both structure and PPCP are matched or calculated based on fingerprints. Incorrect molecular formulas will lead to errors in structural identification and category prediction. From a categorical perspective, some categories have rich features but few structures are matched, or all matched structures have very low similarity. To filter out these classes, an algorithm based on similarity scores is defined. First, the similarity score type is evaluated (by default, P...). simi.score= "Tanimotosimilarity"). Then, set the cutoff value for the similarity score (default is T). simi.score =0.3). All values ​​less than the minimum reach rate (default is T). min.reach Classes with features equal to 0.6 are filtered out. Finally, the remaining classes and related features are collected into MCnebula nebula-index(.MCn.nebula_index).

[0089] Step 107: Cluster the filtered category dataset to obtain multiple clusters.

[0090] Step 108: Generate a molecular network based on the multiple clusters; map the structure dataset to the molecular network; the molecular network is used to visualize the categories and structures of the compounds; the points of the molecular network include the categories and structures of the compounds; the edges of the molecular network are determined by the secondary fragment similarity of the original mass spectrometry data of different compounds.

[0091] Generating a parent nebula is similar to generating a molecular network; the parent nebula consists of node and edge data. Nodes carry feature information or annotations, while edges annotate fragment spectral similarities. To obtain and merge edge and node data into a parent nebula, MCnebula implements the following:

[0092] The secondary fragment similarity of filtered mass spectrometry data between features is evaluated. MCnebula integrates the 'compareSpectra' function of the MSnbaseR package to calculate the cosine similarity between MS / MS spectra. Unlike popular spectral comparison methods, MCnebula does not use the original MS / MS spectra, but instead compiles all noise-filtered MS / MS spectra for comparison. The noise-filtered spectra are from the SIRIUS project space. Different molecular formula candidates for a feature may have their corresponding MS / MS spectra assigned different "valid" or "noise" peaks. Only the "valid" peaks are used to calculate the cosine similarity with the original fragment spectra. To maintain consistency with the algorithm described above, all spectra are acquired based on molecular formulas in the .MCn.formula_set. Furthermore, to reduce computation time, only the same nebula-index(P) is calculated. iden.class Spectral similarity within ); only considering spectral similarity equal to or lower than a specific classification level (T) min.hierarchy =5, by default, i.e., a subclass of ClassyFire). Additionally, if the total number of features exceeds 2000 (by default), the ZODIAC score (default is T) will be adjusted. min.zodiac =0.9) and Tanimoto similarity score (default is T)min.tanimoto =0.4) was used to reduce the number of features that needed to be computed. Then, an edge threshold (default T) was set. edge.filter =0.3) to filter out low similarity. The result is formatted as edge data (.MCn.parentredges).

[0093] Merging multiple datasets. MCnebula merges .MCn.formula_set and .MCn.structural_set into node data (.MCn.parent_nodes).

[0094] The .MCn.parent_nodes and .MCn.parent_edges are integrated into the 'graph' project (.MCn.parent_graph) of the igraph R package. Additionally, the .grahml format file of the parent nebula is exported for interactive exploration in Cytoscape.

[0095] Child nebulas are generated. Based on the .MCn.nebula_index, .MCn.parent_nodes and .MCn.parent_edges are correspondingly partitioned and collected into various "graph" items for generating child nebulas. Simultaneously, a maximum ownership limit (T by default) is defined for each node. max.edges =5), to reduce edges and improve the visualization of sub-nebulae. This means that edges with lower similarity will be cut off first. Finally, all sub-nebula "graphs" are saved to .MCn.child_graph_list and also exported as .graphml format files.

[0096] Visualizing the parent nebula and child nebulae, such as child nebulae... Figure 4 As shown. Various R packages are used for visualization, such as ggplot2, ggrah, etc.

[0097] The analytical methods provided by this invention involve rich datasets of chemical classes, classifications, structures, substructure features, and fragment similarities. Many state-of-the-art techniques and popular methods have been incorporated into the MCnebula workflow to facilitate chemical discovery. MCnebula can be used to explore the classification and structural features of unknown compounds that extend beyond the limitations of spectral libraries. MCnebula is being integrated into the R package for the first time.

[0098] The present invention also provides a compound visualization classification system, comprising:

[0099] The acquisition module is used to acquire the raw mass spectrometry data of the compound.

[0100] The preprocessing module is used to preprocess the raw mass spectrometry data to obtain compound information; the compound information includes the compound molecular formula, molecular formula score, probability of the compound's category, and probability of the compound structure.

[0101] The optimal molecular formula determination module is used to select the molecular formula with the highest molecular formula score among the compound molecular formulas as the optimal molecular formula.

[0102] The structure dataset determination module is used to select the compound structure with the highest probability of the optimal molecular formula as the structure dataset.

[0103] The filtering module is used to filter the category of the compound with the best molecular formula according to the priority parameter, posterior probability, set threshold and probability of the category of the compound, so as to obtain the category dataset.

[0104] The filtering module is used to filter the category dataset according to set conditions to obtain a filtered category dataset; the set conditions include the position of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores.

[0105] The clustering module is used to cluster the filtered category dataset to obtain multiple clusters.

[0106] A generation module is used to generate a molecular network based on multiple said clusters; the structure dataset is mapped to the molecular network; the molecular network is used to visualize the categories and structures of the compounds; the points of the molecular network include the categories and structures of the compounds; the edges of the molecular network are determined by the secondary fragment similarity of the original mass spectrometry data of different compounds.

[0107] As an optional implementation, the preprocessing module specifically includes:

[0108] The format conversion unit is used to convert the original mass spectrometry data into a format that is extensible markup language.

[0109] The feature detection and analysis unit is used to perform feature detection on the extensible markup language using MZmine2 and analysis using SIRIUS to obtain compound information.

[0110] As an optional implementation, the filtering module specifically includes:

[0111] The preliminary screening unit is used to perform preliminary screening of the category of the compound with the optimal molecular formula according to the priority parameter, and obtain preliminary screening results.

[0112] A secondary screening unit is used to perform secondary screening on the preliminary screening results based on the posterior probability, the set threshold, and the probability of the compound belonging to the category, to obtain a category dataset.

[0113] As an optional implementation, the filtering module specifically includes:

[0114] A first-level filtering unit is used to delete categories representing the positions of the chemical functional groups in the category dataset to obtain the first-level filtering result.

[0115] The secondary filtering unit is used to delete the categories with the largest number of category features and the categories with the smallest number of category features from the primary filtering results, thereby obtaining the secondary filtering results.

[0116] The third-level filtering unit is used to delete categories that contain all the same features in the second-level filtering results, thus obtaining the third-level filtering results.

[0117] The calculation unit is used to calculate the similarity score between each pair of categories.

[0118] The filter category dataset determination unit is used to delete categories with similarity scores less than the minimum reach rate from the three-level filtering results, thereby obtaining the filter category dataset.

[0119] The method provided in this invention, called MCnebula, is used for non-targeted LC-MS / MS dataset analysis. MCnebula utilizes state-of-the-art computer prediction techniques—the SIRIUS workflow (SIRIUS, ZODIAC, CSI:fingerID, CANOPUS)—for compound molecular formula prediction, structure retrieval, and class prediction. MCnebula is the first to integrate an abundance-based class selection algorithm into compound annotation. MCnebula also incorporates the advantages of molecular networks, namely intuitive visualization and a large amount of ensembleable information. With MCnebula, researchers can switch from non-targeted to targeted analysis, precisely focusing on compounds or chemical classes of interest. MCnebula has many potential applications, including metabolite identification, classification biomarker tracking, drug discovery, and chemical change exploration.

[0120] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0121] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for visually classifying compounds, characterized in that, include: Obtain the raw mass spectrometry data of the compound; The raw mass spectrometry data were preprocessed to obtain compound information; The compound information includes the compound's molecular formula, molecular formula score, probability of the compound belonging to a category, and probability of the compound's structure. The molecular formula with the highest score among the stated molecular formulas is selected as the optimal molecular formula. The compound structure with the highest probability of selecting the optimal molecular formula is used as the structure dataset. The categories of compounds with the optimal molecular formula are filtered based on priority parameters, posterior probabilities, set thresholds, and the probability of the category to which the compound belongs, to obtain a category dataset; The category dataset is filtered according to set conditions to obtain a filtered category dataset; the set conditions include the position of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores; The filtered category dataset is clustered to obtain multiple clusters; A molecular network is generated based on the multiple clusters described above; The structural dataset is mapped to the molecular network; the molecular network is used to visualize the categories and structures of the compounds; the points of the molecular network include the categories and structures of the compounds; the edges of the molecular network are determined by the secondary fragment similarity of the raw mass spectrometry data of different compounds.

2. The compound visualization classification method according to claim 1, characterized in that, The preprocessing of the raw mass spectrometry data to obtain compound information specifically includes: The raw mass spectrometry data is then converted to a format to obtain Extensible Markup Language (EXPLAIN). The extensible markup language was used for feature detection using MZmine2 and for analysis using SIRIUS to obtain compound information.

3. The compound visualization classification method according to claim 1, characterized in that, The process of filtering the categories of compounds with the optimal molecular formula based on priority parameters, posterior probabilities, set thresholds, and the probability of the compound's category to obtain a category dataset specifically includes: Based on the priority parameters, the categories of compounds with the optimal molecular formula are initially screened to obtain preliminary screening results. The preliminary screening results are then subjected to a secondary screening based on the posterior probability, the set threshold, and the probability of the compound belonging to a certain category, to obtain a category dataset.

4. The compound visualization classification method according to claim 1, characterized in that, The filtering of the category dataset according to the set conditions to obtain the filtered category dataset specifically includes: Delete the categories representing the positions of the chemical functional groups from the category dataset to obtain the first-level filtering result; Delete the categories with the largest number of category features and the categories with the smallest number of category features from the first-level filtering results to obtain the second-level filtering results; The third-level filtering result is obtained by deleting categories that contain all the same features in the second-level filtering result; Calculate the similarity score for each pair of categories; Remove categories from the three-level filtering results whose similarity scores are less than the minimum reach rate to obtain the filtered category dataset.

5. A compound visualization classification system, characterized in that, include: The acquisition module is used to acquire the raw mass spectrometry data of the compound; The preprocessing module is used to preprocess the raw mass spectrometry data to obtain compound information; The compound information includes the compound's molecular formula, molecular formula score, probability of the compound belonging to a category, and probability of the compound's structure. The optimal molecular formula determination module is used to select the molecular formula with the highest molecular formula score among the molecular formulas of the compound as the optimal molecular formula. The structure dataset determination module is used to select the compound structure with the highest probability of the optimal molecular formula as the structure dataset. The filtering module is used to filter the category of the compound with the best molecular formula according to the priority parameter, posterior probability, set threshold and the probability of the category of the compound, so as to obtain the category dataset; The filtering module is used to filter the category dataset according to set conditions to obtain a filtered category dataset; the set conditions include the position of chemical functional groups, the maximum number of category features, the minimum number of category features, identical features, and similarity scores; The clustering module is used to cluster the filtered category dataset to obtain multiple clusters; A generation module is used to generate a molecular network based on the multiple clusters; The structural dataset is mapped to the molecular network; the molecular network is used to visualize the categories and structures of the compounds; the points of the molecular network include the categories and structures of the compounds; the edges of the molecular network are determined by the secondary fragment similarity of the raw mass spectrometry data of different compounds.

6. The compound visualization classification system according to claim 5, characterized in that, The preprocessing module specifically includes: The format conversion unit is used to convert the original mass spectrometry data into a format extensible markup language. The feature detection and analysis unit is used to perform feature detection on the extensible markup language using MZmine2 and analysis using SIRIUS to obtain compound information.

7. The compound visualization classification system according to claim 5, characterized in that, The filtering module specifically includes: The preliminary screening unit is used to perform preliminary screening of the category of the compound with the optimal molecular formula according to the priority parameter, and to obtain preliminary screening results; A secondary screening unit is used to perform secondary screening on the preliminary screening results based on the posterior probability, the set threshold, and the probability of the compound belonging to the category, to obtain a category dataset.

8. The compound visualization classification system according to claim 5, characterized in that, The filtering module specifically includes: A first-level filtering unit is used to delete categories representing the positions of the chemical functional groups in the category dataset to obtain the first-level filtering result; The secondary filtering unit is used to delete the categories with the largest number of category features and the categories with the smallest number of category features from the primary filtering result, thereby obtaining the secondary filtering result; The third-level filtering unit is used to delete categories that contain all the same features in the second-level filtering results, thus obtaining the third-level filtering results; The calculation unit is used to calculate the similarity score between each pair of categories; The filter category dataset determination unit is used to delete categories with similarity scores less than the minimum reach rate from the three-level filtering results, thereby obtaining the filter category dataset.