Structure-driven mass spectrometry feature mining method and device
Patent Information
- Application Number
- CN202611062505.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-08-18
AI Technical Summary
第一,数据处理流程高度碎片化,工具协同性差
(1)全流程一体化,无需多软件切换本发明将数据读取、预处理、结构挖掘、统计分析、分子网络、可视化、结果导出全部集成于同一平台,形成闭环工作流,避免格式转换与信息丢失,大幅提升处理效率与操作便捷性;
Smart Images

Figure CN122594680A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mass spectrometry data analysis technology, and in particular to a structure-driven mass spectrometry feature mining method and apparatus. Background Technology
[0002] High-resolution mass spectrometry (HRMS), with its advantages of ultra-high mass accuracy, high resolution, and high-precision performance, has become a core technology for non-target screening in fields such as environmental monitoring, metabolomics, exocomics, food and drug safety, and bioanalysis. Non-target screening can comprehensively capture and qualitatively analyze unknown compounds, transformation products, novel pollutants, and abnormal metabolites in complex systems, providing crucial technical support for addressing unknown hazardous chemicals and achieving in-depth analysis and scientific discovery. With the continuous improvement of mass spectrometry instrument performance and the increasing data acquisition throughput, a single analysis can generate hundreds of thousands to millions of mass spectrometry signals, resulting in an exponential growth in data volume. How to efficiently, accurately, and automatically extract structurally significant features from massive, noisy, and redundant raw mass spectrometry signals has become a core bottleneck restricting the practical application, standardization, and high-throughput analysis of non-target screening technology.
[0003] Current mass spectrometry feature mining and data analysis techniques still suffer from many unresolved technical shortcomings, specifically in the following aspects: First, the data processing workflow is highly fragmented, and the tools lack interoperability. Existing mass spectrometry data analysis tools have limited functionality and scattered modules. Key steps such as feature extraction, peak alignment, deduplication and noise reduction, structure mining, statistical analysis, and visualization are distributed across different software platforms. Users need to repeatedly switch between multiple tools, convert formats, and reset parameters, which is cumbersome, time-consuming, and labor-intensive, easily introducing human error and making it difficult to achieve fully automated processing. At the same time, the data formats, quality thresholds, and algorithm logic of different tools are incompatible, leading to information loss and feature misalignment during data transmission, seriously affecting the stability and reliability of the analysis results.
[0004] Second, feature mining relies primarily on statistical screening, lacking sufficient structure-driven capabilities. Traditional non-target analysis often uses statistical indicators such as signal intensity, signal-to-noise ratio, and peak shape as the basis for feature selection, lacking in-depth mining from a chemical structure perspective. For structural information with clear chemical significance, such as homologue series, characteristic isotope markers, characteristic fragment ions, neutral loss, and mass defect patterns, existing tools cannot achieve systematic identification and targeted mining. Especially in the screening of new environmental pollutants, a large number of halogenated compounds, perfluorinated and polyfluoroalkyl substances, drugs and their transformation products, and artificial secondary pollutants all possess typical structural features, which traditional tools struggle to effectively capture. This results in the omission of many high-value features, low annotation rates for unknowns, and high false positive rates, failing to meet the needs of in-depth analysis of complex samples.
[0005] Third, inconsistent analysis processes, opaque parameters, and difficulty in reproducing results are problematic. Existing software often employs closed algorithms, with undisclosed parameter setting logic, non-standardized workflows, and untraceable analysis processes. This leads to significant differences in results obtained by different users, different analysis batches, and under different equipment environments, failing to meet the requirements for scientific research comparisons, laboratory certification, and standardized testing. In fields such as environmental monitoring and public health safety, where data reliability is paramount, the lack of reproducibility directly undermines the persuasiveness of analytical conclusions, limiting the widespread application of non-target screening technologies in practical supervision and scientific research.
[0006] Fourth, the lack of advanced analytical capabilities makes it difficult to meet the complex analytical needs of various scenarios. For experimental designs such as grouped controls, time series, and dose-response experiments, existing tools lack comprehensive capabilities for differential feature extraction, trend judgment, cluster analysis, and molecular association network construction. Advanced algorithms such as molecular networks, principal component analysis, feature clustering, and time series trend testing either exist only in specialized academic tools and are difficult to popularize, or cannot be seamlessly integrated with the front-end feature mining process, preventing users from comprehensively interpreting data from multiple perspectives such as structural associations, sample differences, and dynamic changes.
[0007] Fifth, visualization and interactivity are weak, and the learning curve is steep. Most professional mass spectrometry data analysis tools rely on command-line operations or complex parameter panels, lacking graphical, workflow-based, and drag-and-drop interactive interfaces, making them difficult for ordinary experimental personnel to master quickly. Furthermore, the visualization of results is limited, failing to present feature distributions, structural relationships, and inter-group differences in intuitive ways such as network diagrams, heatmaps, trend charts, score plots, and load diagrams. This hinders the rapid identification of key features, the identification of suspicious substances, and the interpretation of biological or environmental significance.
[0008] In summary, existing non-target mass spectrometry data analysis technologies have significant shortcomings in terms of workflow integration, structure mining depth, result reproducibility, scenario adaptability, and ease of use. In particular, there is a lack of a structure-driven mass spectrometry feature mining platform that integrates preprocessing, structure mining, statistical analysis, molecular networks, visualization, and report output.
[0009] Therefore, developing a mass spectrometry feature mining method and device that is fully automated, algorithm-integrated, workflow-standardized, and graphically interfaced is of great practical significance and application value for promoting the development of non-target screening technology and improving the ability to analyze chemical substances in complex systems. Summary of the Invention
[0010] This invention provides a structure-driven mass spectrometry feature mining method and apparatus, which realizes one-stop fully automated processing from raw high-resolution mass spectrometry data to feature screening, structure analysis, statistical analysis, correlation mining, visualization and standardized output.
[0011] The technical solution of the present invention is as follows: A structure-driven mass spectrometry feature mining method includes the following steps: (1) Obtain the high-resolution mass spectrometry raw data of the sample to be tested and perform preprocessing to generate a standardized feature set; (2) Based on chemical structure clues, the feature set is targeted to obtain multi-dimensional structural features. The targeted mining includes suspected target screening, mass defect analysis, homologue detection, feature isotope pattern screening, and feature fragment ion matching with neutral loss. (3) Perform statistical analysis on the feature set to extract inter-group difference features, temporal variation features and similarity clustering features; (4) Based on the fragment information of secondary mass spectrometry, the similarity between features is calculated, and a molecular network is constructed with features as nodes and similarity as edges to realize the association mining of structurally similar compounds and the auxiliary annotation of unknowns; (5) Integrate all mining and analysis results, display them in a visual form, and export them in a common format.
[0012] The method of this invention achieves a one-stop, fully automated processing from raw high-resolution mass spectrometry data to feature screening, structure analysis, statistical analysis, correlation mining, visualization, and standardized output. Driven by chemical structure clues, this method integrates multiple structure mining algorithms and statistical analysis models, significantly improving the identification efficiency and annotation accuracy of unknown compounds, novel pollutants, and characteristic metabolites in complex samples. Simultaneously, it ensures a solidified analytical workflow, configurable parameters, and reproducible results, making it suitable for non-target screening scenarios in multiple fields such as environmental monitoring, metabolomics, exocomics, and food and drug safety.
[0013] In step (1), the preprocessing includes parsing, peak extraction, noise filtering, feature deduplication and standardization of the original high-resolution mass spectrometry data.
[0014] Noise filtering is performed by removing noise signals with intensity below the noise threshold; features from multiple batches and samples are merged and deduplicated according to the set mass-to-charge ratio tolerance and retention time tolerance to eliminate redundant features.
[0015] The feature set includes mass-to-charge ratio, retention time, signal intensity, and secondary mass spectrometry fragment information.
[0016] Preferably, the various directional mining methods in step (2) can be executed independently in parallel or in series.
[0017] A list of suspicious compounds is obtained through suspected target screening, a list of homologue series is obtained through mass defect analysis and homologue detection, and isotopic characteristic clusters are obtained through characteristic isotopic pattern screening.
[0018] The suspected target screening includes: supporting the import of compound lists from local or public databases, matching them based on precise quality and set quality tolerance, obtaining suspicious compounds, and marking their source libraries.
[0019] The mass defect analysis described herein is based on the Kendrick mass defect algorithm, which maps the exact mass to the nominal mass, so that compounds with the same repeating unit exhibit the same mass defect, and identifies a series of compounds with specific elemental composition, functional groups or framework structure.
[0020] The aforementioned homologue detection includes: automatically identifying homologues based on a custom repeating structural unit, according to the mass difference multiple relationship and retention time variation, and setting a minimum member number threshold to exclude false positives.
[0021] The characteristic isotope pattern screening includes: matching and filtering characteristics based on the isotopic mass difference and theoretical abundance ratio of characteristic elements to locate compounds of characteristic elements.
[0022] The matching of characteristic fragment ions and neutral loss includes: performing batch matching of preset characteristic fragment ions or characteristic neutral loss based on secondary mass spectrometry information to identify compounds containing specific functional groups or structural frameworks.
[0023] The preferred step (3) includes: performing case comparison analysis, trend analysis, and cluster analysis on the feature set after structure mining, extracting inter-group difference features, time-series change features, and similar cluster features to form a difference feature set.
[0024] The case-control analysis includes: based on grouping information, using principal component analysis to visualize sample groupings and extract key differential features that significantly contribute to inter-group separation; including: The original feature intensity matrix is centered and standardized. ; in, This is the standardized matrix; This is the original feature intensity matrix, with rows representing samples and columns representing features; The mean of each feature column; The standard deviation of each feature column; The standardized matrix is projected onto the principal component axis to obtain the sample score. ; The importance of features is calculated using the magnitude of the load value vector. The top N features with the highest contribution were extracted as key difference features between groups.
[0025] The trend analysis includes: using the Mann-Kendall trend test and / or quadratic regression fitting method to identify features with significant upward, downward or nonlinear change patterns based on the features after structural mining.
[0026] The clustering analysis includes grouping features based on Euclidean distance, Manhattan distance, cosine distance or Gower distance, using at least one of the algorithms such as KMeans, DBSCAN, HCA to classify similar features.
[0027] Preferably, the molecular network construction in step (4) includes: Normalize the intensity of mass spectrometry peaks; Match the experimental mass spectrometry peaks within the set mass deviation tolerance; Cosine similarity is used as the spectral similarity, and features are used as nodes. Edges are established between nodes based on the spectral similarity between two features to form an undirected weighted graph network. : ; in, For a set of feature nodes; For similarity-based edge connections; For nodes and nodes The edges connecting them.
[0028] The constructed molecular network can visually demonstrate the structural similarity relationships between substances, and can be used for the identification of unknown substances, the discovery of homologs, and the clustering of categories.
[0029] Preferably, in step (5), the visualization forms include one or more of the following: network diagram, heat map, trend curve, PCA score map, PCA load map, and feature list; the output content includes standardized feature table, list of suspected compounds, list of homologous series, isotope feature clusters, differential feature set, molecular network diagram, and statistical analysis results, and supports export in CSV and image formats.
[0030] Further optimized, the output includes text results and visualization charts; the text results include standardized feature tables, lists of suspected compounds, lists of homologues, isotope feature clusters, and sets of differential features; the visualization charts include PCA score plots, PCA load plots, trend curves, clustering heatmaps, dendrograms, and molecular network diagrams; and support for export in CSV and image formats.
[0031] Based on the same inventive concept, the present invention also provides a structure-driven mass spectrometry feature mining device, comprising: The data preprocessing module acquires the high-resolution mass spectrometry raw data of the sample to be tested and preprocesses it to generate a standardized feature set. The feature mining module performs targeted mining of feature sets based on chemical structure clues to obtain multi-dimensional structural annotation information, including a suspected target screening unit, a mass defect analysis unit, a homologue detection unit, a feature isotope pattern screening unit, and a feature fragment ion and neutral loss matching unit. The statistical analysis module performs statistical analysis on the feature set, extracting inter-group difference features, time-series variation features, and similarity clustering features. The molecular network module calculates the similarity between features based on secondary mass spectrometry fragment information, constructs a molecular network with features as nodes and similarity as edges, and realizes the association mining of structurally similar compounds and the auxiliary annotation of unknowns. The visualization and output module integrates all the mining and analysis results, presents them in a visual format, and exports them in a common format.
[0032] Preferably, the visualization and output module provides a graphical interactive interface that supports drag-and-drop workflow construction, real-time parameter configuration, process saving and reuse, and realizes fully automated execution of the entire process.
[0033] The device adopts a graphical, modular, drag-and-drop workflow architecture. Each module can be freely combined, its order can be adjusted, and its parameters can be configured. It supports a flexible analysis mode from independent operation of a single module to one-click execution of the entire process.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The whole process is integrated without switching between multiple software. This invention integrates data reading, preprocessing, structure mining, statistical analysis, molecular network, visualization and result export into the same platform to form a closed-loop workflow, avoid format conversion and information loss, and greatly improve processing efficiency and ease of operation. (2) With structure-driven as the core, the depth of mining is significantly improved, breaking through the limitations of traditional statistical screening. With chemical structural clues such as mass defect, homologues, isotopes, feature fragments, and molecular networks as the core, high-value features are captured in a targeted manner. It is especially suitable for the deep identification of new pollutants, transformation products, and unknown metabolites, significantly improving the annotation rate and reducing false positives. (3) Standardized process and highly reproducible results: Unified parameter system, open algorithm logic, support for workflow saving and reuse, full traceability of the analysis process, stable and consistent results can be obtained by different users, different batches and different laboratories, meeting the needs of scientific research comparison, laboratory certification and standardized testing. (4) It has complete advanced analysis functions and is suitable for multiple research scenarios. It has built-in group control, time series trend, cluster analysis, molecular network and other complete advanced analysis capabilities, supports a variety of experimental designs and application scenarios such as environmental monitoring, metabolomics, exomics, food and drug safety, and has strong scalability. (5) Graphical interaction with low barrier to entry: It adopts a visual drag-and-drop interface, eliminating the need for command-line operation. The parameter settings are intuitive and easy to understand, allowing ordinary experimental personnel to get started quickly and significantly reducing the professional threshold for mass spectrometry data analysis. (6) Standardized output facilitates subsequent verification and publication. Supports export of general formats, with clear results, standardized charts and graphs, and complete feature information. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating the structure-driven mass spectrometry feature mining method proposed in this invention. Figure 2 This is a schematic diagram of the structure-driven mass spectrometry feature mining device proposed in this invention. Detailed Implementation
[0036] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not limit it in any way.
[0037] To overcome the numerous shortcomings of existing non-target mass spectrometry data analysis techniques, such as fragmented workflows, weak structure mining capabilities, non-reproducible results, lack of advanced analytical functions, and high barriers to entry, this invention provides a structure-driven mass spectrometry feature mining method. This method achieves a one-stop, fully automated process from raw high-resolution mass spectrometry data to feature screening, structure analysis, statistical analysis, correlation mining, visualization, and standardized output. Driven by chemical structural clues, this invention integrates multiple structure mining algorithms and statistical analysis models, significantly improving the identification efficiency and annotation accuracy of unknown compounds, novel pollutants, and characteristic metabolites in complex samples. Simultaneously, it ensures a solidified analytical workflow, configurable parameters, and reproducible results, making it suitable for non-target screening scenarios in various fields such as environmental monitoring, metabolomics, exocomics, and food and drug safety.
[0038] like Figure 1 As shown, a structure-driven mass spectrometry feature mining method includes the following steps: Step S1: Mass Spectrometry Data Acquisition and Preprocessing Acquire high-resolution mass spectrometry raw data of the sample to be tested, and perform analysis, peak extraction, noise filtering, feature deduplication, and standardization on the raw data to generate a mass-to-charge ratio (m / z), retention time (RT), signal intensity, and secondary mass spectrometry (MS). 2 A unified feature table for fragment information.
[0039] Specifically, this includes: supporting input in the mzML universal mass spectrometry format; removing low-intensity noise signals based on signal-to-noise ratio (SNR) thresholds; and extracting the mass-to-charge ratio, retention time, peak intensity, isotope distribution, and secondary mass spectrometry fragment information of effective peaks. The raw mass spectrometry data consists of continuous ion current signals. The true compound characteristics are represented by signal peaks at specific SNR and retention times, with the remainder being noise, baseline drift, and redundant signals. Invalid signals can be removed using SNR thresholds; and by using dual matching of SNR and retention time tolerances, features from the same compound in multiple samples can be merged and deduplicated, ensuring feature uniqueness. Features from multiple batches and samples are merged and deduplicated according to the set m / z and retention time tolerances, eliminating redundant features and outputting a standardized feature set without duplication that can be used for subsequent mining.
[0040] Step S2: Structure-driven multi-dimensional feature mining Based on chemical structure clues, standardized features are mined hierarchically, in a targeted and multi-dimensional manner, including: suspected target screening, Kendrick mass defect analysis, homologue detection, characteristic isotope pattern screening, and matching of characteristic fragment ions with neutral loss.
[0041] Suspected target screening: Supports importing compound lists from local or public databases (PubChem, CompTox, MassBank, NIST, etc.), matching them based on precise quality and set quality tolerance, automatically locking suspicious compounds and marking their source libraries; Mass Defect Analysis: Based on the Kendrick mass defect algorithm—the idea of "repeating unit mass normalization"—the precise mass is mapped to the nominal mass, so that compounds with the same repeating units exhibit the same mass defect. This identifies compound series with specific elemental composition, functional groups, or framework structures, thereby quickly identifying homologues and characteristic element series. The calculation formula is as follows: ; ; Nominal mass: sum of integer relative atomic masses of elements; Precise mass: precise molecular weight calculated based on the precise atomic weights of isotopes; Repeating unit: repeating structural unit of homologous series.
[0042] Homologue detection: Based on custom repeating structural units (such as CH2, CF2, C2H4O, etc.), homologue series are automatically identified according to the relationship between mass difference multiples and retention time patterns, and a minimum member number threshold is set to exclude false positives. Homologue mass difference determination: ; in, The number of repeating units ( ); For accurate quality of repeating units; Isotope pattern screening: Based on the isotopic mass difference and theoretical abundance ratio of characteristic elements such as Cl and Br, the features are matched and filtered to quickly locate compounds containing characteristic elements such as halogens. Feature fragment / neutral loss matching: Based on secondary mass spectrometry information, batch matching is performed on preset feature fragment ions or feature neutral loss to identify compounds containing specific functional groups or structural frameworks.
[0043] Step S3: Statistical Analysis and Differential / Trend Feature Mining Statistical analysis is performed on the feature set after structure mining to extract inter-group difference features, temporal change features, and similar clustering features, including case comparison analysis, trend analysis, and cluster analysis.
[0044] Case study analysis: Based on grouping information, principal component analysis (PCA) was used to visualize sample groupings and extract significant differential features contributing to the separation between groups. The specific principles and calculation methods are as follows: The original feature intensity matrix is centered and standardized. ; in, This is the standardized matrix (mean = 0, variance = 1). This is the original feature intensity matrix (rows = samples, columns = features); The mean of each feature column; The standard deviation of each feature column; The standardized matrix is projected onto the principal component axes to obtain the sample scores. : ; in, The characteristic load value represents the first... The feature is related to the first The contribution of each principal component Indicates the first element in the standardized matrix. The intensity value of each feature, Indicates the total number of features.
[0045] The importance of features is calculated using the magnitude of the load value vector. : ; Indicates the first feature to the second feature. The load values of each principal component. Indicates the second feature to the first The load values of each principal component.
[0046] The proportion of principal components that explain the original data : ; in, Indicates the first The eigenvalues corresponding to each principal component Indicates the first The eigenvalues corresponding to each principal component are automatically extracted based on the above formula. These features serve as key differences between groups.
[0047] Trend analysis: For samples such as time series and dose gradients, methods such as the Mann-Kendall trend test and quadratic regression fitting are used to identify features with significant upward, downward, or non-linear change patterns. The specific principles and calculation methods are as follows: Calculation of the Mann-Kendall statistic: ; in, , Represents the first time series / dose gradient. , No. Observations of a sample (sequence data sorted chronologically or by dose from smallest to largest). Kendall's rank correlation coefficient represents the total number of observations (total sample size) for the entire time series / dose gradient sample. ; in, Indicates the strength and direction of the trend. It is on an upward trend. It is in a downward trend.
[0048] Significance determination: when When the trend is determined to be statistically significant.
[0049] Theil-Sen trend slope: ; in, , Represents the first time series / dose gradient. , No. Observations for each sample (sequence data sorted chronologically or by dose from smallest to largest). Trend fitting equation: ; in, This represents the dependent variable (the observed value of the index to be fitted), which is the numerical value of the characteristic index that changes with the independent variable. express Theil-SenThe trend slope calculated by the method represents Follow The average magnitude of change per unit change; The independent variable can be a time series time sequence, a dose gradient, or a gradient variable such as the dose level of a dose gradient. This represents the intercept of the fitted line. , This represents the median of all observed values for the dependent variable. This represents the median of all values for the independent variable.
[0050] Quadratic fitting model: ; in, This represents the fitted predicted value of the dependent variable, which is the index estimate calculated by the quadratic curve model based on the independent variables, and is distinct from the actual observed value. , The representation is consistent with the linear model, and the independent variables are gradient variables such as time and dose gradient. The coefficient of the quadratic term determines the direction of the parabola's opening and the degree of its curvature. It represents the coefficient of the linear term and controls the slope of the curve; The y-intercept of the model is represented by... The predicted value at that time. Nonlinear trend determination: when At that time, the characteristics show a peak trend; when At that time, the feature shows a valley trend.
[0051] Goodness-of-fit calculation: ; in, The coefficient of determination (goodness of fit) ranges from 0 to 1. A value closer to 1 indicates a better fit of the model to the data. Indicates the first The actual observed values of a sample (the measured data of the dependent variable). No. The model fit prediction value for each sample (the estimated value calculated through the regression equation). When The fitting effect was determined to be good.
[0052] The above formula automatically identifies significant nonlinear peak / valley trends, enabling high-throughput mining of dynamic features in non-target mass spectrometry data.
[0053] The above formula automatically identifies features with significant upward or downward trends, making it suitable for feature mining of dynamic samples in non-target mass spectrometry screening.
[0054] Cluster analysis: Based on metrics such as Euclidean distance, Manhattan distance, cosine distance, and Gower distance, algorithms such as KMeans, DBSCAN, and HCA are used to group features, achieving automatic classification of similar features. The specific principles and calculation methods are as follows: (1) Distance metric European distance: ; Manhattan distance: ; Cosine distance: ; Gower distance: ; in, 、 These represent the numbers of the two samples (objects) whose distance is to be calculated; Representative sample With sample The distance between them; Indicates the total number of features (dimensions); Indicates the index of the feature, and iterates through all features. No. The sample at the th The values that can be taken on each feature No. The sample at the th The values that can be taken on each feature; No. The range of a feature (the maximum value of the feature minus the minimum value across all samples).
[0055] (2) Clustering algorithm KMeans clustering: Iterative clustering is performed with the goal of minimizing the within-class squared error: ; in, This indicates the total number of preset clusters. Indicates the cluster number, traversing from 1 to... All clusters, Indicates the first A cluster, Represents a single sample vector in the dataset. Indicates the first Cluster The cluster center (mean vector).
[0056] DBSCAN density clustering: Based on neighborhood radius It identifies high-density connected regions by minimum number of samples, automatically determines the number of clusters, and marks noise points.
[0057] Hierarchical Clustering (HCA): Cluster merging was performed using the Ward variance minimization criterion. ; in, Cluster and cluster The Ward distance between the two clusters is used to measure the cost of merging them. This represents two different clusters whose distance is to be calculated and which are to be merged. Cluster The cluster center (mean vector); Cluster The cluster center (mean vector).
[0058] Feature-level clustering is achieved using a dendrogram, and the classification results are output according to the set number of clusters.
[0059] The above methods enable automatic grouping and similar pattern mining of mass spectrometry features, providing efficient feature classification capabilities for non-target screening.
[0060] Step S4: Molecular network construction and structural association analysis Based on cosine similarity calculations of features using fragment information from secondary mass spectrometry, a molecular network is constructed with features as nodes and similarity as edges. This network automatically identifies groups of structurally similar compounds, transformation product clusters, and homologous families, enabling structural association annotation and visualization of unknown features. The specific principles and calculation methods are as follows: (1) Mass spectrometry data preprocessing Normalize the mass spectrum peak intensities, scaling the intensities to the range of 0-100: ; in, This represents the normalized mass spectrum peak intensity, with values scaled to the range of 0 to 100. This represents the original measured intensity value of a target mass spectrum peak; This represents the maximum original intensity value among all mass spectral peaks in the same mass spectrum.
[0061] (2) Mass spectrometry peak matching: within the set mass deviation tolerance, the peaks of the experimental mass spectrometer and the reference mass spectrometer are matched one by one.
[0062] (3) Calculation of spectral similarity Cosine similarity is used to measure spectral similarity: ; in, This represents the cosine similarity between mass spectra A and B, with a value range of [-1, 1]. In the context of mass spectrometry, it is usually [0, 1]. The closer it is to 1, the higher the similarity between the two spectra. It is the first in graph A The intensity values of each matching peak. It is the first in graph B The intensity values of each matching peak; This indicates the total number of matching mass spectrum peaks obtained after the two spectra are matched.
[0063] A valid similarity is determined only when the number of matching peaks is greater than the minimum number of peaks.
[0064] (4) Molecular network construction When the spectral similarity between two features is greater than a set threshold At that time, establish edges between nodes: ; This ultimately forms an undirected weighted graph network: ; in, For the set of feature nodes, This is the set of edges connected by similarity.
[0065] The molecular networks constructed using the above methods can visually demonstrate the structural similarity relationships between substances, and can be used for the identification of unknown substances, the discovery of homologs, and the clustering of categories.
[0066] Step S5: Results integration, visualization, and standardized output All mining and analysis results are integrated and visualized in the form of network diagrams, heatmaps, trend curves, PCA score plots, PCA loading plots, and feature lists. Outputs include: standardized feature tables, lists of suspected compounds, lists of homologues, isotope feature clusters, differential feature sets, molecular network diagrams, and statistical analysis results. It supports exporting in common formats such as CSV and images for easy use in subsequent analysis.
[0067] like Figure 2 As shown, based on the same inventive concept, this invention also provides a structure-driven mass spectrometry feature mining device, comprising the following five functional modules: 1. Data preprocessing module, used to read and parse raw mass spectrometry data; complete peak extraction, noise filtering, feature extraction, multi-sample feature alignment and redundancy removal; output standardized and deduplicated unified feature table, providing high-quality input for subsequent structure mining and statistical analysis.
[0068] 2. Structure mining module, used to execute a complete set of structure-driven mining algorithms, including: suspected target screening unit, mass defect analysis unit, homologue detection unit, isotope pattern screening unit, feature fragment and neutral loss matching unit, to realize targeted feature mining based on chemical clues.
[0069] 3. Statistical analysis module, used for advanced statistical analysis of feature sets, including: case comparison unit, trend analysis unit, cluster analysis unit, automatically extracting inter-group differences, time series patterns and similar feature groups.
[0070] 4. Molecular network module, used for secondary mass spectrometry similarity calculation, network construction, node layout and visualization, to realize the association mining of structurally similar compounds and auxiliary annotation of unknown substances.
[0071] 5. Visualization and output module, used for chart creation, result integration, report generation and data export; provides a graphical interactive interface, supports drag-and-drop workflow building, real-time parameter configuration, process saving and reuse, and realizes full-process automated execution.
[0072] Furthermore, the device adopts a graphical, modular, drag-and-drop workflow architecture, where modules can be freely combined, their order can be adjusted, and their parameters can be configured, supporting a flexible analysis mode that ranges from independent operation of a single module to one-click execution of the entire process.
[0073] Example 1 This embodiment provides a structure-driven mass spectrometry feature mining method. The overall process includes: mass spectrometry data reading and preprocessing, structure-driven multi-dimensional mining, statistical and trend analysis, molecular network construction, visualization and result output.
[0074] The method of this invention supports the following input data: Raw high-resolution mass spectrometry data, in mzML format; The extracted feature list is in CSV format and includes m / z, retention time RT, intensity, and MS2 information for secondary mass spectrometry.
[0075] This embodiment uses environmental sample mass spectrometry data acquired in DDA mode, with ionization mode supporting positive ion [M+H]. + With negative ion mode [MH] - .
[0076] Step S1: Mass Spectrometry Data Preprocessing (1) Data analysis and peak extraction Use PyMS, PyOpenMS, or pymzML tools to read mzML files and perform peak identification on the raw mass spectrometry data, including: The noise threshold is set to 3000 by default, and noise signals with an intensity lower than this threshold are removed. Feature information is extracted, including m / z, retention time RT, signal intensity, isotope distribution, and secondary mass spectrometry fragments.
[0077] (2) Feature deduplication and standardization Multiple sample features are merged and deduplicated to eliminate redundancy: the m / z tolerance is 5 ppm and the retention time tolerance is 30 seconds; the deduplication rule is to retain only the feature with the highest intensity as the representative feature within the above tolerance range.
[0078] (3) Output results Generate a preprocessed standardized feature table containing the following fields: feature ID, source file name, m / z, RT, RT start, RT end, intensity, MS2 mass-to-charge ratio, MS2 intensity, and precursor m / z.
[0079] Step S2: Structure-driven multi-dimensional feature mining This step includes 5 core mining units, all of which can be run independently or in series.
[0080] (1) Screening of suspected targets Import reference libraries, supporting PubChem, CompTox, MassBank, NIST, or local CSV lists; the mass matching mode is switchable between positive and negative ions; the default mass tolerance is 0.002 Da; the matching rule is: if the deviation of the feature m / z from the precise mass of the reference library is ≤ the set mass tolerance, it is considered a match; output a list of matched suspicious compounds, automatically annotating the source library, CAS number, molecular formula, and matching mass deviation.
[0081] (2) Kendrick quality loss analysis Input repeating units (such as CH2, CF2, C2H4O) to identify series containing specific elements and homologues.
[0082] (3) Detection of homologues Customize the quality of repeating units, such as CH2 at 14.01565 Da and CF2 at 49.99681 Da; the matching tolerance is 5 ppm; the minimum number of series members is ≥ 3; the judgment rule is that the quality difference between features is an integer multiple of the repeating unit, and the retention time increases / decreases regularly; the output includes the homologue series number, repeating unit, number of members, and m / z list.
[0083] (4) Isotope pattern screening The target elements include characteristic isotopes such as Cl and Br; the isotopic mass difference is 1.997 Da for Cl and 1.998 Da for Br; the intensity ratio threshold is M+2, and the intensity target is 0.33 (Cl); the allowable range for intensity ratio error is ≤ 0.2, and the m / z tolerance is 0.001 Da; the output includes a list of features with isotopic labels and isotopic cluster information.
[0084] (5) Feature fragment matching with neutral loss Input method: Manually input fragment m / z or upload fragment list; matching tolerance is 10 ppm, and the judgment rule is: at least 2 characteristic fragments appear in the secondary mass spectrometer and the deviation is within the allowable range; the output content includes a list of compounds containing characteristic functional groups / skeleton structures.
[0085] Step S3: Statistical Analysis and Trend / Cluster Mining (1) Case Comparison Analysis (PCA) Input grouping information CSV (filename and group), use principal component analysis (PCA) algorithm, output the following: PCA score plot (sample grouping), PCA loading plot (feature contribution), and list of top 20 differential features.
[0086] (2) Trend Analysis Two modes are provided: one is the Mann-Kendall trend test, and the other is the significance level. The value is 0.05, indicating a significant increase / decrease in the output; the second method is quadratic regression analysis, with the fitting formula as follows: , The threshold is ≥ 0.6; The threshold value is ≤ 0.05, and the output exhibits significant nonlinear variation characteristics.
[0087] (3) Cluster analysis Supported distance metrics include Euclidean distance, Manhattan distance, cosine distance, and Gower distance; supported algorithms include KMeans, DBSCAN, and hierarchical clustering HCA; outputs clustering results, heatmaps, and dendrograms.
[0088] Step S4: Molecular Network Construction Input preprocessed features and secondary mass spectrometry information, and perform cosine similarity calculation; parameter settings: mass accuracy is 0.01 Da, minimum number of matching peaks is ≥ 2, and similarity threshold is ≥ 0.6; construct a network, with nodes representing mass spectrometry features and edges representing structural similarity between features, and finally output a molecular network diagram, structural similarity groups, and a list of associated features.
[0089] Step S5: Visualization and Result Output The method of this invention supports the following standardized outputs: The text results include a full feature table, a list of suspicious substances, a series of homologous substances, isotopic feature clusters, and a set of differential features; Visual charts, including PCA score plots, PCA load plots, trend curves, cluster heatmaps, dendrograms, and molecular network diagrams; Export formats are CSV (table) and PNG (image).
[0090] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A structure-driven mass spectrometry feature mining method, characterized in that, Includes the following steps: (1) Obtain the high-resolution mass spectrometry raw data of the sample to be tested and perform preprocessing to generate a standardized feature set; (2) Based on chemical structure clues, the feature set is targeted to obtain multi-dimensional structural features. The targeted mining includes suspected target screening, mass defect analysis, homologue detection, feature isotope pattern screening, and feature fragment ion matching with neutral loss. (3) Perform statistical analysis on the feature set to extract inter-group difference features, temporal variation features and similarity clustering features; (4) Based on the fragment information of secondary mass spectrometry, the similarity between features is calculated, and a molecular network is constructed with features as nodes and similarity as edges to realize the association mining of structurally similar compounds and the auxiliary annotation of unknowns; (5) Integrate all mining and analysis results, display them in a visual form, and export them in a common format.
2. The structure-driven mass spectrometry feature mining method according to claim 1, characterized in that, In step (1), the preprocessing includes parsing, peak extraction, noise filtering, feature deduplication and standardization of the original high-resolution mass spectrometry data; the feature set includes mass-to-charge ratio, retention time, signal intensity and secondary mass spectrometry fragment information.
3. The structure-driven mass spectrometry feature mining method according to claim 1, characterized in that, The various directional mining methods in step (2) can be executed independently in parallel or in series.
4. The structure-driven mass spectrometry feature mining method according to claim 1, characterized in that, Step (3) includes: performing case comparison analysis, trend analysis, and cluster analysis on the feature set after structure mining, and extracting inter-group difference features, time series change features, and similar cluster features.
5. The structure-driven mass spectrometry feature mining method according to claim 4, characterized in that, The case-control analysis includes: based on grouping information, using principal component analysis to visualize sample groupings and extract key difference features that significantly contribute to the separation between groups.
6. The structure-driven mass spectrometry feature mining method according to claim 4, characterized in that, The trend analysis includes: using the Mann-Kendall trend test and / or quadratic regression fitting method to identify features with significant upward, downward or nonlinear change patterns based on the features after structural mining.
7. The structure-driven mass spectrometry feature mining method according to claim 4, characterized in that, The clustering analysis includes grouping features based on Euclidean distance, Manhattan distance, cosine distance or Gower distance, using at least one of the algorithms such as KMeans, DBSCAN, HCA to classify similar features.
8. The structure-driven mass spectrometry feature mining method according to claim 1, characterized in that, In step (5), the visualization forms include one or more of the following: network diagram, heat map, trend curve, PCA score map, PCA load map, and feature list; the output includes standardized feature table, list of suspected compounds, list of homologues, isotope feature clusters, differential feature set, molecular network diagram, and statistical analysis results, and supports export in CSV and image formats.
9. A structure-driven mass spectrometry feature mining device, characterized in that, A method for performing the mass spectrometry feature mining method according to any one of claims 1-8, comprising: The data preprocessing module acquires the high-resolution mass spectrometry raw data of the sample to be tested and preprocesses it to generate a standardized feature set. The feature mining module performs targeted mining of feature sets based on chemical structure clues to obtain multi-dimensional structural annotation information, including a suspected target screening unit, a mass defect analysis unit, a homologue detection unit, a feature isotope pattern screening unit, and a feature fragment ion and neutral loss matching unit. The statistical analysis module performs statistical analysis on the feature set, extracting inter-group difference features, time-series variation features, and similarity clustering features. The molecular network module calculates the similarity between features based on secondary mass spectrometry fragment information, constructs a molecular network with features as nodes and similarity as edges, and realizes the association mining of structurally similar compounds and the auxiliary annotation of unknowns. The visualization and output module integrates all the mining and analysis results, presents them in a visual format, and exports them in a common format.
10. The structure-driven mass spectrometry feature mining device according to claim 9, characterized in that, The visualization and output module provides a graphical user interface that supports drag-and-drop workflow building, real-time parameter configuration, process saving and reuse, and enables fully automated execution of the entire process.