A data analysis system based on mass spectrometry detection platform

By building a dual-channel mass spectrometry detection model, combining attention mechanism and wavelet transformation and other technical means, the problems of noise and overlapping peaks in mass spectrometry data are solved, and the accuracy and sensitivity of mass spectrometry data analysis are improved.

CN119936168BActive Publication Date: 2025-08-26RELAIS (HANGZHOU) MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510423920.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-26
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

There are a large number of complex background noise and overlapping peaks in the mass spectrometry data, making it difficult to identify low-abundance molecular or isomer signals. The prior art has limitations in noise suppression and signal separation, which affects the accuracy and sensitivity of the analysis.

Method used

A two-channel mass spectrometry detection model is constructed, the first channel recognizes noise characteristics, and the second channel extracts the signal characteristics of chemical substances, and dynamic baseline correction and signal-to-noise ratio processing are performed in combination with attention mechanism, asymmetric least squares method and wavelet transformation to generate a target mass spectrometry data set.

Benefits of technology

Noise suppression, signal separation and low abundance detection of mass spectrometry data are realized, and the accuracy and sensitivity of mass spectrometry data analysis are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119936168B_ABST
    Figure CN119936168B_ABST
Patent Text Reader

Abstract

The present invention relates to a data analysis system based on a mass spectrometry detection platform, which relates to the field of data processing. A dual-channel mass spectrometry detection model is constructed, wherein a first channel is used to identify noise features in mass spectrometry data, and a second channel is used to extract chemical substance signal features in the mass spectrometry data, and the mass spectrometry data passes through the first and second channels in sequence; an attention mechanism is configured in the second channel to dynamically identify the chemical substance signal features, obtain an attention weight map, and generate a target mass spectrometry data set based on weighted processing of the attention weight map; an asymmetric least squares method is used to dynamically baseline correct the data in the target mass spectrometry data set, and signal-to-noise ratio processing is performed in combination with wavelet transform to obtain mass spectrometry detection results, thereby solving the technical problems of insufficient accuracy and sensitivity of mass spectrometry data analysis during mass spectrometry detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data analysis system based on a mass spectrometry detection platform. Background Art

[0002] As an efficient and sensitive analytical technology, mass spectrometry has been widely used in many fields such as biomedicine, environmental monitoring, and food safety. However, mass spectrometry data often contain a large amount of complex background noise and overlapping peaks. In particular, the signals of low-abundance molecules or isomers are often easily masked or misjudged due to their weak intensity and susceptibility to interference. This poses a huge challenge to the accurate analysis of mass spectrometry data. Traditional mass spectrometry data analysis methods have certain limitations when dealing with noise suppression, signal separation, and low-abundance detection problems. For example, although traditional filtering algorithms can remove noise to a certain extent, they often have the risk of over-suppression of low-abundance signals, resulting in missed detections. In addition, analysis methods for single-dimensional data are usually unable to effectively separate isomers, further limiting the accuracy and reliability of mass spectrometry data analysis. Summary of the Invention

[0003] The present invention aims to solve the technical problem of insufficient accuracy and sensitivity of mass spectrometry data analysis during mass spectrometry detection in the prior art by providing a data analysis system based on a mass spectrometry detection platform.

[0004] The technical solution of the present invention to solve the above technical problems is as follows:

[0005] The present invention provides a data analysis system based on a mass spectrometry detection platform, and the execution steps include: constructing a dual-channel mass spectrometry detection model, wherein the first channel is used to identify noise characteristics in mass spectrometry data, and the second channel is used to extract chemical substance signal characteristics in the mass spectrometry data, and the mass spectrometry data passes through the first and second channels in sequence; configuring an attention mechanism in the second channel to dynamically identify the chemical substance signal characteristics, obtain an attention weight map, and generate a target mass spectrometry data set based on weighted processing of the attention weight map; using an asymmetric least squares method to dynamically baseline correct the data in the target mass spectrometry data set, combining the wavelet transform to perform signal-to-noise ratio processing, and obtain mass spectrometry detection results.

[0006] The beneficial effects of the present invention are: by constructing a dual-channel convolutional neural network, introducing an attention mechanism, and combining dynamic baseline correction technology and wavelet transform, noise suppression, signal separation and low-abundance detection of mass spectrometry data are achieved, thereby improving the accuracy and sensitivity of mass spectrometry data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 This is a schematic flow chart of the execution steps of a data analysis system based on a mass spectrometry detection platform provided by the present invention.

[0008] Figure 2 This is a schematic flow chart of the execution steps of second channel feature extraction in a data analysis system based on a mass spectrometry detection platform provided by the present invention. DETAILED DESCRIPTION

[0009] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0010] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0011] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.

[0012] Example:

[0013] like Figure 1 As shown, the embodiment of the present invention provides a data analysis system based on a mass spectrometry detection platform, and the execution steps include:

[0014] S10: Construct a dual-channel mass spectrometry detection model, wherein the first channel is used to identify noise features in mass spectrometry data, and the second channel is used to extract chemical substance signal features in the mass spectrometry data, and the mass spectrometry data passes through the first and second channels in sequence.

[0015] S20: configuring an attention mechanism in the second channel, dynamically identifying the signal characteristics of the chemical substance, obtaining an attention weight map, and generating a target mass spectrometry data set based on weighted processing of the attention weight map.

[0016] S30: performing dynamic baseline correction on the data in the target mass spectrometry data set using an asymmetric least squares method, performing signal-to-noise ratio processing in combination with wavelet transformation, and obtaining a mass spectrometry detection result.

[0017] For example, when performing data analysis on a mass spectrometry detection platform, mass spectrometry data often contains a large amount of complex background noise and overlapping peaks, especially signals from low-abundance molecules or isomers. Background noise refers to irrelevant signals in the mass spectrum, which may originate from instrument electronic noise, sample matrix interference, environmental noise, etc. The presence of background noise can interfere with the true signal peaks, affecting the accuracy and reliability of mass spectrometry data. Overlapping peaks refer to the overlap of the signal peaks of two or more compounds in the mass spectrum due to similar masses or insufficient instrument resolution. Overlapping peaks can make it difficult to identify the signal peaks, affecting the accuracy of qualitative and quantitative analysis. Low-abundance molecules refer to molecules with low content in the sample, and their signal peaks may be relatively weak in the mass spectrum, or even be obscured by background noise. Isomers, on the other hand, are compounds with the same molecular formula but different structures. In a mass spectrum, the signal peaks of isomers may be difficult to distinguish due to their similar masses. Identifying isomers is important for understanding the structural diversity of compounds and analyzing components in complex samples. Therefore, this application avoids the above complex interference problems and improves the analysis accuracy and reliability of mass spectrometry detection data through technical means such as dual-channel processing, dynamic identification and weighted processing, baseline correction and signal-to-noise ratio optimization.

[0018] Specifically, raw mass spectrometry data is input into the model. This data typically contains a wealth of information, including the signal of the target chemical as well as interfering factors such as background noise. Before entering the dual-channel system, the data typically undergoes preprocessing steps such as smoothing, baseline correction, and normalization to reduce random errors and instrument noise in the data and improve the accuracy of subsequent analysis. The preprocessed mass spectrometry data first enters the first channel, the noise signature identification channel. This channel's primary task is to identify and isolate noise signatures from the data. This is typically achieved through a series of complex algorithms, such as wavelet transforms, principal component analysis (PCA), or machine learning algorithms (such as support vector machines (SVMs) and neural networks). These algorithms analyze the frequency content, statistical properties, or patterns of the data to effectively distinguish noise signals from target chemical signals. In the first channel, identified noise signatures are labeled and removed or attenuated from the raw data to minimize interference with subsequent analysis. After processing in the first channel, the noise-removed or attenuated mass spectrometry data enter the second channel, the chemical signature extraction channel. This channel aims to extract signatures associated with the target chemical. This also relies on advanced algorithms and models, such as feature selection algorithms, pattern recognition algorithms, or deep learning networks. These algorithms can identify specific patterns or features in the data that match the mass spectra of known chemicals. In the second channel, the extracted chemical signal features are further analyzed and processed for subsequent qualitative or quantitative analysis. In a preferred embodiment, the dual channels are constructed using a convolutional neural network (CNN), with each channel learning noise patterns and true signal features respectively. Both channels use synthetic datasets—real mass spectrometry data with simulated noise—to train the models, generating more diverse training samples. This allows the models to better cope with noise interference when faced with new, unknown mass spectrometry data, improving the accuracy and reliability of the analysis. Furthermore, the real mass spectrometry data with simulated noise can be used to optimize noise filtering algorithms, enabling more accurate identification and removal of noise components in the mass spectrometry data, improving the signal-to-noise ratio and clarity of the data, and achieving end-to-end noise filtering.

[0019] Furthermore, in the second channel, a convolutional neural network (CNN) or other deep learning model is used to extract chemical signal features from the mass spectrometry data. These features may include peak position, intensity, and shape, which are crucial for chemical identification and quantitative analysis. To further improve the accuracy and efficiency of feature extraction, an attention mechanism is implemented in the second channel. The attention mechanism is a deep learning technique that mimics the distribution of human visual attention. It dynamically focuses on important components of the input data and ignores irrelevant information. In this second channel, the model dynamically identifies chemical signal features in the mass spectrometry data. This is achieved by calculating an attention weight for each feature element, whose magnitude reflects its importance to the overall identification task. As data flows, the model generates an attention weight map. This map visually displays the distribution of attention across different parts of the data, indicating which components are more important for chemical identification. The resulting attention weight map is then used to weight the raw mass spectrometry data. Specifically, the model weights feature elements according to the weight map, giving greater attention to important features while deemphasizing or ignoring less important ones. This weighted processing can enhance the model's sensitivity to chemical signal characteristics, improving recognition accuracy and efficiency. It also helps reduce the impact of noise and interference on recognition results. After weighted processing, the model generates a target mass spectrometry data set containing optimized and enhanced chemical signal characteristics that are clearer, more accurate, and easier to analyze. This target mass spectrometry data set can be used for subsequent qualitative or quantitative analysis, chemical identification, metabolic pathway analysis, and other tasks. It provides more reliable data support and helps to better understand the chemical components and their content in the sample. In summary, by configuring the attention mechanism in the second channel and dynamically identifying and weighting the chemical signal characteristics, a target mass spectrometry data set can be generated that is more accurate, clear, and easy to analyze.

[0020] Baseline is a crucial concept in mass spectrometry data, representing the instrument's background response in the absence of chemical signal. Baseline instability or drift directly impacts the accuracy and reliability of mass spectrometry data. Therefore, baseline correction is necessary for each data point in the target mass spectrometry data set. Asymmetric least squares (ALS) is an effective method for baseline correction. It automatically finds a smooth curve that approximates the true baseline by analyzing the signal intensity distribution in mass spectrometry data. Compared to traditional least squares methods, ALS assigns different weights to data points on either side of the baseline, better accommodating the asymmetry between signal and noise in mass spectrometry data. During the ALS correction process, the model undergoes iterative optimization until an optimal baseline curve is found. This curve serves as a benchmark for subsequent analysis, removing baseline drift from the raw data. After baseline correction, signal-to-noise ratio (SNR) analysis is performed on the mass spectrometry data to improve data clarity and readability. The SNR, the ratio of signal intensity to noise intensity, is a key metric for assessing mass spectrometry data quality. The wavelet transform is a powerful signal processing tool that decomposes signals into components of varying frequencies and scales, effectively separating signal from noise. In mass spectrometry data analysis, the wavelet transform can be used to decompose mass spectrometry data into multiple wavelet coefficients. The noise component can then be removed by analyzing the distribution and characteristics of these coefficients. Specifically, a threshold can be set, where wavelet coefficients below the threshold are treated as noise and removed, while those above the threshold are retained as signal components. This approach significantly improves the signal-to-noise ratio (SNR) of mass spectrometry data, making the signal clearer and easier to analyze. After dynamic baseline correction and SNR processing, optimized and enhanced mass spectrometry data are obtained. These data are more accurate, clear, and easy to analyze, making them suitable for subsequent qualitative and quantitative analysis, chemical identification, metabolic pathway elucidation, and other tasks. Finally, these processed mass spectrometry data are compiled into a test report and presented to researchers or decision makers. This report contains key information such as the name, content, and structure of the chemical in the sample, providing strong data support for subsequent scientific analysis and decision-making. In summary, by using the asymmetric least squares method for dynamic baseline correction and combining it with wavelet transform for signal-to-noise ratio processing, more accurate and reliable mass spectrometry detection results can be obtained, which realizes noise suppression, signal separation and low-abundance detection of mass spectrometry data, and improves the accuracy and sensitivity of mass spectrometry data analysis.

[0021] In a preferred embodiment, the first channel is used to identify noise features in mass spectrometry data, and the execution steps include: preprocessing the input mass spectrometry data, including data cleaning, format conversion and normalization; based on a machine learning noise recognition algorithm, performing feature extraction and classification on the preprocessed mass spectrometry data to identify and separate the noise features in the data and obtain intermediate mass spectrometry data.

[0022] Specifically, raw mass spectrometry data is input into the first channel. This data may come from different experimental conditions, instrument types, or sample types, and therefore have varying formats and quality. To ensure accuracy and consistency in subsequent analysis, this data requires preprocessing. Preprocessing steps include data cleaning, format conversion, and normalization. Data cleaning aims to remove invalid, missing, or outlier values ​​from the data, ensuring data integrity and accuracy. Format conversion converts the data into a unified format to facilitate subsequent algorithm processing and analysis. Normalization scales the data to a specific range (e.g., between 0 and 1) to eliminate dimensional differences between different data and improve algorithm robustness. After preprocessing, the mass spectrometry data is fed into a machine learning-based noise identification algorithm. The core of this algorithm is feature extraction and classification. Feature extraction is the process of identifying key noise-related features in the mass spectrometry data. These features may include signal intensity, frequency, shape, etc., which can reflect the difference between noise and true signal. Extracting these features provides a strong basis for subsequent noise classification. The classification step uses the extracted features to distinguish noise from true signal in the mass spectrometry data. This is typically achieved by training a classifier (such as a support vector machine, decision tree, or neural network). The classifier learns patterns in noise characteristics and applies these patterns to new data to identify noise. After the classification step, the algorithm outputs a mass spectral dataset labeled with noise characteristics. This dataset clearly separates the noise and true signal components of the original data. Next, these noise characteristics need to be removed from the original data to obtain cleaner intermediate mass spectral data. This is usually achieved through simple data filtering or more complex signal processing techniques. After identifying and separating the noise characteristics, an intermediate mass spectral dataset containing only true signal is obtained. This dataset is clearer and more accurate than the original data, providing a better foundation for subsequent analysis. In summary, the first channel effectively identifies and addresses noise characteristics in mass spectral data through data preprocessing and a machine learning-based noise identification algorithm. This process not only improves data accuracy and consistency but also provides strong support for subsequent qualitative and quantitative analysis, chemical identification, and other tasks, further enhancing the efficiency and accuracy of mass spectrometry data analysis.

[0023] In a preferred embodiment, Figure 2As shown, the second channel is used to extract the signal features of chemical substances in the mass spectrometry data, and the execution steps include: performing signal feature extraction on the intermediate mass spectrometry data to generate a chemical substance signal feature set; performing feature selection, dimensionality reduction and enhancement processing on the chemical substance signal feature set; based on the processed chemical substance signal feature set, using a signal reconstruction algorithm to reconstruct the original mass spectrometry data to generate a target mass spectrometry data set containing chemical substance signal features.

[0024] Furthermore, signal feature extraction techniques are used to extract chemical-related signal features from the intermediate mass spectrometry data. These features may include signal intensity, frequency, shape, and duration, reflecting the unique behavior of different chemicals in mass spectrometry data. Feature extraction generates a chemical signal feature set, encompassing all chemical-related signal features extracted from the data. However, directly deriving subsequent analysis from this feature set can lead to computational overhead and feature redundancy. Therefore, further feature processing is required. Feature selection involves removing features that contribute little to the analysis results or are irrelevant to reduce computational overhead and improve analysis efficiency. Statistical methods, machine learning algorithms, or expert experience are used to select the most important features. Dimensionality reduction transforms a high-dimensional feature space into a low-dimensional space while preserving as much key information as possible from the original data. This helps reduce data complexity and improve algorithm performance. Feature enhancement uses techniques (such as filtering, smoothing, and transformations) to enhance the expressive power of features, making them easier for subsequent algorithms to identify and utilize. After feature processing, an optimized chemical signal feature set is obtained. Next, signal reconstruction algorithms are used to reassemble these features into a new mass spectrometry dataset. The goal of the signal reconstruction algorithm is to restore or enhance the chemical signal in the original mass spectrometry data based on the extracted features. This process may employ mathematical models or optimization algorithms to ensure that the reconstructed data retains the authenticity of the original data while highlighting the chemical signal characteristics. Ultimately, a target mass spectrometry data set containing the chemical signal characteristics is generated. This data set not only removes noise but also highlights the chemical signal, providing a better foundation for subsequent analysis and identification. In summary, the second channel effectively extracts the chemical signal characteristics from the mass spectrometry data through signal feature extraction, feature processing (selection, dimensionality reduction, and enhancement), and signal reconstruction, generating a target mass spectrometry data set containing these features. This process not only improves data analysis efficiency but also provides strong support for subsequent tasks such as chemical identification and quantitative analysis, further enhancing the accuracy and reliability of mass spectrometry data analysis.

[0025] In a preferred embodiment, an attention mechanism is configured in the second channel, and the execution steps include: constructing a convolutional attention mechanism based on the time-mass-to-charge ratio two-dimensional distribution characteristics of the chemical substance signal characteristics, wherein the convolutional attention mechanism includes a spatial attention sublayer and a channel attention sublayer; reshaping the dimension of the generated target mass spectrum data set to form a three-dimensional tensor input of the number of samples × the number of mass-to-charge ratio points × the number of feature channels; generating an attention weight matrix of the same dimension as the input three-dimensional tensor through dual-branch calculation of the attention layer to complete the attention mechanism configuration, and the value range of the matrix elements is [0,1].

[0026] Optionally, an attention mechanism is introduced in the second channel to further enhance the ability to extract chemical signal features. This mechanism strengthens the model's focus on key features while suppressing unimportant ones, thereby improving the accuracy and efficiency of mass spectrometry data analysis. First, a convolutional attention mechanism is constructed based on the two-dimensional time-mass-to-charge ratio distribution of chemical signal features. This mechanism consists of two core sublayers: a spatial attention sublayer and a channel attention sublayer. The spatial attention sublayer uses a 3×3 convolution kernel with a stride of 1 to capture spatial information in the feature map (i.e., the two-dimensional representation of the mass spectrometry data). This sublayer identifies important and unimportant regions in the feature map and generates corresponding spatial attention weights. These weights enhance the features of important regions while suppressing those of unimportant regions. To capture the dependencies between different feature channels, the channel attention sublayer combines global average pooling with full connectivity. First, global average pooling is performed on each channel of the feature map to generate a channel description vector. This vector is then passed through a fully connected layer to generate per-channel attention weights, which reflect the contribution of each channel to the final classification or regression task. Next, the generated target mass spectrometry data set is reshaped. Specifically, the data is reshaped into a three-dimensional tensor consisting of the number of samples × the number of mass-to-charge ratio points × the number of feature channels. This three-dimensional tensor contains both the spatial information of the mass spectrometry data (the number of mass-to-charge ratio points) and the information of the different feature channels. After receiving the three-dimensional tensor input, it is fed into the attention layer for a two-branch computation. This layer considers both spatial and channel attention, generating a spatial attention weight matrix and a channel attention weight matrix through two sublayers, respectively. In the spatial attention sublayer, a 3×3 convolution kernel is used to convolve the input three-dimensional tensor to generate a spatial attention weight matrix, in which each element represents the importance of the feature at the corresponding position. In the channel attention sublayer, global average pooling is first performed on each channel to generate a channel description vector. This vector is then passed through a fully connected layer to generate a channel attention weight matrix, in which each element represents the importance of the corresponding channel. Finally, the spatial attention weight matrix and the channel attention weight matrix are element-wise multiplied to obtain the final attention weight matrix. The elements of this matrix range from [0, 1], indicating the importance of the features at the corresponding position or channel. After obtaining the attention weight matrix, it is applied to the input three-dimensional tensor. This element-wise multiplication enhances important features while suppressing unimportant ones. This completes the configuration of the attention mechanism and produces an attention-weighted three-dimensional tensor output. By introducing the attention mechanism, the second channel can more effectively extract chemical signal features from mass spectrometry data.The combined use of spatial attention sublayer and channel attention sublayer enables the model to simultaneously capture the spatial information in the feature map and the dependencies between channels, thereby improving the accuracy and efficiency of the analysis.

[0027] In a preferred embodiment, the chemical substance signal characteristics are dynamically identified to obtain an attention weight map, and the execution steps include: in the spatial attention sublayer branch, generating a spatial attention map by cross-channel weighting , where H is the number of mass-to-charge ratio points and W is the number of scanning time points; in the channel attention sub-layer branch, the channel weight vector is generated by calculating the autocorrelation between feature channels , C is the number of feature channels; perform tensor outer product of S and C to form a three-dimensional attention weight map , to perform dynamic focusing of time-mass-to-charge ratio-characteristic three-dimensional space.

[0028] In detail, we first focus on the spatial attention sub-layer branch. This branch considers the feature distribution of mass spectrometry data in the two-dimensional space formed by mass-to-charge ratio (H) and scan time points (W). To generate the spatial attention map, a cross-channel weighted approach is employed. Specifically, for the input 3D tensor (number of samples × number of mass-to-charge ratio points H × number of scan time points W × number of feature channels C), spatial features are first extracted independently for each feature channel. This can be achieved through convolution or other spatial feature extraction methods. These extracted spatial features are then weighted and summed across channels to generate a 2D spatial attention map. This map reflects which regions in the 2D space formed by mass-to-charge ratio and scan time point are important and which are not. Next, the channel attention sub-layer branch focuses on the dependencies between different feature channels. To generate the channel weight vector, an autocorrelation method is employed. Specifically, the input 3D tensor is first globally average pooled at each scan time point and mass-to-charge ratio point to obtain a description vector for the feature channel. Next, this description vector is autocorrelated (i.e., its dot product with itself) and then passed through a nonlinear activation function (such as ReLU or Sigmoid) to generate a channel weight vector. This vector reflects the importance of different feature channels. After obtaining the spatial attention map and channel weight vectors, they are combined via tensor outer product to form a three-dimensional attention weight map. This map has the same dimensions as the input three-dimensional tensor (a simplified representation of the number of samples × number of mass-to-charge ratio points H × number of scan time points W × number of feature channels C), but in the attention weight map, the last dimension is replaced by the attention weight. However, it contains richer information. Specifically, for each element in the input tensor, a corresponding weight value is found in the three-dimensional attention weight map. This weight value reflects the importance of that element in the three-dimensional space of time, mass-to-charge ratio, and features. This approach achieves dynamic focusing in the three-dimensional space, enabling the model to more accurately capture key feature information. In summary, by constructing a spatial attention sublayer and a channel attention sublayer and skillfully combining them to generate a three-dimensional attention weight map, dynamic recognition of chemical signal characteristics is achieved. This process not only improves the accuracy of feature extraction, but also enables the model to capture key feature information more accurately, providing strong support for subsequent analysis and recognition tasks.

[0029] In a preferred embodiment, based on the weighted processing of the attention weight map, a target mass spectrum data set is generated, and the execution step includes: performing element-by-element dot multiplication of the three-dimensional attention weight map with the input three-dimensional tensor, and the calculation formula is: , where i is the sample index, j is the number of mass-to-charge ratio points, and k is the number of characteristic channels. A weighted sum is performed along the characteristic channel dimension to generate a two-dimensional enhanced mass spectrum. The calculation formula is: ,in, is a channel weight coefficient optimized through training; local extreme value detection is performed on the enhanced mass spectrum, peak top coordinates are extracted and adjacent peak clusters are merged to form a target mass spectrum data set.

[0030] Specifically, based on the obtained three-dimensional attention weight map, the input three-dimensional tensor is further weighted to generate a more accurate and valuable target mass spectrometry data set. This process combines the advantages of the attention mechanism and significantly improves the analysis quality of mass spectrometry data by enhancing key features and suppressing noise. First, the three-dimensional attention weight map is element-by-element dot multiplication with the input three-dimensional tensor. This step aims to weight each element in the input tensor according to the instructions of the attention weight map, thereby highlighting important feature information and suppressing unimportant information. The specific calculation formula is: , i represents the sample index, j represents the mass-to-charge ratio point number (i.e., m / z value), and k is the number of characteristic channels. This element-by-element dot product yields a weighted three-dimensional tensor. Next, the weighted three-dimensional tensor is weighted summed along the characteristic channel dimension to generate a two-dimensional enhanced mass spectrum. This step aims to fuse the information from different characteristic channels to form a more comprehensive and clear mass spectrum representation. The specific calculation formula is: ,in, The channel weight coefficients are optimized through training and reflect the importance of different feature channels to the final mass spectrum. The result of this weighted summation is a two-dimensional matrix, whose rows represent the number of mass-to-charge ratio points (m / z values) and whose columns represent the corresponding intensity values. After obtaining the two-dimensional enhanced mass spectrum, local extremum detection is performed to extract the peak apex coordinates in the spectrum. These peak apex coordinates correspond to significant features in the mass spectrum and are crucial for subsequent peak identification and qualitative or quantitative analysis of substances. Here, peak apex coordinates are understood as a binary pair of (m / z value, intensity). Since a mass spectrum may contain multiple adjacent peaks, these peaks may originate from different isotopes or different molecular forms of the same chemical. Therefore, after extracting the peak apex coordinates, adjacent peak clusters need to be merged. This step aims to merge peaks with similar positions and intensities into a single peak cluster, thereby simplifying subsequent mass spectrometry data analysis. In summary, the target mass spectrometry data set was generated by weighted processing based on the three-dimensional attention weight map. This process significantly improved the analysis quality and reliability of mass spectrometry data by highlighting key features, suppressing noise, and merging adjacent peak clusters.

[0031] In a preferred embodiment, a confidence assessment is performed on the target mass spectrum data set, and if the peak intensity variation coefficient is greater than 30%, the noise re-check process of the first channel is triggered.

[0032] Before further analysis and application of the target mass spectrometry data set, a confidence assessment step is required to ensure data accuracy and reliability. This step focuses on the stability of peak intensities in the mass spectrum, assessing data confidence by calculating the coefficient of variation (CV) of the peak intensities. First, the CV of the peak intensity is calculated for each peak (or peak cluster) in the target mass spectrometry data set. The CV is a measure of peak intensity fluctuation and is calculated by dividing the standard deviation of the peak intensity by the mean peak intensity. This metric reflects the stability and consistency of peak intensities in the mass spectrum. The calculation process may take into account various factors, such as repeated measurements under different experimental conditions and instrumental measurement errors, to ensure that the CV fully reflects the true state of the data. After the CV is calculated, it is compared with a preset threshold (30% in this example). If the CV of a peak exceeds the threshold, the peak intensity fluctuation is considered large, the data confidence is low, and further review and processing may be required. If the CV of a peak exceeds the threshold, a noise recheck process is triggered for the first channel. The purpose of this step is to perform a more detailed inspection of the first channel to identify and eliminate possible sources of noise or errors. The noise re-inspection process may include multiple steps, such as re-checking experimental conditions, recalibrating instruments, optimizing data processing algorithms, etc. These steps are aimed at improving the accuracy and reliability of the data and ensuring that subsequent analysis and applications can be based on high-quality data. In summary, by performing a confidence assessment on the target mass spectrometry data set and calculating the peak intensity variation coefficient to evaluate the stability and consistency of the data, peaks with low confidence can be discovered and processed in a timely manner. When the peak intensity variation coefficient exceeds the preset threshold, the noise re-inspection process for the first channel is triggered to ensure the accuracy and reliability of the data. This process not only helps to improve the analytical quality of mass spectrometry data, but also provides strong support for subsequent tasks such as substance identification and quantitative analysis.

[0033] In a preferred embodiment, an asymmetric least squares method is used to perform dynamic baseline correction on the data in the target mass spectrometry data set, and the execution steps include: segmenting the mass spectrometry data of each sample, and dividing the high confidence region and the low confidence region based on the attention weight map; using the asymmetric least squares method to fit the baseline in the high confidence region, and switching to an adaptive penalty spline model in the low confidence region to dynamically adjust the smoothing parameters to match the local noise level; connecting each segmented baseline through cubic spline interpolation to generate a global baseline curve and deducting it from the original data.

[0034] For example, in mass spectrometry data analysis, baseline correction is a crucial step that directly impacts the accuracy and reliability of subsequent tasks such as substance identification and quantitative analysis. To dynamically perform baseline correction on the data in the target mass spectrometry dataset, a method combining asymmetric least squares and an adaptive penalized spline model is used. First, the mass spectrometry data of each sample is segmented so that baseline correction can be performed on each segment separately in subsequent steps. Segmentation can be performed based on the characteristics of the mass spectrometry data or the analysis requirements, aiming to divide the data into multiple relatively independent and easily processed parts. Next, each segment is divided into high-confidence and low-confidence regions using the previously generated attention weight map. In this example, high-confidence regions are defined as regions with weights ≥ 0.7, which typically contain relatively stable and distinct mass spectrometry peaks; low-confidence regions are defined as regions with weights < 0.7, which may be significantly affected by factors such as noise and instrument errors. In the high-confidence regions, asymmetric least squares is used for baseline fitting. Asymmetric least squares is a commonly used baseline correction method. It introduces an asymmetric factor into the objective function to prioritize regions of baseline underestimation, thereby avoiding overestimation. In this example, a smoothing coefficient λ = 10³ is set to suppress high-frequency noise interference, and an asymmetry factor p = 0.001 is set to ensure the accuracy and stability of the baseline fit. In low-confidence regions, where noise levels are high and features are less distinct, an adaptive penalized spline model is used for baseline fitting. This adaptive penalized spline model dynamically adjusts the smoothing parameter based on the local noise level, ensuring baseline smoothness while preserving as much useful information as possible in the data. In low-confidence regions, the smoothing parameter is adjusted to match the local noise level to ensure the accuracy and reliability of the baseline fit. Finally, cubic spline interpolation is used to connect the segmented baselines to generate a global baseline curve. Cubic spline interpolation is a commonly used interpolation method that ensures smoothness and continuity of the baseline curve. After generating the global baseline curve, this baseline curve is subtracted from the raw data to obtain the corrected mass spectrometry data. In summary, a method combining asymmetric least squares with an adaptive penalized spline model was used to dynamically correct the baseline of the target mass spectrometry data set. Through steps such as data segmentation, confidence region delineation, baseline fitting and smoothing parameter adjustment, and global baseline curve generation and data subtraction, baseline correction of the mass spectrometry data was achieved, providing high-quality data support for subsequent tasks such as substance identification and quantitative analysis. This process improves the accuracy and reliability of mass spectrometry data analysis.

[0035] In a preferred embodiment, the signal-to-noise ratio processing is performed in combination with wavelet transform, and the execution steps include: selecting the Symlets wavelet basis to perform a 5-layer wavelet decomposition on the baseline-corrected data to separate the high-frequency noise and low-frequency signal components; performing adaptive threshold processing on the detail coefficients of the 1st to 3rd layers, and selectively retaining the approximate coefficients of the 4th to 5th layers based on the attention weight; if the mean of the attention weight corresponding to the peak coordinate interval is ≥0.6, the coefficient is retained; otherwise, the coefficient is attenuated to 20% of the original value; performing wavelet reconstruction to generate a denoised mass spectrum, and outputting the mass spectrum detection results.

[0036] Specifically, to further optimize the baseline-corrected mass spectrometry data, wavelet transform technology was combined with sophisticated decomposition, processing, and reconstruction steps to effectively suppress high-frequency noise and preserve low-frequency signal components. First, the Symlets wavelet basis was selected as the decomposition tool because of its excellent symmetry and regularity, making it suitable for processing mass spectrometry data containing complex features. Next, a five-layer wavelet decomposition was performed on the baseline-corrected mass spectrometry data. This step aims to decompose the raw data into signals with different frequency components, where the high-frequency components primarily contain noise, while the low-frequency components contain useful signal information. This five-layer decomposition yields a low-frequency approximation coefficient (representing the low-frequency signal at the lowest level) and five sets of high-frequency detail coefficients (representing the high-frequency noise components at each level). Adaptive thresholding was performed on the high-frequency detail coefficients of layers 1 through 3. This step aims to remove high-frequency noise from these layers while preserving as much useful signal information as possible. Adaptive thresholding dynamically adjusts the threshold based on the data characteristics, ensuring effective denoising while avoiding signal distortion caused by overprocessing. A selective retention strategy based on attention weights was employed for the low-frequency approximation coefficients of layers 4 and 5. This step aims to finely filter and retain low-frequency signal components based on the information in the attention weight map. Specifically, the mean of the attention weights for the corresponding peak coordinate interval is calculated. If the mean is ≥ 0.6, the signal in that interval is relatively stable and has distinct characteristics, so the approximation coefficient for that interval is retained. If the mean is < 0.6, the signal in that interval may be significantly affected by noise or errors, so the approximation coefficient for that interval is attenuated to 20% of its original value to suppress potential noise interference. After processing the detail and approximation coefficients, a wavelet reconstruction step is performed. This step aims to reassemble the processed coefficients into complete mass spectrometric data, thereby obtaining a denoised mass spectrum. Wavelet reconstruction successfully preserves the useful information in the low-frequency signal components while effectively removing the interference of high-frequency noise. The resulting denoised mass spectrum has a higher signal-to-noise ratio and better signal recognition. Finally, the denoised mass spectrum is used as input for subsequent mass spectrometry detection steps, such as substance identification and quantitative analysis. These steps can be performed based on high-quality data, resulting in more accurate and reliable detection results. In summary, the wavelet transform technology is combined with the sophisticated decomposition, processing and reconstruction steps to achieve the signal-to-noise ratio processing of the mass spectrometry data after baseline correction.

[0037] The data analysis system based on the mass spectrometry detection platform provided by the embodiment of the present invention has at least the following technical effects:

[0038] 1. By building a dual-channel mass spectrometry detection model, the first channel focuses on identifying and separating noise features, while the second channel focuses on extracting chemical signal characteristics. The introduction of an attention mechanism in the second channel dynamically identifies and weights chemical signal characteristics, generating a high-confidence target mass spectrometry data set. This significantly improves the signal-to-noise ratio and signal recognition capabilities of mass spectrometry data, providing a more accurate and reliable data foundation for subsequent analysis.

[0039] 2. Dynamic baseline correction is performed on the target mass spectrometry data set using an asymmetric least squares method. Attention weight maps are used to divide the data into high-confidence and low-confidence regions, enabling precise baseline fitting and adjustment. In the high-confidence region, the asymmetric least squares method effectively suppresses high-frequency noise. In the low-confidence region, an adaptive penalized spline model is used to dynamically adjust the smoothing parameters to match the local noise level. This not only improves the accuracy of baseline correction but also enhances the overall quality of the mass spectrometry data.

[0040] 3. During the signal-to-noise ratio processing stage, wavelet transform technology is combined to separate high-frequency noise from low-frequency signal components through a five-layer wavelet decomposition. Adaptive thresholding is performed on detail coefficients, and approximate coefficients are selectively retained based on attention weights. If the mean attention weight for the corresponding peak coordinate interval is high, the coefficient is retained; otherwise, the coefficient is attenuated to suppress potential noise. This strategy maximizes the preservation of useful information in the mass spectrometry data while ensuring effective denoising, further improving the accuracy and reliability of mass spectrometry results.

[0041] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data analysis system based on a mass spectrometry detection platform, characterized in that: The execution steps include: A dual-channel mass spectrometry detection model was constructed, in which the first channel was used to identify noise features in mass spectrometry data, and the second channel was used to extract chemical substance signal features in mass spectrometry data. The mass spectrometry data passed through the first and second channels in sequence. configuring an attention mechanism in the second channel to dynamically identify the chemical substance signal characteristics, obtain an attention weight map, and generate a target mass spectrometry data set based on weighted processing of the attention weight map; Performing dynamic baseline correction on the data in the target mass spectrometry data set using an asymmetric least squares method, performing signal-to-noise ratio processing in combination with wavelet transform, and obtaining mass spectrometry detection results; A confidence assessment is performed on the target mass spectrometry data set, and if the peak intensity variation coefficient is greater than 30%, the noise re-check process of the first channel is triggered.

2. The system according to claim 1, wherein The first channel is used to identify noise features in mass spectrometry data. The execution steps include: Preprocessing the input mass spectrometry data, including data cleaning, format conversion and normalization; Based on the noise recognition algorithm of machine learning, feature extraction and classification are performed on the preprocessed mass spectrometry data to identify and separate the noise features in the data and obtain intermediate mass spectrometry data.

3. The system according to claim 2, wherein: The second channel is used to extract chemical substance signal characteristics from mass spectrometry data. The execution steps include: performing signal feature extraction on the intermediate mass spectrum data to generate a chemical substance signal feature set; performing feature selection, dimensionality reduction and enhancement processing on the chemical substance signal feature set; Based on the processed chemical substance signal feature set, a signal reconstruction algorithm is used to reconstruct the original mass spectrum data to generate a target mass spectrum data set containing the chemical substance signal features.

4. The system according to claim 3, wherein: Configuring an attention mechanism in the second channel includes the following steps: Based on the time-mass-to-charge ratio two-dimensional distribution characteristics of chemical signal features, a convolutional attention mechanism is constructed, wherein the convolutional attention mechanism includes a spatial attention sublayer and a channel attention sublayer; Reshaping the generated target mass spectrum data set to be determined to form a three-dimensional tensor input of number of samples × number of mass-to-charge ratio points × number of feature channels; Through the dual-branch calculation of the attention layer, an attention weight matrix with the same dimension as the input three-dimensional tensor is generated to complete the attention mechanism configuration, and the value range of the matrix elements is [0, 1].

5. The system according to claim 4, wherein: Dynamically identifying the chemical substance signal characteristics to obtain an attention weight map includes the following steps: In the spatial attention sub-layer branch, a spatial attention map is generated by cross-channel weighting , where H is the mass-to-charge ratio point number and W is the scanning time point number; In the channel attention sub-layer branch, the channel weight vector is generated by calculating the autocorrelation between feature channels. , where C is the number of feature channels; Perform tensor outer product of S and C to form a three-dimensional attention weight map , to perform dynamic focusing of time-mass-to-charge ratio-characteristic three-dimensional space.

6. The system according to claim 5, wherein: Based on the weighted processing of the attention weight map, a target mass spectrometry data set is generated, and the execution steps include: Perform element-wise dot multiplication of the three-dimensional attention weight map and the input three-dimensional tensor. The calculation formula is: , where i is the sample index, j is the number of mass-to-charge ratio points, and k is the number of feature channels; A weighted summation is performed along the characteristic channel dimension to generate a two-dimensional enhanced mass spectrum. The calculation formula is: ,in, is the channel weight coefficient optimized through training; The enhanced mass spectrum is subjected to local extreme value detection, peak top coordinates are extracted, and adjacent peak clusters are merged to form a target mass spectrum data set.

7. The system according to claim 1, wherein: Performing dynamic baseline correction on the data in the target mass spectrometry data set using an asymmetric least squares method, the execution steps comprising: The mass spectrometry data of each sample is segmented into high-confidence areas and low-confidence areas based on the attention weight map; In the high confidence region, an asymmetric least squares method is used to fit the baseline, and in the low confidence region, an adaptive penalized spline model is used to dynamically adjust the smoothing parameters to match the local noise level; The segmented baselines were connected by cubic spline interpolation to generate a global baseline curve which was subtracted from the original data.

8. The system according to claim 7, wherein: The signal-to-noise ratio processing is performed in combination with wavelet transform, and the execution steps include: The Symlets wavelet basis was selected to perform a 5-layer wavelet decomposition on the baseline-corrected data to separate high-frequency noise and low-frequency signal components; Adaptive threshold processing is performed on the detail coefficients of layers 1 to 3, and selective retention based on attention weights is performed on the approximate coefficients of layers 4 to 5; If the mean of the attention weight corresponding to the peak coordinate interval is ≥ 0.6, the coefficient is retained; otherwise, the coefficient is decayed to 20% of the original value; Perform wavelet reconstruction to generate a denoised mass spectrum and output the mass spectrum detection results.

Citation Information

Patent Citations

  • GIS equipment breakdown signal analysis method and system based on hierarchical networking structure

    CN118965107A

  • Sample injection control method and system of mass spectrometer

    CN119936427A