Data analysis system based on mass spectrum detection platform
By building a dual-channel mass spectrometry detection model and introducing attention mechanism, combined with dynamic baseline correction and wavelet transformation processing, the challenges of noise suppression, signal separation and low abundance detection in mass spectrometry data analysis are solved, significantly improving the accuracy and reliability of the analysis.
Patent Information
- Application Number
- CN202510423920.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
There are challenges in mass spectrometry data analysis of noise suppression, signal separation and low abundance detection, resulting in insufficient accuracy and sensitivity of the analysis.
A two-channel mass spectrometry detection model is constructed, the first channel is used to identify noise characteristics, the second channel is used to extract chemical signal characteristics, and an attention mechanism is introduced into the second channel for dynamic identification. Dynamic baseline correction is performed by combining asymmetric least squares method, and signal-to-noise ratio processing is performed through wavelet transformation.
Effectively suppress noise, separate signals, improve the accuracy and sensitivity of low abundance detection, and significantly improve the accuracy and reliability of mass spectrometry data analysis.
Smart Images

Figure CN119936168A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a data analysis system based on a mass spectrometry detection platform. Background Art
[0002] As an efficient and sensitive analytical technology, mass spectrometry has been widely used in many fields such as biomedicine, environmental monitoring, and food safety. However, mass spectrometry data often contains a large amount of complex background noise and overlapping peaks. In particular, the signals of low-abundance molecules or isomers are often easily masked or misjudged due to their weak intensity and susceptibility to interference, which poses a huge challenge to the accurate analysis of mass spectrometry data. Traditional mass spectrometry data analysis methods have certain limitations in dealing with noise suppression, signal separation, and low-abundance detection problems. For example, although traditional filtering algorithms can remove noise to a certain extent, they often have the risk of over-suppression of low-abundance signals, resulting in missed detections. In addition, analysis methods for single-dimensional data are usually unable to effectively separate isomers, further limiting the accuracy and reliability of mass spectrometry data analysis. Summary of the invention
[0003] The present invention aims to solve the technical problem of insufficient accuracy and sensitivity of mass spectrometry data analysis during mass spectrometry detection in the prior art, and provides a data analysis system based on a mass spectrometry detection platform.
[0004] The technical solution of the present invention to solve the above technical problems is as follows:
[0005] The present invention provides a data analysis system based on a mass spectrometry detection platform, and the execution steps include: constructing a dual-channel mass spectrometry detection model, wherein the first channel is used to identify noise characteristics in mass spectrometry data, and the second channel is used to extract chemical substance signal characteristics in the mass spectrometry data, and the mass spectrometry data passes through the first and second channels in sequence; configuring an attention mechanism in the second channel, dynamically identifying the chemical substance signal characteristics, obtaining an attention weight map, and generating a target mass spectrometry data set based on weighted processing of the attention weight map; using an asymmetric least squares method to dynamically perform baseline correction on the data in the target mass spectrometry data set, combining wavelet transform to perform signal-to-noise ratio processing, and obtaining a mass spectrometry detection result.
[0006] The beneficial effects of the present invention are: by constructing a dual-channel convolutional neural network, introducing an attention mechanism, and combining dynamic baseline correction technology and wavelet transform, noise suppression, signal separation and low-abundance detection of mass spectrometry data are achieved, thereby improving the accuracy and sensitivity of mass spectrometry data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 A schematic flow chart of the execution steps of a data analysis system based on a mass spectrometry detection platform provided by the present invention.
[0008] Figure 2 A schematic flow chart of the execution steps of second channel feature extraction in a data analysis system based on a mass spectrometry detection platform provided by the present invention. DETAILED DESCRIPTION
[0009] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0010] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0011] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in the present invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any technician in the field to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present invention.
[0012] Example:
[0013] like Figure 1 As shown, the embodiment of the present invention provides a data analysis system based on a mass spectrometry detection platform, and the execution steps include:
[0014] S10: Construct a dual-channel mass spectrometry detection model, wherein the first channel is used to identify noise features in mass spectrometry data, and the second channel is used to extract chemical substance signal features in mass spectrometry data, and the mass spectrometry data passes through the first and second channels in sequence.
[0015] S20: configuring an attention mechanism in the second channel, dynamically identifying the signal characteristics of the chemical substance, obtaining an attention weight map, and generating a target mass spectrometry data set based on weighted processing of the attention weight map.
[0016] S30: performing dynamic baseline correction on the data in the target mass spectrometry data set by using an asymmetric least squares method, and performing signal-to-noise ratio processing in combination with wavelet transform to obtain a mass spectrometry detection result.
[0017] Exemplarily, when performing data analysis on a mass spectrometry detection platform, mass spectrometry data often contains a large amount of complex background noise and overlapping peaks, especially signals of low-abundance molecules or isomers. Among them, background noise refers to irrelevant signals in the mass spectrogram, which may come from instrument electronic noise, sample matrix interference, environmental noise, etc. The presence of background noise will interfere with the true signal peak and affect the accuracy and reliability of mass spectrometry data. Overlapping peaks refer to signal peaks of two or more compounds in the mass spectrogram that may overlap due to similar mass or insufficient instrument resolution. Overlapping peaks will make it difficult to identify signal peaks, affecting the accuracy of qualitative and quantitative analysis. Low-abundance molecules refer to molecules with low content in the sample, and their signal peaks may be relatively weak in the mass spectrogram, and may even be covered by background noise. Isomers refer to compounds with the same molecular formula but different structures. In the mass spectrogram, the signal peaks of isomers may be difficult to distinguish due to similar mass. The identification of isomers is of great significance for understanding the structural diversity of compounds and analyzing components in complex samples. Therefore, this application avoids the above complex interference problems and improves the analysis accuracy and reliability of mass spectrometry detection data through technical means such as dual-channel processing, dynamic identification and weighted processing, baseline correction and signal-to-noise ratio optimization.
[0018] Specifically, the raw mass spectrometry data are input into the model. These data usually contain a lot of information, including the signal of the target chemical substance and interference factors such as background noise. Before the data enters the dual channel, some preprocessing steps are usually performed, such as smoothing, baseline correction and normalization, to reduce random errors and instrument noise in the data and improve the accuracy of subsequent analysis. The preprocessed mass spectrometry data first enters the first channel, namely the noise feature identification channel. The main task of this channel is to identify and separate the noise features in the data. This is usually achieved through a series of complex algorithms, such as wavelet transform, principal component analysis (PCA) or machine learning algorithms (such as support vector machine SVM, neural network, etc.). These algorithms can analyze the frequency components, statistical characteristics or patterns of the data to effectively distinguish between noise signals and target chemical signals. In the first channel, the identified noise features will be marked and removed or weakened from the raw data to reduce interference with subsequent analysis. After processing in the first channel, the mass spectrometry data with noise removed or weakened enters the second channel, namely the chemical signal feature extraction channel. The goal of this channel is to extract signal features related to the target chemical substance. This also relies on advanced algorithms and models, such as feature selection algorithms, pattern recognition algorithms or deep learning networks. These algorithms can identify specific patterns or features in the data that match the mass spectra of known chemicals. In the second channel, the extracted chemical signal features will be further analyzed and processed for subsequent qualitative or quantitative analysis. In a preferred embodiment, the dual channels are constructed by a convolutional neural network (CNN), and the two channels learn noise patterns and real signal features respectively. And both use synthetic data sets, that is, real mass spectrometry data with simulated noise superposition to train the model, generate more diverse training samples, and when the model faces new and unknown mass spectrometry data, it can better cope with noise interference and improve the accuracy and reliability of the analysis. At the same time, the real mass spectrometry data with simulated noise superposition can also be used to optimize the noise filtering algorithm, which can more accurately identify and remove the noise components in the mass spectrometry data, improve the signal-to-noise ratio and clarity of the data, and achieve end-to-end noise filtering.
[0019] Furthermore, in the second channel, a convolutional neural network (CNN) or other deep learning models are used to extract chemical signal features in the mass spectrometry data. These features may include the position, intensity, shape, etc. of the peak, which are crucial for the identification and quantitative analysis of chemical substances. In order to further improve the accuracy and efficiency of feature extraction, an attention mechanism is configured in the second channel. The attention mechanism is a deep learning technology that simulates the distribution of human visual attention. It can dynamically focus on the important parts of the input data and ignore irrelevant information. In the second channel configured with the attention mechanism, the model can dynamically identify the chemical signal features in the mass spectrometry data. This is achieved by calculating the attention weight of each feature element, and the size of the weight reflects the importance of the element to the overall recognition task. As the data flows, the model generates an attention weight map, which visualizes the attention distribution of different parts of the data, that is, which parts are more important for the identification of chemical substances. After obtaining the attention weight map, it is used to weight the original mass spectrometry data. Specifically, the model weights the feature elements according to the weight map so that important features receive more attention, while unimportant features are weakened or ignored. This weighted processing can enhance the model's sensitivity to the signal characteristics of chemical substances and improve the accuracy and efficiency of recognition. At the same time, it also helps to reduce the impact of noise and interference factors on the recognition results. After weighted processing, the model generates a target mass spectrometry data set, which contains optimized and enhanced chemical signal characteristics, which are clearer, more accurate and easier to analyze. The target mass spectrometry data set can be used for subsequent qualitative or quantitative analysis, chemical identification, metabolic pathway analysis and other tasks. It provides more reliable data support and helps to better understand the chemical components and their contents in the sample. In summary, by configuring the attention mechanism in the second channel and dynamically identifying and weighting the chemical signal characteristics, a more accurate, clear and easy-to-analyze target mass spectrometry data set can be generated.
[0020] Baseline is an important concept in mass spectrometry data, which represents the background response of the instrument when there is no signal of chemical substances. The instability or offset of the baseline will directly affect the accuracy and reliability of mass spectrometry data. Therefore, it is further necessary to perform baseline correction on each data in the target mass spectrometry data set. Asymmetric Least Squares (ALS) is an effective method for baseline correction. It automatically finds a smooth curve close to the true baseline by analyzing the signal intensity distribution in mass spectrometry data. Compared with the traditional least squares method, ALS gives different weights to the data points on both sides of the baseline, so as to better adapt to the asymmetry of signal and noise in mass spectrometry data. During the ALS correction process, the model will be continuously iterated and optimized until an optimal baseline curve is found. This curve will be used as the benchmark for subsequent analysis to remove the baseline offset from the original data. After baseline correction, the mass spectrometry data also needs to be processed for signal-to-noise ratio to improve the clarity and readability of the data. The signal-to-noise ratio refers to the ratio of signal intensity to noise intensity, and is an important indicator for evaluating the quality of mass spectrometry data. Wavelet transform is a powerful signal processing tool that can decompose signals into components of different frequencies and scales, thereby achieving effective separation of signals and noise. In mass spectrometry data analysis, wavelet transform can be used to decompose mass spectrometry data into multiple wavelet coefficients, and then the noise components can be removed by analyzing the distribution and characteristics of these coefficients. Specifically, a threshold can be set, and the wavelet coefficients below the threshold can be regarded as noise and removed, while the wavelet coefficients above the threshold can be retained as signal components. In this way, the signal-to-noise ratio of mass spectrometry data can be significantly improved, making the signal clearer and easier to analyze. After dynamic baseline correction and signal-to-noise ratio processing, optimized and enhanced mass spectrometry data are obtained. These data are more accurate, clear and easy to analyze, and can be used for subsequent qualitative or quantitative analysis, chemical identification, metabolic pathway analysis and other tasks. Finally, these processed mass spectrometry data are organized into a test report and presented to researchers or decision makers. This report contains key information such as the name, content, and structural information of the chemical substances in the sample, providing strong data support for subsequent scientific research analysis and decision-making. In summary, by using asymmetric least squares method for dynamic baseline correction and combining it with wavelet transform for signal-to-noise ratio processing, more accurate and reliable mass spectrometry detection results can be obtained, noise suppression, signal separation and low-abundance detection of mass spectrometry data can be achieved, and the accuracy and sensitivity of mass spectrometry data analysis can be improved.
[0021] In a preferred embodiment, the first channel is used to identify noise features in mass spectrometry data, and the execution steps include: preprocessing the input mass spectrometry data, including data cleaning, format conversion and normalization; based on a machine learning noise identification algorithm, performing feature extraction and classification on the preprocessed mass spectrometry data to identify and separate the noise features in the data and obtain intermediate mass spectrometry data.
[0022] Specifically, the raw mass spectrometry data are input into the first channel. These data may come from different experimental conditions, instrument types or sample types, and therefore have different formats and qualities. In order to ensure the accuracy and consistency of subsequent analysis, these data need to be preprocessed. The preprocessing steps include data cleaning, format conversion and normalization. Data cleaning aims to remove invalid, missing or outliers in the data to ensure the integrity and accuracy of the data. Format conversion is to convert the data into a unified format for subsequent algorithm processing and analysis. Normalization is to scale the data to a specific range (such as between 0 and 1) to eliminate the dimensional differences between different data and improve the robustness of the algorithm. After preprocessing, the mass spectrometry data is fed into the noise identification algorithm based on machine learning. The core of this algorithm is feature extraction and classification. Feature extraction is the process of identifying key features related to noise in mass spectrometry data. These features may include signal intensity, frequency, shape, etc., which can reflect the difference between noise and real signals. By extracting these features, a strong basis can be provided for subsequent noise classification. The classification step is to distinguish the noise from the real signal in the mass spectrometry data based on the extracted features. This is usually achieved by training a classifier (such as a support vector machine, decision tree, or neural network). The classifier learns the patterns of noise features and applies these patterns to new data to identify noise. After the classification step, the algorithm outputs a mass spectrometry data set labeled with noise features, which contains the noise part and the real signal part of the original data, which are clearly separated. Next, these noise features need to be removed from the original data to obtain cleaner intermediate mass spectrometry data. This is usually achieved through simple data filtering or more complex signal processing techniques. After the identification and separation of noise features, an intermediate mass spectrometry data set containing only real signals is obtained. This data set is clearer and more accurate than the original data, providing a better basis for subsequent analysis. In summary, the first channel effectively identifies and processes the noise features in the mass spectrometry data through data preprocessing and noise identification algorithms based on machine learning. This process not only improves the accuracy and consistency of the data, but also provides strong support for subsequent qualitative or quantitative analysis, chemical identification and other tasks, further improving the efficiency and accuracy of mass spectrometry data analysis.
[0023] In a preferred embodiment, Figure 2As shown, the second channel is used to extract signal features of chemical substances in mass spectrometry data, and the execution steps include: performing signal feature extraction on the intermediate mass spectrometry data to generate a chemical substance signal feature set; performing feature selection, dimensionality reduction and enhancement processing on the chemical substance signal feature set; based on the processed chemical substance signal feature set, using a signal reconstruction algorithm to reconstruct the original mass spectrometry data to generate a target mass spectrometry data set to be determined that contains the signal features of chemical substances.
[0024] Furthermore, signal feature extraction technology is used to extract signal features related to chemical substances from intermediate mass spectrometry data. These features may include signal intensity, frequency, shape, duration, etc., which can reflect the unique performance of different chemical substances in mass spectrometry data. Through feature extraction, a set of chemical substance signal features is generated, which contains all signal features related to chemical substances extracted from the data. However, subsequent analysis directly from this feature set may face problems such as large amount of calculation and feature redundancy. Therefore, the features need to be further processed. Feature selection is to remove those features that do not contribute much or are irrelevant to the analysis results to reduce the amount of calculation and improve the efficiency of analysis. Statistical methods, machine learning algorithms or expert experience are used to select the most important features. Dimensionality reduction is to convert high-dimensional feature space into low-dimensional space while retaining the key information in the original data as much as possible. This helps to reduce the complexity of data and improve the performance of the algorithm. Feature enhancement is to enhance the expressiveness of features through some techniques (such as filtering, smoothing, transformation, etc.), making it easier to be recognized and used by subsequent algorithms. After feature processing, an optimized set of chemical substance signal features is obtained. Next, the signal reconstruction algorithm is used to recombine these features into a new mass spectrometry data set. The purpose of the signal reconstruction algorithm is to restore or enhance the chemical signal in the original mass spectrometry data based on the extracted features. In this process, some mathematical models or optimization algorithms may be used to ensure that the reconstructed data not only retains the authenticity of the original data, but also highlights the characteristics of the chemical signal. Finally, a set of target mass spectrometry data containing the signal characteristics of chemical substances is generated. The data in this set not only removes noise, but also highlights the chemical signal, providing a better basis for subsequent analysis and identification. In summary, the second channel effectively extracts the chemical signal characteristics in the mass spectrometry data through the steps of signal feature extraction, feature processing (selection, dimensionality reduction and enhancement) and signal reconstruction, and generates a set of target mass spectrometry data containing these characteristics. This process not only improves the efficiency of data analysis, but also provides strong support for subsequent tasks such as chemical identification and quantitative analysis, and further improves the accuracy and reliability of mass spectrometry data analysis.
[0025] In a preferred embodiment, an attention mechanism is configured in the second channel, and the execution steps include: constructing a convolutional attention mechanism based on the time-mass-to-charge ratio two-dimensional distribution characteristics of the chemical substance signal characteristics, wherein the convolutional attention mechanism includes a spatial attention sublayer and a channel attention sublayer; reshaping the dimension of the generated target mass spectrum data set to form a three-dimensional tensor input of the number of samples × the number of mass-to-charge ratio points × the number of feature channels; generating an attention weight matrix of the same dimension as the input three-dimensional tensor through dual-branch calculation of the attention layer, and completing the attention mechanism configuration, wherein the matrix element value range is [0,1].
[0026] Optionally, in the second channel, in order to further improve the ability to extract chemical signal features, an attention mechanism is introduced. This mechanism can enhance the model's attention to key features while suppressing unimportant features, thereby improving the accuracy and efficiency of mass spectrometry data analysis. First, based on the time-mass-to-charge ratio two-dimensional distribution characteristics of chemical signal features, a convolutional attention mechanism is constructed. This mechanism consists of two core sublayers: a spatial attention sublayer and a channel attention sublayer. Among them, the spatial attention sublayer uses a 3×3 convolution kernel with a step size of 1 to capture the spatial information in the feature map (i.e., the two-dimensional representation of mass spectrometry data). This sublayer can identify which areas in the feature map are important and which areas are unimportant, and generate corresponding spatial attention weights. These weights will enhance the features of important areas while suppressing the features of unimportant areas. In order to capture the dependencies between different feature channels, the channel attention sublayer adopts a combination of global average pooling and full connection. First, global average pooling is performed on each channel of the feature map to obtain a channel description vector. Then, this vector passes through a fully connected layer to generate attention weights for each channel, which reflect the contribution of different channels to the final classification or regression task. Next, the generated target mass spectrometry data set is reshaped. Specifically, the data is reshaped into a three-dimensional tensor of the number of samples × the number of mass-to-charge ratio points × the number of feature channels. This three-dimensional tensor contains both the spatial information of the mass spectrometry data (the number of mass-to-charge ratio points) and the information of different feature channels. After the three-dimensional tensor is input, it is sent to the attention layer for dual-branch calculation. This layer takes into account both spatial attention and channel attention, and generates a spatial attention weight matrix and a channel attention weight matrix through two sublayers. In the spatial attention sublayer, a 3×3 convolution kernel is used to perform a convolution operation on the input three-dimensional tensor to obtain a spatial attention weight matrix, each element of which represents the importance of the feature at the corresponding position. In the channel attention sublayer, global average pooling is first performed on each channel to obtain a channel description vector. Then, this vector passes through a fully connected layer to generate a channel attention weight matrix, each element of which represents the importance of the corresponding channel. Finally, the spatial attention weight matrix and the channel attention weight matrix are element-wise multiplied to obtain the final attention weight matrix. The element value range of this matrix is [0,1], which indicates the importance of the features at the corresponding position or channel. After obtaining the attention weight matrix, it is applied to the input three-dimensional tensor, and important features are enhanced through element-wise multiplication, while unimportant features are suppressed. In this way, the configuration of the attention mechanism is completed, and an attention-weighted three-dimensional tensor output is obtained. By introducing the attention mechanism, the second channel can more effectively extract the signal characteristics of chemical substances in mass spectrometry data.The combined use of the spatial attention sublayer and the channel attention sublayer enables the model to simultaneously capture the spatial information in the feature map and the dependencies between channels, thereby improving the accuracy and efficiency of the analysis.
[0027] In a preferred embodiment, the chemical substance signal feature is dynamically identified to obtain an attention weight map, and the execution steps include: in the spatial attention sublayer branch, generating a spatial attention map by cross-channel weighting , where H is the number of mass-to-charge ratio points and W is the number of scanning time points; in the channel attention sublayer branch, the channel weight vector is generated by calculating the autocorrelation between feature channels. , C is the number of feature channels; perform tensor outer product of S and C to form a three-dimensional attention weight map , to perform dynamic focusing of time-mass-to-charge ratio-characteristic three-dimensional space.
[0028] In detail, we first focus on the spatial attention sub-layer branch. In this branch, the feature distribution of mass spectrometry data in the two-dimensional space composed of mass-to-charge ratio (H) and scanning time point (W) is considered. In order to generate the spatial attention map, a cross-channel weighted method is adopted. Specifically, for the input three-dimensional tensor (number of samples × number of mass-to-charge ratio points H × number of scanning time points W × number of feature channels C), spatial features are first extracted independently on each feature channel. This can be achieved through convolution operations or other spatial feature extraction methods. Then, these extracted spatial features are weighted and summed across channels to generate a two-dimensional spatial attention map, which reflects which areas are important and which areas are not important in the two-dimensional space composed of mass-to-charge ratio and scanning time point. Next, in the channel attention sub-layer branch, attention is paid to the dependencies between different feature channels. In order to generate the channel weight vector, the autocorrelation calculation method is adopted. Specifically, the input three-dimensional tensor is first globally averaged pooled at each scanning time point and mass-to-charge ratio point to obtain a description vector about the feature channel. Then, the autocorrelation of this description vector is calculated, that is, the dot product between it and itself is calculated, and a channel weight vector is generated through a nonlinear activation function (such as ReLU or Sigmoid). This vector reflects the importance of different feature channels. After obtaining the spatial attention map and the channel weight vector, they are combined by tensor outer product to form a three-dimensional attention weight map. The dimension of this map is the same as the input three-dimensional tensor (a simplified representation of the number of samples × the number of mass-to-charge ratio points H × the number of scanning time points W × the number of feature channels C, but in the attention weight map, the last dimension is replaced by the attention weight), but it contains richer information. Specifically, for each element in the input tensor, a corresponding weight value can be found in the three-dimensional attention weight map. This weight value reflects the importance of the element in the three-dimensional space of time-mass-to-charge ratio-feature. In this way, dynamic focusing on the three-dimensional space is achieved, allowing the model to capture key feature information more accurately. In summary, by constructing the spatial attention sublayer and the channel attention sublayer and cleverly combining them to generate a three-dimensional attention weight map, dynamic recognition of chemical signal characteristics is achieved. This process not only improves the accuracy of feature extraction, but also enables the model to capture key feature information more accurately, providing strong support for subsequent analysis and recognition tasks.
[0029] In a preferred embodiment, based on the weighted processing of the attention weight map, a target mass spectrum data set is generated, and the execution step includes: performing element-by-element dot multiplication of the three-dimensional attention weight map and the input three-dimensional tensor, and the calculation formula is: , where i is the sample index, j is the number of mass-to-charge ratio points, and k is the number of characteristic channels; weighted summation is performed along the characteristic channel dimension to generate a two-dimensional enhanced mass spectrum. The calculation formula is: ,in, The channel weight coefficient is optimized through training; local extreme value detection is performed on the enhanced mass spectrum, peak top coordinates are extracted and adjacent peak clusters are merged to form a target mass spectrum data set.
[0030] Specifically, based on the obtained three-dimensional attention weight map, the input three-dimensional tensor is further weighted to generate a more accurate and valuable target mass spectrometry data set. This process combines the advantages of the attention mechanism and significantly improves the analysis quality of mass spectrometry data by enhancing key features and suppressing noise. First, the three-dimensional attention weight map is element-by-element dot multiplied with the input three-dimensional tensor. This step aims to weight each element in the input tensor according to the instructions of the attention weight map, thereby highlighting important feature information and suppressing unimportant information. The specific calculation formula is: , i represents the sample index, j represents the mass-to-charge ratio point number (i.e., m / z value), and k is the number of characteristic channels. A weighted three-dimensional tensor is obtained through this element-by-element dot multiplication. Next, the weighted three-dimensional tensor is weighted and summed along the characteristic channel dimension to generate a two-dimensional enhanced mass spectrum. This step aims to fuse the information of different characteristic channels to form a more comprehensive and clear mass spectrum representation. The specific calculation formula is: ,in, are channel weight coefficients optimized through training, which reflect the importance of different feature channels to the final mass spectrum. The result of weighted summation is a two-dimensional matrix, whose rows represent the number of mass-to-charge ratio points (m / z values) and columns represent the corresponding intensity values. After obtaining the two-dimensional enhanced mass spectrum, local extreme value detection is performed to extract the peak top coordinates in the spectrum. These peak top coordinates correspond to the significant features in the mass spectrum, which are crucial for subsequent peak identification, qualitative or quantitative analysis of substances. Here, the peak top coordinates are understood as a binary of (m / z value, intensity). Since there may be multiple adjacent peaks in the mass spectrum, these peaks may originate from different isotopes or different molecular forms of the same chemical substance. Therefore, after extracting the peak top coordinates, it is also necessary to merge the adjacent peak clusters. This step aims to merge those peaks with similar positions and intensities into a peak cluster, thereby simplifying the subsequent mass spectrometry data analysis process. In summary, the target mass spectrometry data set was generated by weighted processing based on the three-dimensional attention weight map. This process significantly improved the analysis quality and reliability of mass spectrometry data by highlighting key features, suppressing noise, and merging adjacent peak clusters.
[0031] In a preferred embodiment, a confidence assessment is performed on the target mass spectrometry data set, and if the peak intensity variation coefficient is greater than 30%, the noise re-check process of the first channel is triggered.
[0032] Furthermore, before further analysis and application of the target mass spectrometry data set, a confidence assessment step needs to be performed to ensure the accuracy and reliability of the data. This step mainly focuses on the stability of the peak intensity in the mass spectrum, and the confidence of the data is evaluated by calculating the coefficient of variation of the peak intensity. First, the peak intensity variation coefficient is calculated for each peak (or peak cluster) in the target mass spectrometry data set. The peak intensity variation coefficient is an indicator to measure the degree of peak intensity fluctuation, which is calculated by dividing the standard deviation of the peak intensity by the average value of the peak intensity. This indicator can reflect the stability and consistency of the peak intensity in the mass spectrum. In the calculation process, multiple factors may be considered, such as repeated measurements under different experimental conditions, measurement errors of different instruments, etc., to ensure that the calculation of the coefficient of variation can fully reflect the true situation of the data. After obtaining the peak intensity variation coefficient, it is compared with a preset threshold (30% in this case). If the coefficient of variation of a peak is greater than this threshold, it is considered that the intensity fluctuation of this peak is large, the confidence of the data is low, and further inspection and processing may be required. When the intensity variation coefficient of a peak exceeds the threshold, a noise re-check process for the first channel is further triggered. The purpose of this step is to conduct a more detailed inspection of the first channel to identify and eliminate possible sources of noise or errors. The noise re-inspection process may include multiple steps, such as re-checking experimental conditions, recalibrating instruments, optimizing data processing algorithms, etc. These steps are designed to improve the accuracy and reliability of the data and ensure that subsequent analysis and applications can be based on high-quality data. In summary, by performing a confidence assessment on the target mass spectrometry data set and calculating the peak intensity coefficient of variation to evaluate the stability and consistency of the data, peaks with low confidence can be discovered and processed in a timely manner. When the peak intensity coefficient of variation exceeds the preset threshold, the noise re-inspection process for the first channel is triggered to ensure the accuracy and reliability of the data. This process not only helps to improve the analytical quality of mass spectrometry data, but also provides strong support for subsequent tasks such as substance identification and quantitative analysis.
[0033] In a preferred embodiment, an asymmetric least squares method is used to perform dynamic baseline correction on the data in the target mass spectrometry data set, and the execution steps include: segmenting the mass spectrometry data of each sample, and dividing the high confidence region and the low confidence region based on the attention weight map; using the asymmetric least squares method to fit the baseline in the high confidence region, and switching to an adaptive penalized spline model in the low confidence region to dynamically adjust the smoothing parameters to match the local noise level; connecting each segmented baseline through cubic spline interpolation, generating a global baseline curve and subtracting it from the original data.
[0034] For example, in mass spectrometry data analysis, baseline correction is a crucial step, which directly affects the accuracy and reliability of subsequent tasks such as substance identification and quantitative analysis. In order to dynamically correct the data in the target mass spectrometry data set, a method combining asymmetric least squares and adaptive penalized spline model is adopted. First, the mass spectrometry data of each sample is segmented so that each segment can be baseline corrected separately in the subsequent steps. Segmentation can be based on the characteristics or analysis requirements of mass spectrometry data, aiming to divide the data into multiple relatively independent and easy-to-process parts. Then, each segment is divided into high confidence regions and low confidence regions using the attention weight map generated previously. In this example, high confidence regions are defined as regions with weights ≥ 0.7, which usually contain relatively stable and characteristic mass spectrometry peaks; while low confidence regions are regions with weights < 0.7, which may be greatly affected by factors such as noise and instrument errors. In the high confidence region, asymmetric least squares is used for baseline fitting. Asymmetric least squares is a commonly used baseline correction method. It introduces an asymmetric factor in the objective function to preferentially fit the baseline underestimation area, thereby avoiding the problem of overestimation of the baseline. In this example, the smoothing coefficient λ=10³ is set to suppress the interference of high-frequency noise, and the asymmetric factor p=0.001 is set to ensure the accuracy and stability of the baseline fitting. In the low confidence region, due to the high noise level and unclear features, the adaptive penalty spline model is switched to perform baseline fitting. The adaptive penalty spline model can dynamically adjust the smoothing parameters according to the local noise level, thereby ensuring the smoothness of the baseline while retaining the useful information in the data as much as possible. In the low confidence region, the smoothing parameters are adjusted to match the local noise level to ensure the accuracy and reliability of the baseline fitting. Finally, the segmented baselines are connected by cubic spline interpolation to generate a global baseline curve. Cubic spline interpolation is a commonly used interpolation method that can ensure the smoothness and continuity of the baseline curve. After the global baseline curve is generated, this baseline curve is subtracted from the original data to obtain the corrected mass spectrometry data. In summary, a method combining asymmetric least squares method and adaptive penalty spline model is used to perform dynamic baseline correction on the data in the target mass spectrometry data set. Through data segmentation, confidence region division, baseline fitting and smoothing parameter adjustment, global baseline curve generation and data subtraction, the baseline correction of mass spectrometry data is achieved, providing high-quality data support for subsequent tasks such as substance identification and quantitative analysis. This process improves the accuracy and reliability of mass spectrometry data analysis.
[0035] In a preferred embodiment, the signal-to-noise ratio processing is performed in combination with wavelet transform, and the execution steps include: selecting the Symlets wavelet basis to perform a 5-layer wavelet decomposition on the baseline-corrected data to separate high-frequency noise and low-frequency signal components; performing adaptive threshold processing on the detail coefficients of the 1st to 3rd layers, and selectively retaining the approximate coefficients of the 4th to 5th layers based on the attention weight; if the mean attention weight of the corresponding peak coordinate interval is ≥0.6, the coefficient is retained; otherwise, the coefficient is attenuated to 20% of the original value; performing wavelet reconstruction to generate a denoised mass spectrum, and outputting the mass spectrum detection result.
[0036] Specifically, in order to further optimize the mass spectrometry data after baseline correction, combined with wavelet transform technology, through fine decomposition, processing and reconstruction steps, the high-frequency noise is effectively suppressed and the low-frequency signal components are retained. First, the Symlets wavelet basis is selected as the decomposition tool because the Symlets wavelet has good symmetry and regularity and is suitable for processing mass spectrometry data containing complex features. Then, the mass spectrometry data after baseline correction is subjected to a 5-layer wavelet decomposition. This step aims to decompose the original data into signals with different frequency components, in which the high-frequency components mainly contain noise, while the low-frequency components contain useful signal information. Through the 5-layer decomposition, a low-frequency approximation coefficient (representing the low-frequency signal at the bottom layer) and 5 groups of high-frequency detail coefficients (representing the high-frequency noise components of each layer) are obtained. Adaptive threshold processing is performed on the high-frequency detail coefficients of layers 1 to 3. The purpose of this step is to remove the high-frequency noise in these layers while retaining the useful information in the signal as much as possible. Adaptive threshold processing can dynamically adjust the threshold according to the characteristics of the data, thereby ensuring the denoising effect while avoiding signal distortion caused by over-processing. For the low-frequency approximation coefficients of layers 4 to 5, a selective retention strategy based on attention weights is adopted. The purpose of this step is to finely screen and retain the low-frequency signal components based on the information of the attention weight map. Specifically, the mean of the attention weights corresponding to the peak coordinate interval is calculated. If the mean is ≥ 0.6, it means that the signal in this interval is relatively stable and has obvious characteristics, so the approximate coefficient of this interval is retained; if the mean is < 0.6, it means that the signal in this interval may be greatly affected by noise or error, so the approximate coefficient of this interval is attenuated to 20% of the original value to suppress the interference of potential noise. After completing the processing of detail coefficients and approximate coefficients, the wavelet reconstruction step is performed. This step aims to reassemble the processed coefficients into complete mass spectrometry data to obtain a denoised mass spectrum. Wavelet reconstruction successfully retains the useful information in the low-frequency signal components, while effectively removing the interference of high-frequency noise. The denoised mass spectrum obtained in the end has a higher signal-to-noise ratio and better signal recognition ability. Finally, the denoised mass spectrum is used as input to perform subsequent mass spectrometry detection steps, such as substance identification and quantitative analysis. These steps can be performed based on high-quality data, thereby obtaining more accurate and reliable detection results. In summary, the wavelet transform technology is combined with the sophisticated decomposition, processing and reconstruction steps to achieve the signal-to-noise ratio processing of the mass spectrometry data after baseline correction.
[0037] The data analysis system based on the mass spectrometry detection platform provided by the embodiment of the present invention has at least the following technical effects:
[0038] 1. By building a dual-channel mass spectrometry detection model, the first channel focuses on the identification and separation of noise features, while the second channel focuses on the extraction of chemical substance signal features. The introduction of the attention mechanism in the second channel can dynamically identify and weight the chemical substance signal features, generate a high-confidence target mass spectrometry data set, significantly improve the signal-to-noise ratio and signal recognition ability of mass spectrometry data, and provide a more accurate and reliable data basis for subsequent analysis.
[0039] 2. The target mass spectrometry data set is dynamically baseline corrected using the asymmetric least squares method, and the high confidence region and low confidence region are divided into two regions using the attention weight map, achieving accurate fitting and adjustment of the baseline. In the high confidence region, the asymmetric least squares method is used to effectively suppress high-frequency noise; in the low confidence region, the adaptive penalty spline model is switched to dynamically adjust the smoothing parameters to match the local noise level, which not only improves the accuracy of the baseline correction, but also enhances the overall quality of the mass spectrometry data.
[0040] 3. In the signal-to-noise ratio processing stage, the wavelet transform technology is combined to separate high-frequency noise and low-frequency signal components through 5-layer wavelet decomposition. Adaptive threshold processing is performed on the detail coefficients, and the approximate coefficients are selectively retained based on the attention weight. If the mean of the attention weight corresponding to the peak coordinate interval is high, the coefficient is retained; otherwise, the coefficient is attenuated to suppress potential noise. This strategy maximizes the retention of useful information in the mass spectrometry data while ensuring the denoising effect, further improving the accuracy and reliability of the mass spectrometry detection results.
[0041] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data analysis system based on a mass spectrometry detection platform, characterized in that: The implementation steps include: A dual-channel mass spectrometry detection model is constructed, wherein the first channel is used to identify noise features in mass spectrometry data, and the second channel is used to extract chemical substance signal features in mass spectrometry data, and the mass spectrometry data passes through the first and second channels in sequence; An attention mechanism is configured in the second channel to dynamically identify the signal characteristics of the chemical substance to obtain an attention weight map, and a target mass spectrometry data set is generated based on weighted processing of the attention weight map; The asymmetric least square method is used to perform dynamic baseline correction on the data in the target mass spectrometry data set, and the signal-to-noise ratio is processed in combination with wavelet transform to obtain the mass spectrometry detection result.
2. The system according to claim 1, characterized in that The first channel is used to identify noise features in mass spectrometry data, and the execution steps include: Preprocessing the input mass spectrometry data, including data cleaning, format conversion and normalization; Based on the noise recognition algorithm of machine learning, the preprocessed mass spectrometry data is subjected to feature extraction and classification to identify and separate the noise features in the data and obtain intermediate mass spectrometry data.
3. The system according to claim 2, characterized in that The second channel is used to extract the signal characteristics of chemical substances in the mass spectrometry data. The execution steps include: Extracting signal features from the intermediate mass spectrum data to generate a set of chemical substance signal features; Performing feature selection, dimensionality reduction and enhancement processing on the chemical substance signal feature set; Based on the processed chemical substance signal feature set, a signal reconstruction algorithm is used to reconstruct the original mass spectrometry data to generate a target mass spectrometry data set containing chemical substance signal features.
4. The system according to claim 3, characterized in that The attention mechanism is configured in the second channel, and the execution steps include: Based on the time-mass-to-charge ratio two-dimensional distribution characteristics of chemical substance signal characteristics, a convolutional attention mechanism is constructed, wherein the convolutional attention mechanism includes a spatial attention sublayer and a channel attention sublayer; Reshaping the generated target mass spectrum data set to be determined to form a three-dimensional tensor input of number of samples × number of mass-to-charge ratio points × number of feature channels; Through the dual-branch calculation of the attention layer, an attention weight matrix with the same dimension as the input three-dimensional tensor is generated to complete the attention mechanism configuration, and the value range of the matrix elements is [0,1].
5. The system according to claim 4, characterized in that Dynamically identifying the chemical substance signal feature to obtain an attention weight map, the execution steps include: In the spatial attention sublayer branch, a spatial attention map is generated by cross-channel weighting , where H is the mass-to-charge ratio point number, and W is the scanning time point number; In the channel attention sublayer branch, the channel weight vector is generated by calculating the autocorrelation between feature channels. , where C is the number of feature channels; Perform tensor outer product of S and C to form a three-dimensional attention weight map , to perform dynamic focusing of time-mass-to-charge ratio-characteristic three-dimensional space.
6. The system according to claim 5, characterized in that Based on the weighted processing of the attention weight map, a target mass spectrum data set is generated, and the execution steps include: The three-dimensional attention weight map is element-wise multiplied with the input three-dimensional tensor, and the calculation formula is: , where i is the sample index, j is the number of mass-to-charge ratio points, and k is the number of feature channels; A weighted sum is performed along the characteristic channel dimension to generate a two-dimensional enhanced mass spectrum. The calculation formula is: ,in, is the channel weight coefficient optimized through training; The enhanced mass spectrum is subjected to local extrema detection, peak top coordinates are extracted, and adjacent peak clusters are merged to form a target mass spectrum data set.
7. The system according to claim 6, characterized in that The execution step also includes: performing a confidence assessment on the target mass spectrum data set, and if the peak intensity variation coefficient is greater than 30%, triggering a noise re-check process for the first channel.
8. The system of claim 1, wherein: The data in the target mass spectrometry data set are dynamically corrected by using an asymmetric least squares method, and the execution steps include: The mass spectrometry data of each sample is segmented into high confidence areas and low confidence areas based on the attention weight map; In the high confidence region, an asymmetric least squares method is used to fit the baseline, and in the low confidence region, an adaptive penalty spline model is used to dynamically adjust the smoothing parameters to match the local noise level; The segmented baselines were connected by cubic spline interpolation to generate a global baseline curve which was subtracted from the original data.
9. The system according to claim 8, characterized in that The signal-to-noise ratio processing is performed in combination with wavelet transform, and the execution steps include: The Symlets wavelet basis was selected to perform a 5-layer wavelet decomposition on the baseline-corrected data to separate high-frequency noise and low-frequency signal components; Adaptive threshold processing is performed on the detail coefficients of the 1st to 3rd layers, and selective retention based on attention weights is performed on the approximate coefficients of the 4th to 5th layers; If the mean of the attention weight corresponding to the peak coordinate interval is ≥ 0.6, the coefficient is retained; otherwise, the coefficient is decayed to 20% of the original value; Perform wavelet reconstruction to generate a denoised mass spectrum and output the mass spectrum detection results.
Citation Information
Patent Citations
A method for improving the sensitivity of a time-of-flight mass spectrometer
CN109103067A
Multi-mode audio-visual separation method and system fusing two-channel attention mechanism
CN116110423A
Radar signal modulation type identification method based on wavelet transform phase mutation detection
CN116166936A
Communication mode switching method for intelligent dual-mode interphone
CN118042509A
Self-adaptive edge micro-seismic event real-time detection method and system
CN118094208A
Cited By
Water quality element channel positioning and matching method based on mass spectrum peak cluster sequence
CN122598824A
A chromatograph data processing method, system, terminal and storage medium
CN122651939A