An analysis method and system for improving mass spectrum quality

Through multi-scale Gaussian fitting and asymmetric Gaussian function model combined with mass offset correction and convolutional neural network, the problems of inaccurate peak recognition and insufficient resolution in mass spectrometry analysis are solved, and more accurate peak shape description and data resolution improvement are achieved.

CN120121695BActive Publication Date: 2025-08-12TIANJIN ZHIPU INSTR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510599677.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-12
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the existing mass spectrometry analysis methods, inaccurate peak identification, difficulty in fitting asymmetric peaks, and insufficient mass offset and resolution, resulting in inaccurate quantitative analysis.

Method used

Multi-scale Gaussian fitting technology is used to identify potential peak positions, use asymmetric Gaussian function models for peak fitting, and build a mass offset correction model, and combine it with convolutional neural network for super-resolution processing.

Benefits of technology

The measurement accuracy of peak position and shape characteristics is improved, the m/z value deviation is corrected, and the resolution and quantitative analysis capabilities of the mass spectrogram are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120121695B_ABST
    Figure CN120121695B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of mass spectrometry analysis, and more specifically, to an analysis method and system for improving the quality of mass spectrometry. The method comprises the following steps: scanning a sample using a mass spectrometer, collecting raw mass spectrometry data, and preprocessing the raw mass spectrometry data; identifying potential peak positions based on the preprocessed raw mass spectrometry data using a multi-scale Gaussian fitting technique, and screening candidate peaks by calculating the signal-to-noise ratio of each peak; fitting each screened candidate peak using an asymmetric Gaussian function model to describe the candidate peak shape; constructing a mass offset correction model to correct the mass spectrometry data, and performing super-resolution processing on the corrected mass spectrometry data using a convolutional neural network model. The present invention employs a multi-scale Gaussian fitting technique, a signal-to-noise ratio screening mechanism, and an asymmetric Gaussian function model for peak fitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mass spectrometry analysis, and in particular to an analysis method and system for improving the quality of mass spectrometry. Background Art

[0002] In mass spectrometry analysis, accurately distinguishing true signal peaks from background noise is a challenge, especially in low signal-to-noise ratio conditions. Traditional methods have difficulty accurately identifying potential peak positions. Many peaks in mass spectra exhibit asymmetric shapes, and traditional symmetrical Gaussian or other symmetrical models cannot accurately describe these peak morphologies, resulting in inaccurate estimates of peak parameters (height, width, and center position). Due to factors such as instrument drift or environmental changes, the actual measured m / z values may deviate from the theoretical values, affecting the accuracy of quantitative analysis. In addition, the resolution of raw mass spectrometry data is sometimes insufficient to clearly distinguish close peaks or detect weak signals. Therefore, an analytical method and system for improving the quality of mass spectra is provided. Summary of the Invention

[0003] The purpose of the present invention is to provide an analysis method and system for improving the quality of mass spectrometry images, so as to solve the problems of inaccurate peak identification, difficulty in fitting asymmetric peaks, mass offset and insufficient resolution existing in the existing mass spectrometry analysis methods proposed in the above background technology.

[0004] To achieve the above object, the present invention provides an analysis method for improving the quality of a mass spectrum, comprising:

[0005] S1. Scan the sample using a mass spectrometer, collect raw mass spectrum data, and preprocess the raw mass spectrum data;

[0006] S2. Based on the preprocessed raw mass spectrum data, potential peak positions are identified using multi-scale Gaussian fitting technology, and candidate peaks are screened out by calculating the signal-to-noise ratio of each peak;

[0007] S3. For each candidate peak screened, an asymmetric Gaussian function model is used to fit the candidate peak shape, an asymmetric factor is introduced into the asymmetric Gaussian function model, and a progressive optimization process is used to gradually decouple the asymmetric Gaussian function model parameters;

[0008] S4. Construct a mass offset correction model to correct the mass spectrum data, and use a convolutional neural network model to perform super-resolution processing on the corrected mass spectrum data.

[0009] As a further improvement of the present technical solution, in S2, based on the pre-processed original mass spectrum data, identifying potential peak positions by multi-scale Gaussian fitting technology includes the following steps:

[0010] S2.1. Select a scale parameter and use a Gaussian filter to smooth the data for the selected scale.

[0011] S2.2. For each scale of the smoothed data, calculate the first and second derivatives of the smoothed data. The peak position corresponds to the point where the first derivative is zero and the second derivative changes from positive to negative.

[0012] S2.3. Locate potential peak positions by examining the zero crossing points of the first-order derivative and the sign changes of the second-order derivative in its vicinity.

[0013] As a further improvement of the present technical solution, in S2, the candidate peaks are screened out by calculating the signal-to-noise ratio of each peak, comprising the following steps:

[0014] S2.4. For each identified potential peak position, extract the maximum intensity value of the peak as its "peak intensity";

[0015] S2.5. Select a local area window around each potential peak;

[0016] S2.6. Use the standard deviation of the intensity within the selected region as an estimate of the “background noise intensity” corresponding to the peak;

[0017] S2.7. Calculate the signal-to-noise ratio for each potential peak based on the peak intensity and background noise intensity obtained above.

[0018] S2.8. Compare the signal-to-noise ratio of each potential peak with its threshold m. If the signal-to-noise ratio of a peak is greater than or equal to the set threshold m, then the peak is considered a candidate peak.

[0019] Otherwise, exclude it.

[0020] As a further improvement of the present technical solution, in S3, for each candidate peak screened out, an asymmetric Gaussian function model is used for fitting to describe the shape of the candidate peak, including the following steps:

[0021] S3.1. For each candidate peak, select a fitting interval based on the position and width of the candidate peak;

[0022] S3.2. Construct an asymmetric Gaussian function model;

[0023] S3.3, providing initial estimates for the asymmetric Gaussian function model parameters based on the peak intensity, center position, and width of the candidate peak;

[0024] S3.4. Use nonlinear least squares method to fit the asymmetric Gaussian function;

[0025] S3.5. During the fitting process, the model parameters are gradually adjusted. After the fitting is completed, the optimized parameter values are obtained. During the fitting process, dynamic weights are assigned to each data point. The weights of the left data points are additionally subject to the left standard deviation dependency factor. The parameters are gradually decoupled using a progressive optimization process.

[0026] S3.6. Based on the fitting results, extract the key morphological parameters of the candidate peaks.

[0027] As a further improvement of the present technical solution, in S3.2, constructing an asymmetric Gaussian function model includes the following steps:

[0028] S3.21. Calculate the integral area of the left and right half peaks and define the asymmetry factor , adjust the standard deviation ratio on the left and right sides according to the asymmetry factor to generate the initial standard deviation parameter;

[0029] S3.22. Decompose the single standard deviation parameter of the traditional Gaussian function into the left half-peak standard deviation and the right half-peak standard deviation to form a piecewise function structure;

[0030] The piecewise function is:

[0031] ;

[0032] in, Represents the theoretical value of the output of the asymmetric Gaussian function model; represents the height of the peak; Indicates the center position of the peak; represents the standard deviation of the left half peak; represents the standard deviation of the right half peak; Indicates location.

[0033] As a further improvement of the present technical solution, in S3.5, a progressive optimization process is used to gradually decouple parameters, including the following steps:

[0034] S3.51, by fixing the standard deviation of the right half peak, reducing parameter coupling, and prioritizing the optimization of peak height , Peak Center and left half peak standard deviation ;

[0035] S3.52. Based on the first stage, the second stage is carried out, the standard deviation of the left half peak is fixed, and the peak height is further optimized. , Peak Center and right half peak standard deviation ;

[0036] S3.53. Using the results of the first two stages as initial values, a global joint optimization is performed using nonlinear least squares, and a termination criterion driven by physical constraints and residual gradients is introduced.

[0037] Among them, the physical constraints are used to set the upper and lower bounds of the left half peak standard deviation and the right half peak standard deviation;

[0038] The termination criterion driven by residual gradient consists of two parallel conditions: the residual relative change rate and the normalized gradient norm.

[0039] As a further improvement of the present technical solution, in S4, constructing the mass offset correction model involves the following steps:

[0040] S4.1. Construct a dynamic internal standard library;

[0041] S4.2. Align the theoretical scan time of the internal standard in the internal standard library with the actual detection time series using a dynamic time warping algorithm;

[0042] S4.3. With the current scan point as the center, select all internal standard peaks within a ±10 Da window and calculate the offset between the measured m / z value and the theoretical value;

[0043] S4.4. Use robust regression to fit a quadratic polynomial correction model to describe how mass shift changes with m / z value;

[0044] S4.5. Use the local window correction model to perform mass compensation on each detected characteristic peak and generate a corrected mass spectrum.

[0045] As a further improvement of the present technical solution, in S4, super-resolution processing is performed on the corrected mass spectrum data using a convolutional neural network model, comprising the following steps:

[0046] S4.6. extracting a plurality of local window data from the corrected mass spectrum, each window comprising an m / z value and a corresponding intensity value, and preprocessing each intensity value;

[0047] S4.7. Construct a convolutional neural network model to improve the resolution of the data within the window, and define the mean square error as the loss function;

[0048] S4.8. Divide the corrected mass spectrum dataset into a training set, a validation set, and a test set, and train the convolutional neural network model using the training set;

[0049] S4.9. Input the low-resolution mass spectrometry data of each window into the trained convolutional neural network model, and output the corresponding high-resolution data, where data greater than a threshold value a is considered high-resolution, and data less than a threshold value a is considered low-resolution;

[0050] S4.10. Perform denormalization on the generated high-resolution data to restore the original intensity range, and splice the high-resolution data of all windows into a complete mass spectrum.

[0051] As a further improvement of the present technical solution, the convolutional neural network model constructed in S4.7 is used to improve the resolution of the data in the window, including the following steps:

[0052] S4.71. The input layer is designed to have low-resolution window data in the shape of ,in, Indicates the height of the window. Indicates the width of the window, Indicates the number of channels;

[0053] S4.72. Use multiple convolutional layers to extract features from low-resolution data.

[0054] S4.73, use convolutional layers to map low-resolution features to high-resolution feature space;

[0055] S4.74, use deconvolution operation to enlarge the feature map to the target resolution;

[0056] S4.75, the last layer uses a 1×1 convolution kernel to generate the final high-resolution output, the shape is ,in, Indicates the height of the target resolution, Indicates the width of the target resolution, Indicates the number of channels for output data.

[0057] On the other hand, the present invention provides an analysis system for improving the quality of mass spectrum, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein the processor executes the computer program to implement any one of the above-mentioned analysis methods for improving the quality of mass spectrum.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. This analytical method and system for improving mass spectrometry quality utilizes multi-scale Gaussian fitting technology, a signal-to-noise ratio screening mechanism, and an asymmetric Gaussian function model for peak fitting, enabling more accurate identification and description of potential peak positions and shape characteristics in mass spectra. In particular, the introduction of an asymmetry factor and a progressive optimization process effectively address the inability of traditional symmetric Gaussian models to accurately describe asymmetric peaks. This improves the measurement accuracy of peak position, height, width, and asymmetry, while reducing noise interference and the likelihood of false positive peak misidentification.

[0060] 2. This analytical method and system for improving mass spectrometry quality utilizes a mass offset correction model combined with a robust regression algorithm to correct for m / z value deviations, and employs a convolutional neural network (CNN) for super-resolution processing. This not only corrects mass spectrometry data errors caused by instrument drift or environmental changes, but also significantly improves mass spectrometry data resolution. This enables clearer differentiation of overlapping peaks and detection of weak signals in complex sample analysis, enhancing the ability to qualitatively and quantitatively analyze sample components, and providing more accurate and reliable mass spectrometry data support for scientific research and practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION

[0062] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0063] Example 1: Please refer to Figure 1 As shown, this embodiment provides an analysis method for improving the quality of a mass spectrum, comprising the following steps:

[0064] S1. Scan the sample using a mass spectrometer, collect raw mass spectrum data, and preprocess the raw mass spectrum data;

[0065] S2. Based on the preprocessed raw mass spectrum data, potential peak positions are identified using multi-scale Gaussian fitting technology, and candidate peaks are screened out by calculating the signal-to-noise ratio of each peak;

[0066] In this embodiment, based on the pre-processed raw mass spectrum data, potential peak positions are identified by multi-scale Gaussian fitting technology, including the following steps:

[0067] S2.1. Select a scale parameter (the parameter determines the width of the Gaussian function; each scale corresponds to a peak of a different width in the data). For the selected scale, smooth the data using a Gaussian filter (this step helps highlight peaks of a width corresponding to that scale at that scale).

[0068] S2.2. For each scale of the smoothed data, calculate the first and second derivatives of the smoothed data. The peak position corresponds to the point where the first derivative is zero and the second derivative changes from positive to negative.

[0069] S2.3. Locate potential peak positions by checking the zero crossing point of the first-order derivative and the sign change of the second-order derivative in its neighboring area. In addition, a threshold can be set to filter out peaks with too low intensity to avoid detecting artifacts caused by noise.

[0070] Furthermore, candidate peaks are screened out by calculating the signal-to-noise ratio of each peak, including the following steps:

[0071] S2.4. For each potential peak position identified, extract the maximum intensity value of the peak as its "peak intensity". This step directly determines the molecular part in the subsequent signal-to-noise ratio calculation;

[0072] S2.5. Select a local area window around each potential peak (the area around the peak is defined as the range of 5 mass units on both sides of the peak), and try to avoid including other significant peaks in this area.

[0073] S2.6. Use the standard deviation of the intensity within the selected region as the estimated "background noise intensity" corresponding to the peak. This step provides data support for the denominator part of the signal-to-noise ratio calculation;

[0074] S2.7. Calculate the signal-to-noise ratio for each potential peak based on the peak intensity and background noise intensity obtained above.

[0075] S2.8. Compare the signal-to-noise ratio of each potential peak with its threshold m. If the signal-to-noise ratio of a peak is greater than or equal to the set threshold m, then the peak is considered a candidate peak.

[0076] Otherwise, it is excluded and considered as a noise artifact.

[0077] S3. For each candidate peak screened, an asymmetric Gaussian function model is used to fit the candidate peak shape, an asymmetric factor is introduced into the asymmetric Gaussian function model, and a progressive optimization process is used to gradually decouple the asymmetric Gaussian function model parameters;

[0078] In this embodiment, for each candidate peak screened out, an asymmetric Gaussian function model is used for fitting to describe the shape of the candidate peak, including the following steps:

[0079] S3.1. For each candidate peak, select a fitting interval (a certain range around the peak center, ±5 Da) based on the position and width of the candidate peak. This interval contains the complete peak shape and avoids interference from adjacent peaks.

[0080] S3.2. Construct an asymmetric Gaussian function model;

[0081] The construction of an asymmetric Gaussian function model includes the following steps:

[0082] S3.21. Calculate the integral area of the left and right half peaks and define the asymmetry factor , adjust the standard deviation ratio on the left and right sides according to the asymmetry factor to generate the initial standard deviation parameter;

[0083] Among them, the asymmetry factor for: , the integrated areas of the left and right half-peaks were extracted by numerical integration method;

[0084] Adjust the standard deviation ratio on the left and right sides according to the asymmetry factor to generate the initial standard deviation parameter:

[0085] , ;

[0086] in, Represents the initial value of the symmetric standard deviation of the multi-scale Gaussian fitting output. The ratio of the left and right standard deviations is dynamically adjusted by the difference in the integrated area to avoid the initial parameter deviation caused by relying solely on the half-height width conversion; represents the initial left half peak standard deviation; represents the initial right half peak standard deviation;

[0087] S3.22. Decompose the single standard deviation parameter of the traditional Gaussian function into the left half-peak standard deviation and the right half-peak standard deviation, forming a piecewise function structure. The piecewise function is the asymmetric Gaussian function model.

[0088] The piecewise function is:

[0089] ;

[0090] in, Represents the theoretical value of the output of the asymmetric Gaussian function model; represents the height of the peak; Indicates the center position of the peak; represents the standard deviation of the left half peak; represents the standard deviation of the right half peak; Indicates location;

[0091] S3.3. Provide initial estimates of the asymmetric Gaussian function model parameters based on the peak intensity, center position, and width of the candidate peak (model parameters include: maximum intensity of the candidate peak, center position of the candidate peak (determined by multi-scale Gaussian fitting), and standard deviations to the left and right of the peak);

[0092] S3.4. Fit an asymmetric Gaussian function using nonlinear least squares (Levenberg-Marquardt algorithm) to minimize the sum of squared errors between the experimental data and the fitted curve. (Nonlinear least squares is an optimization technique for fitting model parameters to observed data. It finds the best parameter estimate by minimizing the sum of squared errors between the model predictions and the actual observations.)

[0093] S3.5. During the fitting process, the model parameters are gradually adjusted. After the fitting is completed, the optimized parameter values are obtained. These parameters can directly reflect the characteristics of the candidate peaks. Dynamic weights are given to each data point during the fitting process. ( , Represents the residual sensitivity coefficient (default value , represents the intensity of the experimentally measured data point, Indicates Position, the theoretical intensity value calculated by the asymmetric Gaussian function model, Represents a specific m / z value in the mass spectrum, represents the data point index), and applies an additional left standard deviation dependency factor to the weight of the left data point (this mechanism is an adaptive residual weight allocation mechanism that prioritizes optimizing the fitting accuracy in areas with significant peak asymmetry while suppressing the interference of noise points), and uses a progressive optimization process to gradually decouple the parameters;

[0094] Among them, the weight of the left data point Additional left standard deviation dependence factor = (The left standard deviation dependency factor is a factor used to adjust the weight of the data points on the left side of the asymmetric peak. By dynamically adjusting the weight of the left data points, the fitting process of the asymmetric peak can be more effectively optimized, especially the fitting quality of the left part of the peak can be improved):

[0095] ;

[0096] Furthermore, a progressive optimization process is used to gradually decouple the parameters, including the following steps:

[0097] In an asymmetric Gaussian function model, there are complex interdependencies between multiple parameters (peak height, center position, left-half peak standard deviation, and right-half peak standard deviation). Progressive optimization allows certain parameters to be fixed at different stages, simplifying the optimization problem and avoiding local minima or overfitting caused by strong coupling between parameters. For complex asymmetric peak shapes, optimizing all parameters simultaneously at once can lead to instability, especially when the initial parameters are poorly chosen. Progressive optimization, by gradually adjusting the parameters, makes each optimization step relatively simple and clear, thereby improving the stability and convergence of the overall optimization process.

[0098] S3.51, by fixing the standard deviation of the right half peak, reducing parameter coupling, and prioritizing the optimization of peak height , Peak Center and left half peak standard deviation , which can better capture the details on the left side of the peak;

[0099] Specifically: Fixed (The initial value is generated by the dynamic asymmetry factor);

[0100] Using the nonlinear least squares algorithm, the following objective function is optimized:

[0101] ;

[0102] Among them, the weight Generated by the residual weight adaptive allocation mechanism; represents the intensity of the experimentally measured data points;

[0103] Iterate optimization until the residual change rate is less than the preset threshold (<1%);

[0104] Initially optimized , and record the fitting error and termination condition status;

[0105] S3.52. Based on the first stage, the second stage is carried out, the standard deviation of the left half peak is fixed, and the peak height is further optimized. , Peak Center and right half peak standard deviation , to ensure that the entire peak shape can be accurately described;

[0106] Specifically: Fixed Optimize results for the first phase;

[0107] Again, the nonlinear least squares method is used to optimize the following objective function, adjusting :

[0108] ;

[0109] in, , Represents the weight of the data point on the right;

[0110] Iterative optimization is performed until the residual change rate is less than a stricter threshold (<0.5%);

[0111] Further optimized , and record the fitting error and termination condition status;

[0112] S3.53. Using the results of the first two stages as initial values, a global joint optimization is performed using nonlinear least squares, and a termination criterion driven by physical constraints and residual gradients is introduced.

[0113] Physical constraints are used to set upper and lower bounds for the left and right half-peak standard deviations, preventing extremely narrow or wide peaks that cannot be detected by the instrument. Limiting the left / right half-peak standard deviation ratio reflects the detector's actual response limit to peak broadening in different directions. By limiting the parameter range, the likelihood of the optimization algorithm falling into a local minimum can be reduced, thereby improving optimization stability and convergence speed.

[0114] The residual gradient-driven termination criterion is a dynamic condition used to determine whether the optimization algorithm has reached convergence. It determines when to terminate the iteration by jointly monitoring the objective function residual (fitting error) and its gradient. It includes two parallel conditions: the relative rate of change of the residual (the L2 norm of the residual vector (fitting error) between two adjacent iterations is less than a threshold, indicating that the optimization has entered a stable plateau) and the normalized gradient norm (the overall gradient of the objective function with respect to the parameter approaches zero, indicating that the extreme value has been approached). This dynamic termination criterion can flexibly adjust the termination timing based on actual data and fitting conditions, making it more intelligent and efficient than termination methods with a fixed number of iterations.

[0115] Specifically: take the results of the first two stages as the initial values, that is: ;

[0116] in, The optimization results from the first and second stages are taken respectively;

[0117] All parameters ( ) to perform joint optimization, the goal is to minimize the sum of squared errors:

[0118] ;

[0119] After each iteration, the boundary conditions embedded in the physical constraints ( ) to perform projection correction on the parameters, represents the initial standard deviation;

[0120] Iterate the optimization until the termination criterion driven by the residual gradient is met:

[0121] (Condition on the relative rate of change of residuals),

[0122] (Normalized gradient norm condition);

[0123] in, Represents the threshold value of the residual relative change rate condition, , represents the threshold of the normalized gradient norm condition, , Indicates the The residual vector of the iteration, Indicates the number of iterations;

[0124] Finalized parameters These are the key morphological parameters of the candidate peaks;

[0125] S3.6. Based on the fitting results, extract the key morphological parameters of the candidate peaks (including peak height, peak position, peak width and asymmetry).

[0126] S4. Constructing a mass offset correction model to correct the mass spectrum data, and using a convolutional neural network (CNN) model to perform super-resolution processing on the corrected mass spectrum data;

[0127] In this embodiment, constructing a mass offset correction model involves the following steps:

[0128] S4.1. Inject reference compounds with known m / z values (mass-to-charge ratios) before mass spectrometry detection to construct a dynamic internal standard library. Each internal standard should meet the following requirements: absolute mean mass error ≤ 2 ppm, peak intensity coefficient of variation ≤ 5%;

[0129] S4.2. Align the theoretical scan times of the internal standards in the internal standard library with the actual detection time series using the dynamic time warping algorithm (extract the theoretical scan times and actual detection time series of the internal standards, use the dynamic time warping (DTW) algorithm to calculate the optimal alignment path between the theoretical and actual time series, and adjust the actual detection time series based on the alignment results to align it with the theoretical time series).

[0130] S4.3. Centered on the current scan point, select all internal standard peaks within a ±10 Da window and calculate the offset between the measured m / z value (mass-to-charge ratio) and the theoretical value (from the calibrated mass spectrum, extract data for multiple local windows according to step S4.5, where each window contains the m / z value and corresponding intensity value).

[0131] The deviation between the measured m / z value (mass-to-charge ratio) and the theoretical value for:

[0132] ;

[0133] in, Indicates that the internal standard is indexed;

[0134] S4.4. Use robust regression (RANSAC algorithm) to fit a quadratic polynomial correction model to describe how mass shift changes with m / z value;

[0135] The quadratic polynomial correction model is:

[0136] ;

[0137] in, Represents the regression residual, iterative elimination ; Represents the coefficients of the quadratic polynomial correction model, which is used to describe the variation of mass shift with m / z value;

[0138] S4.5. Use the local window correction model to perform mass compensation on each detected characteristic peak and generate a corrected mass spectrum.

[0139] Furthermore, a convolutional neural network (CNN) model is used to perform super-resolution processing on the corrected mass spectrum data, including the following steps:

[0140] S4.6. Extract multiple local window data (±10 Da window) from the corrected mass spectrum, where each window contains the m / z value and the corresponding intensity value, and perform preprocessing (normalization and data enhancement) on each intensity value.

[0141] S4.7. Construct a convolutional neural network (CNN) model to improve the resolution of the data within the window and define the mean squared error (MSE) as the loss function (MSE is a statistic that measures the difference between the predicted value and the true value. It evaluates the accuracy of the model by calculating the average of the squares of the differences between the predicted and the true values).

[0142] The convolutional neural network (CNN) model constructed to improve the resolution of data within the window includes the following steps:

[0143] S4.71. The input layer is designed to have low-resolution window data in the shape of ,in, Indicates the height of the window. Indicates the width of the window, Indicates the number of channels (in this example, it is 1 because the mass spectrometry data is a single channel);

[0144] S4.72. Use multiple convolutional layers to extract features from low-resolution data.

[0145] S4.73, use deeper convolutional layers to map low-resolution features to high-resolution feature space;

[0146] S4.74, use deconvolution operation to enlarge the feature map to the target resolution;

[0147] S4.75, the last layer uses a 1×1 convolution kernel to generate the final high-resolution output, the shape is ,in, Indicates the height of the target resolution, Indicates the width of the target resolution, Indicates the number of channels of the output data, that is, the number of feature channels contained in each pixel;

[0148] S4.8. Divide the corrected mass spectrum dataset into a training set, a validation set, and a test set. Use the training set to train a convolutional neural network (CNN) model to ensure that the model can accurately predict high-resolution data.

[0149] S4.9. Input the low-resolution mass spectrometry data of each window into a trained convolutional neural network (CNN) model, and output the corresponding high-resolution data, including finer m / z values and intensity values. Values greater than a threshold a are considered high-resolution, while values less than a threshold a are considered low-resolution.

[0150] S4.10. Perform denormalization on the generated high-resolution data to restore the original intensity range, and splice the high-resolution data of all windows into a complete mass spectrum.

[0151] Example 2: This example provides an analysis system for improving the quality of mass spectrometry, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement any one of the above-described analysis methods for improving the quality of mass spectrometry.

[0152] The basic principles, main features, and advantages of the present invention are shown and described above. It should be understood by those skilled in the art that the present invention is not limited to the above-described embodiments. The above-described embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention claimed.

Claims

1. An analytical method for improving the quality of a mass spectrum, characterized in that: The following steps are involved: S1. Scan the sample using a mass spectrometer, collect raw mass spectrum data, and preprocess the raw mass spectrum data; S2. Based on the preprocessed raw mass spectrum data, potential peak positions are identified using multi-scale Gaussian fitting technology, and candidate peaks are screened out by calculating the signal-to-noise ratio of each peak; S3. For each candidate peak screened, an asymmetric Gaussian function model is used to fit the candidate peak shape, an asymmetric factor is introduced into the asymmetric Gaussian function model, and a progressive optimization process is used to gradually decouple the asymmetric Gaussian function model parameters; S4. constructing a mass offset correction model to correct the mass spectrum data, and using a convolutional neural network model to perform super-resolution processing on the corrected mass spectrum data; In S3, for each candidate peak screened out, an asymmetric Gaussian function model is used for fitting to describe the shape of the candidate peak, including the following steps: S3.

1. For each candidate peak, select a fitting interval based on the position and width of the candidate peak; S3.

2. Construct an asymmetric Gaussian function model; S3.3, providing initial estimates for the asymmetric Gaussian function model parameters based on the peak intensity, center position, and width of the candidate peak; S3.

4. Use nonlinear least squares method to fit the asymmetric Gaussian function; S3.

5. During the fitting process, the model parameters are gradually adjusted. After the fitting is completed, the optimized parameter values are obtained. During the fitting process, dynamic weights are assigned to each data point. The weights of the left data points are additionally subject to the left standard deviation dependency factor. The parameters are gradually decoupled using a progressive optimization process. S3.

6. Extract key morphological parameters of candidate peaks based on the fitting results; In S3.5, a progressive optimization process is used to gradually decouple parameters, including the following steps: S3.51, by fixing the standard deviation of the right half peak, reducing parameter coupling, and prioritizing the optimization of peak height , Peak Center and left half peak standard deviation ; S3.

52. Based on the first stage, the second stage is carried out, the standard deviation of the left half peak is fixed, and the peak height is further optimized. , Peak Center and right half peak standard deviation ; S3.

53. Using the results of the first two stages as initial values, a global joint optimization is performed using nonlinear least squares, and a termination criterion driven by physical constraints and residual gradients is introduced. Among them, the physical constraints are used to set the upper and lower bounds of the left half peak standard deviation and the right half peak standard deviation; The termination criterion driven by residual gradient consists of two parallel conditions: the residual relative change rate and the normalized gradient norm.

2. The method for improving the quality of mass spectra according to claim 1, wherein: In S2, based on the pre-processed raw mass spectrum data, potential peak positions are identified by multi-scale Gaussian fitting technology, including the following steps: S2.

1. Select a scale parameter and use a Gaussian filter to smooth the data for the selected scale. S2.

2. For each scale of the smoothed data, calculate the first and second derivatives of the smoothed data. The peak position corresponds to the point where the first derivative is zero and the second derivative changes from positive to negative. S2.

3. Locate potential peak positions by examining the zero crossing points of the first-order derivative and the sign changes of the second-order derivative in its vicinity.

3. The method for improving the quality of mass spectra according to claim 1, wherein: In S2, candidate peaks are screened out by calculating the signal-to-noise ratio of each peak, including the following steps: S2.

4. For each identified potential peak position, extract the maximum intensity value of the peak as its "peak intensity"; S2.

5. Select a local area window around each potential peak; S2.

6. Use the standard deviation of the intensity within the selected region as an estimate of the "background noise intensity" corresponding to the peak; S2.

7. Calculate the signal-to-noise ratio for each potential peak based on the peak intensity and background noise intensity obtained above. S2.

8. Compare the signal-to-noise ratio of each potential peak with its threshold m. If the signal-to-noise ratio of a peak is greater than or equal to the set threshold m, then the peak is considered a candidate peak. Otherwise, exclude it.

4. The method for improving the quality of mass spectra according to claim 1, wherein: In S3.2, constructing an asymmetric Gaussian function model includes the following steps: S3.

21. Calculate the integral area of the left and right half peaks and define the asymmetry factor , adjust the standard deviation ratio on the left and right sides according to the asymmetry factor to generate the initial standard deviation parameter; S3.

22. Decompose the single standard deviation parameter of the traditional Gaussian function into the left half-peak standard deviation and the right half-peak standard deviation to form a piecewise function structure; The piecewise function is: ; in, Represents the theoretical value of the output of the asymmetric Gaussian function model; represents the height of the peak; Indicates the center position of the peak; represents the standard deviation of the left half peak; represents the standard deviation of the right half peak; Indicates location.

5. The method for improving the quality of mass spectra according to claim 1, wherein: In S4, constructing the mass offset correction model involves the following steps: S4.

1. Construct a dynamic internal standard library; S4.

2. Align the theoretical scan time of the internal standard in the internal standard library with the actual detection time series using a dynamic time warping algorithm; S4.

3. With the current scan point as the center, select all internal standard peaks within a ±10 Da window and calculate the offset between the measured m / z value and the theoretical value; S4.

4. Use robust regression to fit a quadratic polynomial correction model to describe how mass shift changes with m / z value; S4.

5. Use the local window correction model to perform mass compensation on each detected characteristic peak and generate a corrected mass spectrum.

6. The method for improving the quality of mass spectra according to claim 1, wherein: In S4, super-resolution processing is performed on the corrected mass spectrum data using a convolutional neural network model, comprising the following steps: S4.

6. extracting a plurality of local window data from the corrected mass spectrum, each window comprising an m / z value and a corresponding intensity value, and preprocessing each intensity value; S4.

7. Construct a convolutional neural network model to improve the resolution of the data within the window, and define the mean square error as the loss function; S4.

8. Divide the corrected mass spectrum dataset into a training set, a validation set, and a test set, and train the convolutional neural network model using the training set; S4.

9. Input the low-resolution mass spectrometry data of each window into the trained convolutional neural network model, and output the corresponding high-resolution data, where data greater than a threshold value a is considered high-resolution, and data less than a threshold value a is considered low-resolution; S4.

10. Perform denormalization on the generated high-resolution data to restore the original intensity range, and splice the high-resolution data of all windows into a complete mass spectrum.

7. The method for improving the quality of mass spectra according to claim 6, wherein: In S4.7, the convolutional neural network model constructed to improve the resolution of the data in the window includes the following steps: S4.

71. The input layer is designed to have low-resolution window data in the shape of ,in, Indicates the height of the window. Indicates the width of the window, Indicates the number of channels; S4.

72. Use multiple convolutional layers to extract features from low-resolution data. S4.73, use convolutional layers to map low-resolution features to high-resolution feature space; S4.74, use deconvolution operation to enlarge the feature map to the target resolution; S4.75, the last layer uses a 1×1 convolution kernel to generate the final high-resolution output, the shape is ,in, Indicates the height of the target resolution, Indicates the width of the target resolution, Indicates the number of channels for output data.

8. An analysis system for improving the quality of a mass spectrum, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the analysis method for improving the quality of mass spectrum as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for extracting radioactive energy spectrum heavy peak area through asymmetric Gaussian function

    CN118671817A