A mass spectrometry detection method and system based on data twin
By adopting a data twin-based mass spectrometry detection method in the mass spectrometry detection system, multi-stage intelligent preprocessing and machine learning model training, the problems of complex and low efficiency of data processing in the existing mass spectrometry detection system are solved, and efficient and accurate mass spectrometry data analysis is achieved.
Patent Information
- Application Number
- CN202510104949.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The data processing and analysis of existing mass spectrometry detection systems have problems such as large amount of data and complex processing, resulting in large resource investment and low efficiency.
The mass spectrometry detection method based on data twins is adopted to obtain raw data through mass spectrometry detection equipment, perform multi-stage intelligent preprocessing, build a data twin model, and use machine learning algorithms for training and analysis.
It improves the quality and reliability of mass spectrometry data, improves the accuracy and efficiency of data analysis, and realizes real-time, accurate and efficient mass spectrometry data analysis.
Smart Images

Figure CN119538016B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mass spectrometry data processing, and in particular to a mass spectrometry detection method and system based on data twins. Background Art
[0002] Mass spectrometry is an important technology widely used in chemical analysis, environmental monitoring, biomedicine and other fields. The data processing and analysis of existing mass spectrometry detection systems have problems such as large data volume and complex processing. It requires a lot of manpower, material or financial resources, and has relatively large limitations. Therefore, it is urgent to provide a real-time, accurate and efficient mass spectrometry data analysis method to improve the utilization efficiency and analysis effect of mass spectrometry detection data. Summary of the invention
[0003] In view of the problems existing in the above-mentioned prior art, the present invention provides a mass spectrometry detection method and system based on data twins. The technical solution is as follows:
[0004] In a first aspect, a mass spectrometry detection method based on data twin is provided, comprising the following steps:
[0005] Obtaining original mass spectrum data of the sample through a mass spectrometry detection device;
[0006] Perform multi-stage intelligent mass spectrometry data preprocessing on raw mass spectrometry data;
[0007] Using a machine learning algorithm to train the processed mass spectrometry data to construct a data twin model, wherein the data twin model is used to obtain mass spectrometry analysis results;
[0008] New mass spectrometry data is input into the trained data twin model to obtain mass spectrometry analysis results.
[0009] In some embodiments, the multi-stage intelligent mass spectrometry data preprocessing of the raw mass spectrometry data includes:
[0010] The original mass spectrometry data is subjected to multi-level and refined denoising through wavelet transform and adaptive filtering;
[0011] Eliminate instrument and sample differences through dual standardization and intelligent calibration methods.
[0012] In some embodiments, the multi-level and refined denoising process of the raw mass spectrometry data by wavelet transform and adaptive filtering includes:
[0013] Processing the original mass spectrum data by wavelet transform to obtain first preprocessed data;
[0014] The first preprocessed data is processed by a least mean square adaptive filtering algorithm to obtain second preprocessed data; in the least mean square adaptive filtering algorithm, the update formula for each weight of the filter is: w(n+1)=w(n)+μe(n)x(n), wherein w(n) is the filter weight, μ is the learning rate, e(n) is the error, and x(n) is the input signal.
[0015] In some embodiments, the dual standardization and intelligent correction method is used to eliminate the differences between instruments and samples, including: using a pre-trained model to automatically identify and correct outliers in the data, specifically including:
[0016] Performing global and local standardization on the mass spectrometry data, and comprehensively determining the standardization result of the mass spectrometry data based on the double standardization data;
[0017] Use intelligent correction to further adjust the standardized data to eliminate the influence of systematic errors and instrument drift;
[0018] Standardized processing methods for mass spectrometry data, including:
[0019] The mass spectrometry data of different samples or different groups are locally standardized respectively; the mass spectrometry data of all samples or all groups are globally standardized to obtain local standardized data and global standardized data corresponding to the mass spectrometry data of the same sample or the same group, wherein the standardized method is calculated based on the difference between the current data and the minimum value of the data and the ratio of the difference between the maximum and minimum values of the data;
[0020] Analyze the difference between the local standardized data and the global standardized data of the mass spectrometry data of the same sample or the same group to obtain the double standardized data difference data of the mass spectrometry data of each sample or each group;
[0021] Based on the statistical distribution data of the double-standardized data difference data corresponding to the mass spectral data of all samples or all groups, the standardization result of the mass spectral data of each sample or group is determined.
[0022] In some embodiments, the data twin model includes a spatiotemporal feature extraction module, a frequency domain feature calculation unit, and a machine learning algorithm integration module, wherein the spatiotemporal feature extraction module and the frequency domain feature calculation unit are connected in parallel and synchronously receive input mass spectrometry data, and the spatiotemporal feature extraction module and the frequency domain feature calculation unit are connected in parallel and then cascaded to the machine learning algorithm integration module;
[0023] The training process of the data twin model includes:
[0024] The spatiotemporal features are extracted from the input mass spectrometry data using a spatiotemporal feature extraction module, and the frequency domain features are extracted using a frequency domain feature calculation unit through fast Fourier transform FFT; the network model parameters in the spatiotemporal feature extraction module are to be trained, and are continuously iterated and updated during the training process of the data twin model, and the frequency domain feature calculation unit does not contain parameters to be trained;
[0025] Concatenate the spatiotemporal features and frequency domain features of the mass spectrometry data to form mass spectrometry features;
[0026] The mass spectrometry features are input into the machine learning algorithm integration module for comprehensive decision-making to obtain the mass spectrometry data analysis results. The network model parameters in the machine learning algorithm integration module are to be trained and continuously iterated and updated during the training process of the data twin model.
[0027] In some embodiments, in the data twin model, the machine learning algorithm integration module includes at least one machine learning model and a weight distribution model; at least one machine learning model synchronously receives input data of the machine learning algorithm integration module in parallel, and at least one machine learning model is cascaded with the weight distribution model after being connected in parallel; the network model parameters of each of the machine learning models are to be trained, and the network model parameters of the weight distribution model are to be trained, and are continuously iterated and updated during the training process of the data twin model.
[0028] In some embodiments, during the training of the data twin model, the machine learning algorithm integration module performs the following process:
[0029] Obtain the mass spectrometry data analysis results output by each machine learning model;
[0030] Calculating the error between the mass spectrometry data analysis results of each output and the theoretical mass spectrometry data analysis results; and reversely updating the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis results;
[0031] Calculate the theoretical value of the weight of each machine learning model based on the error corresponding to each machine learning model, and obtain the theoretical value of the weight vector for all machine learning models in the integrated module;
[0032] Inputting the mass spectrometry data analysis results output by each machine learning model into the weight distribution model to obtain the weight vector output by the weight distribution model;
[0033] The network model parameters of the weight allocation model are reversely updated based on the error between the output weight vector and the theoretical value of the weight vector.
[0034] In some embodiments, during the training of the data twin model, the machine learning algorithm integration module performs the following process:
[0035] Obtain the mass spectrometry data analysis results output by each machine learning model;
[0036] Calculating the error between the mass spectrometry data analysis results of each output and the theoretical mass spectrometry data analysis results; and reversely updating the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis results;
[0037] The mass spectrometry data analysis results output by each machine learning model are input into the weight allocation model for weighted fusion, and the summarized mass spectrometry data analysis results are output; the network model parameters of the weight allocation model are reversely updated based on the error between the summarized mass spectrometry data analysis results and the theoretical mass spectrometry data analysis results.
[0038] In some embodiments, the spatiotemporal feature extraction module is based on a cascade of a convolutional neural network (CNN) and a long short-term memory (LSTM) network; the data processing steps of the spatiotemporal feature extraction module include: mass spectrum data is input through a data input layer, enters a convolutional layer to extract local spatiotemporal features of the mass spectrum data, enters a pooling layer for feature dimension reduction, enters an LSTM layer to extract temporal features of the mass spectrum data, generates a feature vector of the mass spectrum data through a fully connected layer, and outputs through an output layer;
[0039] The frequency domain features include: frequency component, power spectrum density, spectrum entropy, center frequency, bandwidth, harmonic component, wavelet coefficient, and instantaneous frequency.
[0040] In a second aspect, a mass spectrometry detection system based on data twin is provided, the system comprising:
[0041] A raw data acquisition unit, used to acquire raw mass spectrometry data of a sample through a mass spectrometry detection device;
[0042] A preprocessing unit, used for performing multi-stage intelligent mass spectrometry data preprocessing on raw mass spectrometry data;
[0043] A model training unit, used to train the processed mass spectrometry data using a machine learning algorithm to construct a data twin model, wherein the data twin model is used to obtain mass spectrometry analysis results;
[0044] The data analysis unit is used to input new mass spectrometry data into the trained data twin model to obtain mass spectrometry analysis results.
[0045] According to a third aspect, an electronic device is provided, the electronic device comprising:
[0046] processor;
[0047] a memory for storing processor-executable instructions;
[0048] Wherein, the processor implements the mass spectrometry detection method described in the first aspect above by running the executable instructions.
[0049] In a fourth aspect, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the mass spectrometry detection method described in the first aspect are implemented.
[0050] A mass spectrometry detection method and system based on data twins of the present invention have the following beneficial effects: improving the quality and reliability of data, improving the accuracy of data, and analyzing mass spectrometry data through a data twin model to improve the efficiency of mass spectrometry data analysis by performing multi-stage intelligent mass spectrometry data preprocessing on raw mass spectrometry data. The data standardization process is comprehensively determined by local data distribution characteristics and global data distribution characteristics. When a certain local data distribution feature or some local data distribution features appear less frequently in the global data, the effect of its local features is increased to obtain preferred standardized data. If a certain local data distribution feature or some local data distribution features appear more frequently in the global data, the importance of its local features is relatively small. The standardization method in the embodiment of the present application fully retains the personalized characteristics of mass spectrometry data on the basis of global standardization, and enhances the characterization ability of the features extracted in the feature extraction stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of a flow chart of a mass spectrometry detection method based on data twins provided in an embodiment of the present application;
[0052] Figure 2 Schematic diagram of the mass spectrometry data preprocessing process in the embodiment of the present application;
[0053] Figure 3 It is a structural diagram of the data twin model in the embodiment of the present application;
[0054] Figure 4 It is a flow chart of the training process of the machine learning algorithm integration module in the embodiment of the present application;
[0055] Figure 5 It is another flowchart of the training process of the machine learning algorithm integration module in the embodiment of the present application;
[0056] Figure 6 It is a structural schematic diagram of a mass spectrometry detection system based on data twins provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0058] See also Figure 1, a mass spectrometry detection method based on data twins provided in an embodiment of the present application comprises the following steps:
[0059] Step 1, obtaining original mass spectrum data of the sample through a mass spectrometry detection device;
[0060] Step 2, performing multi-stage intelligent mass spectrometry data preprocessing on the raw mass spectrometry data;
[0061] Step 3, using a machine learning algorithm to train the processed mass spectrometry data to build a data twin model, wherein the data twin model is used to obtain mass spectrometry analysis results;
[0062] Step 4: Input the new mass spectrometry data into the trained data twin model to obtain the mass spectrometry analysis results.
[0063] In the embodiments of the present application, mass spectrometry data is identified and classified, and specific chemical components in the sample are identified and quantitatively analyzed. The quality and reliability of the data are improved by performing multi-stage intelligent mass spectrometry data preprocessing on the raw mass spectrometry data, and the accuracy of the data is improved. The mass spectrometry data is analyzed through a data twin model to improve the efficiency of mass spectrometry data analysis.
[0064] See also Figure 2 In one embodiment, in the above step 2, the raw mass spectrum data is subjected to multi-stage intelligent mass spectrum data preprocessing, including:
[0065] Step 21, performing multi-level and refined denoising processing on the original mass spectrum data through wavelet transform and adaptive filtering;
[0066] Step 22, eliminate the differences between instruments and samples through double standardization and intelligent calibration methods.
[0067] In one embodiment, the above step 21, performing multi-level and refined denoising processing on the original mass spectrum data by wavelet transform and adaptive filtering, includes:
[0068] Step 211, using wavelet transform to process the original mass spectrum data to obtain first preprocessed data;
[0069] Step 212, the first preprocessed data is processed by a least mean square adaptive filtering algorithm to obtain second preprocessed data; in the least mean square adaptive filtering algorithm, the update formula of each filter weight is: w(n+1)=w(n)+μe(n)x(n), wherein w(n) is the filter weight, μ is the learning rate, e(n) is the error, and x(n) is the input signal.
[0070] In the embodiment of the present application, a wavelet transform combined with an adaptive filtering method is used to remove noise from mass spectrometry data and improve the signal-to-noise ratio. A multi-level noise removal algorithm is introduced, wavelet transform is combined with adaptive filtering technology to remove noise at different scales, and signal features are enhanced through nonlinear filtering. Compared with traditional simple filters, this method can better retain the key features in mass spectrometry signals while reducing noise.
[0071] In the embodiment of the present application, the wavelet transform formula is as follows:
[0072]
[0073] Where x(t) is the input signal, is the wavelet function, a is the scale parameter, b is the translation parameter, represents the complex conjugate of the wavelet function.
[0074] x(t): In a mass spectrometry system, x(t) represents the raw data output by the mass spectrometer. These data are usually time series signals that reflect the response intensity of the sample at different time points.
[0075] : Mother wavelet. The mother wavelet is used to generate the basis function of other wavelets. It generates a complete set of basis functions through scaling and translation. Common mother wavelets include Morlet wavelet, Mexican Hat wavelet and Daubechies wavelet. In mass spectrometry data processing, choosing a suitable mother wavelet function is very important for extracting local features of the signal and denoising.
[0076] a: Scale parameter. The scale parameter controls the expansion and contraction of the wavelet function. Small scale corresponds to the high frequency part (details), and large scale corresponds to the low frequency part (overall trend). In mass spectrometry data analysis, by adjusting the scale parameter, the characteristics of the signal in different frequency ranges can be analyzed to help identify different chemical components and noise.
[0077] b: Translation parameter. The translation parameter controls the position of the wavelet function on the time axis. By translating the wavelet, the entire signal can be analyzed locally. In mass spectrometry data processing, the translation parameter is used to locate signal features, such as peak position and noise position.
[0078] W(a,b): Wavelet coefficients. Wavelet coefficients are the wavelet transform results of the original signal at a specific scale and position, reflecting the characteristics of the signal at that scale and position. In mass spectrometry data analysis, by analyzing the distribution of wavelet coefficients, important features and abnormal points in the signal can be identified.
[0079] In the embodiment of the present application, in the least mean square adaptive filtering algorithm, w(n) is the filter coefficient vector, w(n+1) is the updated filter coefficient vector, and the updated filter coefficient vector is the new coefficient adjusted according to the input signal and the error signal, indicating the state of the filter after the current iteration. In mass spectrometry data processing, the updated filter coefficient vector is used for the next iteration to gradually optimize the filter performance to improve the effect of signal processing.
[0080] In one embodiment, the above step 22, through the dual standardization and intelligent correction method, eliminates the differences between instruments and samples, including: using a pre-trained model to automatically identify and correct outliers in the data, specifically including:
[0081] Step 221, performing global standardization and local standardization on the mass spectrum data, and comprehensively determining the standardization result of the mass spectrum data based on the double standardization data;
[0082] Step 222, using intelligent calibration to further adjust the standardized data to eliminate the influence of systematic errors and instrument drift; the calibration method uses linear regression, uses standard samples or known data sets to train the calibration model, and the calibration model uses a linear regression model y=mx+b. Then use the trained calibration model to calibrate the new data.
[0083] It should be noted that, in the embodiments of the present application, before performing wavelet transform and adaptive filtering, it also includes: selecting an appropriate preliminary processing method according to the type of mass spectrometry data, for example, for protein mass spectrometry data, using baseline correction and peak detection; for metabolite mass spectrometry data, using smoothing filtering and peak recognition.
[0084] The above step 221, the method for standardizing the mass spectrum data, includes:
[0085] Step 2211, performing local standardization on mass spectral data of different samples or different groups respectively; performing global standardization on mass spectral data of all samples or all groups, to obtain local standardized data and global standardized data corresponding to mass spectral data of the same sample or the same group, wherein the standardization method is calculated based on the difference between the current data and the minimum value of the data and the ratio of the difference between the maximum and minimum values of the data;
[0086] Step 2212, analyzing the difference between the local standardized data and the global standardized data of the mass spectrum data of the same sample or the same group, and obtaining double standardized data difference data of the mass spectrum data of each sample or each group;
[0087] Step 2213, based on the statistical distribution data of the double-standardized data difference data corresponding to the mass spectral data of all samples or all groups, determine the standardization result of the mass spectral data of each sample or group.
[0088] Specifically, the method determines the standardized result of the mass spectral data of each sample or group based on the statistical distribution data of the double-standardized data difference data corresponding to the mass spectral data of all samples or all groups, including: determining the standardized result of the mass spectral data of each sample or group based on the size distribution of the double-standardized data difference data and the number of samples or groups belonging to different segments of the difference data.
[0089] Specifically, the method of determining the standardized result of the mass spectral data of each sample or group based on the statistical distribution data of the double standardized data difference data corresponding to the mass spectral data of all samples or all groups includes:
[0090] Step A1, arranging the difference data in order of size, segmenting the difference data based on the maximum and minimum values of the difference data, and obtaining a plurality of continuous non-overlapping difference data segments;
[0091] Step A2, counting the number of samples or groups corresponding to different difference data segments;
[0092] Step A3, obtaining a reference weight of local standardized data and a reference weight of global standardized data preset under a reference condition of differential data distribution, wherein the reference condition is that the differential data segments corresponding to the mass spectrometry data of all samples or groups are consistent;
[0093] Step A4, determining the weight correction coefficient of the local standardized data of the mass spectrum data for the sample or group corresponding to the single difference data segment based on the stability of the number of samples or groups corresponding to different difference data segments, the number of samples or groups corresponding to a single difference data segment, and the mean of the difference values of the single difference data segment under actual conditions; the smaller the stability of the number of samples or groups corresponding to the different difference data segments, the greater the weight of the local standardized data of the corresponding mass spectrum data; the smaller the number of samples or groups corresponding to the single difference data segment, the greater the weight of the local standardized data of the corresponding mass spectrum data; the greater the mean of the difference values of the single difference data segment, the greater the weight of the local standardized data of the corresponding mass spectrum data;
[0094] Step A5, weighted fusion based on local standardization and global standardization is used as the preferred standardized data.
[0095] It can be understood that the difference data segments corresponding to the mass spectrometry data of all samples or groups under the reference conditions are consistent. At this time, the difference between the local standardized data and the global standardized data of each mass spectrometry data is small. At this time, the weight of the local standardized data is small. In extreme cases, the weight of the local standardized data can be 0, and the preferred standardized data of the global standardized data corresponding to the mass spectrometry data can be directly used. When the distribution of the difference data of the double standardized data is uneven, it is necessary to increase the weight of the local standardized data to highlight the characteristics of the individual mass spectrometry data. At this time, the increase in the weight of the local standardized data is related to the stability of the number of samples or groups corresponding to different difference data segments, the number of samples or groups corresponding to a single difference data segment, and the mean of the difference value of a single difference data segment. The smaller the stability of the number of samples or groups corresponding to different difference data segments, the greater the weight of the local standardized data of the corresponding mass spectrometry data; the smaller the number of samples or groups corresponding to a single difference data segment, the greater the weight of the local standardized data of the corresponding mass spectrometry data; the greater the mean of the difference value of a single difference data segment, the greater the weight of the local standardized data of the corresponding mass spectrometry data.
[0096] In the embodiment of the present application, the data standardization process is comprehensively determined by the local data distribution characteristics and the global data distribution characteristics. When a certain local data distribution characteristic or some local data distribution characteristics appear less frequently in the global data, the effect of its local characteristics is increased to obtain the preferred standardized data. If a certain local data distribution characteristic or some local data distribution characteristics appear more frequently in the global data, the importance of its local characteristics is relatively small. The standardization method in the embodiment of the present application fully retains the personalized characteristics of the mass spectrometry data on the basis of global standardization, and enhances the characterization ability of the characteristics extracted in the feature extraction stage.
[0097] See also Figure 3 In one embodiment, in the above step 3, the data twin model includes a spatiotemporal feature extraction module, a frequency domain feature calculation unit and a machine learning algorithm integration module, wherein the spatiotemporal feature extraction module and the frequency domain feature calculation unit are connected in parallel and synchronously receive the input mass spectrometry data, and the spatiotemporal feature extraction module and the frequency domain feature calculation unit are connected in parallel and then cascaded to the machine learning algorithm integration module;
[0098] The training process of the data twin model includes:
[0099] The spatiotemporal features are extracted from the input mass spectrometry data using a spatiotemporal feature extraction module, and the frequency domain features are extracted using a frequency domain feature calculation unit through fast Fourier transform FFT; the network model parameters in the spatiotemporal feature extraction module are to be trained, and are continuously iterated and updated during the training process of the data twin model, and the frequency domain feature calculation unit does not contain parameters to be trained;
[0100] Concatenate the spatiotemporal features and frequency domain features of the mass spectrometry data to form mass spectrometry features;
[0101] The mass spectrometry features are input into the machine learning algorithm integration module for comprehensive decision-making to obtain the mass spectrometry data analysis results. The network model parameters in the machine learning algorithm integration module are to be trained and continuously iterated and updated during the training process of the data twin model.
[0102] In the embodiment of the present application, the data twin model includes a spatiotemporal feature extraction module and a frequency domain feature calculation unit connected in parallel in front, and a machine learning algorithm integration module cascaded in the back. When extracting features, spatiotemporal feature extraction and frequency domain feature extraction based on neural networks are used to realize the mining of explicit features and implicit features of mass spectrometry data. When performing mass spectrometry data analysis, machine learning algorithm integration technology is used, and a variety of machine learning algorithms are used for comprehensive decision-making. Considering that different machine learning algorithms are adapted to mass spectrometry data of different types of characteristics, the advantages of a variety of machine learning algorithms are used to complement each other, and the mass spectrometry data analysis results facing mass spectrometry data of different characteristics are improved. Figure 3 In the figure, the modules in the dashed box do not contain parameters to be trained.
[0103] In one embodiment, in the data twin model of step 3 above, the machine learning algorithm integration module includes at least one machine learning model and a weight distribution model; at least one machine learning model synchronously receives input data of the machine learning algorithm integration module in parallel, and at least one machine learning model is cascaded to the weight distribution model after being connected in parallel; the network model parameters of each of the machine learning models are to be trained, and the network model parameters of the weight distribution model are to be trained, and are continuously iterated and updated during the training process of the data twin model.
[0104] In an embodiment of the present application, a weight allocation model is additionally designed in the machine learning algorithm integration module to adaptively weight the output decision results of each machine learning model. That is, the weight for each machine learning model is not a fixed value, but when the mass spectrometry data received by the data twin model is different, the weight will be adaptively determined. During the training process, the weight allocation model continuously learns how to assign weights to each machine learning model when faced with different mass spectrometry data output result distributions of each machine learning model.
[0105] The parameters of the weight allocation model are learnable parameters. Based on a large amount of sample data, the weight assignment scheme is intelligently learned for the output results of multiple machine learning models after inputting mass spectrometry data of different samples, so as to achieve intelligent and efficient weight allocation for multiple algorithm models in the machine learning algorithm integration module. The machine learning models in the machine learning algorithm integration module include but are not limited to random forest (RF), support vector machine (SVM), convolutional neural network (CNN), and long short-term memory network (LSTM).
[0106] See also Figure 4In one embodiment, in the above step 3, during the training process of the data twin model, the machine learning algorithm integration module performs the following process:
[0107] Step 301, obtaining the mass spectrometry data analysis results output by each machine learning model;
[0108] Step 302, calculating the error between each mass spectrometry data analysis result and the theoretical mass spectrometry data analysis result based on each output; and reversely updating the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis result;
[0109] Step 303, calculating the weight theoretical value of each machine learning model based on the error corresponding to each machine learning model, and obtaining the weight vector theoretical value for all machine learning models in the integrated module;
[0110] Step 304, inputting the mass spectrometry data analysis results output by each machine learning model into the weight distribution model to obtain a weight vector output by the weight distribution model;
[0111] Step 305: reversely update the network model parameters of the weight allocation model based on the error between the output weight vector and the theoretical value of the weight vector.
[0112] In an embodiment of the present application, the network model parameter training process of each machine learning model in the machine learning algorithm integration module and the network model parameter training process of the weight distribution model can be performed synchronously or asynchronously.
[0113] When the training processes of the machine learning model and the weight distribution model are synchronized: the error between the mass spectrometry data analysis results output by each machine learning model and the theoretical mass spectrometry data analysis results is calculated; the network model parameters of each machine learning model are reversely updated based on the error of the mass spectrometry data analysis results; at the same time, the mass spectrometry data analysis results output by each machine learning model are input into the weight distribution model to obtain the weight vector output by the weight distribution model; the network model parameters of the weight distribution model are reversely updated based on the error between the output weight vector and the theoretical value of the weight vector.
[0114] When the training processes of the machine learning model and the weight distribution model are not carried out synchronously: calculate the error between the mass spectrometry data analysis results output by each machine learning model and the theoretical mass spectrometry data analysis results; reversely update the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis results; until each machine learning model reaches the preset accuracy and error conditions, that is, when the training process of each machine learning model reaches a certain level, start to synchronously train the weight distribution model: input the mass spectrometry data analysis results output by each machine learning model into the weight distribution model to obtain the weight vector output by the weight distribution model; reversely update the network model parameters of the weight distribution model based on the error between the output weight vector and the theoretical value of the weight vector.
[0115] See also Figure 5 In another embodiment, in the above step 3, during the training process of the data twin model, the machine learning algorithm integration module performs the following process:
[0116] Step 311, obtaining the mass spectrometry data analysis results output by each machine learning model;
[0117] Step 312, calculating the error between each output mass spectrometry data analysis result and the theoretical mass spectrometry data analysis result; and reversely updating the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis result;
[0118] Step 313, input the mass spectrometry data analysis results output by each machine learning model into the weight allocation model for weighted fusion, and output the summarized mass spectrometry data analysis results; based on the error between the summarized mass spectrometry data analysis results and the theoretical mass spectrometry data analysis results, reversely update the network model parameters of the weight allocation model.
[0119] In the embodiment of the present application, the weight parameters and linear combinations of each machine learning model are implemented through layer-by-layer data processing within the weight allocation model, and the weighted results of each machine learning model are not explicitly output to the outside. Instead, the mass spectrometry data analysis results output by each machine learning model are summarized and fused to obtain the final mass spectrometry data analysis result.
[0120] The error based on the mass spectrometry data analysis result reversely updates the network model parameters of each machine learning model, and the error based on the summarized mass spectrometry data analysis result and the theoretical mass spectrometry data analysis result reversely updates the network model parameters of the weight distribution model. The network model parameter updates of the machine learning model and the weight distribution model are respectively carried out based on different error functions.
[0121] In one embodiment, in the data twin model, the spatiotemporal feature extraction module is based on a cascade of a convolutional neural network (CNN) and a long short-term memory (LSTM) network; the data processing steps of the spatiotemporal feature extraction module include: mass spectrum data is input through a data input layer, enters a convolutional layer to extract local spatiotemporal features of the mass spectrum data, enters a pooling layer for feature dimension reduction, enters an LSTM layer to extract temporal features of the mass spectrum data, generates a feature vector of the mass spectrum data through a fully connected layer, and outputs through an output layer;
[0122] Frequency domain features include: frequency components, power spectrum density, spectrum entropy, center frequency, bandwidth, harmonic components, wavelet coefficients, and instantaneous frequency.
[0123] In the embodiments of the present application, based on the frequency domain features extracted by Fourier transform, spatiotemporal features are further extracted through convolutional neural network CNN and long short-term memory network LSTM, so as to improve the recognition ability of complex patterns in mass spectrometry data and capture deeper features.
[0124] In the embodiment of the present application, in the frequency domain features:
[0125] Frequency components: The main frequency components in the mass spectrometry signal, which correspond to the characteristic peaks of different chemical substances in the sample and are used to identify and quantify specific chemical components in the sample.
[0126] Power spectral density: Power spectral density describes how the power of a signal is distributed across different frequencies, helping to identify noise and signal characteristics. It is used to analyze the noise characteristics of mass spectrometry signals, optimize signal processing and filtering algorithms, and improve the signal-to-noise ratio of the signal.
[0127] Spectral Entropy: Spectral entropy measures the complexity and disorder of the signal spectrum. A higher spectral entropy indicates that the signal is more complex or contains more random components. It is used for pattern recognition and classification of mass spectrometry signals to help distinguish different types of samples.
[0128] Centroid Frequency: The centroid frequency represents the average frequency of the signal and reflects the concentration position of the main frequency components in the signal. It is used to describe and compare the frequency characteristics of different samples and identify the main frequency range of the signal.
[0129] Bandwidth: Bandwidth describes the range of frequency components contained in a signal. A wider bandwidth indicates that the signal contains more frequency components. It is used to evaluate the frequency extension range of mass spectrometry signals and optimize signal processing and filtering strategies.
[0130] Harmonics: includes the fundamental frequency of the mass spectrometry signal and its integer multiple frequency components. Harmonics reflect the periodicity and repetitive characteristics of the signal and help identify specific patterns in the mass spectrometry signal. It is used to analyze the periodic characteristics of the mass spectrometry signal and optimize the signal processing and filtering algorithms.
[0131] Wavelet Coefficients: Wavelet coefficients reflect the frequency components of mass spectrometry signals at different scales, helping to identify local features and mutations in the signal. They are used to analyze mass spectrometry signals at multiple scales and optimize signal processing and filtering algorithms.
[0132] Instantaneous Frequency: Instantaneous frequency reflects the frequency variation characteristics of non-stationary signals and helps identify the rapidly changing components in the signal. It is used to analyze non-stationary mass spectrometry signals, identify the characteristics of signal frequency variation over time, and optimize signal processing and filtering algorithms.
[0133] In the embodiment of the present application, in the face of the task of processing mass spectrometry data with complex spatiotemporal features, a network model combining a convolutional layer and an LSTM layer is innovatively used when extracting deep features. The convolutional layer can effectively extract local features in the mass spectrometry data through convolution operations. In particular, in mass spectrometry data, the convolutional layer can detect local feature patterns in the data, such as peaks, frequency components, etc. Through multiple convolutional layers and pooling layers, the convolutional layer can gradually capture the low-level to high-level features of the data and integrate these features into a feature map. In mass spectrometry data analysis, LSTM can further process the dependencies in the time dimension after receiving the feature map extracted by the convolutional layer, thereby identifying the time series patterns and trends in the data. The feature map extracted by the convolutional layer contains local spatial information, and LSTM can further model the changes in these spatial information over time to form a deep understanding of spatiotemporal features. Compared with directly using a fully connected network to process mass spectrometry data, the combination of convolutional layers and LSTM layers can greatly reduce the amount of model parameters, improve training efficiency, and avoid overfitting. This combination can better understand the complex characteristics of mass spectrometry data and improve the analysis accuracy and generalization ability of the model. Specifically, the convolution layer in the present application is a convolution network containing multiple layers of convolution operations, which sequentially includes input layer → convolution layer 1 → activation layer 1 → pooling layer 1 → convolution layer 2 → activation layer 2 → pooling layer 2; the LSTM layer in the present application is a multi-layer network, which includes at least one LSTM processing layer.
[0134] See also Figure 6 , an embodiment of the present application provides a mass spectrometry detection system based on data twins, the system comprising:
[0135] A raw data acquisition unit, used to acquire raw mass spectrometry data of a sample through a mass spectrometry detection device;
[0136] A preprocessing unit, used for performing multi-stage intelligent mass spectrometry data preprocessing on raw mass spectrometry data;
[0137] A model training unit, used to train the processed mass spectrometry data using a machine learning algorithm to construct a data twin model, wherein the data twin model is used to obtain mass spectrometry analysis results;
[0138] The data analysis unit is used to input new mass spectrometry data into the trained data twin model to obtain mass spectrometry analysis results.
[0139] For specific limitations on the mass spectrometry detection system, please refer to the limitations on the mass spectrometry detection method above, which will not be repeated here.
[0140] An embodiment of the present application provides an electronic device, the electronic device comprising:
[0141] processor;
[0142] a memory for storing processor-executable instructions;
[0143] The processor implements the mass spectrometry detection method described in the above method embodiment by running the executable instructions.
[0144] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0145] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory is used to store at least one instruction, which is used to be executed by the processor to implement the mass spectrometry detection method provided in the method embodiment of the present application.
[0146] In some embodiments, the electronic device may optionally include: an input interface and an output interface. The processor, memory and the input interface and output interface may be connected via a bus or a signal line. Each peripheral device may be connected to the input interface and the output interface via a bus, a signal line or a circuit board. The input interface and the output interface may be used to connect at least one input / output related peripheral device to the processor and the memory. In some embodiments, the processor, the memory and the input interface and the output interface are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor, the memory and the input interface and the output interface may be implemented on a separate chip or circuit board, which is not limited in the embodiments of the present application.
[0147] An embodiment of the present application provides a computer-readable storage medium having computer instructions stored thereon, and when the instructions are executed by a processor, the steps of the mass spectrometry detection method described in the above method embodiment are implemented.
[0148] A person of ordinary skill in the art will appreciate that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned computer-readable storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0149] Those skilled in the art should be aware that the functions described in the embodiments of the present application can be implemented with hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that a general or special-purpose computer can access.
[0150] The present invention is not limited to the above-mentioned specific implementation modes. Various changes made by ordinary technicians in this field based on the above-mentioned concepts without creative work are all within the protection scope of the present invention.
Claims
1. A mass spectrometry detection method based on data twins, characterized in that: The steps include: Obtaining original mass spectrum data of the sample through a mass spectrometry detection device; The raw mass spectrometry data is preprocessed in multiple stages with intelligent technology, including: multi-level and refined denoising of the raw mass spectrometry data through wavelet transform and adaptive filtering; dual standardization and intelligent correction methods are used to eliminate the differences between instruments and samples; Using a machine learning algorithm to train the processed mass spectrometry data to construct a data twin model, wherein the data twin model is used to obtain mass spectrometry analysis results; Input new mass spectrometry data into the trained data twin model to obtain mass spectrometry analysis results; The dual standardization and intelligent correction method is used to eliminate the differences between instruments and samples, including: using a pre-trained model to automatically identify and correct outliers in the data, specifically including: Performing global and local standardization on the mass spectrometry data, and comprehensively determining the standardization result of the mass spectrometry data based on the double standardization data; Use intelligent correction to further adjust the standardized data to eliminate the influence of systematic errors and instrument drift; Standardized processing methods for mass spectrometry data, including: The mass spectrometry data of different samples are locally standardized respectively; the mass spectrometry data of all samples are globally standardized to obtain local standardized data and global standardized data corresponding to the mass spectrometry data of the same sample, wherein the standardization method is calculated based on the difference between the current data and the minimum value of the data and the ratio of the difference between the maximum and minimum values of the data; Analyze the difference between the local standardized data and the global standardized data of the mass spectrum data of the same sample to obtain double standardized data difference data of the mass spectrum data of each sample; Determine the standardized result of the mass spectral data of each sample based on the statistical distribution data of the double standardized data difference data corresponding to the mass spectral data of all samples; The data twin model includes a machine learning algorithm integration module, which includes at least one machine learning model and a weight distribution model; in the machine learning algorithm integration module, the following process is performed: Obtain the mass spectrometry data analysis results output by each machine learning model; Calculating the error between the mass spectrometry data analysis results of each output and the theoretical mass spectrometry data analysis results; and reversely updating the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis results; Calculate the theoretical value of the weight of each machine learning model based on the error corresponding to each machine learning model, and obtain the theoretical value of the weight vector for all machine learning models in the integrated module; Inputting the mass spectrometry data analysis results output by each machine learning model into the weight distribution model to obtain the weight vector output by the weight distribution model; The network model parameters of the weight allocation model are reversely updated based on the error between the output weight vector and the theoretical value of the weight vector.
2. The mass spectrometry detection method based on data twinning according to claim 1, characterized in that: The multi-level and refined denoising process of the original mass spectrum data by wavelet transform and adaptive filtering includes: Processing the original mass spectrum data by wavelet transform to obtain first preprocessed data; The first preprocessed data is processed by a least mean square adaptive filtering algorithm to obtain second preprocessed data; in the least mean square adaptive filtering algorithm, the update formula for each weight of the filter is: w(n+1)=w(n)+μe(n)x(n), wherein w(n) is the filter weight, μ is the learning rate, e(n) is the error, and x(n) is the input signal.
3. The mass spectrometry detection method based on data twinning according to claim 1, characterized in that: The data twin model includes a spatiotemporal feature extraction module, a frequency domain feature calculation unit and a machine learning algorithm integration module, wherein the spatiotemporal feature extraction module and the frequency domain feature calculation unit are connected in parallel and synchronously receive input mass spectrometry data, and the spatiotemporal feature extraction module and the frequency domain feature calculation unit are connected in parallel and then cascaded to the machine learning algorithm integration module; The training process of the data twin model includes: The spatiotemporal features are extracted from the input mass spectrometry data using a spatiotemporal feature extraction module, and the frequency domain features are extracted using a frequency domain feature calculation unit through fast Fourier transform FFT; the network model parameters in the spatiotemporal feature extraction module are to be trained, and are continuously iterated and updated during the training process of the data twin model, and the frequency domain feature calculation unit does not contain parameters to be trained; Concatenate the spatiotemporal features and frequency domain features of the mass spectrometry data to form mass spectrometry features; The mass spectrometry features are input into the machine learning algorithm integration module for comprehensive decision-making to obtain the mass spectrometry data analysis results. The network model parameters in the machine learning algorithm integration module are to be trained and continuously iterated and updated during the training process of the data twin model.
4. The mass spectrometry detection method based on data twinning according to claim 3 is characterized in that: In the data twin model, the machine learning algorithm integration module includes at least one machine learning model and a weight distribution model; at least one machine learning model synchronously receives input data of the machine learning algorithm integration module in parallel, and at least one machine learning model is cascaded with the weight distribution model after being connected in parallel; the network model parameters of each machine learning model are to be trained, and the network model parameters of the weight distribution model are to be trained, and are continuously iterated and updated during the training process of the data twin model.
5. The mass spectrometry detection method based on data twinning according to claim 3, characterized in that: The spatiotemporal feature extraction module is based on a cascade of a convolutional neural network (CNN) and a long short-term memory (LSTM) network; the data processing steps of the spatiotemporal feature extraction module include: mass spectrum data is input through a data input layer, enters a convolutional layer to extract local spatiotemporal features of the mass spectrum data, enters a pooling layer to perform feature dimension reduction, enters an LSTM layer to extract temporal features of the mass spectrum data, generates a feature vector of the mass spectrum data through a fully connected layer, and outputs the data through an output layer; The frequency domain features include: frequency component, power spectrum density, spectrum entropy, center frequency, bandwidth, harmonic component, wavelet coefficient, and instantaneous frequency.
6. A mass spectrometry detection system based on data twins, characterized in that: include: A raw data acquisition unit, used to acquire raw mass spectrometry data of a sample through a mass spectrometry detection device; The preprocessing unit is used to perform multi-stage intelligent preprocessing of the raw mass spectrometry data, including: multi-level and refined denoising of the raw mass spectrometry data through wavelet transform and adaptive filtering; eliminating the differences between instruments and samples through double standardization and intelligent correction methods; The dual standardization and intelligent correction method is used to eliminate the differences between instruments and samples, including: using a pre-trained model to automatically identify and correct outliers in the data, specifically including: Performing global and local standardization on the mass spectrometry data, and comprehensively determining the standardization result of the mass spectrometry data based on the double standardization data; Use intelligent correction to further adjust the standardized data to eliminate the influence of systematic errors and instrument drift; Standardized processing methods for mass spectrometry data, including: The mass spectrometry data of different samples are locally standardized respectively; the mass spectrometry data of all samples are globally standardized to obtain local standardized data and global standardized data corresponding to the mass spectrometry data of the same sample, wherein the standardization method is calculated based on the difference between the current data and the minimum value of the data and the ratio of the difference between the maximum and minimum values of the data; Analyze the difference between the local standardized data and the global standardized data of the mass spectrum data of the same sample to obtain double standardized data difference data of the mass spectrum data of each sample; Determine the standardized result of the mass spectral data of each sample based on the statistical distribution data of the double standardized data difference data corresponding to the mass spectral data of all samples; A model training unit, used to train the processed mass spectrometry data using a machine learning algorithm to construct a data twin model, wherein the data twin model is used to obtain mass spectrometry analysis results; The data twin model includes a machine learning algorithm integration module, which includes at least one machine learning model and a weight distribution model; in the machine learning algorithm integration module, the following process is performed: Obtain the mass spectrometry data analysis results output by each machine learning model; Calculating the error between the mass spectrometry data analysis results of each output and the theoretical mass spectrometry data analysis results; and reversely updating the network model parameters of each machine learning model based on the error of the mass spectrometry data analysis results; Calculate the theoretical value of the weight of each machine learning model based on the error corresponding to each machine learning model, and obtain the theoretical value of the weight vector for all machine learning models in the integrated module; Inputting the mass spectrometry data analysis results output by each machine learning model into the weight distribution model to obtain the weight vector output by the weight distribution model; Reversely updating the network model parameters of the weight allocation model based on the error between the output weight vector and the theoretical value of the weight vector; The data analysis unit is used to input new mass spectrometry data into the trained data twin model to obtain mass spectrometry analysis results.
Citation Information
Patent Citations
Transformer fault diagnosis method and equipment based on oil chromatography time-frequency domain information and residual attention network
CN113889198A
Power transformer defect diagnosis method based on multi-mode sound image fusion
CN118779807A
GIS equipment breakdown signal analysis method and system based on hierarchical networking structure
CN118965107A