Original mass spectrum data classification method based on multi-channel embedded representation
Through the multi-channel embedding representation method, the problems of low classification efficiency and low accuracy in mass spectrometry data analysis are solved. Through the joint optimization training of binning, normalization and multi-channel embedding modules, the multi-channel embedding representation is generated, which improves the classification accuracy and effectiveness of mass spectrometry data and improves the generalization ability of the model.
Patent Information
- Application Number
- CN202510701564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The prior art has problems with low classification efficiency and low accuracy in mass spectrometry data analysis, especially in high-dimensional and large-scale mass spectrometry data, which are difficult to fully capture multi-dimensional correlations. Traditional methods need to rely on complex preprocessing processes and are prone to information loss.
The multi-channel embedding representation method is adopted, and standardized intensity data is obtained through binning and normalization processing, and a multi-channel embedding module and preset classification model are built for joint optimization training. The encoder submodule, channel embedding submodule and channel splicing operations are used to generate a multi-channel embedding representation, and feature extraction and classification are performed through the end-to-end architecture.
It significantly improves the accuracy and effectiveness of mass spectrometry data classification, improves the generalization level of the model, reduces the calculation cost, enhances the feature expression ability and information density, and achieves a more comprehensive data characterization.
Smart Images

Figure CN120524271A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mass spectrometry data analysis and deep learning technology, and in particular relates to a method for classifying raw mass spectrometry data based on multi-channel embedding representation. Background Art
[0002] Mass spectrometry, as an efficient and versatile detection method, can accurately identify, analyze, and quantify a variety of target substances based on their mass-to-charge ratio (m / z). While the efficiency and structural complexity of mass spectrometry data have significantly increased with technological advancements, its high dimensionality and large volume present significant challenges for data analysis, particularly in applications requiring extremely high precision, such as cancer screening.
[0003] In the field of mass spectrometry data analysis, traditional machine learning algorithms (such as support vector machines, random forests, and XGBoost) are widely adopted and used for benchmarking, but they rely on complex preprocessing procedures. For example, steps such as denoising, baseline correction, peak identification, and peak alignment are required to eliminate interference factors during signal acquisition, thereby enhancing the stability of subsequent multivariate analysis and the interpretability of the results. These preprocessing operations are not only time-consuming but can also lead to the loss of critical information due to manual intervention.
[0004] In recent years, deep learning technology, owing to its advantages in automatic feature extraction, has gradually become a core method for data analysis. Continuous innovations in architectures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformers have significantly expanded the boundaries of data processing capabilities. However, a core challenge in mass spectrometry data analysis lies in constructing a high-information-density feature representation. To balance computational efficiency and model performance, existing methods typically map high-dimensional mass spectral vectors into a low-dimensional space. However, such strategies are often limited to single-channel representations and fail to fully capture the multidimensional correlations of the data. In contrast, the field of image classification has demonstrated the superiority of multi-channel representations—different channels can capture complementary information such as the target's spectrum and texture, thereby providing a more comprehensive description. This approach provides important insights for mass spectrometry analysis: through multi-channel collaborative modeling, features at different levels in mass spectrometry data (such as global intensity distribution and local peak patterns) can be more flexibly represented, thereby generating more discriminative embedding vectors. Therefore, developing an efficient and robust multi-channel representation method is crucial to improving classification accuracy and optimizing computational resource utilization.
[0005] In this context, convolutional neural networks (CNNs) have become an ideal choice for multi-channel modeling due to their unique feature extraction mechanism. CNNs can mine potential features from raw mass spectrometry signals through the spatial invariance properties of convolution kernels (such as translation invariance), and further enhance the identification of key structures through pooling operations. More importantly, CNNs do not need to rely on manually designed feature engineering and can autonomously learn multi-scale feature expressions directly from data. By combining stacked convolutional layers with nonlinearity, CNNs can dynamically construct multi-channel feature maps while capturing global trends and local details in mass spectrometry data. This multi-level feature fusion not only enhances the representational capabilities of the model, but also greatly improves its adaptability to high-dimensional data. Experiments have shown that even in the presence of noise or signal offset, CNN-based multi-channel models can maintain a high classification accuracy. At present, this type of method has demonstrated significant advantages in the direct classification task of raw mass spectrometry data and has become an important technical path in this field. Summary of the Invention
[0006] In response to the above-mentioned deficiencies in the prior art, the present invention provides a raw mass spectrometry data classification method based on multi-channel embedding representation, which solves the problems of low efficiency in raw mass spectrometry data classification and low accuracy of classification tasks.
[0007] In order to achieve the above objectives, the technical solution adopted by the present invention is: a method for classifying raw mass spectrometry data based on multi-channel embedding representation, comprising the following steps:
[0008] S1. Obtaining raw mass spectrometry data, performing binning and normalization processing on the raw mass spectrometry data to obtain standardized intensity data;
[0009] S2. Construct a multi-channel embedding module, and use the multi-channel embedding module as a feature representation layer to structurally connect it with the preset classification model to obtain a connected preset classification model;
[0010] S3. Performing joint optimization training on the parameters of the connected preset classification models based on the normalized intensity data, using an early stopping mechanism to obtain a trained preset classification model;
[0011] S4. Use the trained preset classification model to classify the standardized intensity data to obtain a classification result.
[0012] The beneficial effects of the present invention are as follows: the present invention proposes an innovative multi-channel embedding representation method to construct a highly targeted multi-channel embedding module, aiming to explore and utilize the mutual correlation between different representation channels within the data, so as to generate features with stronger expressive ability and higher information density; and by jointly optimizing and training the multi-channel embedding module with the preset classification model, the accuracy and efficiency of the classification task are significantly improved, and the generalization level of the preset classification model used is improved.
[0013] Furthermore, the S1 includes the following steps:
[0014] S101, obtaining raw mass spectrum data, and classifying the raw mass spectrum data according to data characteristics of different mass spectrum preparation technologies to obtain first mass spectrum data and second mass spectrum data;
[0015] S102, binning the first mass spectrum data according to a user-defined mass-to-charge ratio to form binning intervals, and taking, based on the binning intervals, mass spectrum data in the first mass spectrum data having a total ion intensity greater than a preset threshold as first training samples;
[0016] S103, binning the second mass spectrum data according to the mass-to-charge ratio defined by the user to form binning intervals, performing signal alignment on the second mass spectrum data based on the binning intervals and using the user-defined retention time as a window, and superimposing and averaging the signal intensity values in the same binning interval to obtain a second training sample;
[0017] S104 , normalizing the total number of ions in each first training sample and each second training sample to obtain standardized intensity data.
[0018] Furthermore, the first mass spectrum data is water-assisted laser desorption or ionization technology mass spectrum data;
[0019] The second mass spectrum data is liquid chromatography-mass spectrometry mass spectrum data.
[0020] The beneficial effect of the above further scheme is as follows: the present invention eliminates the noise caused by instrument resolution differences, removes the interference of low-quality data, eliminates the absolute intensity differences between sample data, and improves the quality of the original data by binning and normalizing the original mass spectrometry data.
[0021] Furthermore, the S2 includes the following steps:
[0022] S201, constructing a multi-channel embedding module using the encoder submodule, the channel embedding submodule, and the channel splicing operation;
[0023] S202. Based on the end-to-end architecture, the multi-channel embedding module is used as a feature representation layer and structurally connected with the preset classification model to obtain a connected preset classification model.
[0024] The beneficial effects of the above further scheme are as follows: the present invention constructs a multi-channel embedding module through an encoder sub-module, a channel embedding sub-module and a channel splicing operation, and structurally connects the multi-channel embedding module with a preset classification model based on an end-to-end architecture, thereby constructing a learnable deep learning front-end module, which significantly improves the classification performance and reduces the computational cost.
[0025] Furthermore, the S3 includes the following steps:
[0026] S301, using a stratified random sampling strategy to divide the standardized intensity data into a training set, a validation set, and a test set;
[0027] S302: The mass spectrometry data in the training set is used as input data and input into the multi-channel embedding module for processing to generate an optimal multi-channel embedding representation;
[0028] S303: Input the optimal multi-channel embedding representation into the preset classification model for joint optimization training, verify it through the validation set, and adopt an early stopping mechanism to test the classification model with the optimal parameters using the test set to obtain a trained preset classification model.
[0029] The beneficial effects of the above further scheme are as follows: the present invention realizes data-driven adaptive feature optimization by integrating advanced deep learning classification models for end-to-end joint training, and converts the mass spectrometry data in the training set into a multi-channel embedded expression to simultaneously capture the macro-global features and micro-local features of the data, providing the classification model with better quality and more comprehensive input information.
[0030] Furthermore, the S302 includes the following steps:
[0031] S3021. Input the mass spectrometry data in the training set as input data to the multi-channel embedding module, use the first fully connected layer in the encoder submodule to compress the input data to an intermediate dimension, and use layer normalization, rectified linear units, and random dropout units to perform feature extraction to obtain an initial latent embedding representation;
[0032] S3022. Using the second fully connected layer in the encoder submodule, compress the input data from the intermediate dimension to the embedding dimension and generate a global embedding representation;
[0033] S3023, reshaping the global embedding representation into a three-dimensional tensor, performing feature extraction using the first one-dimensional convolutional layer in the channel embedding submodule to obtain an initial first multi-channel embedding representation with half the preset number of channels, processing the initial first multi-channel embedding representation through batch normalization, rectified linear unit activation function, and random dropout to obtain an enhanced first multi-channel embedding representation;
[0034] S3024. Perform feature extraction on the enhanced first multi-channel embedding representation using the second one-dimensional convolutional layer in the channel embedding submodule to obtain an initial second multi-channel embedding representation with a preset number of channels, and obtain an enhanced second multi-channel embedding representation through batch normalization, rectified linear unit activation function, and random dropout.
[0035] S3025. Using the channel residual connection submodule, the global embedding representation reshaped into a three-dimensional tensor is concatenated with the enhanced second multi-channel embedding representation to obtain the optimal multi-channel embedding representation.
[0036] Furthermore, the expression for concatenating the global embedding representation reshaped into a three-dimensional tensor and the enhanced second multi-channel embedding representation is as follows:
[0037] O=Concat(E′,C2)∈R B×(1+C)×d
[0038] Where O represents the optimal multi-channel embedding representation, Concat(·) represents the concatenation operation, E′ represents the global embedding representation reshaped into a three-dimensional tensor, and E′∈R B×1×d , C2 represents the enhanced second multi-channel embedding representation, R B×(1+C)×d Represents the dimensions of the three-dimensional tensor, B represents the batch size, 1 represents dimension 1, d represents the embedding dimension, and C represents the number of channels defined by the user.
[0039] The beneficial effects of the above further solution are as follows: the present invention utilizes the encoder submodule to convert the original high-dimensional mass spectrum vector into a lower-dimensional space through nonlinear mapping, ensuring the preservation of key information and removing redundant data, forming a more refined feature representation, performing dimensionality reduction at an early stage, significantly reducing the computational burden of subsequent processing, and improving overall efficiency;
[0040] The channel embedding submodule uses two identical one-dimensional convolutional layers to maintain data consistency in the length dimension, avoiding changes in feature scale. It uses convolution operations to effectively identify the local arrangement structure and spatial correlation within the input data, and transforms the embedding representation of a single channel into an embedding representation containing multiple channels, which is more conducive to the execution of subsequent classification tasks.
[0041] By utilizing the channel splicing operation, the output features can retain both global information and local details, achieving effective fusion and mutual enhancement of features, and making the fused features more diverse and informative, capable of efficiently capturing key discriminative information in mass spectrometry data. The generated multi-channel embedding representation can be seamlessly connected and directly used as input to the downstream classification model, forming a complete, end-to-end learning system.
[0042] Furthermore, the S303 includes the following steps:
[0043] S3031. Input the multi-channel embedding representation into a preset classification model, and update the preset classification model parameters using an adaptive moment estimation optimizer based on a staged optimization strategy;
[0044] S3032. Dynamically adjust the learning rate of the adaptive moment estimation optimizer using a preset learning rate scheduler based on the validation set. In response to the validation set loss not decreasing for consecutive preset attenuation rounds, decay the learning rate by a preset attenuation multiple.
[0045] S3033. Calculate the validation set accuracy and F1 score, save the model weight of the best performance, reversely distribute the loss function weights based on the training set class frequency, and calculate the validation loss using the weighted cross entropy loss function based on the validation set. Use an early stopping mechanism. If the validation loss does not decrease after a preset number of early stopping rounds, stop training and proceed to step S3035.
[0046] S3034: Determine whether the preset training period has been reached. If not, return to step S3031 to perform joint optimization training again. If so, proceed to step S3035.
[0047] S3035. Based on the optimal performance model weights saved in the last round, the classification model with the optimal parameters is tested using the test set to obtain a trained preset classification model.
[0048] Furthermore, the preset classification model includes: a convolutional neural network, a long short-term memory network, and a converter model;
[0049] The convolutional neural network takes the multi-channel embedding representation as multi-channel input and uses hierarchical convolution and pooling operations to gradually fuse the global distribution and local peak details in the multi-channel to complete the classification;
[0050] The LSTM network uses multi-channel embedding representations as temporal sequence inputs to the model. Each time step corresponds to a feature vector of a channel. The LSTM layer is used to capture the temporal pattern across channels. The fully connected layer is used to map the hidden state of the last time step to the classification label to complete the classification.
[0051] The converter model expands the multi-channel embedding representation according to the embedding dimension and uses it as a multi-sequence input. It adds a learnable preset marker at the beginning of the sequence to obtain the input sequence, and uses a multi-layer multi-head self-attention module to form an encoder to capture the cross-channel dependencies of the input sequence. It extracts the hidden state corresponding to the preset marker and uses a fully connected layer for mapping to obtain the classification result.
[0052] The beneficial effects of the above further scheme are as follows: the present invention realizes the collaborative learning of feature representation and classification decision through the joint optimization of the multi-channel embedding module and the preset classification model parameters, and adopts a phased optimization strategy to balance efficiency and stability; by adopting weighted cross entropy loss in the loss function, the data imbalance problem is alleviated, and in order to prevent overfitting, an early stopping mechanism is introduced to improve the training efficiency of the preset classification model.
[0053] To achieve the above object, according to a second aspect of the present invention, a system for classifying raw mass spectrum data based on multi-channel embedded representation is provided, which is used to perform the above-mentioned method for classifying raw mass spectrum data based on multi-channel embedded representation, and is characterized by comprising:
[0054] Data preprocessing subsystem, used to perform binning and normalization processing on raw mass spectrometry data;
[0055] The multi-channel embedding subsystem is used to construct a multi-channel embedding module and use it as a feature representation layer to structurally connect it with the preset classification model.
[0056] The classification subsystem is used to jointly optimize and train the parameters of the connected preset classification models and perform classification prediction steps.
[0057] The beneficial effects of the above scheme are as follows: the present invention provides a raw mass spectrometry data classification system based on multi-channel embedded representation, which breaks through the limitations of single-channel representation through multi-channel collaborative feature extraction and the joint design of global dimensionality reduction and local convolution, and realizes multi-level semantic mining of mass spectrometry data; dynamic feature fusion mechanism, channel residual connection retains the original information while strengthening feature interaction, avoiding information loss and improving model generalization ability; end-to-end training process, modular design takes into account both performance and efficiency, and provides a scalable solution for high-dimensional mass spectrometry data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Flow chart of the method of the present invention.
[0059] Figure 2 This is a reference diagram for the multi-channel embedding representation generation stage in this embodiment. DETAILED DESCRIPTION
[0060] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0061] Before describing this embodiment, the following terms are explained:
[0062] SpiderMass: Water-assisted laser desorption or ionization technology;
[0063] LC-MS: liquid chromatography-mass spectrometry;
[0064] TIC: total ion count normalized;
[0065] LayerNorm: layer normalization;
[0066] BatchNorm: batch normalization;
[0067] ReLU: rectified linear unit;
[0068] Dropout: random inactivation unit;
[0069] Transformer model: converter model;
[0070] Adam: an optimization method with adaptive learning rate.
[0071] CLS_token: Adds a special token to the beginning of the input sequence for sequence-level downstream tasks;
[0072] ReduceLROnPlateau: A learning rate scheduler for deep learning model training.
[0073] FLOPS: Floating-point operations per second.
[0074] Example
[0075] In this embodiment, Figure 1 As shown in the figure, the classification process includes the following stages: data preprocessing stage, in which the original mass spectrometry data is binned and normalized to generate standardized intensity data; multi-channel embedding feature generation stage, in which the preprocessed data is input into the multi-channel embedding module to generate a multi-channel embedding representation; module and deep learning classification model collaborative optimization training stage, in which the multi-channel embedding representation is input into the classification model, and the classification results are collaboratively optimized and output during the training process.
[0076] like Figure 1 As shown, the present invention provides a method for classifying raw mass spectrometry data based on multi-channel embedding representation, and its implementation method is as follows:
[0077] S1. Obtain raw mass spectrometry data, perform binning and normalization on the raw mass spectrometry data to obtain standardized intensity data. The specific steps are as follows:
[0078] S101, obtaining raw mass spectrum data, and classifying the raw mass spectrum data according to data characteristics of different mass spectrum preparation technologies to obtain first mass spectrum data and second mass spectrum data;
[0079] S102, binning the first mass spectrum data according to a user-defined mass-to-charge ratio to form binning intervals, and taking, based on the binning intervals, mass spectrum data in the first mass spectrum data having a total ion intensity greater than a preset threshold as first training samples;
[0080] S103, binning the second mass spectrum data according to the mass-to-charge ratio defined by the user to form binning intervals, performing signal alignment on the second mass spectrum data based on the binning intervals and using the user-defined retention time as a window, and superimposing and averaging the signal intensity values in the same binning interval to obtain a second training sample;
[0081] S104 , normalizing the total number of ions in each first training sample and each second training sample to obtain standardized intensity data.
[0082] In this embodiment, in the data preprocessing stage, raw mass spectrum data is obtained, and the raw mass spectrum data is classified according to the data characteristics of different mass spectrum preparation technologies (such as LC-MS and SpiderMass) to obtain first mass spectrum data (SpiderMass mass spectrum data) and second mass spectrum data (LC-MS mass spectrum data);
[0083] For SpiderMass mass spectrometry data, binning was performed with a fixed width of 0.1 Da according to the user-defined mass-to-charge ratio (m / z). For example, each mass spectral vector was divided into 12,000 bins in the range of 400–1,600 Da. The intensity values in each bin were accumulated to eliminate the noise caused by instrument resolution differences and to filter by thresholds (e.g., when the total number of ions was less than 1 × 10 4 The mass spectrum of the data is regarded as low-quality data) and the interference of low-quality data is eliminated to obtain the first training sample;
[0084] For LC-MS mass spectrometry data, binning was performed with a fixed width of 0.1 Da according to the user-defined mass-to-charge ratio (m / z), and further aligned with a 10-second window according to the retention time (RT). The signal intensity values within the same binning interval were superimposed and averaged to obtain the second training sample;
[0085] Total ion count normalization (TIC) was performed on each sample (the first training sample and the second training sample) to scale the total intensity to 1.0 and eliminate the absolute intensity differences between samples.
[0086] S2. Construct a multi-channel embedding module and use it as a feature representation layer to structurally connect it with the preset classification model to obtain a connected preset classification model. The specific steps are as follows:
[0087] S201, constructing a multi-channel embedding module using the encoder submodule, the channel embedding submodule, and the channel splicing operation;
[0088] S202: Based on the end-to-end architecture, the multi-channel embedding module is used as the feature representation layer and structurally connected with the preset classification model to obtain the connected preset classification model.
[0089] In this embodiment, the multi-channel embedding module includes: an encoder submodule, a channel embedding submodule, and a channel splicing operation for realizing inter-channel information fusion;
[0090] The encoder submodule includes: a two-layer fully connected neural network, with layer normalization, rectified linear units, and random dropout units connected after the first fully connected layer;
[0091] The channel embedding submodule includes: two one-dimensional convolutional layers, and each convolution layer is followed by BatchNorm, ReLU activation function and Dropout unit, where the Dropout probability is configurable;
[0092] The channel splicing operation is used to perform a splicing operation on the channel dimension.
[0093] S3. Based on the normalized intensity data, jointly optimize and train the parameters of the connected preset classification model, using an early stopping mechanism to obtain a trained preset classification model. The specific steps are as follows:
[0094] S301. Divide the standardized intensity data into a training set, a validation set, and a test set using a stratified random sampling strategy.
[0095] In this embodiment, after the preprocessing stage, a stratified random sampling strategy is used to divide the standardized intensity data into a training set (70%), a validation set (20%), and a test set (10%) to ensure that the category distribution of each subset is consistent with the original dataset and to address the category imbalance problem.
[0096] S302: The mass spectrometry data in the training set is used as input data and input into the multi-channel embedding module for processing to generate the optimal multi-channel embedding representation, as follows:
[0097] S3021. Input the mass spectrometry data in the training set as input data to the multi-channel embedding module, use the first fully connected layer in the encoder submodule to compress the input data to an intermediate dimension, and use layer normalization, rectified linear units, and random dropout units to perform feature extraction to obtain an initial latent embedding representation;
[0098] S3022. Utilize the second fully connected layer in the encoder submodule to compress the input data from the intermediate dimension to the embedding dimension and generate a global embedding representation.
[0099] In this embodiment, Figure 2As shown in the figure, the multi-channel embedding module uses the encoder submodule to reduce the dimensionality of the high-dimensional mass spectrum vector, reducing the computational complexity while retaining key features. Subsequently, through the stacking of convolutional layers in the channel embedding submodule, the reduced single-channel mass spectrum vector is embedded into multiple complementary channel representations. This multi-channel embedding method can capture the feature information at different levels in the original mass spectrum data and establish deep associations between the channels, thereby effectively improving the classification performance of the subsequent classification model. The multi-channel embedding submodule serves as the front feature extraction layer of the classification model and achieves deep coupling with the downstream model through an end-to-end architecture.
[0100] Specifically: the preprocessed mass spectrometry data X∈R in the training set B×D As input data (where B represents the batch size and D is the original dimension of each mass spectral vector), it is input to the multi-channel embedding module, and the fully connected encoder performs nonlinear dimensionality reduction. The first fully connected layer is used to compress the dimension of the mass spectral data from the original dimension D of the mass spectral vector to the preset intermediate dimension 2048. The intermediate dimension mass spectral data is processed by LayerNorm, ReLU activation function and random inactivation to obtain the initial potential embedding representation. The second fully connected layer is further used to reduce the dimension to the embedding dimension d preset to 1024, generating a low-dimensional global embedding representation E∈R B×d ;
[0101] The mathematical process of generating a low-dimensional global embedding representation can be summarized as follows:
[0102] H1=ReLU(LayerNorm(XW1+b1))
[0103] H′1=Dropout(H1)
[0104] E=H1′W2+b2
[0105] Where H1 represents the intermediate dimension mass spectrum data after LayerNorm processing and ReLU activation function processing, ReLU(·) represents the ReLU activation function processing process, LayerNorm(·) represents LayerNorm processing, X represents the preprocessed mass spectrum data, W1 represents the weight matrix of the first fully connected layer, and W1∈R D×2048 , which linearly projects the mass spectrum vector X of dimension D into an intermediate feature space of 2048, b1 represents the bias vector of the first fully connected layer, and b1∈R 2048 , after the linear transformation of XW1, the bias vector b1 is added to each dimension of the 2048-dimensional intermediate result, H1′ represents the initial potential embedding representation, Dropout(·) represents the random inactivation process, E represents the global embedding representation, W2 represents the weight matrix of the second fully connected layer, and W2∈R 2048×d, b2 represents the bias vector of the second fully connected layer, and b2∈R d .
[0106] S3023, reshaping the global embedding representation into a three-dimensional tensor, performing feature extraction using the first one-dimensional convolutional layer in the channel embedding submodule to obtain an initial first multi-channel embedding representation with half the preset number of channels, processing the initial first multi-channel embedding representation through batch normalization, rectified linear unit activation function, and random dropout to obtain an enhanced first multi-channel embedding representation;
[0107] S3024. Perform feature extraction on the enhanced first multi-channel embedding representation using the second one-dimensional convolutional layer in the channel embedding submodule to obtain an initial second multi-channel embedding representation with a preset number of channels, and obtain an enhanced second multi-channel embedding representation through batch normalization, rectified linear unit activation function, and random dropout.
[0108] S3025. Using the channel residual connection submodule, the global embedding representation reshaped into a three-dimensional tensor is concatenated with the enhanced second multi-channel embedding representation to obtain the optimal multi-channel embedding representation.
[0109] In this embodiment, the global embedding representation is reshaped into a three-dimensional tensor E′∈R B×1×d , input to the channel embedding submodule, and then pass through two layers of one-dimensional convolution (kernel size 3, stride 1, padding 1). After each layer of convolution, BatchNorm, ReLU activation function and Dropout are executed in turn to enhance generalization ability. The Dropout probability is configurable. The first layer of one-dimensional convolution layer outputs half of the preset number of channels. The first multi-channel embedding representation is processed by the second one-dimensional convolutional layer, and a second multi-channel embedding representation with a preset number of channels C = 256 is output. After enhanced generalization, an enhanced multi-channel embedding representation is obtained. The global embedding representation reshaped into a three-dimensional tensor is spliced with the enhanced second multi-channel embedding representation (i.e., multi-channel local features) through a channel splicing operation to obtain a fused feature. The channel dimension of the fused feature is expanded to 257 (original channels + 256 convolution channels), so that the fused feature can be directly adapted to a variety of classification models.
[0110] The expression for reshaping the global embedding representation into a three-dimensional tensor is as follows:
[0111]
[0112] Among them, E′ represents the global embedding representation reshaped into a three-dimensional tensor, which is composed of the global embedding representation plus a channel dimension, and E represents the global embedding representation. Represents the expansion in the channel dimension;
[0113] The expression for feature extraction in the channel embedding submodule is as follows:
[0114]
[0115] Where C1 represents the first multi-channel embedding representation after batch normalization and ReLU activation, and BatchNorm(·) represents batch normalization, Conv1D(·) represents one-dimensional convolution, K1 represents the kernel parameter of the first one-dimensional convolution layer, and C2 represents the enhanced second multi-channel embedding representation, and C2∈R B×C×d , Dropout(·) represents random dropout processing, K2 represents the kernel parameter of the second one-dimensional convolutional layer, and
[0116] The expression for concatenating the global embedding representation reshaped into a three-dimensional tensor and the enhanced second multi-channel embedding representation is as follows:
[0117] O=Concat(E′,C2)∈R B×(1+C)×d
[0118] Where O represents the optimal multi-channel embedding representation, Concat(·) represents the concatenation operation, E′ represents the global embedding representation reshaped into a three-dimensional tensor, and E′∈R B×1×d , C2 represents the global embedding representation after the second augmentation and reshaping into a three-dimensional tensor, R B ×(1+C)×d Represents the dimensions of the three-dimensional tensor, B represents the batch size, 1 represents dimension 1, d represents the embedding dimension, and C represents the number of channels defined by the user.
[0119] S303: Input the multi-channel embedding representation into the preset classification model for joint optimization training, verify it through the validation set, and adopt an early stopping mechanism to test the classification model with the optimal parameters using the test set to obtain a trained preset classification model. The specific steps are as follows:
[0120] S3031. Input the multi-channel embedding representation into a preset classification model, and update the preset classification model parameters using an adaptive moment estimation optimizer based on a staged optimization strategy;
[0121] S3032. Dynamically adjust the learning rate of the adaptive moment estimation optimizer using a preset learning rate scheduler based on the validation set. In response to the validation set loss not decreasing for consecutive preset attenuation rounds, decay the learning rate by a preset attenuation multiple.
[0122] S3033. Calculate the validation set accuracy and F1 score, save the model weight of the best performance, reversely distribute the loss function weights based on the training set class frequency, and calculate the validation loss using the weighted cross entropy loss function based on the validation set. Use an early stopping mechanism. If the validation loss does not decrease after a preset number of early stopping rounds, stop training and proceed to step S3035.
[0123] S3034: Determine whether the preset training period has been reached. If not, return to step S3031 to perform joint optimization training again. If so, proceed to step S3035.
[0124] S3035. Based on the optimal performance model weights saved in the last round, the classification model with the optimal parameters is tested using the test set to obtain a trained preset classification model.
[0125] In this embodiment, the parameters of the multi-channel embedding module and the parameters of the preset classification model are optimized together during joint optimization training;
[0126] The preset classification model training adopts a phased optimization strategy to balance efficiency and stability; first, the Adam optimizer is used (the initial learning rate is set to 1e -3 , the decay rate is set to 1e -5 ) Update the parameters and dynamically adjust the learning rate in combination with the ReduceLROnPlateau scheduler: If the validation set loss does not decrease for 5 consecutive rounds (preset attenuation rounds), the learning rate is attenuated to 0.1 times the current value, with a minimum limit of 1e -32 The preset training cycle is fixed at 64 rounds. After each round, the validation set accuracy and F1 score are calculated, and only the model weight with the best performance is saved. The loss function uses weighted cross-entropy loss to calculate the validation loss. The weights are inversely distributed according to the frequency of the training set categories (for example, the weight of the minority class is the total number of samples or the number of samples in that class) to alleviate the data imbalance problem. To prevent overfitting, an early stopping mechanism is introduced. If the validation loss does not decrease for 10 consecutive rounds (the preset number of early stopping rounds), the training is terminated. In terms of training acceleration, a parallel computing strategy is adopted to improve data throughput in a multi-GPU environment.
[0127] After one round of training, determine whether the preset training cycle of 64 rounds has been reached. If not, perform joint optimization training again. If so, use the test set to test the classification model with the optimal parameters based on the optimal performance model weights saved in the last round to obtain the trained preset classification model.
[0128] In this embodiment, during the collaborative optimization training phase, the present invention combines the multi-channel embedding module with a variety of mainstream deep learning model architectures, including convolutional neural networks (CNNs) such as ResNet, EfficientNet, and DenseNet, as well as sequence modeling networks such as LSTM and Transformer. Rather than adopting an ensemble learning strategy, the present invention uses the multi-channel embedding module as a feature representation layer and directly connects it to the classification model to optimize feature input.
[0129] The preset classification models include: convolutional neural network, long short-term memory network and converter model;
[0130] For CNN (such as ResNet-50, EfficientNetB0, DenseNet121 and other models), O can be directly used as multi-channel input, and the global distribution and local peak details in the multi-channel features can be gradually integrated through hierarchical convolution and pooling operations;
[0131] For the Transformer, O is used as a multi-sequence input. The embedding dimension of the multi-channel embedding module is the sequence dimension of the model. To meet the requirements of the classification task, a learnable preset tag CLS_token (B, 1, d) is added to the beginning of the sequence. Then, it is input into the encoder of the Transformer, which consists of multiple layers of multi-head self-attention modules. Each layer captures cross-channel dependencies through the self-attention mechanism. Finally, the hidden state (B, d) corresponding to the CLS_token is extracted and mapped using a fully connected layer to obtain the classification result.
[0132] For the LSTM model, the multi-channel features output by the multi-channel embedding module are directly input into the model as a time series. Each time step corresponds to the feature vector of a channel. The LSTM layer captures the cross-channel temporal pattern, and the hidden state of the last time step is mapped to the classification label through the fully connected layer to complete the classification.
[0133] S4. Use the trained preset classification model to classify the standardized intensity data to obtain a classification result.
[0134] In this embodiment, the deep learning classification model without the multi-channel embedding module has large values in FLOPs and model size dimensions, indicating that its computational overhead is high. After the module is integrated, the FLOPs of all models are significantly reduced, the model size is also greatly reduced, and the classification accuracy is greatly improved. This performance improvement is due to the characteristics of the multi-channel embedding module. It first performs a dimensionality reduction operation on the high-dimensional data before channel embedding, which significantly reduces the data dimension, so that subsequent convolution operations no longer require excessive convolution calculations to capture the structural information of the data, reducing unnecessary computational overhead in sparse high-dimensional space. In addition, although the computational requirements of sequence models such as LSTM and Transformer increase (FLOPs) after the module is integrated, the training stability of the model is significantly improved, avoiding the collapse phenomenon during training;
[0135] The multi-channel embedding representation classification method proposed in this paper significantly enhances feature expression capabilities by converting raw mass spectral vectors into multi-channel embedding representations. Compared with traditional feature engineering methods, this method is more adaptable and can maintain robust performance across different datasets and tasks. Overall, this study provides a novel deep feature representation method for mass spectrometry data analysis, expands the application prospects of deep representation learning in the field of mass spectrometry, and provides new technical support for mass spectrometry data analysis.
Claims
1. A method for classifying raw mass spectrometry data based on multi-channel embedding representation, characterized in that: The following steps are involved: S1. Obtaining raw mass spectrometry data, performing binning and normalization processing on the raw mass spectrometry data to obtain standardized intensity data; S2. Construct a multi-channel embedding module, and use the multi-channel embedding module as a feature representation layer to structurally connect it with the preset classification model to obtain a connected preset classification model; S3. Performing joint optimization training on the parameters of the connected preset classification models based on the normalized intensity data, using an early stopping mechanism to obtain a trained preset classification model; S4. Use the trained preset classification model to classify the standardized intensity data to obtain a classification result.
2. The method for classifying raw mass spectrometry data based on multi-channel embedded representation according to claim 1, characterized in that: Said S1 comprises the following steps: S101, obtaining raw mass spectrum data, and classifying the raw mass spectrum data according to data characteristics of different mass spectrum preparation technologies to obtain first mass spectrum data and second mass spectrum data; S102, binning the first mass spectrum data according to a user-defined mass-to-charge ratio to form binning intervals, and taking, based on the binning intervals, mass spectrum data in the first mass spectrum data having a total ion intensity greater than a preset threshold as first training samples; S103, binning the second mass spectrum data according to the mass-to-charge ratio defined by the user to form binning intervals, performing signal alignment on the second mass spectrum data based on the binning intervals and using the user-defined retention time as a window, and superimposing and averaging the signal intensity values in the same binning interval to obtain a second training sample; S104 , normalizing the total number of ions in each first training sample and each second training sample to obtain standardized intensity data.
3. The method for classifying raw mass spectrometry data based on multi-channel embedded representation according to claim 2, characterized in that: The first mass spectrum data is water-assisted laser desorption or ionization technology mass spectrum data; The second mass spectrum data is liquid chromatography-mass spectrometry mass spectrum data.
4. The method for classifying raw mass spectrometry data based on multi-channel embedded representation according to claim 1, characterized in that: The S2 comprises the following steps: S201, constructing a multi-channel embedding module using the encoder submodule, the channel embedding submodule, and the channel splicing operation; S202. Based on the end-to-end architecture, the multi-channel embedding module is used as a feature representation layer and structurally connected with the preset classification model to obtain a connected preset classification model.
5. The method for classifying raw mass spectrometry data based on multi-channel embedded representation according to claim 4, characterized in that: The S3 includes the following steps: S301, using a stratified random sampling strategy to divide the standardized intensity data into a training set, a validation set, and a test set; S302: The mass spectrometry data in the training set is used as input data and input into the multi-channel embedding module for processing to generate an optimal multi-channel embedding representation; S303: Input the optimal multi-channel embedding representation into the preset classification model for joint optimization training, verify it through the validation set, and adopt an early stopping mechanism to test the classification model with the optimal parameters using the test set to obtain a trained preset classification model.
6. The method for classifying raw mass spectrometry data based on multi-channel embedded representation according to claim 5, characterized in that: The S302 includes the following steps: S3021. Input the mass spectrometry data in the training set as input data to the multi-channel embedding module, use the first fully connected layer in the encoder submodule to compress the dimension of the input data to an intermediate dimension, and use layer normalization, rectified linear units, and random dropout units to perform feature extraction to obtain an initial latent embedding representation; S3022. Using the second fully connected layer in the encoder submodule, compress the input data from the intermediate dimension to the embedding dimension and generate a global embedding representation; S3023, reshaping the global embedding representation into a three-dimensional tensor, performing feature extraction using the first one-dimensional convolutional layer in the channel embedding submodule to obtain an initial first multi-channel embedding representation with half the preset number of channels, processing the initial first multi-channel embedding representation through batch normalization, rectified linear unit activation function, and random dropout to obtain an enhanced first multi-channel embedding representation; S3024. Perform feature extraction on the enhanced first multi-channel embedding representation using the second one-dimensional convolutional layer in the channel embedding submodule to obtain an initial second multi-channel embedding representation with a preset number of channels, and obtain an enhanced second multi-channel embedding representation through batch normalization, rectified linear unit activation function, and random dropout. S3025. Use the channel residual connection submodule to splice the global embedding representation reshaped into a three-dimensional tensor and the enhanced second multi-channel embedding representation to obtain the optimal multi-channel embedding representation.
7. The method for classifying raw mass spectrometry data based on multi-channel embedded representation according to claim 6, characterized in that: The expression for concatenating the global embedding representation reshaped into a three-dimensional tensor and the enhanced second multi-channel embedding representation is as follows: O=Concat(E′,C2)∈R B×(1+C)×d Where O represents the optimal multi-channel embedding representation, Concat(·) represents the concatenation operation, E′ represents the global embedding representation reshaped into a three-dimensional tensor, and E′∈R B×1×d , C2 represents the enhanced second multi-channel embedding representation, R B×(1+C)×d Represents the dimensions of the three-dimensional tensor, B represents the batch size, 1 represents dimension 1, d represents the embedding dimension, and C represents the number of channels defined by the user.
8. The method for classifying raw mass spectrum data based on multi-channel embedded representation according to claim 5, characterized in that: The S303 includes the following steps: S3031. Input the multi-channel embedding representation into a preset classification model, and update the preset classification model parameters using an adaptive moment estimation optimizer based on a staged optimization strategy; S3032. Dynamically adjust the learning rate of the adaptive moment estimation optimizer using a preset learning rate scheduler based on the validation set. In response to the validation set loss not decreasing for consecutive preset attenuation rounds, decay the learning rate by a preset attenuation multiple. S3033. Calculate the validation set accuracy and F1 score, save the model weight of the best performance, reversely distribute the loss function weights based on the training set class frequency, and calculate the validation loss using the weighted cross entropy loss function based on the validation set. Use an early stopping mechanism. If the validation loss does not decrease after a preset number of early stopping rounds, stop training and proceed to step S3035. S3034: Determine whether the preset training period has been reached. If not, return to step S3031 to perform joint optimization training again. If so, proceed to step S3035. S3035. Based on the optimal performance model weights saved in the last round, the classification model with the optimal parameters is tested using the test set to obtain a trained preset classification model.
9. The method for classifying raw mass spectrum data based on multi-channel embedded representation according to claim 5, characterized in that: The preset classification models include: convolutional neural network, long short-term memory network and converter model; The convolutional neural network takes the multi-channel embedding representation as multi-channel input and uses hierarchical convolution and pooling operations to gradually fuse the global distribution and local peak details in the multi-channel to complete the classification; The LSTM network uses multi-channel embedding representations as temporal sequence inputs to the model. Each time step corresponds to a feature vector of a channel. The LSTM layer is used to capture the temporal pattern across channels. The fully connected layer is used to map the hidden state of the last time step to the classification label to complete the classification. The converter model expands the multi-channel embedding representation according to the embedding dimension and uses it as a multi-sequence input. It adds a learnable preset marker at the beginning of the sequence to obtain the input sequence, and uses a multi-layer multi-head self-attention module to form an encoder to capture the cross-channel dependencies of the input sequence. It extracts the hidden state corresponding to the preset marker and uses a fully connected layer for mapping to obtain the classification result.
10. A system for classifying raw mass spectrum data based on multi-channel embedded representation, for executing the method for classifying raw mass spectrum data based on multi-channel embedded representation according to any one of claims 1 to 9, characterized in that: include: Data preprocessing subsystem, used to perform binning and normalization processing on raw mass spectrometry data; The multi-channel embedding subsystem is used to construct a multi-channel embedding module and use it as a feature representation layer to structurally connect it with the preset classification model. The classification subsystem is used to jointly optimize and train the parameters of the connected preset classification models and perform classification prediction steps.
Citation Information
Patent Citations
Sample Mass Spectrum Analysis
CN107683476A
Digital image coding method based on metabonomics mass spectrum data
CN116183796A
Mass spectrum classification method and system based on deep learning, medium and equipment
CN117034017A
Mass spectrum data classification method suitable for nephropathy identification and electronic equipment
CN117975102A
Deep learning-based mass spectrum classification method and biomarker discovery method and device
CN118609668A
Cited By
Multi-source mass spectrum metadata fusion method and system
CN121935847A