Mass spectrum ion fragment peak region identification method and device based on deep learning

Through a two-branch convolutional neural network model based on deep learning, combining charge state and additive ion information, peak region segmentation and charge state prediction are optimized, and the accuracy and chemical rationality of fragment peak recognition in the mass spectrogram are solved, and fragment peak recognition with high accuracy and low misjudgment rate is achieved.

CN120068004AActive Publication Date: 2025-05-30TIANJIN ZHIPU INSTR CO LTD

Patent Information

Application Number
CN202510543595.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Complex background noise and artifact interference in mass spectra, charge state and added ion impact, and effective extraction difficulties of local and global features, resulting in insufficient accuracy and chemical rationality of fragment peak recognition.

Method used

A two-branch convolutional neural network model based on deep learning is adopted, combining charge state limitation and added ion probability weights, a multi-task loss function is constructed, peak region segmentation, charge state prediction and added ion matching is optimized, local and global features are extracted through one-dimensional and two-dimensional convolutional layers, and multi-scale information fusion technology and spatial attention module are used to improve model performance.

Benefits of technology

It significantly improves the recognition accuracy and chemical rationality of fragment peaks in the mass spectrogram, reduces interference from background noise or artifacts, ensures detection sensitivity and significantly reduces the misjudgment rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068004A_ABST
    Figure CN120068004A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fragment peak region recognition, in particular to a mass spectrum ion fragment peak region recognition method and device based on deep learning. The method comprises the following steps: collecting mass spectrogram data, and preprocessing the mass spectrogram data; and constructing a double-branch convolutional neural network model based on the preprocessed mass spectrogram data in combination with charge state limitation and adduct ion probability weight for identifying a fragment peak region in the mass spectrogram, a multi-task loss function is constructed by combining optimization peak region segmentation, charge state prediction and adduct ion matching so as to optimize the process of constructing the double-branch convolutional neural network model; and extracting an optimal fragment peak region from the output of the double-branch convolutional neural network model. According to the invention, by adopting the double-branch convolutional neural network model (including one-dimensional and two-dimensional convolutional layers), local and global features in a mass spectrum can be effectively extracted, so that the recognition precision of fragment peaks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fragment peak region recognition, and more specifically, to a method and device for identifying ion fragment peak regions in mass spectra based on deep learning. Background Art

[0002] Mass spectra usually contain complex background noise and non-target signals (artifacts), which makes it difficult to accurately identify true ion fragment peaks from raw data. Traditional denoising and baseline correction methods are not sufficient to completely eliminate these interfering factors; due to variations in instruments and sample preparation processes, there are significant differences in intensity ranges between different samples, and this variation can lead to deviations in analysis results, affecting subsequent data comparison and interpretation; in mass spectrometry analysis, ions carry different charges and can combine with different adduct ions. Correctly identifying the charge states of these ions and the adduct ions they form is crucial for analyzing complex mixtures, but this is also a technical challenge; in order to accurately identify fragment peaks in a mass spectrum, it is necessary to consider both local details (such as fine structure changes near specific mass numbers) and global information (such as features across multiple neighboring mass numbers) simultaneously. Traditional methods often have difficulty meeting the requirements of both aspects. Therefore, a method and device for identifying ion fragment peak regions in mass spectra based on deep learning are provided. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and device for identifying ion fragment peak regions in mass spectra based on deep learning, so as to solve the problems of complex background noise and artifact interference, the influence of charge states and adduct ions, and the effective extraction of local and global features proposed in the above background art.

[0004] To achieve the above purpose, the present invention aims to provide a method for identifying ion fragment peak regions in mass spectra based on deep learning, including the following steps: S1. Collect mass spectrum data and preprocess the mass spectrum data; S2. Based on the preprocessed mass spectrum data and combined with charge state constraints and adduct ion probability weights, construct a dual-branch convolutional neural network model for identifying fragment peak regions in the mass spectrum, and jointly optimize peak region segmentation, charge state prediction, and adduct ion matching to construct a multi-task loss function to optimize the process of constructing the dual-branch convolutional neural network model; S3. Extract the optimal fragment peak region from the output of the dual-branch convolutional neural network model.

[0005] As a further improvement of this technical solution, in S2, based on the dual-branch convolutional neural network model to identify fragment peak regions in the mass spectrum, it includes the following steps: S2.1. Receive the preprocessed mass spectrometry data as input data, and embed charge state and adduct ion rule information in the input data to generate bimodal input data; S2.2. Construct a one-dimensional convolutional layer and a two-dimensional convolutional layer in the dual-branch convolutional neural network model; Among them, the one-dimensional convolutional layer is used to capture local peak shape details; The two-dimensional convolutional layer is used to capture the global noise distribution pattern; S2.3. Jointly optimize peak region segmentation, charge state prediction, and adduct ion matching to construct a multi-task loss function; S2.4. Use the bimodal input data to pre-train the dual-branch convolutional neural network model. After training, the dual-branch convolutional neural network model outputs a probability heat map, a charge state, and an adduct ion matching score matrix.

[0006] As a further improvement of this technical solution, in S2.1, receiving the preprocessed mass spectrometry data as input data and embedding charge state and adduct ion rule information in the input data to generate bimodal input data includes the following steps: S2.11. Predict the charge state of candidate peaks according to the isotope peak spacing in the mass spectrometry signal; S2.12. For each candidate peak, calculate the mass-to-charge ratio of the formed adduct ion according to its mass-to-charge ratio, and use a Gaussian distribution to assign probability values to each formation result to construct an adduct ion probability weight matrix; S2.13. For each candidate peak, calculate the matching probability between its mass-to-charge ratio and the theoretical mass-to-charge ratio of the adduct ion; S2.14. Combine the original mass spectrometry signal with the adduct ion matching probability and the charge state as the input of the dual-branch convolutional neural network model.

[0007] As a further improvement of this technical solution, in S2.2, constructing a one-dimensional convolutional layer includes the following steps: S2.21. Select convolutional kernels, including small-scale kernels, medium-scale kernels, and large-scale kernels; S2.22. Parallelly use convolutional kernels of different scales to extract local features in different ranges; S2.23. Concatenate the convolutional results at the three scales along the channel dimension to form a multi-channel feature representation, and introduce skip connections to optimize the multi-channel feature representation and retain the original signal features.

[0008] As a further improvement of this technical solution, in S2.2, constructing a two-dimensional convolutional layer includes the following steps: S2.24. Use multi-level 2D convolutional layers to gradually expand the receptive field and capture the global noise pattern; S2.25. Use convolution with a hole rate in the deep layer to expand the receptive field without increasing the number of parameters. S2.26. Use pooling windows of different sizes in parallel to capture multi-scale noise distributions, and splice the multi-scale pooling results into a fixed-length feature vector. S2.27. Add a spatial attention module to enable the network to focus on high-frequency noise regions. S2.28. Preprocess the spectrogram using a high-pass filter to highlight the noise regions. S2.29. Generate a two-dimensional probability map and label the regions with high noise confidence.

[0009] As a further improvement of this technical solution, in S2.3, jointly optimize peak region segmentation, charge state prediction, and adduct ion matching to construct a multi-task loss function, including the following steps: S2.31. Construct a multi-task loss function to simultaneously optimize three subtasks: fragment peak region segmentation, charge state prediction, and adduct ion matching. Among them, the fragment peak region segmentation task uses a weighted Dice loss function as the main loss term; the charge state prediction task uses a cross-entropy loss function; the adduct ion matching task uses an adduct ion contrast loss function based on contrast learning, and optimizes adduct ion matching based on the feature space distance. S2.32. For each candidate peak, construct a weight factor according to its matching probability with the theoretical isotope distribution and adduct ion offset; and use the isotope spacing matching degree and adduct similarity as differentiable gradient adjustment factors to automatically enhance the peak region response that conforms to the rules during backpropagation; at the same time, impose continuous penalties on predictions where the charge state conflicts with the isotope spacing. S2.33. Adopt an adaptive task weighting method based on uncertainty modeling to dynamically adjust the weights of the losses of each subtask according to the uncertainty of the prediction results of each subtask; and use an LSTM network to dynamically generate task weights. The conventional multi-task loss updates the main model parameters, and generates chemical rule-driven meta-gradients through a small-scale MLP to maximize the rule compliance of the validation set.

[0010] As a further improvement of this technical solution, in S3, extract the optimal fragment peak region from the output of the dual-branch convolutional neural network model and optimize the output result, including the following steps: S3.1. Set a confidence threshold m. S3.2. Traverse each pixel point in the segmentation probability map, mark the region with a probability value higher than the confidence threshold m as the "high-confidence fragment peak region", and consider the region with a probability value lower than the confidence threshold m as the background and exclude it. S3.3. Combine the segmentation probability map, charge state prediction, and adduct ion matching probability to define a comprehensive confidence score formula, and further optimize the selection of high-confidence fragment peak regions; S3.4. Re-evaluate the confidence of each fragment peak region according to the comprehensive confidence score; S3.5. Output a list containing all fragment peaks with a confidence greater than m; S3.6. Sort the fragment peak list according to the comprehensive confidence score, and preferentially display the fragment peaks with a confidence greater than m.

[0011] As a further improvement of this technical solution, in S3.3, combining the segmentation probability map, charge state prediction, and adduct ion matching probability to define a comprehensive confidence score formula includes the following steps: S3.31. Utilize the peak region probability heat map output by the dual-branch convolutional neural network model, and directly use the probability values output by the dual-branch convolutional neural network model; S3.32. Calculate the matching score matrix of candidate peaks through a preset adduct ion template, and select the maximum matching score; S3.33. Calculate the charge consistency score according to the deviation between the measured isotope spacing and the theoretical value; S3.34. Perform product fusion on the probability value, the maximum matching score, and the consistency score to calculate the comprehensive confidence score.

[0012] On the other hand, the present invention provides a device for identifying ion fragment peak regions in a mass spectrum based on deep learning, including a sensor, a storage, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for identifying ion fragment peak regions in a mass spectrum based on deep learning described in any one of the above; On the other hand, the device for identifying ion fragment peak regions in a mass spectrum based on deep learning includes a mass spectrum data acquisition module, a model construction and identification module, and a fragment peak extraction module; Among them, the mass spectrum data acquisition module is used to collect mass spectrum data and preprocess the mass spectrum data; The model construction and identification module is used to construct a dual-branch convolutional neural network model based on the preprocessed mass spectrum data in combination with charge state constraints and adduct ion probability weights for identifying fragment peak regions in the mass spectrum; The fragment peak extraction module is used to extract the optimal fragment peak region from the output of the dual-branch convolutional neural network model.

[0013] Compared with the prior art, the beneficial effects of the present invention: 1. In the method and device for identifying ion fragment peak regions in mass spectra based on deep learning, by adopting a dual-branch convolutional neural network model (including one-dimensional and two-dimensional convolutional layers), this method can effectively extract local and global features in mass spectra, thereby improving the recognition accuracy of fragment peaks. In addition, by combining charge state constraints and adduct ion probability weights, and using a differentiable chemical rule engine (DCR) and meta-gradient optimization for deep fusion to achieve multi-task joint optimization, it not only improves the accuracy of fragment peak recognition, but also enhances the chemical rationality of the recognition results, avoiding the recognition of pseudo-peaks without physical meaning.

[0014] 2. In the method and device for identifying ion fragment peak regions in mass spectra based on deep learning, multi-scale information fusion technology is used, including the parallel use of convolutional kernels of different scales, the application of a spatial attention module (CBAM), and high-pass filter preprocessing, etc., enabling the model to better understand and process complex mass spectrometry data. At the same time, by setting a confidence threshold to screen high-confidence fragment peak regions, and further optimizing the selection by combining the segmentation probability map, charge state prediction, and adduct ion matching probability, the interference of background noise or artifacts is effectively reduced, and the quality of the output results is improved. This method significantly reduces the false positive rate while ensuring detection sensitivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is the overall method flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Embodiment 1: Please refer to Figure 1 As shown, this embodiment provides a method for identifying ion fragment peak regions in mass spectra based on deep learning, including the following steps: S1. Collect mass spectrometry data, and preprocess the mass spectrometry data, perform denoising, baseline correction, and normalization on the original mass spectrometry data to eliminate instrument noise and background interference, and unify the intensity ranges of different samples; S2. Based on the preprocessed mass spectrometry data and combined with charge state constraints and adduct ion probability weights, construct a dual-branch convolutional neural network model for identifying fragment peak regions in mass spectra, and jointly optimize peak region segmentation, charge state prediction, and adduct ion matching to construct a multi-task loss function to optimize the process of constructing the dual-branch convolutional neural network model; In this embodiment, identifying the fragment peak region in a mass spectrum based on a dual-branch convolutional neural network model includes the following steps: S2.1. Receive the preprocessed mass spectrum data as input data, and embed charge state and adduct ion rule information in the input data to generate dual-modal input data; Among them, receiving the preprocessed mass spectrum data as input data and embedding charge state and adduct ion rule information in the input data to generate dual-modal input data includes the following steps: S2.11. According to the isotope peak spacing in the mass spectrum signal ( , where is the number of charges), predict the charge state of the candidate peak. For each candidate peak, find its nearby isotope peaks and calculate their mass differences. Based on these mass differences, use a predefined threshold or automatically determine the most likely charge state through an algorithm; S2.12. For each candidate peak, calculate the mass-to-charge ratio of the formed adduct ion according to its mass-to-charge ratio, and use a Gaussian distribution to assign probability values to each formation result to construct an adduct ion probability weight matrix; Specifically: For each candidate peak, based on its mass-to-charge ratio, calculate the mass-to-charge ratio of the formed adduct ion and assign a probability value based on a Gaussian distribution to each possibility to form an adduct ion probability weight matrix; S2.13. For each candidate peak, calculate the matching probability (such as Gaussian kernel similarity) between its mass-to-charge ratio and the theoretical mass-to-charge ratio of the adduct ion; S2.14. Combine the original mass spectrum signal with the adduct ion matching probability and the charge state (such as charge state encoding, adduct ion matching probability) as the input of the dual-branch convolutional neural network model; S2.2. Construct a one-dimensional convolutional layer and a two-dimensional convolutional layer in the dual-branch convolutional neural network model. The purpose is to effectively extract local and global features in the mass spectrum, improve the accuracy of fragment peak recognition, and combine the local features output by the one-dimensional convolutional layer with the global features obtained by the two-dimensional convolutional layer to form a more comprehensive feature representation; Among them, the one-dimensional convolutional layer is used to capture local peak shape details. These layers can effectively identify fine structure changes near specific mass numbers in the mass spectrum and help locate potential fragment peak positions; The two-dimensional convolutional layer is used to capture the global noise distribution pattern. Although the mass spectrum is essentially a one-dimensional signal, by converting it into a two-dimensional representation (for example, time series to image conversion), the two-dimensional convolutional layer can be used to more effectively extract features across multiple adjacent mass numbers, thereby helping to distinguish real fragment peaks from background noise or artifacts; Further, constructing a one-dimensional convolutional layer includes the following steps: S2.21. Select convolution kernels, which include small-scale kernels ((kernel size = 3): detect steep peak edges (such as rising edges and falling edges)), medium-scale kernels ((kernel size = 5): identify peak shapes of medium width), and large-scale kernels ((kernel size = 9): capture the overall contour of wide peaks or overlapping peaks); S2.22. Use convolution kernels of different scales in parallel to extract local features in different ranges, and configure an appropriate number of filters for each scale of convolution operation (for example, set 16 or 32 respectively) to ensure that rich feature representations can be obtained at each scale; S2.23. Concatenate the convolution results at the three scales along the channel dimension to form a multi-channel feature representation. This step allows subsequent layers to integrate information from different scales, thereby enhancing the model's ability to understand complex patterns. Use a 1x1 convolution to reduce the number of channels (i.e., dimensionality reduction) to reduce computational costs and prevent overfitting. This step also helps fuse multi-scale information, generate a more compact feature representation, and introduce skip connections to optimize the multi-channel feature representation, retain the original signal features, and avoid the problem of gradient disappearance.

[0018] Furthermore, construct a two-dimensional convolutional layer, including the following steps: S2.24. Use multi-level 2D convolutional layers to gradually expand the receptive field and capture global noise patterns; S2.25. Use convolutions with a dilation rate (2 in this embodiment) in the deeper layers to expand the receptive field without increasing the number of parameters. Introduce dilated convolutions in the deeper layers and set the dilation rate to 2, so that each convolution kernel can cover a larger area without increasing the number of parameters. Dilated convolutions allow the network to obtain information in a larger range without reducing the resolution, and are particularly suitable for tasks that need to consider larger neighborhood relationships; S2.26. Use pooling windows of different sizes (4x4, 8x8, 16x16) in parallel to capture multi-scale noise distributions, and concatenate the multi-scale pooling results into a fixed-length feature vector; S2.27. Add a spatial attention module (CBAM) to enable the network to focus on high-frequency noise regions (introduce CBAM as a spatial attention module, which includes two parts: channel attention and spatial attention. CBAM first generates a channel attention map through global average pooling, then uses convolution operations to generate a spatial attention map, and finally multiplies these two attention maps with the original feature map to strengthen the feature representation of key regions); S2.28. Preprocess the spectrogram using a high-pass filter (Laplacian operator) to highlight the noise regions (before the input data enters the convolutional neural network, apply the Laplacian operator as a high-pass filter to preprocess the spectrogram. The Laplacian operator can detect rapid changes (i.e., high-frequency components) in the image, thereby highlighting the noise and other fine structures and improving the model's sensitivity to noise). S2.29. Generate a two-dimensional probability map and label the high-noise confidence regions (using the feature map processed through the above steps, generate a two-dimensional probability map through a fully connected layer or an additional convolutional layer, assign a noise confidence score to each pixel position to form a complete noise probability map, and label the high-noise regions where the confidence exceeds the threshold n). S2.3. Jointly optimize peak region segmentation, charge state prediction, and adduct ion matching to construct a multi-task loss function ; S2.4. Pre-train a dual-branch convolutional neural network model using bimodal input data (including mass spectra of charge states and adduct ions). After training, the dual-branch convolutional neural network model outputs a probability heat map, charge states, and an adduct ion matching score matrix.

[0019] Jointly optimize peak region segmentation, charge state prediction, and adduct ion matching to construct a multi-task loss function, including the following steps: This multi-task loss function combines a differentiable chemical rule engine (DCR) with meta-gradient optimization; S2.31. Construct a multi-task loss function to simultaneously optimize three sub-tasks: fragment peak region segmentation, charge state prediction, and adduct ion matching; Among them, for the fragment peak region segmentation task, a weighted Dice loss function is used as the main loss term to optimize the accuracy of predicting the fragment peak region; for the charge state prediction task, a cross-entropy loss function is used to train the model to correctly classify the charge states of each candidate peak; for the adduct ion matching task, an adduct ion contrast loss function based on contrastive learning is used, and the adduct ion matching is optimized based on the feature space distance to measure the matching degree between the candidate peak and the theoretical mass-to-charge ratio of the adduct ion, thereby enhancing the model's ability to model the adduct ion pattern; Weighted Dice loss function is: ; In the formula, represents the weight factor; represents the true label (0 or 1) of the i-th sample; represents the predicted probability of the i-th sample; Cross-entropy loss function is: ; In the formula, Denote the true charge state label of the th sample; Denote the probability distribution of the predicted charge state of the th sample; Denote the sample index; Adduct ion matching optimization based on feature space distance is: ; In the formula, denotes the feature vector of the anchor point (true adduct peak), denotes the feature vector of the positive sample (peaks of the same type), denotes the feature vector of the negative sample (noise peak); denotes the boundary value, which controls the minimum distance difference between positive and negative samples; S2.32. For each candidate peak, according to its matching probability with the theoretical isotope distribution (isotope spacing , adduct shift +21.9819 Da) and the adduct ion shift, construct a weight factor, which is used to amplify the contribution of this peak in the loss function, making the model more inclined to learn chemically reasonable fragment peak patterns; Among them, according to its matching probability with the theoretical isotope distribution and the adduct ion shift, the constructed weight factor is (the weight factor guides the model to focus on regions that conform to chemical rules by amplifying the contributions of these peaks with high chemical rationality in the loss function, avoiding learning pseudo-peaks without physical meaning. In multi-task learning, the loss functions of different sub-tasks (such as peak region segmentation, charge state prediction, adduct ion matching) need to be jointly optimized. The role of the weight factor is to dynamically adjust the importance of each candidate peak in the loss function; for candidate peaks with high chemical rationality (such as isotope peaks or adduct ion peaks with high matching degrees), assign a larger weight to make their impact on the model greater during training; by introducing the weight factor, the model can more accurately locate the region of true fragment peaks while suppressing the interference of background noise and artifacts): ; In the formula, denotes the isotope spacing matching degree (calculated by Gaussian kernel, differentiable), which reflects the consistency between the candidate peak and its theoretical isotope distribution. If the candidate peak highly matches the theoretical isotope distribution, it indicates that it is more likely to be a true fragment peak; denotes the adduct ion template convolution similarity (differentiable), which reflects the matching degree between the candidate peak and the known adduct ion pattern. A high similarity indicates that this peak may be formed by common adduct ions; denotes the adjustable coefficient, which is used to control the influence of different matching degrees on the weight; Using the isotope spacing matching degree and the adduct similarity as differentiable gradient adjustment factors, the response of the peak region that conforms to the rules is automatically enhanced during backpropagation (by performing backpropagation using these differentiable gradient adjustment factors, the model can pay more attention to the data points that conform to the expected chemical patterns during training. This means that when updating the model parameters, the algorithm tends to increase the confidence of the predicted values near the correct peak positions while reducing the influence of background noise or irrelevant signals. In this way, the true fragment peak region can be more accurately located and identified); At the same time, continuous penalties are imposed on the predictions where the charge state conflicts with the isotope spacing , forcing the model to output structural features that conform to the mass spectrometry fragmentation rules, avoiding the identification results of pseudo-peaks without physical meaning. By explicitly constraining the physical rationality of the model output, the prediction results are more in line with the mass spectrometry fragmentation mechanism. The continuous penalty negatively constrains the model from the physical rule level to avoid contradictions between the charge state and the peak spacing, acting independently on the total loss through a fixed coefficient to emphasize the non-compromisability of physical rules, while the weights of other tasks are dynamically adjusted to optimize the multi-task balance; ; In the formula, represents the measured isotope spacing; represents the observed charge state; represents the predicted charge state; represents the penalty intensity coefficient; S2.33. Adopt an adaptive task weighting method based on uncertainty modeling to dynamically adjust the weights of the losses of each subtask according to the uncertainty of the prediction results of each subtask. In the initial stage, set the peak segmentation task as the main task (e.g., the initial weight is 0.6), and the charge state and adduct ion matching tasks each account for a secondary weight (e.g., 0.2 each). As the training progresses, automatically adjust the weights of each subtask according to the prediction confidence and loss convergence situation, making the optimization direction more balanced and avoiding a single task from dominating the model learning; and use an LSTM network to dynamically generate task weights. Update the main model parameters with the conventional multi-task loss, generate chemical rule-driven meta-gradients through a small-scale MLP, and maximize the rule compliance of the validation set (in multi-task learning, the difficulties and convergence speeds of different subtasks (such as peak segmentation, charge state prediction, and adduct ion matching) may vary. If the loss function values of some tasks are too large or too small, it may cause these tasks to dominate during the training process while the learning of other tasks is ignored. The initial task weight allocation (the peak segmentation task accounts for 0.6, and the charge state and adduct ion matching each account for 0.2) is to ensure that the model preferentially learns the main task (such as peak segmentation) while taking into account the secondary tasks. As the training progresses, dynamically adjust the weights according to the prediction confidence and loss convergence situation of each task, enabling the model to optimize all tasks more evenly and avoiding the excessive influence of a single task on the overall learning process; using an LSTM network to dynamically generate task weights can quantify the difficulty of each task in real time (such as the IoU of the segmentation task, the charge classification error, the mean contrast loss, etc.) and automatically adjust the task weights according to these indicators. This dynamic adjustment mechanism enables the model to flexibly adapt to changes in task difficulty and maintain the balance of the optimization direction; by introducing chemical rule-driven meta-gradients (generated by a small-scale MLP), the model's ability to understand the mass spectrometry fragmentation rules can be further enhanced on the basis of multi-task learning).

[0020] Among them, use an LSTM network to dynamically generate task weights: ; Among them, represents the difficulty of the segmentation task (1 - batch average IoU); represents the charge classification error; represents the mean contrast loss; represents the th iteration of the th task's weight; Introduce a task discriminator Optimize the shared features (in multi-task learning, different tasks may conflict with the shared features. The peak segmentation task may focus more on local details, while the charge state prediction task may rely on global information. This conflict may make it difficult for the model to optimize multiple tasks simultaneously. By introducing an adversarial task discriminator , it is possible to optimize the shared feature representation so that the shared features can not only meet the requirements of individual tasks but also take into account the overall goals of multiple tasks, effectively alleviating the conflicts between tasks and improving the overall performance of the model): ; In the formula, represents the shared feature; Among them, the main model parameters are updated by the conventional multi-task loss as: ; In the formula, represents the weight of each sub-task loss; represents the weight of the adversarial loss; represents the adversarial loss; Generate chemically rule-driven meta-gradients through a small-scale MLP , and the maximum validation set rule compliance is: ; Parameter update: ; In the formula, represents the gradient of the multi-task loss with respect to the model parameter ; represents the current model parameter; represents the learning rate; Training process: Phase 1 (rule pre-training): Pre-train the DCR module on synthetic data to enforce learning of chemical rules; Phase 2 (dynamic curriculum learning): Gradually increase the complexity of the adduct ion type, and the LSTM dynamically adjusts the task weights; Phase 3 (meta-gradient fine-tuning): Optimize the meta-learner based on real data to improve the rule generalization.

[0021] This method realizes multi-task joint optimization through the deep integration of a differentiable chemical rule engine (DCR) and meta-gradient optimization. Specifically: First, construct a ternary loss function of weighted Dice loss (main task), cross-entropy loss (charge classification), and contrast loss (adduct matching). Among them, the weight factor dynamically fuses the isotope spacing matching degree (calculated by Gaussian kernel) and the adduct template convolution similarity, and enhances the peak region response that conforms to chemical rules through backpropagation; at the same time, design a continuous penalty term Enforce the physical consistency between the predicted charge state and the isotope spacing; secondly, introduce a dynamic curriculum learning mechanism: based on the LSTM network, the difficulty of the real-time quantization segmentation task (1-IoU), the charge classification error, and the mean contrast loss are used to generate adaptive task weights, and the shared features are optimized through an adversarial task discriminator to avoid multi-task conflicts; further, a two-layer meta-learning framework is adopted. The inner layer updates the main model parameters through the conventional multi-task loss and the outer layer generates meta-gradients through the MLP to maximize the compliance of the validation set with chemical rules and achieve the global alignment of the gradient direction with the mass spectrometry fragmentation mechanism; finally, it is trained in three stages: rule pre-training (forced learning of isotope / adduct rules with multi-modal data), dynamic curriculum learning (gradually increasing the complexity of adduct ions and adjusting weights), and meta-gradient fine-tuning (optimizing the generalization ability with real data). This method synchronously improves the fragmentation peak recognition accuracy and chemical rationality through the differentiable encoding of chemical rules, dynamic task collaboration, and meta-learning optimization.

[0022] S3. Extract the optimal fragmentation peak region from the output of the dual-branch convolutional neural network model; In this embodiment, the optimal fragmentation peak region is extracted from the output of the dual-branch convolutional neural network model and the output result is optimized, including the following steps: S3.1. Set a confidence threshold m for screening high-confidence fragmentation peak regions; S3.2. Traverse each pixel point in the segmentation probability map, and mark the region with a probability value higher than the confidence threshold m as the "high-confidence fragmentation peak region". For the region with a probability value lower than the confidence threshold m, it is regarded as background or noise and excluded; S3.3. Combine the segmentation probability map, the predicted charge state, and the adduct ion matching probability to define a comprehensive confidence score formula to further optimize the selection of high-confidence fragmentation peak regions. For a certain region, if its segmentation probability is high, but the corresponding predicted charge state and adduct ion matching probability are low, then its confidence score is reduced; Furthermore, combining the segmentation probability map, the predicted charge state, and the adduct ion matching probability to define a comprehensive confidence score formula includes the following steps: S3.31. Use the peak region probability heat map output by the dual-branch convolutional neural network model, and directly use the probability value output by the dual-branch convolutional neural network model. The probability value directly reflects the model's segmentation ability for the peak region and is the basic index of the comprehensive confidence score; S3.32. Calculate the matching score matrix of candidate peaks through a preset adduct ion template, and select the maximum matching score. Specifically, for each candidate peak, calculate its matching score with each adduct ion template to form a matching score matrix, and select the maximum matching score to represent the most likely adduct ion type; S3.33. Calculate the charge consistency score according to the deviation between the measured isotope spacing and the theoretical value; S3.34. Multiply and fuse the probability value, the maximum matching score, and the consistency score to calculate the comprehensive confidence score. By multiplying and fusing the probability value, the adduct matching score, and the charge consistency score, the comprehensive confidence score aims to strictly screen the fragment peaks that simultaneously conform to the model prediction, the adduct rule, and the charge physical law. Its advantage is to significantly reduce the false detection rate through a non-linear penalty mechanism and ensure that high-confidence peaks need to meet multi-dimensional chemical constraints.

[0023] The comprehensive confidence score is: ; where represents the comprehensive confidence score of the th candidate peak, represents the probability value; represents the maximum matching score; represents the charge consistency score, represents the normalized isotope spacing (i.e., the mass difference between adjacent isotope peaks divided by the charge state); S3.4. Re-evaluate the confidence of each fragment peak region according to the comprehensive confidence score; S3.5. Output a list of all fragment peaks with a confidence greater than m. The information of each peak includes: mass-to-charge ratio (peak center position), intensity (peak height or integral area), charge state (predicted value), adduct ion type (matching result), and comprehensive confidence score; S3.6. Sort the fragment peak list according to the comprehensive confidence score and preferentially display the fragment peaks with a confidence greater than m.

[0024] Example 2: This example provides a device for identifying ion fragment peak regions in a mass spectrum based on deep learning, including a sensor, a storage, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for identifying ion fragment peak regions in a mass spectrum based on deep learning.

[0025] This example provides a device for identifying ion fragment peak regions in a mass spectrum based on deep learning, including a mass spectrum data acquisition module, a model construction and identification module, and a fragment peak extraction module; Among them, the mass spectrum data acquisition module is used to collect mass spectrum data and preprocess the mass spectrum data; The model construction and identification module is used to construct a dual-branch convolutional neural network model based on the preprocessed mass spectrum data and combined with charge state constraints and adduct ion probability weights (isotope distribution, adduct ion rules) to identify the fragment peak regions in the mass spectrum; The fragment peak extraction module is used to extract the optimal fragment peak region from the output of the dual-branch convolutional neural network model.

[0026] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for identifying ion fragment peak regions in mass spectra based on deep learning, characterized in that: The following steps are involved: S1, collecting mass spectrum data and preprocessing the mass spectrum data; S2. Based on the preprocessed mass spectrum data and combined with charge state restrictions and adduct ion probability weights, a two-branch convolutional neural network model is constructed to identify the fragment peak region in the mass spectrum, and a multi-task loss function is constructed by jointly optimizing peak region segmentation, charge state prediction and adduct ion matching to optimize the process of constructing the two-branch convolutional neural network model; S3. Extract the optimal fragmentation peak region from the output of the two-branch convolutional neural network model.

2. The method for identifying ion fragment peak regions of a mass spectrum based on deep learning according to claim 1, characterized in that: In S2, identifying the fragment peak region in the mass spectrum based on a dual-branch convolutional neural network model comprises the following steps: S2.1, receiving the preprocessed mass spectrum data as input data, and embedding charge state and adduct ion rule information into the input data to generate bimodal input data; S2.

2. Construct a one-dimensional convolutional layer and a two-dimensional convolutional layer in a two-branch convolutional neural network model; Among them, the one-dimensional convolution layer is used to capture the local peak details; The 2D convolutional layer is used to capture the global noise distribution pattern; S2.3, jointly optimize peak region segmentation, charge state prediction and adduct ion matching to construct a multi-task loss function; S2.

4. Use bimodal input data to pre-train a two-branch convolutional neural network model. After training, the two-branch convolutional neural network model outputs a probability heat map, charge state, and adduct ion matching score matrix.

3. The method for identifying ion fragment peak regions of a mass spectrum based on deep learning according to claim 2, characterized in that: In S2.1, the pre-processed mass spectrum data is received as input data, and the charge state and adduct ion rule information are embedded in the input data to generate bimodal input data, including the following steps: S2.11, predicting the charge state of the candidate peak based on the isotope peak spacing in the mass spectrometry signal; S2.

12. For each candidate peak, the mass-to-charge ratio of the formed adduct ion is calculated according to its mass-to-charge ratio, and a probability value is assigned to each formation result using Gaussian distribution to construct an adduct ion probability weight matrix; S2.13, for each candidate peak, calculating the probability of its mass-to-charge ratio matching the theoretical adduct ion mass-to-charge ratio; S2.

14. The original mass spectrometry signal is combined with the adduct ion matching probability and the charge state as the input of the two-branch convolutional neural network model.

4. The method for identifying ion fragment peak regions of a mass spectrum based on deep learning according to claim 2, characterized in that: In S2.2, constructing a one-dimensional convolutional layer includes the following steps: S2.

21. Select a convolution kernel, which includes a small-scale kernel, a medium-scale kernel, and a large-scale kernel; S2.22, use convolution kernels of different scales in parallel to extract local features of different ranges; S2.

23. Concatenate the convolution results at three scales along the channel dimension to form a multi-channel feature representation, and introduce skip connections to optimize the multi-channel feature representation to retain the original signal characteristics.

5. The method for identifying ion fragment peak regions of a mass spectrum based on deep learning according to claim 2, characterized in that: In S2.2, constructing a two-dimensional convolutional layer includes the following steps: S2.24, use multiple levels of 2D convolutional layers to gradually expand the receptive field and capture the global noise pattern; S2.25, use dilated convolution in deep layers to expand the receptive field without increasing the number of parameters; S2.26, using pooling windows of different sizes in parallel to capture multi-scale noise distribution, and concatenating the multi-scale pooling results into a fixed-length feature vector; S2.27, add a spatial attention module to make the network focus on high-frequency noise areas; S2.28, use a high-pass filter to preprocess the spectrum to highlight the noise area; S2.

29. Generate a two-dimensional probability map and mark the high noise confidence area.

6. The method for identifying ion fragment peak regions of mass spectra based on deep learning according to claim 2, characterized in that: In S2.3, the multi-task loss function is constructed by jointly optimizing peak region segmentation, charge state prediction and adduct ion matching, including the following steps: S2.31, construct a multi-task loss function to simultaneously optimize the three subtasks of fragment peak region segmentation, charge state prediction and adduct ion matching; Among them, the weighted Dice loss function is used as the main loss term for the fragment peak region segmentation task; the cross entropy loss function is used for the charge state prediction task; the adduct ion matching task adopts the adduct ion contrast loss function based on contrastive learning, and the adduct ion matching is optimized based on the feature space distance; S2.

32. For each candidate peak, a weight factor is constructed based on the probability of matching it with the theoretical isotope distribution and the adduct ion offset; the isotope spacing matching and adduct similarity are used as differentiable gradient adjustment factors to automatically enhance the peak region response that meets the rules during back propagation; at the same time, a continuous penalty is imposed on the prediction of the charge state and the isotope spacing being inconsistent; S2.

33. Adopt an adaptive task weighting method based on uncertainty modeling to dynamically adjust the weight of each task loss according to the uncertainty of the prediction results of each subtask; and use the LSTM network to dynamically generate task weights, update the main model parameters with conventional multi-task losses, and generate chemical rule-driven meta-gradients through small-scale MLP to maximize the compliance of the verification set rules.

7. The method for identifying ion fragment peak regions of a mass spectrum based on deep learning according to claim 1, characterized in that: In S3, the optimal fragment peak region is extracted from the output of the dual-branch convolutional neural network model, and the output result is optimized, including the following steps: S3.1, set the confidence threshold m; S3.2, traverse each pixel in the segmentation probability map, mark the area with probability value higher than the confidence threshold m as "high confidence fragment peak area", and regard the area below the confidence threshold m as background and exclude it; S3.3, combining the segmentation probability map, charge state prediction and adduct ion matching probability to define a comprehensive confidence scoring formula to further optimize the selection of high-confidence fragment peak regions; S3.

4. Re-evaluate the confidence of each fragment peak region based on the comprehensive confidence score; S3.5, output a list containing all fragment peaks with confidence greater than m; S3.

6. Sort the fragment peak list according to the comprehensive confidence score, and give priority to displaying the fragment peaks with a confidence score greater than m.

8. The method for identifying ion fragment peak regions of a mass spectrum based on deep learning according to claim 7, characterized in that: In S3.3, the segmentation probability map, charge state prediction and adduct ion matching probability are combined to define a comprehensive confidence scoring formula, including the following steps: S3.31, using the peak area probability heat map output by the two-branch convolutional neural network model, and directly using the probability value output by the two-branch convolutional neural network model; S3.32, calculating the matching score matrix of the candidate peaks by using the preset adduct ion template, and selecting the maximum matching score; S3.33, calculate the charge consistency score based on the deviation between the measured isotope spacing and the theoretical value; S3.

34. Multiply the probability value, the maximum matching score and the consistency score to obtain a comprehensive confidence score.

9. A mass spectrum ion fragmentation peak region recognition device based on deep learning, comprising a sensor, a storage, a processor, and a computer program stored in the storage and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for identifying ion fragment peak regions in a mass spectrum based on deep learning as described in any one of claims 1 to 8 are implemented.

10. The mass spectrum ion fragment peak region identification device based on deep learning according to claim 9, characterized in that: It includes a mass spectrum data acquisition module, a model building and recognition module, and a fragment peak extraction module; Wherein, the mass spectrum data acquisition module is used to collect mass spectrum data and pre-process the mass spectrum data; The model building recognition module is used to build a two-branch convolutional neural network model based on the preprocessed mass spectrum data and combined with the charge state restriction and the adduct ion probability weight to identify the fragment peak area in the mass spectrum; The fragment peak extraction module is used to extract the optimal fragment peak region from the output of the two-branch convolutional neural network model.

Citation Information

Patent Citations

  • Mass spectrum image super-resolution reconstruction method based on deep learning

    CN108062744A

  • Vocs mass spectrum ion fragment peak region identification method and device based on deep learning

    CN116500118A

  • Zero sample learning-based mass spectrum image super-resolution reconstruction method

    CN118014843A

  • Mass spectrum processing apparatus and method

    EP3683823A1

  • Systems and methods for patient-specific identification of neoantigens by de novo peptide sequencing for personalized immunotherapy

    US20200243164A1

Cited By

  • Mycobacterium tuberculosis drug-resistant substance screening method based on mass spectrometric detection

    CN120254021A

  • Marine ecological environment influence tracking and monitoring method based on offshore wind power engineering

    CN120410274A

  • Mycotoxin non-targeted screening mass spectrum data processing method based on deep learning

    CN121617459A

  • Deep learning based mycotoxin non-targeted screening mass spectrometry data processing method

    CN121617459B