Transform-based mass spectrum data noise reduction method, apparatus and device, and medium

By using a mass spectrometry data denoising model based on Transformer neural network and combining it with a cosine-weighted contrast loss function, the problem of unsatisfactory performance of existing mass spectrometry data denoising algorithms is solved, achieving more efficient noise removal and retention of useful information, and improving the accuracy and adaptability of mass spectrometry detection.

CN121071306APending Publication Date: 2025-12-05SHENZHEN HYMSON LASER INTELLIGENT EQUIP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511183742.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing mass spectrometry data denoising algorithms are not ideal, failing to effectively remove noise and affecting the accuracy and reliability of mass spectrometry detection.

Method used

A denoising model based on Transformer neural network is adopted and trained with cosine weighted contrast loss function. A multi-head self-attention mechanism is used to capture long-distance dependencies in mass spectrometry data, and class imbalance is handled by cosine similarity and class weight.

Benefits of technology

It improves the accuracy and reliability of mass spectrometry data denoising, effectively removes global noise, retains useful information to the greatest extent, and is suitable for various types of mass spectrometry data and complex scenarios, reducing human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071306A_ABST
    Figure CN121071306A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing and artificial intelligence, and discloses a Transform-based mass spectrum data noise reduction method, device, equipment and medium, and the method comprises the steps: inputting a data set into a pre-training noise reduction model, and training the pre-training noise reduction model based on a cosine weighted contrast loss function, a cosine weighted comparison loss function is constructed based on cosine similarity and category weight; and carrying out noise reduction processing on the mass spectrum data to be subjected to noise reduction by utilizing the trained noise reduction model. According to the method, the long-distance dependency relationship in the mass spectrum data can be effectively captured, more useful information is reserved to the maximum extent, global noise is effectively removed, and the noise reduction effect is improved; loss calculation is carried out in combination with cosine similarity and class imbalance, so that the training precision of a noise reduction model can be improved, and the accuracy and reliability of mass spectrum detection noise reduction are further improved; the method can be suitable for automatic noise reduction processing of various mass spectrum data and various complex scenes, and manual intervention is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing and artificial intelligence, and particularly relates to a mass spectrum data denoising method and device based on a Transformer, equipment and a medium. BACKGROUND

[0002] Mass spectrum detection technology is an analysis technology for determining the composition and structure of a substance by measuring the mass-to-charge ratio of ions, and is widely used in the fields of chemistry, biology, medicine, environment, etc. However, in the mass spectrum detection process, due to the noise of the instrument itself, external interference and other factors, the mass spectrum data collected often contains a large amount of noise, which will affect the subsequent processing such as identification and quantitative analysis of mass spectrum peaks, and reduce the accuracy and reliability of mass spectrum detection.

[0003] At present, common mass spectrum data denoising algorithms mainly include wavelet transform-based denoising algorithms, Fourier transform-based denoising algorithms, etc. The wavelet transform-based denoising algorithm removes noise by wavelet decomposition and threshold processing of the signal, but the selection of the threshold greatly affects the denoising effect, and it is difficult for humans to grasp the selection of the threshold, and the denoising effect is uncertain; the Fourier transform-based denoising algorithm performs filtering processing by converting the signal to the frequency domain, but the denoising effect on non-stationary signals is not ideal. SUMMARY

[0004] The main purpose of the present application is to provide a mass spectrum data denoising method and device based on a Transformer, equipment and a medium, aiming to solve the technical problems that the denoising effect of the prior art mass spectrum data is difficult to control and is not ideal.

[0005] The first aspect of the present application provides a mass spectrum data denoising method based on a Transformer, comprising: inputting a data set into a pre-trained denoising model, and training the pre-trained denoising model based on a cosine weighted contrast loss function to obtain a trained denoising model, wherein the data set contains multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data, the pre-trained denoising model is constructed based on a Transformer neural network, and the cosine weighted contrast loss function is constructed based on cosine similarity and category weight; performing denoising processing on the to-be-denoised mass spectrum data by using the trained denoising model to obtain target denoised mass spectrum data.

[0006] The present application also provides a mass spectrum data denoising device based on a Transformer, comprising: a model training module configured to input a data set into a pre-trained denoising model and train the pre-trained denoising model based on a cosine weighted contrastive loss function to obtain a trained denoising model, wherein the data set contains multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data, the pre-trained denoising model is constructed based on a Transformer neural network, and the cosine weighted contrastive loss function is constructed based on a cosine similarity and a class weight; a denoising processing module configured to perform denoising processing on to-be-denoised mass spectrum data by using the trained denoising model to obtain target denoised mass spectrum data.

[0007] The third aspect of the present application provides a computer device, comprising a memory and at least one processor, the memory stores instructions; the at least one processor invokes the instructions in the memory to enable the computer device to execute the above-mentioned Transformer-based mass spectrum data denoising method.

[0008] The fourth aspect of the present application provides a computer readable storage medium, which stores instructions, when running on a computer, enables the computer to execute the above-mentioned Transformer-based mass spectrum data denoising method.

[0009] The present application provides a Transformer-based mass spectrum data denoising method, which processes the noisy data collected by a mass spectrum detector by constructing a denoising model based on a Transformer neural network, can effectively capture the long-distance dependency relationship in the mass spectrum data, retain more useful information to the greatest extent, effectively remove global noise, and improve the denoising effect; the present application combines cosine similarity and class imbalance for loss calculation, which can improve the training accuracy of the denoising model, and further improve the accuracy and reliability of mass spectrum detection denoising; the denoising algorithm of the present application can be applied to automatic denoising processing of various mass spectrum data and various complex scenes, and greatly reduces the manual intervention. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 a first embodiment flowchart of the Transformer-based mass spectrum data denoising method in the embodiments of the present application; Figure 2 a connection diagram between an encoder and a decoder in the denoising model in the embodiments of the present application; Figure 3 a functional module diagram of one embodiment of the Transformer-based mass spectrum data denoising device in the embodiments of the present application; Figure 4 a diagram of one embodiment of the computer device in the embodiments of the present application. DETAILED DESCRIPTION

[0011] The terms "first", "second", "third", "fourth" and the like in the description and claims of this application and in the above Summary, if any, are used for distinguishing between similar objects and are not necessarily used to describe a particular sequential or chronological order. It is to be understood that the use of these terms herein is to be construed to cover a non-sequential process unless expressly indicated otherwise. Further, the use of terms such as "including", "comprising" or "having" and variations thereof herein is intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises a list of steps or elements is not necessarily limited to those steps or elements that are expressly identified, but can include other steps or elements not expressly listed or inherent to such process, method, product or apparatus.

[0012] Mass spectrometry is an analytical technique that determines the composition and structure of substances by measuring the mass-to-charge ratio of ions, and is widely used in the fields of chemistry, biology, medicine, environment, etc. However, in the process of mass spectrometry, due to the noise of the instrument itself, external interference and other factors, the mass spectrometry data collected often contains a large amount of noise, which will affect the identification of mass spectrometry peaks, quantitative analysis and other subsequent processing, and reduce the accuracy and reliability of mass spectrometry.

[0013] At present, the commonly used mass spectrometry data denoising algorithms mainly include wavelet transform based denoising algorithm, Fourier transform based denoising algorithm, Kalman filter based denoising algorithm, etc. The wavelet transform based denoising algorithm removes noise by wavelet decomposition and threshold processing of the signal, but the selection of the threshold greatly affects the denoising effect, and it is difficult for people to grasp the selection of the threshold, and the denoising effect is uncertain; the Fourier transform based denoising algorithm converts the signal to the frequency domain for filtering processing, but the denoising effect on non-stationary signals is not ideal; the Kalman filter based denoising algorithm needs to establish an accurate and complex system model, which is difficult to apply in complex scenes.

[0014] With the development of deep learning technology, some neural network based denoising algorithms are applied to mass spectrometry data processing, such as convolutional neural network (CNN) based denoising algorithm. CNN has the ability to extract local features and can effectively remove local noise, but it is difficult to capture long-distance dependencies in mass spectrometry data, and the denoising effect needs to be further improved.

[0015] Based on this, the present application provides a mass spectrometry data denoising scheme based on Transformer.

[0016] Reference Figure 1 The present application provides a mass spectrometry data denoising method based on Transformer, which comprises: S100: input a data set into a pre-trained denoising model, and train the pre-trained denoising model based on a cosine weighted contrastive loss function to obtain a trained denoising model, wherein the data set contains multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data, the pre-trained denoising model is constructed based on a Transformer neural network, and the cosine weighted contrastive loss function is constructed based on cosine similarity and a category weight.

[0017] Specifically, sample noisy mass spectrum data (i.e., mass spectrum data containing noise) and its corresponding sample denoised mass spectrum data as a true label are obtained. A data set is constructed with multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data.

[0018] Before obtaining the data set, data acquisition and preprocessing can also be included, that is, collecting noisy mass spectrum data and clean mass spectrum data output by a mass spectrum detector, and preprocessing the data to obtain standardized sample noisy mass spectrum data and sample denoised mass spectrum data.

[0019] The preprocessing includes data normalization and data division. The data normalization is to map the mass spectrum data to the interval [0, 1] or [-1, 1] to eliminate the influence of different magnitudes of data. The data division is to divide the preprocessed data into a training set, a validation set and a test set; wherein the training set is used for model training, the validation set is used for adjusting model hyperparameters, and the test set is used for evaluating model performance.

[0020] A pre-trained denoising model based on a Transformer neural network is constructed, which includes an input layer, an encoder, a decoder and an output layer.

[0021] The input layer converts the preprocessed sample data as input data into a feature vector that can be processed by the model.

[0022] The encoder adopts the encoder structure of the Transformer, which is used for feature extraction of the input data, and the decoder adopts the decoder structure of the Transformer, which is used for generation of the denoising signal according to the features extracted by the encoder.

[0023] In one specific embodiment, the encoder includes a multi-head self-attention mechanism layer and a feedforward neural network layer. The multi-head self-attention mechanism layer can capture long-distance dependencies in the input data from different angles at the same time, and the feedforward neural network layer is used for nonlinear transformation of the features output by the multi-head attention mechanism layer to enhance the expression ability of the model.

[0024] The decoder comprises a multi-head self-attention mechanism layer, an encoder-decoder attention mechanism layer and a feedforward neural network layer, the encoder-decoder attention mechanism layer is used for making the decoder pay attention to key features extracted by the encoder, and improving the generation quality of the denoised signal.

[0025] The output layer converts the feature vector output by the decoder into final denoised mass spectrum data.

[0026] In the model training process, the preprocessed sample noisy mass spectrum data and the corresponding clean mass spectrum data (i.e., sample denoised mass spectrum data) are used as training samples to train the constructed denoising model, the model parameters are adjusted through the back propagation algorithm, the error between the denoised data output by the model and the sample denoised mass spectrum data as the label is minimized, and a trained denoising model is obtained.

[0027] In the training process, the embodiment uses a cosine-weighted contrastive loss function (Cosine-Weighted Contrastive Loss), which is constructed based on cosine similarity and class weight. More specifically, the cosine similarity of the feature vectors of any two samples in the same batch is calculated, and the class weight of each class is calculated; the cosine-weighted contrastive loss function is calculated according to the cosine similarity and the class weight corresponding to all samples in the same batch. The cosine similarity can measure the similarity of the feature vectors, which is conducive to learning discriminative features by the model; the class weight can handle the problem of class imbalance.

[0028] Based on the cosine-weighted contrastive loss function and the back propagation algorithm, the model parameters of the pre-trained denoising model are updated by gradient, and a trained denoising model is obtained.

[0029] During the training process, the performance of the model is monitored by the validation set, and when the performance of the model on the validation set no longer improves, the training is stopped and the trained model parameters are saved.

[0030] In one specific embodiment, the random gradient descent algorithm, Adam algorithm or RMSprop algorithm is used for optimization during model training.

[0031] S200: Denoising the to-be-denoised mass spectrum data by using the trained denoising model to obtain target denoised mass spectrum data.

[0032] Specifically, the mass spectrum data to be denoised is input into the trained denoising model, and the model output is the target denoised mass spectrum data.

[0033] The embodiment adopts a denoising model based on a Transformer neural network. The multi-head self-attention mechanism in the Transformer can effectively capture long-distance dependencies in the mass spectrum data, and can better understand the global features of the mass spectrum data compared with traditional denoising algorithms, thereby improving the denoising effect.

[0034] The denoising algorithm of the embodiment can maximize the retention of useful signals while removing noise, thereby avoiding the loss of useful information and facilitating subsequent mass spectrum peak identification, quantitative analysis and other processing.

[0035] The cosine weighted contrast loss function used in the embodiment can measure the similarity of the feature vectors, which is conducive to the learning of discriminative features by the model. The class weight can process the class imbalance problem, so that the denoising model training is more accurate.

[0036] The denoising algorithm of the embodiment has strong adaptability and generalization ability, and can be applied to the denoising of mass spectrum data of different types of mass spectrum detectors and different detection scenarios. It can denoise both stationary signals and non-stationary signals, and does not require manual threshold setting and other interventions.

[0037] The embodiment constructs a denoising model based on a Transformer neural network to process the noisy data collected by the mass spectrum detector. The model can effectively capture long-distance dependencies in the mass spectrum data, maximize the retention of more useful information, effectively remove global noise, and improve the denoising effect. The embodiment combines cosine similarity and class imbalance for loss calculation, which can improve the training accuracy of the denoising model, thereby improving the accuracy and reliability of mass spectrum detection and denoising. The denoising algorithm of the embodiment can be applied to automatic denoising processing of various types of mass spectrum data and various complex scenarios, and greatly reduces manual intervention.

[0038] In one embodiment, the cosine weighted contrast loss function is shown in the following formula 1: Formula 1 Wherein, N is the batch size; is the true class of sample i, is the total number of classes; is the weight of class . is the cosine similarity of the feature vector of sample i and the feature vector of sample . is the scaling factor of the same class sample, which is used to control the punishment intensity of the similarity of the same class sample, and .​ is a scaling factor for different class samples, used to control the penalty strength for the similarity of different class samples, and ; is a threshold parameter, used to define the maximum acceptable similarity of different class samples.

[0039] Specifically, N is the batch size, specifically the number of samples in a batch.

[0040] For sample mass spectrum data, if it contains the same substance composition, it belongs to the same class sample; if it contains different substance compositions, it belongs to different class samples.

[0041] is the weight of the class , used to handle class imbalance.

[0042] .

[0043] Specifically, the cosine similarity of the feature vector of sample i in the same batch (batch) and the feature vector of sample .

[0044] In one embodiment, the same class sample or different class sample can be indicated by indicating the vector , specifically as follows:

[0045] This embodiment directly measures the similarity of the feature vectors by using the cosine similarity, which is conducive to learning discriminative features, and the class weight can handle the class imbalance problem. For the same class sample, the lower the similarity, the greater the penalty; for different class samples, only when the similarity exceeds the threshold , the penalty is applied.

[0046] In addition, in another embodiment, the attention degree to the same class and different class samples can also be balanced by adjusting .

[0047] In one embodiment, the scaling factor of the same class sample can be dynamically adjusted in different iteration rounds of the model training process; specifically as follows: Based on the average cosine similarity of the same class sample pair, the expected average cosine similarity of the same class sample, the value of the scaling factor of the same class sample in the last iteration round, the adjustment coefficient for controlling the adjustment range of the scaling factor of the same class sample, the maximum value (upper limit) and the minimum value (lower limit) allowed for the scaling factor of the same class sample, the value of the scaling factor of the same class sample in the next iteration round is calculated. ​​

[0048] In one embodiment, the scaling factor of the same class sample can be dynamically adjusted in different iteration rounds of the model training process. The iterative relationship of the scaling factor of the same class sample is shown in the following formula 2: Formula 2 Wherein, is the scaling factor of the same class sample in the (t+1)th iteration round, is the scaling factor of the same class sample in the tth iteration round; is the average value of the cosine similarity of all same class sample pairs in the tth iteration round of training; is the expected average cosine similarity of the same class sample; k is an adjustment coefficient for controlling the adjustment range of the scaling factor of the same class sample; λ max , λ min are the upper and lower limit values of λ, respectively.

[0049] Specifically, ∈[−1,1]. In actual scenarios, if the similarity of the same class sample is a non-negative value, ∈[0,1].

[0050] t is greater than or equal to 1.

[0051] k>0, for controlling the adjustment range of λ. In one specific embodiment, k=0.5. Of course, the value of k can be set according to the actual scenario, which is not limited in the present application.

[0052] Setting the upper and lower limit values of λ max , λ min can avoid extreme values.

[0053] In one specific embodiment, λ min =0.1, λ max =10.0.

[0054] Of course, λ max , λ min can be set according to the actual scenario, which is not limited in the present application.

[0055] The clip() function is a clipping function for limiting the value within a specified range, ensuring that it is between the specified minimum and maximum values.

[0056] In one embodiment, the expected average cosine similarity of the same class sample can be updated with the iteration rounds by the following steps: Calculate the cosine similarity of all same class sample pairs in each iteration round of training; Take the p% quantile of all cosine similarities as the expected average cosine similarity of the same sample period corresponding to the iteration round.

[0057] Specifically, each iteration training uses a batch of samples, and the expected average cosine similarity of the same sample is Updated with the iteration training round.

[0058] The value of p may be, for example, 90, 85, 80, etc.

[0059] In one specific embodiment, by calculating the cosine similarity of all pairs of same samples in an iteration round, taking the 90% quantile as the initial value, and with each iteration, the similarity distribution of the current same sample is recalculated, and the latest 90% quantile is updated to keep the expected same sample feature vector consistent.

[0060] In one embodiment, the scaling factor of different class samples can be dynamically adjusted in different iteration rounds of the model training process; Specifically as follows: Based on the proportion of sample pairs with cosine similarity exceeding threshold parameter τ among all different class sample pairs, the maximum proportion of different class sample pairs with cosine similarity exceeding threshold parameter τ among acceptable different class sample pairs, the adjustment coefficient for controlling the adjustment range of the scaling factor of different class samples, the maximum value (upper limit) and minimum value (lower limit) allowed for the scaling factor of different class samples, and the value of the scaling factor of different class samples in the last iteration round, the value of the scaling factor of different class samples in the next iteration round is calculated.

[0061] In one embodiment, the scaling factor of different class samples can be dynamically adjusted in different iteration rounds of the model training process; The iteration relationship of the scaling factor of different class samples is shown in the following formula 3: Formula 3 Wherein, is the scaling factor of different class samples in the (t+1) iteration round, is the scaling factor of different class samples in the t iteration round; is the proportion of sample pairs with cosine similarity exceeding threshold parameter τ among all different class sample pairs in the t iteration training; is the maximum proportion of different class sample pairs with cosine similarity exceeding threshold parameter τ among acceptable different class sample pairs; m is the adjustment coefficient for controlling the adjustment range of the scaling factor of different class samples; μ max , μ min are the upper and lower limit values of μ, respectively.

[0062] Specifically, the ratio of the sample pairs whose cosine similarity exceeds the threshold parameter τ in all different-class sample pairs ranges from 0 to 1.

[0063] m>0, used to control the adjustment range of μ, such as m=0.5. Of course, the value of m can be set according to the actual scene, and the present application does not limit this.

[0064] The upper and lower limit values of μ are set as μ max , μ min The extreme value can be avoided.

[0065] μ max , μ min The value of can be set according to the actual scene, and the present application does not limit this.

[0066] In one embodiment, the maximum ratio of different-class sample pairs whose cosine similarity exceeds the threshold parameter τ that can be accepted can be updated with each iteration round by the following steps: Calculate the cosine similarity of all different-class sample pairs in each iteration training, and calculate the ratio of sample pairs whose cosine similarity exceeds the threshold parameter τ; Take the q% quantile of all sample pairs whose cosine similarity exceeds the threshold parameter τ as the maximum ratio of different-class sample pairs whose cosine similarity exceeds the threshold parameter τ that can be accepted in the corresponding iteration round .

[0067] Specifically, one batch of samples is used for each iteration training, and the ratio of sample pairs whose cosine similarity exceeds the threshold parameter τ in different-class sample pairs is updated with each iteration round.

[0068] The value of q can be, for example, 10, 15, 20, etc.

[0069] In one specific embodiment, the ratio of sample pairs whose cosine similarity exceeds τ in all different-class sample pairs is calculated, the 10% quantile is taken as the initial value, and then the ratio of different-class samples exceeding the threshold value is recalculated every round. The latest 10% quantile is updated, indicating that only 10% or less of different-class sample pairs are expected to exceed τ.

[0070] In one embodiment, the value of the threshold parameter τ can be adaptively and dynamically adjusted, or, receive and respond to user adjustment instructions for the threshold parameter τ to adjust the value of the threshold parameter.

[0071] More specifically, τ defines the "tolerance upper limit" of the similarity of different class sample features (τ∈[-1, 1]). When there is a certain reasonable similarity between different class samples (such as the peak structure of some substances is close), τ needs to be increased to avoid excessive punishment; when it is necessary to strictly distinguish different class samples (such as high-precision detection scenarios), τ needs to be reduced to strictly limit the similarity.

[0072] For different scenarios, τ needs to be adjusted, such as in high-precision detection scenarios, different class samples must be strictly distinguished (such as toxic substances and ordinary substances), even if si,j=0.3 needs to be vigilant, τ needs to be reduced (such as adjusted to 0.2), and the similarity exceeding 0.2 is punished to strengthen the distinction. When the reasonable feature similarity of different class samples (such as substances with completely opposite properties, whose mass spectrum signal features naturally have strong anti-correlation) is negative, τ needs to be set to a negative value.

[0073] In one embodiment, if the scaling factor of the same class sample , the scaling factor of the different class sample and the threshold parameter τ are updated with the iteration round, the above formula 1 can be updated to the following formula 1-1: Formula 1-1 Of course, it needs to be pointed out that the scaling factor of the same class sample , the scaling factor of the different class sample and the threshold parameter τ can be partially updated, partially not updated, and the three are not required to be updated synchronously; according to the actual application scenario setting, this application does not limit it.

[0074] In one embodiment, the pre-trained denoising model comprises an unequal number of encoders and decoders, and the number of encoders is not less than the number of decoders, the encoders are sequentially connected in order, the decoders are sequentially connected in order, the output end of the last encoder is connected to the input end of the first decoder, and the input end of part of other decoders is respectively connected to the output end of one other encoder. The other encoder is an encoder other than the last encoder, and the other decoder is a decoder other than the first decoder.

[0075] Specifically, the output end of the previous encoder is connected to the input end of the next encoder, and the output end of the previous decoder is connected to the input end of the next decoder.

[0076] In addition to the last encoder, the output end of part of the other encoders can be connected to the output end of one other decoder. In this way, the transmission of shallow, middle and deep features can be increased, the multi-scale features can be fused by the decoder, the more comprehensive data information can be captured, and the loss of small signals caused by single deep features or the insufficient noise suppression caused by single shallow features can be avoided.

[0077] In one embodiment, the other encoders are the encoders other than the first encoder and the last encoder, the other decoders are the decoders other than the first decoder and the last decoder, one encoder is connected to at most one decoder, one decoder is connected to at most one encoder, and the other encoders connected to the decoders are not adjacent.

[0078] In order to avoid repeated reception of information and data redundancy and increase the amount of calculation, in this embodiment, the output ends of some other encoders are respectively connected to the input ends of different other decoders, and the other encoders connected to the decoders are not adjacent. For example, if the second encoder is connected to a decoder, then the first encoder and the third encoder cannot be connected to any decoder, and so on. In this way, multi-scale feature fusion can be achieved, as much useful information as possible can be retained, and data redundancy and calculation amount can be reduced.

[0079] The trained denoising model has the same structure as the pre-trained denoising model.

[0080] Figure 2 FIG. 1 is a schematic diagram of the connection between the encoders and the decoders in the denoising model according to an embodiment of the present application. Figure 2 For example, there are 12 encoders and 6 decoders. The second encoder is connected to the fifth decoder, the fourth encoder is connected to the fourth decoder, the sixth encoder is connected to the third decoder, the eighth encoder is connected to the second decoder, and the twelfth encoder is connected to the first decoder. This connection mode enhances the transmission of features in the shallow layer, the middle layer and the deep layer. The decoders fuse these multi-scale features to avoid small signal loss caused by single deep layer features or insufficient noise suppression caused by single shallow layer features. Considering the redundancy of deep layer features, the feature transmission of the tenth layer encoder is not performed. The outputs of the ninth to twelfth layer encoders have deeply aggregated global information, and a transmission will be performed after the twelfth layer. In order to avoid repeated reception of information by the decoders and increase the amount of calculation, the outputs of the ninth to twelfth layer encoders are not input to other decoders.

[0081] FIG. 1 is a schematic diagram of the connection between the encoders and the decoders in the denoising model according to an embodiment of the present application. Figure 3 The application further provides a Transformer-based mass spectrum data denoising device, which comprises: The model training module 100 is configured to input a data set into a pre-trained denoising model and train the pre-trained denoising model based on a cosine weighted contrastive loss function to obtain a trained denoising model, wherein the data set contains multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data, the pre-trained denoising model is constructed based on a Transformer neural network, and the cosine weighted contrastive loss function is constructed based on a cosine similarity and a category weight. The denoising processing module 200 is configured to perform denoising processing on to-be-denoised mass spectrum data by using the trained denoising model to obtain target denoised mass spectrum data.

[0082] In one embodiment, the cosine weighted contrastive loss function is as shown in the following formula 1. Formula 1 Wherein N is a batch size. is a true category of sample i, is a total number of categories. is a weight of category . is a feature vector of sample i and a cosine similarity between the feature vector of sample . is a scaling factor of same-category samples, used to control the penalty strength of the similarity of same-category samples, and . is a scaling factor of different-category samples, used to control the penalty strength of the similarity of different-category samples, and . is a threshold parameter, used to define the maximum acceptable similarity of different-category samples.

[0083] In one embodiment, the scaling factor of same-category samples can be dynamically adjusted in different iteration rounds of the model training process. The iteration relationship of the scaling factor of same-category samples is as shown in the following formula 2. Formula 2 Wherein is the scaling factor of same-category samples in the (t+1)th iteration round, is the scaling factor of same-category samples in the tth iteration round. is an average value of the cosine similarity of all pairs of same-category samples in the tth iteration round. ​an expected average cosine similarity of the same class samples; k is an adjustment coefficient for controlling the adjustment range of the scaling factor of the same class samples; λ max , λ min are upper and lower limit values of λ, respectively.

[0084] In an embodiment, the Transformer-based mass spectrometry data denoising device further comprises: a same class sample expected average cosine similarity calculation and updating module configured to calculate the cosine similarity of all same class sample pairs in each iteration of training, and take the p% quantile of all cosine similarities as the expected average cosine similarity of the same class samples in the corresponding iteration round.

[0085] In an embodiment, the scaling factor of different class samples can be dynamically adjusted in different iteration rounds of the model training process. The iteration relationship of the scaling factor of different class samples is shown in the following formula 3: Formula 3 wherein, is the scaling factor of different class samples in the (t+1)th iteration round, is the scaling factor of different class samples in the tth iteration round; is the proportion of sample pairs whose cosine similarity exceeds the threshold parameter in all different class sample pairs in the tth iteration of training; is the maximum proportion of acceptable different class sample pairs whose cosine similarity exceeds the threshold parameter; m is an adjustment coefficient for controlling the adjustment range of the scaling factor of different class samples; μ max , μ min are upper and lower limit values of μ, respectively.

[0086] In an embodiment, the Transformer-based mass spectrometry data denoising device further comprises a maximum proportion calculation and updating module configured to calculate the cosine similarity of all different class sample pairs in each iteration of training, and calculate the proportion of sample pairs whose cosine similarity exceeds the threshold parameter; take the q% quantile of all sample pairs whose cosine similarity exceeds the threshold parameter as the maximum proportion of acceptable different class sample pairs whose cosine similarity exceeds the threshold parameter in the corresponding iteration round.

[0087] In one embodiment, the noise reduction model comprises an unequal number of encoders and decoders, and the number of encoders is not less than the number of decoders, the encoders are sequentially connected in order, the decoders are sequentially connected in order, the output end of the last encoder is connected to the input end of the first decoder, the input end of part of other decoders is respectively connected to the output end of one other encoder, the other encoder is an encoder other than the last encoder, and the other decoder is a decoder other than the first decoder.

[0088] In one embodiment, the other encoder is an encoder other than the first encoder and the last encoder, and the other decoder is a decoder other than the first decoder and the last decoder. One encoder is connected to at most one decoder, and one decoder is connected to at most one encoder. The other encoders connected to the decoders are not adjacent.

[0089] Compared with the prior art, the noise reduction model based on the Transformer neural network is adopted, the multi-head self-attention mechanism in the Transformer can effectively capture the long-distance dependency relationship in the mass spectrum data, can better understand the global features of the mass spectrum data compared with the traditional CNN neural network, and improve the noise reduction effect.

[0090] The cosine weighted contrast loss function adopted in the application can measure the similarity of the feature vectors, is beneficial to the model learning of discriminative features, can process the class imbalance problem through the class weight, and makes the noise reduction model training more accurate.

[0091] The noise reduction algorithm of the application can maximize the retention of useful signals while removing noise, avoids the loss of useful information, and is beneficial to subsequent mass spectrum peak identification, quantitative analysis and other processing.

[0092] The algorithm of the application has strong adaptability and generalization ability, and can be applied to different types of mass spectrum detectors and different detection scenes.

[0093] Figure 4is a structural schematic diagram of a computer device provided by an embodiment of the present application. The computer device 7000 can have great differences due to different configurations or performances, and can include one or more processors (central processing units, CPUs) 710 (for example, one or more processors) and a memory 720, one or more storage media 730 (for example, one or more mass storage devices) storing application programs 733 or data 732. The memory 720 and the storage medium 730 can be temporary storage or persistent storage. The programs stored in the storage medium 730 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations on the computer device 7000. Further, the processor 710 can be configured to communicate with the storage medium 730 and execute a series of instruction operations in the storage medium 730 on the computer device 7000.

[0094] The computer device 7000 can also include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input and output interfaces 760, and / or one or more operating systems 731, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that the computer device 7000 can include more or fewer components than those shown in the figure, or combine some components, or arrange different components. Figure 4 The computer device structure shown does not constitute a limitation on the computer device, and can include more or fewer components than those shown in the figure, or combine some components, or arrange different components.

[0095] The present application also provides a computer device including a memory and a processor, the memory storing computer readable instructions, and the computer readable instructions being executed by the processor to make the processor execute the steps of the mass spectrum data denoising method based on the Transformer in each of the above embodiments. The present application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores instructions, and the instructions make the computer execute the steps of the mass spectrum data denoising method based on the Transformer when the instructions run on the computer.

[0096] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0097] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0098] The above-described embodiments are merely used to illustrate the technical solutions of the present application, rather than limit the same; even though the present application has been described in detail with reference to the foregoing embodiments, those ordinarily skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for denoising mass spectrometry data based on Transformer, characterized in that, The Transformer-based mass spectrum data denoising method comprises: inputting a data set into a pre-trained denoising model, and training the pre-trained denoising model based on a cosine weighted contrast loss function to obtain a trained denoising model, wherein the data set contains multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data, the pre-trained denoising model is constructed based on a Transformer neural network, and the cosine weighted contrast loss function is constructed based on a cosine similarity and a category weight; performing denoising processing on to-be-denoised mass spectrum data by using the trained denoising model to obtain target denoised mass spectrum data.

2. The Transformer-based mass spectrometry data denoising method of claim 1, wherein, The cosine weighted contrast loss function is shown in the following formula 1: Equation 1 wherein N is a batch size; true class of the sample i, total number of classes; For the category of weights; feature vector for sample i cosine similarity with feature vector of sample feature vector of sample i cosine similarity with feature vector of sample a scaling factor for the same class samples, used to control the intensity of the penalty for the similarity of the same class samples, and ; scaling factors for different classes of samples, for controlling the strength of the penalty for similarity of different classes of samples, and ; Threshold parameter used to define the maximum similarity acceptable for different class samples.

3. The Transformer-based mass spectrometry data denoising method of claim 2, wherein, The scaling factor of the same type of samples can be dynamically adjusted in different iteration rounds in the model training process; specifically by the following steps: Based on the average cosine similarity of the same type of sample pairs, the expected average cosine similarity of the same type of samples, the value of the scaling factor of the same type of samples in the last iteration round, the adjustment coefficient for controlling the adjustment range of the scaling factor of the same type of samples, the maximum value and the minimum value of the scaling factor of the same type of samples, the value of the scaling factor of the same type of samples in the next iteration round is calculated.

4. The Transformer-based mass spectrometry data denoising method of claim 3, wherein, The expected average cosine similarity of the same type of samples can be updated with the iteration rounds by the following steps: Calculate the cosine similarity of all same type of sample pairs in each iteration training; Take the p% quantile of all cosine similarities as the expected average cosine similarity of the same type of samples in the corresponding iteration round.

5. The Transformer-based mass spectrometry data denoising method according to any one of claims 2-4, characterized in that, The scaling factor of different types of samples can be dynamically adjusted in different iteration rounds in the model training process; specifically by the following steps: Based on the proportion of sample pairs with cosine similarity exceeding the threshold parameter in all different type of sample pairs, the maximum proportion of sample pairs with cosine similarity exceeding the threshold parameter in the acceptable different type of sample pairs, the adjustment coefficient for controlling the adjustment range of the scaling factor of the different type of samples, the maximum value and the minimum value of the scaling factor of the different type of samples, the value of the scaling factor of the different type of samples in the last iteration round, the value of the scaling factor of the different type of samples in the next iteration round is calculated.

6. The Transformer-based mass spectrometry data denoising method of claim 5, wherein, The maximum proportion of sample pairs with cosine similarity exceeding the threshold parameter in the acceptable different type of sample pairs can be updated with the iteration rounds by the following steps: Calculate the cosine similarity of all different type of sample pairs in each iteration training, and calculate the proportion of sample pairs with cosine similarity exceeding the threshold parameter in the different type of sample pairs; Take the q% quantile of all sample pairs with cosine similarity exceeding the threshold parameter as the maximum proportion of sample pairs with cosine similarity exceeding the threshold parameter in the corresponding iteration round of the acceptable different type of sample pairs.

7. The Transformer-based mass spectrometry data denoising method according to any one of claims 1-4, 6, characterized in that, The pre-trained denoising model contains an unequal number of encoders and decoders, and the number of encoders is not less than the number of decoders, the encoders are sequentially connected in order, the decoders are sequentially connected in order, the output end of the last encoder is connected to the input end of the first decoder, the input end of part of other decoders is respectively connected to the output end of one other encoder, the other encoder is an encoder other than the last encoder, and the other decoder is a decoder other than the first decoder.

8. A Transformer-based mass spectrometry data denoising apparatus, characterized in that, The Transformer-based mass spectrum data denoising device comprises: a model training module configured to input a data set into a pre-trained denoising model, train the pre-trained denoising model based on a cosine weighted contrast loss function, and obtain a trained denoising model, wherein the data set contains multiple groups of sample noisy mass spectrum data and corresponding sample denoised mass spectrum data, the pre-trained denoising model is constructed based on a Transformer neural network, and the cosine weighted contrast loss function is constructed based on a cosine similarity and a category weight; a denoising processing module configured to perform denoising processing on to-be-denoised mass spectrum data by using the trained denoising model, and obtain target denoised mass spectrum data.

9. A computer device, comprising: The computer device comprises a memory and at least one processor, and the memory stores instructions; The at least one processor invokes the instructions in the memory, so that the computer device performs the Transformer-based mass spectrum data denoising method in any one of claims 1-7.

10. A computer-readable storage medium having stored thereon instructions, the instructions comprising, The instructions are executed by the processor to implement the Transformer-based mass spectrum data denoising method in any one of claims 1-7.

Citation Information

Cited By

  • Target identification method, system and device based on binocular vision and medium

    CN121544872A