Peptide de novo sequencing model based on non-autoregressive Transform and dual learning
By using non-autoregressive Transformer and dual learning models in de novo peptide sequencing, the problem of insufficient accuracy and efficiency of de novo peptide sequencing in the prior art is solved, and more efficient and accurate mass spectrometry data analysis is achieved.
Patent Information
- Application Number
- CN202510121330.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-03
AI Technical Summary
Existing deep learning models have limitations in accuracy and efficiency in de novo peptide sequencing, making it difficult to effectively process complex mass spectrometry data.
The peptide de novo sequencing model based on non-autoregressive Transformer and dual learning was adopted to extract the potential characteristics of mass spectrometry data through multi-scale processing and the self-attention mechanism of the Transformer encoder, and directed acyclic graph inference peptide sequences were constructed using the non-autoregressive Transformer decoder, while optimizing the mass spectrometry data reconstruction through the dual learning paradigm.
It significantly improves the accuracy and operation efficiency of peptide de novo sequencing, can effectively process large-scale mass spectrometry data, and exhibits high robustness in complex samples.
Smart Images

Figure CN120089196A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and specifically to a peptide de novo sequencing model based on non-autoregressive Transformer and dual learning, which is used to infer peptide sequences from mass spectrometry data and improve the accuracy and efficiency of mass spectrometry data analysis. Background Art
[0002] Peptide de novo sequencing is an important technique in bioinformatics and is widely used in the research of proteomics and metabolomics. Different from traditional peptide sequence analysis methods, de novo sequencing aims to directly infer the amino acid sequence of a peptide from mass spectrometry data without a known reference sequence. This process is particularly important for studying unknown proteins, variants, and new biomarkers. Traditional peptide de novo sequencing algorithms usually rely on algorithm models, such as spectral comparison, dynamic programming, etc. In recent years, the rapid development of deep learning technology has provided new solutions for peptide de novo sequencing. Nevertheless, the application of existing deep learning models in peptide de novo sequencing still faces limitations in accuracy and efficiency. Summary of the Invention
[0003] The purpose of the present invention is to provide a peptide de novo sequencing model based on non-autoregressive Transformer and dual learning, which can effectively improve the accuracy and running efficiency of the peptide de novo sequencing model.
[0004] The technical solution adopted by the present invention is as follows: A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning, comprising the following steps:
[0005] Step 1: Taking the original mass spectrometry data and peptide precursor information as inputs, performing segmentation of the original mass spectrometry data at different resolutions through multi-scale processing to generate embedded representation information of the mass spectrometry;
[0006] Step 1.1: Performing multi-scale segmentation on the input mass spectrometry data X to generate mass spectrometry vectors at multiple resolutions;
[0007] Step 1.2: Embedding and projecting the mass spectrometry vectors at different scales respectively, and then splicing them to obtain the embedded representation information of the mass spectrometry.
[0008] Step 2: Using the self-attention mechanism of the Transformer encoder to extract the latent features of the original mass spectrometry and capture the multi-scale feature dependence relationships in the mass spectrometry data;
[0009] Step 2.1: By using the Transformer encoder, using its self-attention mechanism to extract features from the embedded representation information of the mass spectrometry to capture the latent multi-scale feature dependence relationships in the mass spectrometry data;
[0010] Step 2.2: In the encoder, use a multi-layer self-attention structure so that each layer dynamically models the complex interactions between features of different scales by calculating the self-attention weights of the input information;
[0011] Step 2.3: Finally, the latent features of the original mass spectrum are output at the last layer of the encoder, and this feature is the comprehensive information of the mass spectrometry data at multiple scales.
[0012] Step 3: Based on the latent features of the original mass spectrum and the peptide precursor information, use a non-autoregressive Transformer decoder to construct a directed acyclic graph and infer the peptide sequence;
[0013] Step 3.1: Based on the latent features of the original mass spectrum and the peptide precursor information, use the Transformer decoder to construct an amino acid distribution matrix and an amino acid transition matrix. Each node represents an amino acid, and the edge represents the probability of transferring from the current amino acid to the next amino acid;
[0014] Step 3.2: Construct a directed acyclic graph, use the Beam Search algorithm to decode the directed acyclic graph, infer the peptide sequence in parallel, and generate an amino acid sequence;
[0015] Step 3.3: During the decoding process, adjust the generation of each amino acid according to the constraint conditions of the precursor peptide information, and finally output the complete peptide sequence.
[0016] Step 4: Use the paradigm of dual learning to reconstruct the mass spectrometry data through the Transformer encoder;
[0017] Step 4.1: Re-project and splice the amino acid distribution matrix and the amino acid transition matrix to obtain the latent representation of the reconstructed mass spectrum;
[0018] Step 4.2: Generate the reconstructed mass spectrometry data through the multi-head attention mechanism.
[0019] Step 5: Train the de novo peptide sequencing model based on non-autoregressive Transformer and dual learning, establish an association between the original mass spectrum and the inferred peptide sequence through a joint loss function, minimize the negative log-likelihood loss value of the generated peptide sequence and the mass spectrum reconstruction loss value, and obtain the final de novo peptide sequencing model based on non-autoregressive Transformer and dual learning.
[0020] Step 5.1: Calculate the loss value between the predicted peptide sequence and the true peptide sequence;
[0021] Step 5.2: Calculate the loss value between the predicted mass spectrometry data and the true mass spectrometry data;
[0022] Step 5.3: During the loss calculation process, balance the above two loss values through a dynamic weighting strategy.
[0023] The beneficial effects of the present invention are as follows:
[0024] The present invention utilizes a non-autoregressive Transformer architecture to support parallel inference of peptide sequences. Compared with the traditional step-by-step generation method, it significantly improves the inference speed and is suitable for large-scale mass spectrometry data analysis.
[0025] By processing mass spectrometry data at multiple scales, the present invention can integrate information at different resolutions, improve the modeling ability of the complexity of mass spectrometry data, and ensure reliable results in diverse samples.
[0026] Through the paradigm of dual learning, the present invention simultaneously optimizes mass spectrometry data reconstruction and peptide sequence generation during the training process, enhances the robustness of the model to data noise and missing information, and ensures its effectiveness in complex samples. Description of the Drawings
[0027] Figure 1 Model flow chart;
[0028] Figure 2 Construction process diagram of the amino acid distribution matrix and the amino acid transfer matrix;
[0029] Figure 3 Mass spectrometry data reconstruction flow chart;
[0030] Figure 4 Line chart of performance comparison between the embodiments of the present invention and other methods on a cross-species dataset. Detailed Embodiments
[0031] The present invention proposes a de novo peptide sequencing model based on non-autoregressive Transformer and dual learning:
[0032] Step 1: The mass spectrometry data is represented as a vector X, which consists of pairs (m i , I i ), where X = {(m 1 , I 1 ), (m 2 , I 2 ), …, (m N , I N )}, and each (m i , I i ) represents a mass spectrometry peak.
[0033] Step 2: The peptide sequence is represented as Y = {y 1 , y 2 , y 3 , …, yL}, where y i represents an amino acid, and L represents the length of the peptide sequence.
[0034] Step 3: Peptide precursor information Z = (m pre , c pre ), where represents the mass of the precursor peptide, and c pre ∈{1, 2, 3, …, 10} represents the charge information.
[0035] Step 4: As shown in Figure 1 . This model takes mass spectrometry data X and peptide precursor information Z as inputs, and outputs peptide sequence Y′ and reconstructed mass spectrometry X′. First, the m / z values and intensity values are embedded separately. For the m / z values, sine encoding is used to effectively capture the m / z values and make the model focus on the changes between peaks. The dimension of the peak embedding is set to 512. For the intensity value part, linear embedding is used to project it into a 512-dimensional space.
[0036] Step 5: The mass spectrometry data X is divided into N p patches, and this segmentation helps to capture the global features in the spectrum, thereby facilitating the effective decoding of the peptide sequence. After each patch is separately projected into a 512-dimensional space and then position-encoded, it is concatenated with the embedded mass spectrometry data X to form the embedded representation information of the mass spectrometry.
[0037] Step 6: This information is separately input into different Transformer encoding layers for calculation. The Transformer encoding layer includes multiple self-attention modules and linear modules. After being processed by inputs of different scales, combined with layer normalization, linear layers, and residual connections, the latent feature representation of the original mass spectrometry is finally obtained.
[0038] Step 7: By applying learnable position information embedding to the fully masked sequence G = {g 1 , g 2 , g 3 , …, g s} and peptide precursor information Z to represent the vertex indices. Then, the encoded sequence is input into the Transformer decoding layer, which uses self-attention and cross-attention mechanisms similar to those in the Transformer encoder to generate the graph latent representation D. Subsequently, as shown in Figure 2 , the graph latent representation D undergoes Softmax function, weighted summation processing, and multiplication with learnable matrices W q and W kThe operation is performed to finally obtain the amino acid transfer matrix E. Meanwhile, the graph latent representation D is processed by the Softmax function and a linear layer to obtain the amino acid distribution matrix P. The amino acid transfer matrix E and the amino acid distribution matrix P passing through the lower triangular mask form a directed acyclic graph, where each node represents an amino acid and the edge represents the probability of transferring from the current amino acid to the next amino acid.
[0039] Step 8: Direct global search for the constructed directed acyclic graph may bring high computational complexity, and greedy search may miss the optimal solution. Therefore, the Beam Search algorithm is adopted to generate multiple possible paths, and these paths form peptide sequences from vertices. The path with the highest score is selected as the final output.
[0040] Step 9: The amino acid distribution matrix P and the amino acid transfer matrix E are concatenated and input into the projection layer to obtain the latent representation H of the peptide. e Subsequently, the peptide precursor information Z and the latent representation H of the peptide e are concatenated to generate the feature vector H for reconstructing the mass spectrum. H is then processed through multiple standard Transformer encoding layers, and the latent representation of the reconstructed mass spectrum is obtained through a multi-layer perceptron (MLP) with a sigmoid activation function, thereby obtaining the discretized reconstructed mass spectrum data. Specifically, this step is only used in model training.
[0041] Effectiveness verification:
[0042] Through comparative experiments, the performance of the present invention is evaluated on cross-species datasets respectively. The comparison results of the present invention with other methods are as Figure 4 shown, and the coverage curves of the present invention are higher than those of other methods.
Claims
1. A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning, characterized in that: The following steps are involved: Step 1: Take the raw mass spectrometry data and peptide precursor information as input, segment the raw mass spectrometry data at different resolutions through multi-scale processing, and generate embedded representation information of the mass spectrum; Step 2: Use the self-attention mechanism of the Transformer encoder to extract the potential features of the original mass spectrum and capture the multi-scale feature dependencies in the mass spectrometry data; Step 3: Based on the potential features of the original mass spectrum and the peptide precursor information, a non-autoregressive Transformer decoder is used to construct a directed acyclic graph and infer the peptide sequence; Step 4: Use the dual learning paradigm to reconstruct the mass spectrometry data through the Transformer encoder; Step 5: Train the peptide de novo sequencing model based on non-autoregressive Transformer and dual learning, establish an association between the original mass spectrum and the inferred peptide sequence through a joint loss function, minimize the negative log-likelihood loss value of the generated peptide sequence and the mass spectrum reconstruction loss value, and finally obtain the peptide de novo sequencing model based on non-autoregressive Transformer and dual learning.
2. A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning according to claim 1, characterized in that: The specific steps in step 1 are: Step 1.1: Perform multi-scale segmentation on the input mass spectrum data X to generate mass spectrum vectors with multiple resolutions; Step 1.2: Embed and project the mass spectrum vectors of different scales respectively, and then concatenate them to obtain the embedded representation information of the mass spectrum.
3. A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning according to claim 1, characterized in that: The specific steps in step 2 are: Step 2.1: Use the Transformer encoder to extract features from the embedded representation of mass spectrometry using its self-attention mechanism to capture the potential multi-scale feature dependencies in mass spectrometry data; Step 2.2: In the encoder, use a multi-layer self-attention structure so that each layer dynamically models the complex interactions between features of different scales by calculating the self-attention weights of the input information; Step 2.3: Finally, the last layer of the encoder outputs the latent features of the original mass spectrum, which is the comprehensive information of the mass spectrum data at multiple scales.
4. A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning according to claim 1, characterized in that: The specific steps in step 3 are: Step 3.1: Based on the potential features of the original mass spectrum and the peptide precursor information, the Transformer decoder is used to construct the amino acid distribution matrix and the amino acid transfer matrix. Each node represents an amino acid, and the edge represents the probability of transferring from the current amino acid to the next amino acid. Step 3.2: Construct a directed acyclic graph, decode the directed acyclic graph using the Beam Search algorithm, infer the peptide sequence in parallel, and generate the amino acid sequence; Step 3.3: During the decoding process, the generation of each amino acid is adjusted according to the constraints of the peptide precursor information, and the complete peptide sequence is finally output.
5. A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning according to claim 4, characterized in that: The specific steps in step 4 are: Step 4.1: reproject and concatenate the amino acid distribution matrix and the amino acid transfer matrix to obtain the potential representation of the reconstructed mass spectrum; Step 4.2: Generate reconstructed mass spectrometry data through a multi-head attention mechanism.
6. A peptide de novo sequencing model based on non-autoregressive Transformer and dual learning according to claim 1, characterized in that: The specific steps in step 5 are: Step 5.1: Calculate the loss value between the predicted peptide sequence and the true peptide sequence; Step 5.2: Calculate the loss value between the predicted mass spectrum data and the actual mass spectrum data; Step 5.3: During the loss calculation process, the two loss values mentioned above are balanced through a dynamic weighting strategy.