A deep learning-based top-down mass spectrometry protein identification method

CN122658415APending Publication Date: 2026-08-28SHENZHEN MSU-BIT UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611141513.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]本发明的目的在于针对现有Top-Down质谱深度学习方法无法直接处理完整蛋白质长序列和复杂碎片模式、需要拆分序列导致精度下降的技术缺陷,提供一种端到端的专用深度学习算法,该方法通过对前体单同位素质量及多电荷态分布的联合嵌入形成先验约束,融合氨基酸序列和理化特征嵌入提升泛化能力,采用自适应分段束搜索有效缓解长序列解码的组合爆炸问题,引入扩散模型对低置信区域进行多步迭代精修,并在去噪过程中施加全局质量一致性、局部碎片匹配度及修饰位点兼容性的联合约束,从而在无需拆分、充分利用谱图全部信息的前提下,实现对完整蛋白质形式的高精度、高可信度识别和翻译后修饰位点的准确标注,满足高通量Top-Down蛋白质组学的应用需求

Benefits of technology

[0016]Compared with existing technologies, the deep learning-based Top-Down mass spectrometry protein identification method provided in this invention directly identifies sequences from fragmented spectra of intact proteins, avoiding information loss and splicing errors caused by enzyme digestion. It employs an end-to-end deep learning model to generate sequences from spectra, eliminating reliance on prior databases and demonstrating stronger discovery capabilities for unknown proteins and post-translational modifications. By jointly embedding the precursor monoisotope mass and multi-charge state distribution to form a prior constraint vector, and integrating amino acid sequence embedding with four types of physicochemical feature embeddings—hydrophobicity, charge, molecular weight, and secondary structure tendency—the model can utilize the evolutionary patterns and inherent properties of amino acids to identify unseen peptides and modification sites. Accurate prediction and significantly improved generalization performance; adaptive segmentation of single isotope mass based on peak energy distribution and allocation of segment mass budget, combined with segmented bundle search, effectively alleviates the combinatorial explosion problem in decoding long complete protein sequences, significantly shortening analysis time; the diffusion model only iteratively refines low-confidence regions and is simultaneously guided by global mass consistency, local fragment matching degree, and modification site compatibility during denoising, ensuring that prediction results are protected in terms of mass conservation, spectral matching, and modification rationality. The output includes the complete protein form with post-translational modification site annotations and its confidence score, which can be widely applied to scenarios such as unknown protein discovery, antibody characterization, structural protein analysis, and biopharmaceutical quality control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658415A_ABST
    Figure CN122658415A_ABST
Patent Text Reader

Abstract

The application discloses a kind of Top-Down mass spectrum protein identification methods based on deep learning, comprising: the Top-Down mass spectrum data of the complete protein form to be identified is preprocessed and encoded to obtain spectrum vector representation;By peak encoder, its encoding obtains spectrum latent feature representation;The precursor monoisotopic mass and the distribution of multiple charge state are jointly embedded to obtain prior constraint vector;With spectrum peak energy distribution, monoisotopic mass is adaptively segmented to obtain sub-section mass budget;Sequence decoder performs beam search constrained by mass budget in each subsection and splices into skeleton sequence;By diffusion model, low confidence area is iteratively refined in multiple steps, and the denoising process is jointly guided by global mass consistency, local fragment matching degree and modification site compatibility, and outputs complete protein form containing post-translational modification site annotation and confidence score thereof.The application does not need to be cut and database dependent, can identify long sequence and post-translational modification, significantly improve recognition accuracy and generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of proteomics and deep learning technologies, specifically to a deep learning-based Top-Down mass spectrometry protein identification method. Background Technology

[0002] De novo protein sequencing refers to the direct inference of protein sequences from mass spectrometry spectra obtained without a reference database. Currently, most research focuses on bottom-up mass spectrometry data, where proteins are first digested into peptides using enzymes before mass spectrometry analysis and sequence identification. However, enzyme digestion destroys the complete structural information of proteins and results in insufficient peptide coverage, making it difficult to assemble the full-length sequence. Top-down mass spectrometry, which directly analyzes intact proteins, preserves post-translational modifications and isoform information, representing an important direction for the development of high-throughput proteomics. Existing top-down identification methods mainly fall into two categories: one is based on database search, which enumerates a large number of candidate sequences and calculates the matching degree between theoretical fragment spectra and experimental spectra to screen for correct sequences. However, this method is computationally intensive and time-consuming, making it difficult to apply to large-scale top-down spectra. The other category, based on deep learning-based de novo sequencing algorithms, has achieved significant results on bottom-up data. However, its network structure and loss function are designed for short peptide sequences and cannot directly handle the complex fragmentation patterns of complete proteins. Furthermore, due to the scarcity of training data, the complexity of fragment types, and strict quality constraints, proteins usually need to be split into several short tags and then recombined, leading to a decrease in precision and recall. Therefore, developing end-to-end dedicated deep learning algorithms for high-resolution top-down mass spectrometry data, which can fully utilize all the information in the spectra while meeting precursor quality constraints, is key to overcoming the shortcomings of existing technologies. Summary of the Invention

[0003] The purpose of this invention is to address the technical shortcomings of existing Top-Down mass spectrometry deep learning methods, which cannot directly handle long, intact protein sequences and complex fragment patterns, and require sequence splitting leading to decreased accuracy. This invention provides an end-to-end dedicated deep learning algorithm. This method establishes prior constraints through the joint embedding of precursor monoisotope mass and multi-charge state distribution, enhances generalization ability by fusing amino acid sequence and physicochemical feature embedding, effectively alleviates the combinatorial explosion problem in long sequence decoding by employing adaptive segmented bundle search, introduces a diffusion model for multi-step iterative refinement of low-confidence regions, and applies joint constraints of global mass consistency, local fragment matching degree, and modification site compatibility during denoising. Thus, without requiring splitting and fully utilizing all spectral information, it achieves high-precision, high-confidence identification of intact protein forms and accurate labeling of post-translational modification sites, meeting the application requirements of high-throughput Top-Down proteomics.

[0004] To achieve the above objectives, this invention provides a deep learning-based Top-Down mass spectrometry protein identification method, the method comprising the following steps: S1: Preprocess and encode the Top-Down mass spectrometry data of the intact protein form to be identified to obtain a spectral vector representation; S2: The peak encoder encodes the spectral vector representation to obtain the spectral latent feature representation; S3: Jointly embed the single isotopic mass and multiple charge state distribution of the complete protein precursor to obtain the prior constraint vector; S4: Adaptively segment the mass of the single isotope according to the energy distribution of the spectral peaks in the latent feature representation of the spectrum to obtain several sub-segment mass budgets; S5: The latent feature representation of the spectrum, the prior constraint vector, the amino acid sequence embedding vector, and the amino acid physicochemical feature embedding vector are fused and input into the sequence decoder. The sequence decoder performs a bundle search within each segment, constrained by the quality budget of the corresponding segment, and splices the outputs of each segment into a backbone sequence. S6: The diffusion model performs multi-step iterative refinement on regions in the backbone sequence whose position confidence is lower than the preset position confidence threshold. The denoising process is simultaneously guided by global quality consistency, local fragment matching degree and modification site compatibility, and outputs the complete protein form containing amino acid residue sequences and post-translational modification site annotations and its confidence score.

[0005] Optionally, the preprocessing and encoding in step S1 specifically involves: performing baseline correction, denoising, and peak extraction on the original spectrum of the Top-Down mass spectrometry data to obtain a preprocessed spectrum; transforming the mass-to-charge ratio of each spectral peak in the preprocessed spectrum using sine and cosine functions of different frequencies to obtain a position encoding vector; and concatenating the position encoding vector with the corresponding intensity information to form the spectrum vector representation.

[0006] Optionally, the peak encoder is a multi-layer Transformer encoding layer, each layer containing a multi-head self-attention sub-layer and a feedforward network sub-layer, with residual connections and layer normalization set between the sub-layers.

[0007] Optionally, step S3 specifically involves: extracting several charge states and their relative abundances of the complete protein precursor from the preprocessed spectrum to form a charge state abundance distribution; projecting the charge state abundance distribution and the single isotope mass obtained by deconvolution onto the same vector space and summing them to obtain the prior constraint vector.

[0008] Optionally, the amino acid sequence embedding vector is obtained by adding the pre-trained amino acid vector to the position encoding; the amino acid physicochemical feature embedding vector obtains four types of values ​​for each amino acid—hydrophobicity, charge, molecular mass, and secondary structure tendency—by looking up a table, and then linearly projects them onto the dimension of the amino acid sequence embedding vector and adds them element by element.

[0009] Optionally, step S4 specifically involves: identifying N energy peak clusters based on the spectral peak energy distribution in the latent feature representation of the spectrum; dividing the single isotope mass into N sub-segments using the energy peak clusters as segment boundaries; and allocating the mass budget for each sub-segment according to the following formula:

[0010] in, B i For the first i The sub-segment quality budget, in units of Da; M p The mass of the single isotope is in the dimension Da; E i For the first i The sum of the squares of the intensities of all spectral peaks within an energy peak cluster represents the energy value corresponding to that segment; ε is the preset mass budget redundancy interval, with dimensions Da; and the mass budget of each segment satisfies:

[0011] Optionally, the position confidence in step S6 is calculated based on the posterior probability of each amino acid position output by the sequence decoder; the region where the position confidence of M consecutive positions is lower than the preset position confidence threshold is determined as a low confidence region; the diffusion model only applies noise to the tokens in the low confidence region and performs iterative denoising, while the remaining positions in the backbone sequence remain unchanged during the iteration process.

[0012] Optionally, the joint guidance in step S6 adds the three constraint terms to the noise term predicted by the diffusion model at time step t according to the following formula:

[0013] in, ε θ ( x t , t ) represents the original noise term predicted by the diffusion model; M ( x t ) represents the current sequence x t Theoretical cumulative mass, and M pBoth are units of Da, and the difference between them is... M p After normalization, it becomes a dimensionless quality deviation term; S match ( x t The cosine similarity between the theoretical fragment spectrum corresponding to the current sequence and the spectral vector representation is [0,1]. P PTM ( x t The value represents the incompatibility ratio between the modified token and the amino acid residue type at its current position in the current sequence, and its value ranges from [0,1]. λ 1 , λ 2 , λ 3 Dimensionless weighted coefficients; corrected noise term ( x t , t ) used for driving x t arrive x t-1 Inverse denoising.

[0014] Optionally, the vocabulary corresponding to the amino acid sequence embedding vector is pre-expanded with modification tokens corresponding to common post-translational modification types. Each modification token carries the molecular weight of the corresponding modification group and the type of amino acid residue that can be modified. In step S6, the modification site compatibility is calculated based on the matching relationship between the type of amino acid residue that can be modified recorded by the modification token at the current position and the amino acid residue at the current position, and a penalty term is applied to mismatched positions.

[0015] Optionally, the model training step is also included: constructing a labeled dataset based on known complete protein form samples, using a target decoy strategy to take correct spectral sequence pairs as positive samples, and using decoy sequences generated from positive samples by flipping or random perturbation as negative samples; and performing end-to-end joint training of the peak encoder, the sequence decoder, and the diffusion model by weighted summation of sequence prediction cross-entropy loss, mean square error loss between diffusion model predicted noise and real noise, and modification site prediction cross-entropy loss.

[0016] Compared with existing technologies, the deep learning-based Top-Down mass spectrometry protein identification method provided in this invention directly identifies sequences from fragmented spectra of intact proteins, avoiding information loss and splicing errors caused by enzyme digestion. It employs an end-to-end deep learning model to generate sequences from spectra, eliminating reliance on prior databases and demonstrating stronger discovery capabilities for unknown proteins and post-translational modifications. By jointly embedding the precursor monoisotope mass and multi-charge state distribution to form a prior constraint vector, and integrating amino acid sequence embedding with four types of physicochemical feature embeddings—hydrophobicity, charge, molecular weight, and secondary structure tendency—the model can utilize the evolutionary patterns and inherent properties of amino acids to identify unseen peptides and modification sites. Accurate prediction and significantly improved generalization performance; adaptive segmentation of single isotope mass based on peak energy distribution and allocation of segment mass budget, combined with segmented bundle search, effectively alleviates the combinatorial explosion problem in decoding long complete protein sequences, significantly shortening analysis time; the diffusion model only iteratively refines low-confidence regions and is simultaneously guided by global mass consistency, local fragment matching degree, and modification site compatibility during denoising, ensuring that prediction results are protected in terms of mass conservation, spectral matching, and modification rationality. The output includes the complete protein form with post-translational modification site annotations and its confidence score, which can be widely applied to scenarios such as unknown protein discovery, antibody characterization, structural protein analysis, and biopharmaceutical quality control. Attached Figure Description

[0017] Figure 1 A schematic diagram of the overall process of the deep learning-based Top-Down mass spectrometry protein identification method provided by the present invention; Figure 2 This is a schematic diagram of the spectrum preprocessing and multi-scale sine function encoding process provided by the present invention; Figure 3 This is a schematic diagram of the peak encoder structure provided by the present invention; Figure 4 A schematic diagram illustrating the prior constraint vector generation process provided by this invention; Figure 5 This is a schematic diagram of the adaptive segmentation and sub-segment quality budget allocation process provided by the present invention; Figure 6 This is a schematic diagram illustrating the process of multi-step iterative refinement of the low-confidence region using the diffusion model provided by the present invention. It should be noted that, Figure 5 The number of energy peak clusters and sub-segments shown are illustrative examples. Figure 6 The specific position numbers, amino acid residue types, and modification types shown are illustrative examples, and the method described in this invention is not limited to these specific examples. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature; in the description of this application, unless otherwise stated, "multiple" means two or more.

[0020] To more clearly illustrate the technical solution of the present invention, the present invention will be described in detail below with reference to specific embodiments, but should not be construed as limiting the scope of protection of the present invention. For ease of review, the hyperparameters used in the various embodiments of the present invention are uniformly listed as follows: the number of effective spectral peaks L is 2000, the model hidden dimension dmodel is 512, the number of peak encoder layers L' is 9, the number of attention heads is 8, the hidden dimension of the feedforward network is 2048; the maximum possible number of charge states of the precursor Zmax is 30; the number of adaptive segments N is 12, the mass budget redundancy interval ε is 0.5 Da; the beam search width K is 16; the back diffusion time step T is 15; the position confidence threshold θc is 0.6, the number of consecutive positions M is 3; the joint guidance weight coefficients λ1, λ2, and λ3 are 1.0, 0.5, and 2.0; the training loss weighting coefficients α, β, and γ are 1.0, 0.5, and 2.0; the training batch size is 32, and the initial learning rate is 1×10⁻⁶. -4 .

[0021] Furthermore, in the various embodiments of this invention, "linear projection," "projection to the same vector space," "fully connected embedding layer," and "linear projection layer" all refer to the same type of mathematical operation. That is, by using a learnable weight matrix and a learnable bias vector, the input vector is transformed by first multiplying it with the weight matrix and then adding it to the bias vector, mapping the input vector from its original dimension to a target dimension vector space. The number of rows in the weight matrix corresponds to the dimension of the input vector, and the number of columns corresponds to the dimension of the target vector space. The dimension of the bias vector is equal to the dimension of the target vector space. Both the weight matrix and the bias vector are learnable parameters, optimized together with the peak encoder, sequence decoder, and diffusion model through the end-to-end joint training process described in Embodiment 10. The function of linear projection is to unify the dimensions of tensors from different sources, enabling multiple tensors to perform subsequent operations such as summation, concatenation, or attention calculation within the same vector space. All linear projections involved in the various embodiments of this invention are implemented according to the above definition and will not be repeated below.

[0022] In some embodiments, such as Figure 1 As shown, the Top-Down mass spectrometry protein identification method of the present invention will be described below with reference to a specific embodiment. In this embodiment, the complete human myoglobin protein with a molecular weight of approximately 17 kDa is used as the target protein. The complete protein is electrospray ionized and then enters a high-resolution Fourier transform ion cyclotron resonance mass spectrometer to acquire the raw Top-Down mass spectrometry data, and sequence identification is completed according to the following steps.

[0023] Step S1: Preprocess and encode the Top-Down mass spectrometry data of the intact protein form to be identified to obtain the spectral vector representation.

[0024] The acquired raw spectra are first processed using a rolling ball algorithm to subtract baseline drift, then smoothed using Savitzky-Golay filtering to remove high-frequency noise. Subsequently, peak detection is performed based on a signal-to-noise ratio threshold to extract effective spectral peaks with an intensity greater than three times the background noise. Each effective spectral peak is represented by a mass-to-charge ratio... m i and intensity I i Two attributes represent, i =1,2,…, L In this embodiment L Taking 2000 as an example, the mass-to-charge ratio of each spectral peak covers a range of 500 to 3000 Th. A multi-scale sine function transform is applied to the mass-to-charge ratio of each spectral peak to obtain the position encoding vector. P i Then P i and I iAfter splicing, the spectral vector representation is formed, which simultaneously carries the position and intensity information of the spectral peaks, facilitating in-depth modeling by the subsequent peak encoder.

[0025] Step S2: The peak encoder encodes the spectral vector representation to obtain the spectral latent feature representation.

[0026] The peak encoder receives the spectral vector representation obtained in the previous step and models the relationship between each spectral peak, outputting a dimension of... L × d model The latent feature representation of the spectrum h 1 , h 2 … h L In this embodiment d model The value is 512. The peak encoder consists of multiple stacked coding layers. Each coding layer includes a self-attention mechanism and a feedforward transform, enabling the features at each spectral peak position to gather contextual information from other relevant spectral peak positions in the entire spectrum, thereby extracting high-order semantic features that reflect the fragmentation patterns of the complete protein form. The specific number of layers in the peak encoder can be set as needed; the specific value in this embodiment is described in the corresponding embodiment below. The latent feature representation of the spectrum serves both as the decoding context for the subsequent bundle search step and as global reference information for the iterative refinement stage of the diffusion model.

[0027] Step S3: Jointly embed the single isotopic mass and multiple charge state distribution of the complete protein precursor to obtain the prior constraint vector.

[0028] The precursor protein in its complete form was extracted from the raw Top-Down mass spectrometry data, showing parent ion peaks in multiple charge states. In this embodiment, the myoglobin precursor exhibited distinct isotopic peak clusters in the charge state range of +12 to +22. The single isotopic mass was obtained after deconvolution processing. M p The precursor mass is approximately 16951.5 Da, and the relative abundance of the precursor in each charge state is recorded to form the charge state distribution characteristics. The single isotopic mass and the multi-charge state distribution are mapped to the same vector space as the latent feature representation of the spectrum through an embedding layer, and then jointly obtained to obtain the prior constraint vector. The specific implementation of the embedding layer can be found in the corresponding embodiment below. The prior constraint vector carries both the overall mass information of the precursor and the statistical characteristics of the charge state distribution within the source, serving as global conditions in the decoding and diffusion refinement stages, helping to ensure that the prediction results remain consistent in terms of mass conservation and the reasonableness of charge distribution.

[0029] Step S4: Adaptively segment the mass of the single isotope according to the peak energy distribution in the latent feature representation of the spectrum to obtain several sub-segment mass budgets.

[0030] Energy distribution statistics are performed on the potential feature representation of the spectrum to identify several energy peak clusters. Each energy peak cluster corresponds to an amino acid region in the intact protein form that concentrates the generation of fragment ions. In this embodiment, the identified peaks are... N =12 energy peak clusters, and then the mass of the single isotope is distributed to each sub-segment according to the relative proportion of the energy carried by each energy peak cluster, to obtain N Sub-segment quality budget B 1 , B 2 … B N It should be noted that a strict physical correspondence is not required between the energy peak clusters and the continuous amino acid regions in the complete protein form. The essence of this adaptive segmentation is to redistribute the entire single-isotope mass budget according to the density of fragment information, allowing the search interval rich in fragment information to obtain a wider mass budget, thereby supporting beam search to explore more candidate amino acid combinations. The specific assignment of the sequence among each segment is automatically determined by the subsequent beam search stage under the constraint of the mass budget, and the redundancy interval ε tolerates boundary deviations between adjacent segments. To allow for slight deviations in the assignment of boundary amino acids between adjacent segments, a preset mass budget redundancy interval ε is superimposed on the mass budget of each segment; in this embodiment, ε is 0.5 Da. Compared with uniform segmentation, this adaptive segmentation method can align the segment boundaries with the high-incidence areas of fragment ions, thereby more accurately constraining the candidate amino acid combinations within each segment during the decoding stage.

[0031] S5: The latent feature representation of the spectrum, the prior constraint vector, the amino acid sequence embedding vector, and the amino acid physicochemical feature embedding vector are fused and input into the sequence decoder. The sequence decoder performs a bundle search within each sub-segment, constrained by the quality budget of the corresponding sub-segment, and splices the outputs of each sub-segment into a backbone sequence.

[0032] The amino acid sequence embedding vector provides semantic information about the type and position of amino acids, while the amino acid physicochemical feature embedding vector provides inherent property information such as hydrophobicity, charge, molecular weight, and secondary structure tendency of amino acids. Both, along with the spectral latent feature representation and the prior constraint vector, are mapped to the same vector space through a fusion layer and then input into the sequence decoder. The sequence decoder is a structure combining a Transformer decoding layer and a diffusion model, performing beam search in parallel within each segment. In this embodiment, the beam width... K Choose 16, and for each candidate path, when expanding amino acid residues, the cumulative mass must ensure that it does not exceed the corresponding segment's mass budget.B i Paths that do not meet the constraints are pruned using dynamic programming. After each sub-segment bundle search is completed, the candidate sequence with the highest cumulative probability for each segment is selected and spliced ​​together in the order of the sub-segments to form the skeleton sequence. In this embodiment, a skeleton sequence with a length of 153 is obtained.

[0033] S6: The diffusion model performs multi-step iterative refinement on regions in the backbone sequence whose position confidence is lower than the preset position confidence threshold. The denoising process is simultaneously guided by global quality consistency, local fragment matching degree and modification site compatibility, and outputs the complete protein form containing amino acid residue sequences and post-translational modification site annotations and its confidence score.

[0034] For each position in the skeleton sequence, a position confidence score is calculated. This score is obtained from the posterior probability of the corresponding position output by the sequence decoder. Regions where the position confidence scores of consecutive positions are all below a preset position confidence threshold are identified as low-confidence regions. In this embodiment, the threshold is set to 0.6, identifying three low-confidence regions in the skeleton sequence. The remaining high-confidence positions remain unchanged in subsequent iterations. The diffusion model only adds noise to the tokens within the low-confidence regions and distributes it according to a preset time step sequence. t = T Reverse denoising to t =0, in this embodiment T Take 15. The denoising process at each time step is simultaneously guided by three constraints: the global quality consistency constraint, the cumulative quality of the current sequence, and the single isotope quality. M p The deviation between the two sequences is used as the basis for the local fragment matching degree constraint, which is based on the similarity between the theoretical fragment spectrum corresponding to the current sequence and the spectrum vector representation. The modification site compatibility constraint is based on the matching relationship between the modification token and the amino acid residue type at its position. The final output is a complete protein form with a length of 153, containing multiple post-translational modification site annotations, where lysine at position 37 is labeled as acetylation and threonine at position 81 is labeled as phosphorylation, with an overall confidence score of 0.94.

[0035] like Figure 2 As shown, the preprocessing and encoding process will be described below with reference to a specific embodiment. This embodiment uses a set of Top-Down mass spectrometry raw data obtained from bovine hemoglobin in its intact protein form as an example. The raw data was acquired by an Orbitrap high-resolution mass spectrometer at a resolution of 120,000.

[0036] The raw spectra of the Top-Down mass spectrometry data are subjected to baseline correction, noise reduction, and peak extraction to obtain preprocessed spectra.

[0037] This embodiment first uses the rolling ball algorithm to estimate and subtract the baseline signal of the original spectrum to eliminate baseline fluctuations caused by ion source contamination and detector drift. Then, a Savitzky-Golay filter with a window width of 11 and a polynomial order of 3 is used to smooth the baseline-subtracted signal, suppressing high-frequency electronic noise while preserving the peak shape. The smoothed signal is then subjected to local maxima algorithm for peak detection, retaining only valid peaks with a signal-to-noise ratio greater than 3. Each valid peak is represented by a mass-to-charge ratio... m i and intensity I i Two attributes represent, i =1,2,…, L In this embodiment, a total of L =1800 valid spectral peaks constitute the preprocessed spectrum.

[0038] The mass-to-charge ratio of each spectral peak in the preprocessed spectrum is transformed using sine and cosine functions of different frequencies to obtain the position encoding vector.

[0039] Mass-to-charge ratio for each valid spectral peak m i A multi-scale sine function transform is applied, with sine functions used for even-numbered dimensions and cosine functions for odd-numbered dimensions. The frequencies corresponding to different dimensions decay exponentially. Specifically, the second... k 2nd peacekeeping mission k +1 dimension position encoding respectively according to PE( m i ,2 k )=sin( m i / 10000 2k / d_model ) and PE ( m i ,2 k +1)=cos( m i / 10000 2k / d_model )calculate, k =0,1,…, d model / 2−1, so that the position encoding vector reflects coarse-grained position information of mass-to-charge ratio in a low dimension and fine-grained position information of mass-to-charge ratio in a high dimension. The dimension of the position encoding vector in this embodiment is... d model Taking 512, each valid spectral peak yields a corresponding position encoding vector. P i The position encoding vectors of all spectral peaks in the entire spectrum together constitute L × d modelA positional encoding matrix.

[0040] The position encoding vector is concatenated with the corresponding intensity information to form the spectral vector representation.

[0041] For each valid spectral peak i , and its position encoding vector P i Corresponding strength I i The feature vector of the spectral peak is obtained by concatenating the feature vectors of all peaks in the entire spectrum. The feature vectors of all peaks in the spectrum are arranged according to the peak index to form the spectrum vector representation, with a dimension of . L ×( d model +1), then mapped back through a linear projection layer L × d model To adapt to the input requirements of the subsequent peak encoder, the intensity information is normalized to the [0,1] interval according to the maximum value of each peak intensity before splicing, to avoid high-abundance peaks excessively dominating the collection of contextual information in subsequent attention modeling. The final spectral vector representation carries both the position and intensity information of each valid peak, and is used as input to the peak encoder to extract high-order semantic features of the fragmentation pattern of the complete protein form.

[0042] like Figure 3 As shown, the specific structure of the peak encoder will be described below with reference to a specific embodiment. This embodiment follows the spectral vector representation obtained in the preceding steps, and the dimension of the spectral vector representation is... L × d model ,in L =2000 represents the number of valid spectral peaks. d model =512 is the hidden dimension of the model.

[0043] The peak encoder is a multi-layer Transformer encoding layer.

[0044] The peak encoder described in this embodiment is composed of L 'The Transformer encoding layers are stacked sequentially.' L Take 9. The spectral vector represents the input of the first coding layer and the output of each coding layer. z i l ( i =1,…, L ) as input to the next coding layer z i l , No. LThe output of the layer coding layer is the latent feature representation of the spectral graph. h 1 , h 2 … h L Each coding layer shares the same structure but maintains its own independent parameters.

[0045] Each layer contains a multi-head self-attention sublayer and a feedforward network sublayer.

[0046] The multi-head self-attention sublayer will input features z i l-1 The values ​​are mapped to query vector Q, key vector K, and value vector V, respectively. Scaling dot product attention is computed in parallel to model the correlation between any two spectral peaks within the entire spectrum. In this embodiment, the number of attention heads is 8, and each head has a dimension of 64. The feedforward network sublayer consists of two linear transformations and an intermediate nonlinear activation function. In this embodiment, the intermediate hidden dimension is 2048, and the activation function is GELU. A nonlinear transformation is performed on each positional feature output by the multi-head self-attention sublayer to improve the model's ability to fit fragmentation patterns.

[0047] Residual connections and layer normalization are set between sub-layers.

[0048] This embodiment uses a Pre-LN topology, and the following operations are performed sequentially within each coding layer: First, the input... z i l-1 Layer normalization is performed to obtain normalized features. These normalized features are then fed into the multi-head self-attention sub-layer to obtain the attention output. Finally, the attention output is compared with the original input. z i l-1 The first residual connection is formed by adding the bits one by one; then, the sum is normalized again, and the normalized features are fed into the feedforward network sublayer to obtain the feedforward output. The feedforward output is then added bit by bit to the result of the first residual connection to form the second residual connection, thus obtaining the final output of the layer. z i l The layer normalization mentioned above applies to each location feature along... d model The mean and variance of the dimension are calculated and normalized. The residual connections ensure that deep gradient backpropagation does not result in vanishing gradients. L After layer stacking, the final output dimension of the peak encoder is L × d model The latent feature representation of the spectrum serves as global reference information for subsequent decoding and diffusion refinement stages.

[0049] like Figure 4 As shown, the specific generation process of the prior constraint vector is explained below with reference to a specific embodiment. This embodiment follows the aforementioned steps to obtain the raw Top-Down mass spectrometry data of the complete protein precursor, and further processes the isotope peak clusters of the precursor in multiple charge states to obtain the prior constraint vector.

[0050] Several charge states and their relative abundances of the complete protein precursor are extracted from the preprocessed spectrum to form a charge state abundance distribution.

[0051] In this embodiment, the intact protein to be identified is myoglobin. Its precursor undergoes multi-charge ionization within the source, resulting in a series of equally spaced isotope peak clusters in the pre-processed spectrum. Each isotope peak cluster corresponds to a charge state. z k Peak fitting is performed on each isotope peak cluster, and the intensities of all isotope peaks within it are summed to obtain the abundance of the precursor in that charge state. A k In this embodiment, the identification is... z k There are 11 charge states ∈{+12,+13,…,+22}. The abundance of each charge state is... A k Normalized to the [0,1] interval by its maximum value, the charge state abundance distribution {( z k , A k )| k =1,2,…,11}. The charge state abundance distribution reflects the statistical characteristics of the charge state of the precursor in the source. It has different distribution patterns for complete protein forms with different molecular weights and structures, and can be used as auxiliary prior information for sequence identification.

[0052] The prior constraint vector is obtained by projecting the charge state abundance distribution and the single isotope mass obtained by deconvolution onto the same vector space and summing them.

[0053] Deconvolution is performed on the isotope peak clusters corresponding to each charge state in the charge state abundance distribution to map the multi-charge parent ion peak clusters onto the neutral mass axis, thereby extracting the single isotope mass corresponding to the single isotope peak after deconvolution. M p In this embodiment, we obtain... M p =16951.5 Da. The charge state abundance distribution is represented as a length of... Z max sparse vectorsv charge ,in Z max The maximum possible number of charge states of the precursor is taken in this embodiment. Z max =30, v charge The Middle z k The value of the bit is A k The remaining positions are set to zero; the mass of the single isotope is... M p Represented as a scalar and normalized according to a preset mass scaling factor, it is obtained. v mass .Will v charge and v mass Each feature is mapped to the same dimension as the latent feature representation of the spectral map through two independent fully connected embedding layers. d model From the vector space of 512, we obtain the charge state embedding vector. e charge and quality embedding vector e mass Finally, e charge and e mass The prior constraint vector is obtained by adding the bits one by one. c prior Its dimensions are d model The prior constraint vector carries both the overall mass information of the precursor and the statistical characteristics of the charge state distribution within the source. It serves as the global condition input for the subsequent sequence decoder at each decoding step, ensuring that the prediction results remain consistent in terms of mass conservation and the rationality of charge distribution.

[0054] In some embodiments, the generation process of the amino acid sequence embedding vector and the amino acid physicochemical feature embedding vector is described. This embodiment, based on the spectral latent feature representation and the prior constraint vector obtained in the preceding steps, constructs the two types of amino acid feature embeddings for the input side of the sequence decoder.

[0055] The amino acid sequence embedding vector is obtained by adding the pre-trained amino acid vector to the position encoding.

[0056] This embodiment constructs a vocabulary list for 20 standard amino acids, with each amino acid corresponding to one... d modelThe pre-trained amino acid vectors are obtained through self-supervised training on a large-scale protein sequence database using a masked amino acid prediction task. In this embodiment... d model The dataset is set to 512, and the pre-training database is UniRef50. For each position in the skeleton sequence... p ( p =1,2,…, n , n amino acid residues (length of sequence) a p The corresponding pre-trained amino acid vector is retrieved from the vocabulary. u a_p The position is calculated using a sine / cosine position encoding method. p Corresponding positional encoding vector pos p The two are added position by position to obtain the amino acid sequence embedding vector at that position. e seq p = u a_p + pos p The amino acid sequence embedding vector carries both semantic information about the type of amino acid and semantic information about its position in the backbone sequence.

[0057] The physicochemical feature embedding vector of the amino acid obtains four types of values ​​for each amino acid—hydrophobicity, charge, molecular mass, and secondary structure tendency—by looking up a table. After linear projection onto the dimension of the amino acid sequence embedding vector, these values ​​are added element by element.

[0058] This embodiment pre-constructs a 20-row, 4-column physicochemical characteristic lookup table. Each row corresponds to one amino acid, and the four columns record four types of values ​​for the amino acid: hydrophobicity, charge, molecular weight, and secondary structure tendency. Hydrophobicity is represented by the Kyte-Doolittle hydrophobicity index, charge by the net charge of the side chain at pH 7.0, molecular weight by the monoisotopic mass of the amino acid residue in Da, and secondary structure tendency by the Chou-Fasman α-helical tendency coefficient. For each position... p amino acid residues a p The corresponding four-dimensional physicochemical feature vector is obtained from the physicochemical feature lookup table. f a_p The values ​​of each dimension are normalized to the [0,1] interval according to their respective value ranges, and then subjected to a 4× d model The linear projection layer is mapped to the same vector as the amino acid sequence embedding vector. d model=512-dimensional space, to obtain the amino acid physicochemical feature embedding vector at this position. e phys p The amino acid sequence is embedded into the vector. e seq p Embedded vectors of the physicochemical features of the amino acids e phys p In each location p The amino acid input characteristics of the fusion are obtained by adding elements one by one position. e in p = e seq p + e phys p All positions e in p The sequence features are arranged in sequence to form the input feature matrix of the sequence decoder. The amino acid physicochemical feature embedding vectors supplement the inherent properties of amino acids not covered by the pure sequence embedding, enabling the sequence decoder to maintain strong predictive ability for unseen peptides and post-translational modification sites. The vocabulary can be further expanded to include modification tokens corresponding to post-translational modifications; specific expansion methods are described in the corresponding embodiments below.

[0059] like Figure 5 As shown, the adaptive segmentation and sub-segment quality budget allocation process is further explained. This embodiment follows the spectral latent feature representation obtained from the aforementioned steps. h 1 , h 2 … h L and the single isotopic mass obtained by deconvolution M p =16951.5Da.

[0060] Based on the peak energy distribution in the potential feature representation of the spectrum, N energy peak clusters are identified, and the single isotope mass is divided into N sub-segments using the energy peak clusters as segment boundaries.

[0061] This embodiment describes each position in the latent feature representation of the spectrum. i eigenvectors h i The energy value at that location can be obtained by calculating its L2 norm. e i , i =1,2,…, L For the energy value sequencee 1 , e 2 … e L A smooth energy curve is obtained by averaging a sliding window with a width of 50 Th along the mass-to-charge ratio axis. Then, a local maximum algorithm is used to identify peak positions on the curve, retaining only peaks with heights greater than 1.2 times the global energy mean as energy peak cluster centers. The midpoint between adjacent energy peak cluster centers is used as the sub-segment boundary, resulting in... N =12 energy peak clusters and corresponding 12 sub-segments. The energy peak clusters correspond to the amino acid regions on the complete protein backbone where fragment ions are concentrated. Using these as segment boundaries helps to align the sub-segment boundaries of subsequent beam searches with the fragmentation-prone areas.

[0062] Allocate the quality budget for each segment according to the following formula:

[0063] For the first i Each energy peak cluster covers all spectral peaks within a given range, ranked by their intensity. I k The sum of squares is used to calculate the energy value corresponding to the segment. E i =Σ I k 2 In this embodiment E i The dimension is the square unit of intensity. E i Divide by all N The sum of the energy values ​​of each segment Σ E j The dimensionless energy percentage of this segment is obtained, and then multiplied by the mass of the single isotope. M p =16951.5Da, thus obtaining the basic mass budget allocated to this segment according to its energy proportion, with the dimension Da. A preset mass budget redundancy interval ε is superimposed on the basic mass budget as the allowable deviation for amino acid assignment at the boundary of adjacent segments; in this embodiment, ε is taken as 0.5Da. In this embodiment, the segment with the largest energy proportion is identified. E i / Σ E j =0.142, allocated segment quality budget B i =0.142×16951.5+0.5=2408.6Da; the segment with the smallest energy percentage corresponds to E i / Σ E j=0.041, the allocated segment quality budget B i =695.5Da. All N =The sum of the quality budgets of the 12 segments satisfies Σ B i = M p + N ε = 16951.5 + 12 × 0.5 = 16957.5 Da, ensuring that the sum of the mass budgets for each segment and the mass of the single isotope are strictly conserved within the allowable redundancy range. The segment mass budget... B 1 , B 2 … B N As a hard constraint for subsequent sequence decoders to perform beam search within each segment, the cumulative quality of each candidate path within a segment must not exceed the corresponding... B i This effectively reduces the search space for decoding long sequences of complete proteins.

[0064] It should be noted that this embodiment adopts... E i / Σ E j The allocation weights for the quality budget of each segment are a heuristic allocation strategy based on the energy distribution of the spectrum, and do not require... E i The intensity of the peak is strictly proportional to the actual mass fraction of the corresponding sequence segment. However, in Top-Down mass spectrometry, peak intensity is influenced by various factors such as fragmentation efficiency, ion transport, and detection response, and does not have a strictly linear relationship with sequence fragment quality. This embodiment uses... E i The significance of weighted engineering lies in the fact that regions with concentrated energy peak clusters often correspond to segments rich in fragment information, thus granting these segments a relatively broad mass budget. B i This allows the beam search to explore more candidate amino acid combinations within the sub-segment; while the information in lower-energy sub-segments is relatively sparse, allocating a narrower mass budget can effectively reduce the search space. This strategy, combined with a preset mass redundancy interval ε, allows for slight deviations in the boundary amino acid assignments between adjacent sub-segments, and the sum of the mass budgets for all sub-segments satisfies the global conservation constraint Σ. B i = M p + N ·ε, therefore even individual segments B i Slightly deviating from the true mass share of the sequence, the final output is still affected by the mass of a single isotope.M p Strict constraints ensure that the overall quality consistency of the recognition results is not affected.

[0065] It should be further clarified that the energy peak clusters described in this embodiment are statistical descriptions of the local aggregation patterns of spectral peak energies along the mass-to-charge ratio axis in the latent feature representation of the spectrum, and not physical mappings of continuous amino acid regions in the form of a complete protein. Protein fragmentation sites in a mass spectrometer are discretely distributed, ion types are complex, and the same peptide bond may generate multiple fragments. Therefore, there is no universally accepted, strict one-to-one correspondence between energy peak clusters and sequence amino acid regions. The engineering value of using energy peak clusters as segmentation anchors in this embodiment lies in the fact that the search interval where energy peak clusters are concentrated usually corresponds to a mass window rich in fragmentation information, allowing for a wider sub-segment mass budget within this window. B i This allows beam search to explore more candidate amino acid combinations, improving recognition accuracy; energy-sparse search intervals are allocated narrower sub-segment quality budgets to compress the search space. Regardless of each sub-segment... B i How should the quality budgets of all sub-segments be allocated so that the sum of their quality budgets satisfies the global conservation constraint Σ? B i = M p + N •ε, the specific assignment and boundaries of the sequence among each sub-segment are automatically learned and determined by the bundle search stage under the mass budget constraint, and further corrected by the subsequent multi-step iterative refinement process of the diffusion model. Therefore, the adaptive segmentation method described in this invention does not rely on the physical assumption that energy peak clusters and sequence segments are strictly corresponding.

[0066] like Figure 6 As shown, this embodiment illustrates the process of calculating the location confidence and identifying low-confidence regions. It can be understood that this embodiment follows the skeleton sequence obtained from the preceding steps, and the length of the skeleton sequence... n =153, corresponding to the full length of the complete protein form.

[0067] The position confidence is calculated based on the posterior probability of each amino acid position output by the sequence decoder.

[0068] In this embodiment, the sequence decoder searches for each position on the skeleton sequence during the beam search phase. p ( p =1,2,…, n Each outputs a length of V posterior probability distribution vec p =( P ( a 1 |·),P ( a 2 |·), …, P ( a V |·)), where V The vocabulary size is specified in this embodiment. The vocabulary contains 20 standard amino acid residues and several modification tokens. V Take 40. The skeleton sequence is at position p The actual selected amino acid residue is denoted as â p , and its corresponding posterior probability P ( â p |·) is directly used as the location confidence level for that position. c p , c p ∈[0,1], the closer to 1, the more certain the sequence decoder's prediction of that position. In this embodiment, a position confidence sequence of length 153 is obtained in the manner described above. c 1 , c 2 … c n .

[0069] The region where the location confidence of M consecutive locations is lower than the preset location confidence threshold is defined as a low confidence region.

[0070] This embodiment presets a position confidence threshold θ c Take 0.6, number of consecutive positions M Take 3. For the location confidence sequence from location... p =1 to start sequential scanning, if consecutive... M The location confidence at each position satisfies c p <θ c , c p+1 <θ c … c p+M-1 <θ c Then [ p , p + M -1] is marked as a candidate low-confidence interval; then the scan continues to extend backward, marking all subsequent consecutive intervals that satisfy... c q <θ c Location q Merge into the candidate low-confidence interval until the first one is encountered. c q ≥θc The scanning operation is repeated until all positions have been traversed. In this embodiment, three low-confidence regions were finally identified, located at positions [34,39], [78,84], and [121,128] in the skeleton sequence, respectively. The remaining positions remained unchanged in subsequent iterations.

[0071] The diffusion model only applies noise to the tokens within the low-confidence region and performs iterative denoising.

[0072] For the amino acid tokens in the three low-confidence regions, the time step is scheduled according to a preset noise schedule. t = T In this embodiment, noise is initially applied and denoised progressively according to a reverse time step sequence. T Take 15, the preset noise scheduling adopts a cosine scheduling form, so that the noise variance β t With time step t Increasing according to the cosine curve. At each time step t The diffusion model only updates the token values ​​within the low-confidence region, and the position confidence in the skeleton sequence. c p ≥θ c The high-confidence location tokens remain frozen throughout the iteration. This strategy of iteratively refining only the low-confidence regions preserves the high-quality predictions of the sequence decoder for most locations while avoiding the additional computational overhead of adding noise to the entire sequence and the risk of erroneous rewriting of high-confidence locations, thus achieving a balance between accuracy and efficiency in the final identification of the complete protein form.

[0073] Furthermore, such as Figure 6 As shown, the three low-confidence regions identified in the preceding steps and the single isotopic mass obtained by deconvolution are... M p =16951.5Da, for each inverse denoising time step t The original noise term predicted by the diffusion model described above is corrected using multiple constraints.

[0074] The joint guidance superimposes the three constraint terms onto the noise term predicted by the diffusion model at time step t according to the following formula.

[0075]

[0076] It should be noted that the three constraint terms in the formula... , and

[0077] Numerically, they are all scalars, while the original noise term ε θ ( x t , t ) and the current sequence x t Same dimension. This embodiment uses a classifier-guided approach to transform the three scalar constraints into a similar dimension. x t Superimposing gradient tensors of the same dimension: Specifically, for each time step t, first calculate the scalar values ​​of the three constraint terms according to the formula definition, and then apply them to the current sequence. x t Finding the gradient yields three... x t Gradient tensor of the same dimension , , Finally, the three gradient tensors are multiplied by the weighting coefficient λ respectively. 1 , λ 2 , λ 3 The corrected noise term is then superimposed onto the original noise term. ( x t , t The aforementioned formula expresses this superposition process in a compact form, and its strict tensor form is:

[0078] The gradients are calculated using an automatic differentiation framework (such as PyTorch's autograd) without requiring manual derivation. The scalar values ​​of the three constraint terms and their corresponding gradient tensors reflect the degree of deviation of the current sequence from the constraints and the direction in which correction can best reduce the deviation, respectively. The former is used to monitor the constraint satisfaction status, and the latter is used to drive inverse denoising.

[0079] Furthermore, in this embodiment, the diffusion model is implemented according to a reverse time step sequence from... t =15 Step-by-step noise reduction to t =0, at each time step t The diffusion model first determines the current sequence. x t The latent features of the spectrum represent the output of the original noise term. ε θ ( x t , t Then, calculate the three constraint terms separately and then add the corrections.

[0080] For global quality consistency constraints, based on the current sequence x tThe theoretical cumulative mass of the current sequence is obtained by sequentially adding up the monoisotopic masses carried by all amino acid tokens and modification tokens. M ( x t The dimension is Da; compare it with the mass of the single isotope. M p Take the absolute value after subtraction, then divide by. M p Normalization yields the dimensionless mass deviation term. M ( x t )- M p | / M p For example in t At time 8, the theoretical cumulative mass of the current sequence in this embodiment is... M ( x t =16949.2Da, corresponding to the quality deviation term is |16949.2-16951.5| / 16951.5≈1.36×10 -4 .

[0081] For the local fragment matching degree constraint, according to the current sequence x t The theoretical mass-to-charge ratio of b / y ions is enumerated by the amino acid residue composition. All enumerated theoretical mass-to-charge ratios of b / y ions are binned at 50 Th intervals and encoded using single-thermal encoding to form theoretical fragment spectral vectors of the same dimension as the spectral vector representation. t x_t If theoretical b / y ions are present in each sub-compartment, the corresponding value is set to 1; otherwise, it is set to 0. t x_t With the spectral vector representation v spec After projecting onto the same feature space, the cosine similarity is calculated, resulting in... S match ( x t )= t x_t · v spec / (|| t x_t ||·|| v spec ||)∈[0,1]. This embodiment is in... t =Calculated at time 8 S match ( x t =0.83, corresponding to constraint term 1- Smatch ( x t =0.17.

[0082] For the modification site compatibility constraint, for the current sequence x t For each position predicted as a modification token, the matching relationship between the modifiable amino acid residue type recorded by that modification token and the amino acid residue at that position is checked. The incompatibility ratio is obtained by dividing the number of incompatible positions by the total number of modification positions. P PTM ( x t ()∈[0,1], and the specific calculation method is described in the corresponding embodiment below. This embodiment is in t At time 8, the current sequence predicts a total of 3 modification positions, namely position... p =37 acetylated token, position p =81 phosphorylated token and its position p =124 glycosylated tokens, where adjacent amino acid residues at the position of the glycosylated token are incompatible, corresponding to P PTM ( x t )≈0.33.

[0083] The three constraint terms are weighted by λ. 1 , λ 2 , λ 3 Weighted superposition to the original noise term ε θ ( x t , t The corrected noise term is obtained on ) ( x t , t In this embodiment, λ 1 Take 1.0, λ 2 Take 0.5, λ 3 The weight is set to 2.0, and the weights are obtained through parameter tuning via a grid search on the labeled dataset. The corrected noise term... ( x t , t Used for driving according to preset noise. x t arrive x t-1 The reverse denoising method. The joint guidance ensures that the denoising result of each step simultaneously approaches the cumulative quality. M pThe theoretical fragment spectrum approximates the measured spectrum, and the modification sites approximate the chemical rationality in a convergent direction. The final output of the complete protein form is guaranteed in terms of mass conservation, spectrum matching, and modification rationality.

[0084] In some embodiments, the construction of the modified token vocabulary and the calculation process of the modification site compatibility are described. Understandably, this embodiment follows the steps described above regarding the joint guidance process of the amino acid sequence embedding vector on the input side of the sequence decoder and the diffusion model on the output side, and describes the specific construction, attribute recording, and compatibility determination method of the modified token.

[0085] The vocabulary corresponding to the amino acid sequence embedding vector is pre-expanded with modification tokens corresponding to common post-translational modification types. Each modification token carries the molecular weight of the corresponding modification group and the type of amino acid residue that can be modified.

[0086] In this embodiment, in addition to the 20 standard amino acid residue tokens, the vocabulary corresponding to the amino acid sequence embedding vector is further expanded with 20 modification tokens corresponding to common post-translational modification types. These modification tokens include, but are not limited to, phosphorylation, acetylation, methylation, dimethylation, trimethylation, ubiquitination, glycosylation, nitration, oxidation, and amidation. The total vocabulary size is expanded from 20 to 40, with each modification token corresponding to an independent index, which is maintained in the embedding layer. d model =512-dimensional learnable embedding vectors. Each modification token records its attributes in a triplet attribute table, including the modification token name, the monoisotopic molecular mass Δ of the corresponding modification group. m (Dimensions in Da) and the set of types of amino acid residues that can be modified. R mod For example, the Δ of phosphorylated tokens m =79.9663Da R mod ={S, T, Y}; Δ of acetylated token m =42.0106Da R mod ={K, N-terminus}; Δ of methylated token m =14.0157Da R mod ={K, R};Δ of oxidized token m =15.9949Da R mod ={M, W, C}; Δ of glycosylated tokens m =162.0528Da Rmod ={N, S, T}. The attribute table is used as a constant lookup table during the model inference phase. In this embodiment, the modification token is inserted into the token sequence as an independent token immediately following the modified amino acid residue token. For example, the token sequence fragment corresponding to the acetylation of lysine at position 37 is "...,K, Acetyl, ...". The sequence decoder outputs not only the posterior probability for the 20 amino acid residue tokens at each decoding position, but also the posterior probability for the 20 modification tokens, thereby achieving end-to-end unified modeling of amino acid residue identification and post-translational modification site identification.

[0087] The modification site compatibility is calculated based on the matching relationship between the types of modifiable amino acid residues recorded by the modification token at the current position and the amino acid residues at the current position, and a penalty term is applied to mismatched positions.

[0088] At each time step of the diffusion model inverse denoising t For the current sequence x t Perform compatibility checks on all positions predicted to modify the token. For each modification position... p First, identify the type of the modified token at that position and then look up the set of amino acid residue types that can be modified in the attribute table. R mod Then read the current sequence. x t The modification of the amino acid residue preceding the token as described in the text a p (If the modified token is located at the beginning of the sequence, then) a p (Take the first amino acid residue of the sequence). Judgment a p Does it belong to the set of modifiable amino acid residue types? R mod :like a p ∈ R mod The modified position is then marked as compatible, with a corresponding position weight. w p =0; if a p ∉ R mod The modified position is then marked as incompatible, and the corresponding position weight is... w p =1. This embodiment is in... t At time 8, the current sequence predicts three modification positions, all of which fall within the three low-confidence regions identified in Example 7.p =37 falls into the low confidence region [34, 39], and is predicted to be an acetylated token. a p =K, belonging to R mod ={K,N end}, marked as compatible; position p =81 falls into the low confidence region [78, 84], and is predicted to be a phosphorylated token. a p =T, belonging to R mod ={S, T, Y}, marked as compatible; position p =124 falls into the low confidence region [121, 128], and is predicted to be a glycosylated token. a p =L, does not belong to R mod ={N, S, T}, marked as incompatible.

[0089] After determining the compatibility of each modification location, the incompatibility ratio is... P PTM ( x t )=Σ w p / n mod ,in n mod For the current sequence x t The total number of modified tokens mentioned in the example, in this embodiment n mod =3、Σ w p =1, corresponding to P PTM ( x t = 1 / 3 ≈ 0.33. This incompatibility ratio is used as the value of the modification site compatibility constraint term, multiplied by the preset weight coefficient λ. 3 This is then superimposed onto the original noise term predicted by the diffusion model at that time step, and a penalty term is applied to incompatible positions, making the modification tokens more likely to appear at chemically reasonable amino acid residues during subsequent denoising, thereby improving the chemical rationality of post-translational modification site annotations in the complete protein form.

[0090] In some embodiments, the peak encoder, the sequence decoder, and the diffusion model used in the inference phase need to undergo end-to-end joint training before being put into use. The training process is based on a labeled dataset constructed from known complete protein form samples, and the parameters of each model are jointly optimized on the dataset.

[0091] A labeled dataset is constructed based on known complete protein samples. A target decoy strategy is adopted to use correct spectral sequence pairs as positive samples and decoy sequences generated by flipping or randomly perturbing positive samples as negative samples.

[0092] The real spectrum-sequence pairs in this embodiment are derived from the publicly available proteomics database ProteomeXchange and its associated Top-Down datasets (such as the Top-Down datasets in MassIVE and PRIDE). Spectra meeting the following criteria were selected: instrument resolution of at least 60,000; precursor monoisotope mass determinable by deconvolution; and the corresponding complete protein form confirmed by database alignment using publicly available tools such as TopPIC, ProSight, or MASH Suite, and verified against a 1% FDR (false detection rate) threshold. Approximately 20,000 real-labeled spectra were obtained based on these criteria. These were then expanded to approximately 80,000 spectra using three data augmentation methods: Gaussian noise perturbation, random peak discarding, and slight mass-to-charge ratio shift. Each spectrum carries the corresponding real amino acid residue sequence and post-translational modification site annotations. The post-translational modification site annotations are uniformly coded according to the modification type registry number in the Unimod database, forming a positive sample set. For each positive sample, decoy sequences are generated in two ways: first, sequence reversal, where the correct sequence is reversed from C-terminus to N-terminus; second, random perturbation, where amino acid residues are replaced or modified at 10% random positions in the correct sequence to remove the token. Two decoy sequences are generated for each positive sample, resulting in approximately 160,000 negative samples. All positive and negative samples are divided into training and validation sets in a 4:1 ratio, forming the labeled dataset.

[0093] The peak encoder, the sequence decoder, and the diffusion model are jointly trained end-to-end by weighting and summing the sequence prediction cross-entropy loss, the mean square error loss between the diffusion model's predicted noise and the actual noise, and the modification site prediction cross-entropy loss to form a total loss function.

[0094] In this embodiment, a batch of training samples is extracted from the labeled dataset at each training step, with a batch size of 32. For each sample within the batch, the peak encoder and the sequence decoder first perform forward propagation on its spectrum to obtain the predicted amino acid residue distribution of the backbone sequence, and then calculate the sequence prediction cross-entropy loss with the actual amino acid residue sequence of that sample. L seqThen, Gaussian noise is added to the real sequence of the sample according to a preset noise schedule to obtain a noisy sequence, which is then input into the diffusion model. The diffusion model predicts the added noise, and the mean squared error loss is calculated by comparing it with the actual added noise. L diff Finally, the predicted cross-entropy loss of the modified tokens output by the sequence decoder is calculated between the probability distribution of the modified tokens and the actual modified site annotations of the sample. L PTM The three losses are calculated according to the formula. L total =α· L seq +β· L diff +γ· L PTM The total loss function is formed by weighted summation; in this embodiment, α is 1.0, β is 0.5, and γ is 2.0. The AdamW optimizer is used to backpropagate and update the gradients of the total loss function with respect to all parameters of the peak encoder, the sequence decoder, and the diffusion model. The initial learning rate is 1×10⁻⁶. -4 Cosine annealing was used for scheduling decay, with a total of 200 training rounds. During training, the sequence recognition accuracy and modification site recognition recall were evaluated on the validation set every 5 rounds. The model parameters with the best overall performance on the validation set were used as the final model for the inference phase.

[0095] The above description is merely an exemplary embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's specification and drawings under the technical concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A deep learning-based top-down mass spectrometry protein identification method, characterized in that, include: S1: Preprocess and encode the Top-Down mass spectrometry data of the intact protein form to be identified to obtain a spectral vector representation; S2: The peak encoder encodes the spectral vector representation to obtain the spectral latent feature representation; S3: Jointly embed the single isotopic mass and multi-charge state distribution of the complete protein precursor to obtain the prior constraint vector; S4: Adaptively segment the mass of the single isotope according to the energy distribution of the spectral peaks in the latent feature representation of the spectrum to obtain several sub-segment mass budgets; S5: The latent feature representation of the spectrum, the prior constraint vector, the amino acid sequence embedding vector, and the amino acid physicochemical feature embedding vector are fused and input into the sequence decoder. The sequence decoder performs a bundle search within each sub-segment, constrained by the quality budget of the corresponding sub-segment, and splices the outputs of each sub-segment into a backbone sequence. S6: The diffusion model performs multi-step iterative refinement on regions in the backbone sequence whose position confidence is lower than the preset position confidence threshold. The denoising process is simultaneously guided by global quality consistency, local fragment matching degree and modification site compatibility, and outputs the complete protein form containing amino acid residue sequences and post-translational modification site annotations and its confidence score.

2. The method according to claim 1, characterized in that, The preprocessing and encoding in step S1 specifically involves: performing baseline correction, denoising, and peak extraction on the original spectrum of the Top-Down mass spectrometry data to obtain a preprocessed spectrum; The mass-to-charge ratio of each spectral peak in the preprocessed spectrum is transformed using sine and cosine functions of different frequencies to obtain the position encoding vector; The position encoding vector is concatenated with the corresponding intensity information to form the spectral vector representation.

3. The method according to claim 1, characterized in that, The peak encoder is a multi-layer Transformer encoding layer, each layer containing a multi-head self-attention sub-layer and a feedforward network sub-layer, with residual connections and layer normalization set between the sub-layers.

4. The method according to claim 1, characterized in that, Step S3 is as follows: Several charge states and their relative abundances of the complete protein precursor were extracted from the preprocessed spectrum to form a charge state abundance distribution; The prior constraint vector is obtained by projecting the charge state abundance distribution and the single isotope mass obtained by deconvolution onto the same vector space and summing them.

5. The method according to claim 1, characterized in that, The amino acid sequence embedding vector is obtained by adding the pre-trained amino acid vector and the position code; the amino acid physicochemical feature embedding vector obtains four types of values ​​for each amino acid—hydrophobicity, charge, molecular mass, and secondary structure tendency—by looking up a table, and then linearly projects them onto the dimension of the amino acid sequence embedding vector and adds them element by element.

6. The method according to claim 1, characterized in that, Step S4 specifically involves: identifying N energy peak clusters based on the spectral peak energy distribution in the latent feature representation of the spectrum; dividing the single isotope mass into N sub-segments using the energy peak clusters as segment boundaries; and allocating the mass budget for each sub-segment according to the following formula: in, B i For the first i The sub-segment quality budget, in units of Da; M p The mass of the single isotope is in the dimension Da; E i For the first i The sum of the squares of the intensities of all spectral peaks within an energy peak cluster represents the energy value corresponding to that segment; ε is the preset mass budget redundancy interval, with dimensions Da; and the mass budget of each segment satisfies: 。 7. The method according to claim 1, characterized in that, The position confidence in step S6 is calculated based on the posterior probability of each amino acid position output by the sequence decoder; the region where the position confidence of M consecutive positions is lower than the preset position confidence threshold is determined as a low confidence region; The diffusion model only applies noise to the tokens within the low-confidence region and performs iterative denoising, while the remaining positions in the skeleton sequence remain unchanged during the iteration process.

8. The method according to claim 1, characterized in that, In step S6, the joint guidance adds the three constraint terms to the noise term predicted by the diffusion model at time step t according to the following formula: in, ε θ ( x t , t ) represents the original noise term predicted by the diffusion model; M ( x t ) represents the current sequence x t Theoretical cumulative mass, and M p Both are units of Da, and the difference between them is... M p After normalization, it becomes a dimensionless quality deviation term; S match ( x t The cosine similarity between the theoretical fragment spectrum corresponding to the current sequence and the spectral vector representation is [0,1]. P PTM ( x t The value represents the incompatibility ratio between the modified token and the amino acid residue type at its current position in the current sequence, and its value ranges from [0,1]. λ 1 , λ 2 , λ 3 Dimensionless weighted coefficients; corrected noise term ( x t , t ) used for driving x t arrive x t-1 Inverse denoising.

9. The method according to claim 1, characterized in that, The vocabulary corresponding to the amino acid sequence embedding vector is pre-expanded with modification tokens corresponding to common post-translational modification types. Each modification token carries the molecular weight of the corresponding modification group and the type of amino acid residue that can be modified. In step S6, the modification site compatibility is calculated based on the matching relationship between the type of amino acid residue that can be modified recorded by the modification token at the current position and the amino acid residue at the current position. A penalty term is applied to mismatched positions.

10. The method according to claim 1, characterized in that, It also includes a model training step: constructing a labeled dataset based on known complete protein form samples, using a target decoy strategy to use correct spectral sequence pairs as positive samples, and using decoy sequences generated from positive samples by flipping or random perturbation as negative samples; using a weighted summation of sequence prediction cross-entropy loss, mean square error loss between diffusion model predicted noise and real noise, and modification site prediction cross-entropy loss to form a total loss function, and performing end-to-end joint training on the peak encoder, the sequence decoder, and the diffusion model.