Peptide fragment secondary spectrogram prediction method for introducing protein large language model coding
By introducing protein large language model encoding and index single-hot encoding methods, the peptide information is encoded, which solves the problems of low efficiency and low accuracy of secondary spectrogram prediction of peptides in the prior art, and achieves more efficient and accurate secondary spectrogram prediction of peptides.
Patent Information
- Application Number
- CN202510422815.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
The existing secondary peptide spectrogram prediction software has low operating efficiency, and the prediction accuracy of peptide spectrograms with post-translation modifications is poor, which cannot meet the current data processing requirements.
The peptide information is encoded by protein large language model encoding, index encoding and single-hot encoding. The fragment ion intensity of the peptide is obtained through the encoding vector and spectrum prediction model, and the secondary spectrum is constructed.
The accuracy and efficiency of secondary spectrogram prediction of peptides is improved, especially the prediction accuracy of peptides with post-translational modifications.
Smart Images

Figure CN120279981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of biotechnology and proteomics, and in particular to a method for predicting the secondary spectrum of peptide segments encoded by a protein large language model. Background Art
[0002] Currently, proteomics often uses mass spectrometry technology for analysis. The steps of one of the mainstream algorithms are as follows: First, digest proteins with enzymes to form peptide segments; then send the peptide segments into a mass spectrometer to draw a primary spectrum to reflect the process of ion signals from emergence to disappearance; then perform matching within the retention time window based on the relevant information of the primary spectrum to obtain the peptide segment information in the corresponding chromatographic interval from another running program; finally, batch process the incoming components according to the above steps to obtain the corresponding mass spectrometry data set, and then deduce the types and contents of proteins through algorithms. Among them, the abscissa of the primary spectrum is the mass-to-charge ratio (m / z) of the spectral peak signal, the ordinate is the spectral peak intensity (Intensity), and the isotopes of different molecules of the same peptide segment form a multi-isotope pattern. However, proteins cannot be accurately identified only through the signals of the primary spectrum. Therefore, researchers draw a secondary spectrum for the fragment ions of the peptide segment. The secondary spectrum reflects the process of peptide segment fragment ions from emergence to disappearance. The abscissa and ordinate of the secondary spectrum are the same as those of the primary spectrum, and each fragment ion has a unique spectral peak signal in the secondary spectrum.
[0003] Existing peptide secondary spectrum prediction software generally has the problems of low running efficiency, only targeting a small number of post-translational modifications, or poor prediction accuracy for spectra of peptide segments with post-translational modifications, so it cannot meet the current requirements for data processing. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for predicting the secondary spectrum of peptide segments encoded by a protein large language model to improve the accuracy of predicting the secondary spectrum of peptide segments.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] In the first aspect, the present invention provides a method for predicting the secondary spectrum of peptide segments encoded by a protein large language model, including: obtaining the peptide segments of a protein, and encoding the peptide segment information of the peptide segments using a preset encoding algorithm to obtain an encoding vector of the peptide segments; wherein, the encoding algorithm includes: encoding based on a protein large language model, index encoding, and one-hot encoding; obtaining the fragment ion intensity of the peptide segments based on the encoding vector and a spectrum prediction model; and converting the fragment ion intensity into the secondary spectrum of the peptide segments.
[0007] Optionally, the peptide information includes at least: amino acid sequence, post-translational modification, charge, normalized collision energy, and instrument; the peptide information of the peptide is encoded using a preset encoding algorithm to obtain a peptide encoding vector, including: encoding the amino acid sequence of the peptide according to a preset amino acid index to obtain a peptide sequence encoding vector; performing structure encoding on the amino acid sequence based on a protein large language model to obtain a peptide structure feature encoding vector; adding the post-translational modification to the corresponding peptide sequence encoding vector, and performing dimensionality reduction on the peptide sequence encoding vector after adding the modification based on a linear neural network to obtain a peptide modification encoding vector; performing one-hot encoding on the charge, and replicating the encoding of the charge according to the sequence length of the peptide to obtain a charge encoding vector; performing one-hot encoding on the normalized collision energy after multiplying it by a preset energy coefficient, and replicating the encoding of the normalized collision energy according to the sequence length of the peptide to obtain a normalized fragmentation energy encoding vector; obtaining the identifier corresponding to the instrument, performing one-hot encoding on the identifier, and replicating the encoding of the identifier according to the sequence length of the peptide to obtain an instrument encoding vector.
[0008] Optionally, performing structure encoding on the amino acid sequence based on a protein large language model to obtain a peptide structure feature encoding vector, including: inputting the amino acid sequence into the pre-trained model of the protein large language model to obtain a structure feature matrix of the amino acid sequence; performing dimensionality reduction on the structure feature matrix based on a fully connected neural network to obtain a peptide structure feature encoding vector.
[0009] Optionally, based on the encoding vector and the spectrum prediction model, the fragment ion intensity of the peptide is obtained, including: splicing the encoding vectors to obtain a feature vector of the peptide information; wherein, the encoding vectors include: peptide sequence encoding vector, peptide structure feature encoding vector, peptide modification encoding vector, charge encoding vector, normalized fragmentation energy encoding vector, and instrument encoding vector; inputting the feature vector of the peptide information into the spectrum prediction model to obtain the fragment ion intensity of the peptide.
[0010] Optionally, the spectrum prediction model includes: an encoding layer and a decoding layer; inputting the feature vector of the peptide information into the spectrum prediction model to obtain the fragment ion intensity of the peptide, including: converting the feature vector of the peptide information into a feature matrix through the encoding layer; inputting the feature matrix into the decoding layer to obtain the fragment ion intensity of the peptide.
[0011] Optionally, converting the fragment ion intensity into a secondary spectrum of the peptide, including: normalizing the fragment ion intensity; using the mass-to-charge ratio of the fragment ion as the abscissa and the normalized fragment ion intensity as the ordinate to construct the secondary spectrum of the peptide.
[0012] Second aspect, the present invention provides a device for predicting the secondary spectrum of a peptide segment encoded by a protein large language model, including: an encoding module, configured to obtain the peptide segments of a protein and encode the peptide segment information of the peptide segments using a preset encoding algorithm to obtain an encoded vector of the peptide segments; wherein, the encoding algorithm includes: encoding based on a protein large language model, index encoding, and one-hot encoding; a prediction module, configured to obtain the fragment ion intensity of the peptide segments based on the encoded vector and a spectrum prediction model; a secondary spectrum drawing module, configured to convert the fragment ion intensity into the secondary spectrum of the peptide segments.
[0013] Optionally, the peptide segment information at least includes: amino acid sequence, post-translational modification, charge, normalized collision energy, and instrument; the encoding module is specifically configured to: encode the amino acid sequence of the peptide segment according to a preset amino acid index to obtain a peptide segment sequence encoded vector; perform structural encoding on the amino acid sequence based on a protein large language model to obtain a peptide segment structure feature encoded vector; add the post-translational modification to the corresponding peptide segment sequence encoded vector, and perform dimensionality reduction on the peptide segment sequence encoded vector after adding the modification based on a linear neural network to obtain a peptide segment modification encoded vector; perform one-hot encoding on the charge, and replicate the encoding of the charge according to the sequence length of the peptide segment to obtain a charge encoded vector; multiply the normalized collision energy by a preset energy coefficient and then perform one-hot encoding, and replicate the encoding of the normalized collision energy according to the sequence length of the peptide segment to obtain a normalized fragmentation energy encoded vector; obtain the identifier corresponding to the instrument, perform one-hot encoding on the identifier, and replicate the encoding of the identifier according to the sequence length of the peptide segment to obtain an instrument encoded vector.
[0014] Third aspect, the present invention provides an electronic device, including a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any method provided in the first aspect above.
[0015] Fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of any method provided in the first aspect above.
[0016] The present invention brings the following beneficial effects:
[0017] The above-mentioned method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model provided by the present invention first obtains the peptide segments of a protein, and encodes the peptide segment information of the peptide segments using a preset encoding algorithm to obtain the encoded vector of the peptide segments; wherein, the encoding algorithm includes: encoding based on a protein large language model, index encoding, and one-hot encoding; then, based on the encoded vector and the spectrum prediction model, the fragment ion intensity of the peptide segment is obtained; finally, the fragment ion intensity is converted into the secondary spectrum of the peptide segment. The above method encodes the peptide segment information of the peptide segments through encoding algorithms such as encoding based on a protein large language model, index encoding, and one-hot encoding to obtain the encoded vector of the peptide segments, obtains the fragment ion intensity of the peptide segment according to the encoded vector and the spectrum prediction model, and finally constructs the secondary spectrum according to the fragment ion intensity of the peptide segment, improving the accuracy of predicting the secondary spectrum of the peptide segment.
[0018] Other features and advantages of the present invention will be described in the subsequent description, and, in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, claims, and drawings.
[0019] To make the above-mentioned objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, is described in detail as follows. Description of the Drawings
[0020] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of a method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model provided by an embodiment of the present invention;
[0022] Figure 2 It is a schematic diagram of predicting the secondary spectrum of a peptide segment provided by an embodiment of the present invention;
[0023] Figure 3 It is a schematic diagram of a sequence encoding method provided by an embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of a protein large language model outputting a peptide segment structure feature vector provided by an embodiment of the present invention;
[0025] Figure 5 It is a schematic diagram of a peptide segment experimental information encoding method provided by an embodiment of the present invention;
[0026] Figure 6 This is a flowchart of a method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model provided by an embodiment of the present invention;
[0027] Figure 7 This is a schematic structural diagram of a device for predicting the secondary spectrum of a peptide segment encoded by a protein large language model provided by an embodiment of the present invention;
[0028] Figure 8 This is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] Currently, existing peptide secondary spectrum prediction software generally has the problems of low operating efficiency, only targeting a small number of post-translational modifications, or poor prediction accuracy for spectra of peptides with post-translational modifications, and thus cannot meet the requirements of current data processing. Specifically, the existing prediction software mainly has the following disadvantages: (1) The existing software has low sequence coding efficiency for peptide segments, resulting in low prediction efficiency. (2) The existing software has low prediction accuracy for spectra of peptide segments with post-translational modifications and does not have a spectral prediction method for multiple post-translational modifications.
[0031] Based on this, a method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model provided by an embodiment of the present invention can improve the accuracy of predicting the secondary spectrum of peptide segments.
[0032] To facilitate the understanding of this embodiment, first, a method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model disclosed in an embodiment of the present invention will be introduced in detail. This method can be executed by an electronic device, such as a smart phone, a computer, a tablet computer, etc. Refer to Figure 1 The flowchart of a method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model shown, which shows that the method mainly includes the following steps S101 to step S103:
[0033] Step S101: Obtain the peptide segments of a protein, and encode the peptide segment information of the peptide segments using a preset encoding algorithm to obtain the encoding vectors of the peptide segments.
[0034] In one embodiment, first, peptides are obtained by digesting proteins, and peptide information of the peptides is acquired. The peptide information includes at least: amino acid sequence, post-translational modification, charge, normalized collision energy, and instrument (i.e., the instrument used to obtain peptide-related information). Then, a preset encoding algorithm is used to encode the peptide information of the peptides to obtain an encoded vector of the peptides. The encoding algorithms used in this embodiment include: encoding based on a protein large language model, index encoding, one-hot encoding, etc. For example: index encoding is used to encode the amino acid sequence, the protein large language model is used to perform structural encoding on the peptides, and one-hot encoding is used to encode the charge, normalized collision energy, and instrument.
[0035] Step S102: Based on the encoded vector and the spectral prediction model, obtain the fragment ion intensity of the peptide.
[0036] In one embodiment, during the training of the spectral prediction model, the feature vector of the peptide information of each peptide can be used as the input, and the fragment ion intensity of the peptide can be used as the output to make a training set, and the spectral prediction model is trained. During the training process, the activation function of the network uses ReLU and Softmax, the loss function uses the L1 loss function, as shown in formula (1), the optimizer uses SGD, the hyperparameters are adjusted, and then the training is carried out until the loss curve drops to convergence.
[0037] l(x,y) = L = {l1, l2... l n} T , l n = ||x n - y n | (1)
[0038] Based on this, in the embodiment of the present invention, after obtaining the encoded vector, the encoded vector can be passed through an embedding layer to obtain respective feature vectors, all the feature vectors are concatenated to obtain the feature vector of the peptide information, and then the feature vector is input into the spectral prediction model to obtain the fragment ion intensity of the peptide.
[0039] Step S103: Convert the fragment ion intensity into a secondary spectrum of the peptide.
[0040] In one embodiment, first, the fragment ion intensity is normalized, and then, with the mass-to-charge ratio of the fragment ions as the abscissa and the normalized fragment ion intensity as the ordinate, a secondary spectrum of the peptide is constructed. Among them, for predicting the secondary spectrum of the peptide, reference can be made to Figure 2 as shown.
[0041] In specific implementation, the fragment ion intensity is output in the form of relative intensity. For all the fragment ion intensities of each peptide segment, the highest fragment ion intensity is taken, and the remaining fragment ions are normalized by dividing them by the highest fragment ion intensity to obtain the relative intensity. Among them, the method of normalizing the fragment ion intensity is shown in formula (2):
[0042]
[0043] Among them, Input Intensity is the normalized fragment ion intensity, MaxIntensity is the highest fragment ion intensity, and FragIntensity is the intensity of the i-th fragment ion of the peptide segment.
[0044] A method for predicting the secondary spectrum of a peptide segment encoded by a protein large language model provided by an embodiment of the present invention encodes the peptide segment information of the peptide segment through encoding algorithms such as protein large language model encoding, index encoding, and one-hot encoding to obtain an encoded vector of the peptide segment, and obtains the fragment ion intensity of the peptide segment according to the encoded vector and the spectrum prediction model. Finally, a secondary spectrum is constructed based on the fragment ion intensity of the peptide segment, improving the accuracy of predicting the secondary spectrum of the peptide segment.
[0045] In one implementation manner, for the foregoing step S101, that is, when encoding the peptide segment information of the peptide segment by using a preset encoding algorithm to obtain an encoded vector of the peptide segment, the following manners may be adopted, including but not limited to:[[]]
[0046] (1) Encode the amino acid sequence of the peptide segment according to a preset amino acid index to obtain a peptide segment sequence encoded vector.
[0047] In specific implementation, when encoding the amino acid sequence, an index (including a starting bit set to 0) may be set for each amino acid. As Figure 3 shown, among them, the index of alanine (A) is set to 1, the index of cysteine (C) is set to 3, and so on, thereby converting the amino acid sequence from a string structure to an encoded vector structure. By means of index encoding, the situation of redundant matrices during encoding can be avoided, and the complexity of the encoded vector E(I) of the amino acid sequence is reduced from O(n 2 ) to O(n). Specifically, the peptide segment sequence encoded vector E(I) is shown in formula (3), where I represents the amino acid index.
[0048]
[0049] (2) Perform structural encoding on the amino acid sequence based on the protein large language model to obtain a peptide segment structural feature encoded vector.
[0050] In specific implementation, first, the amino acid sequence is input into the pre-trained model of the protein large language model to obtain the structural feature matrix of the amino acid sequence; then, based on the fully connected neural network, the dimensionality of the structural feature matrix is reduced to obtain the peptide structural feature encoding vector.
[0051] Specifically, as shown in Figure 4 When using the protein large language model to output the peptide structural feature encoding, first, the amino acid sequence of the peptide is input into the pre-trained model of the protein large language model, and then the pre-trained model of the protein large language model will output a structural feature matrix of the amino acid sequence. After that, the structural feature matrix is reduced in dimensionality through a fully connected neural network to obtain the peptide structural feature encoding.
[0052] (3) Add the post-translational modification to the corresponding peptide sequence encoding vector, and based on the linear neural network, reduce the dimensionality of the peptide sequence encoding vector added with the modification to obtain the peptide modification encoding vector.
[0053] In specific implementation, when encoding the peptide with post-translational modification, first attach each modified C, H, N, O, S, P atom to the corresponding peptide sequence encoding vector, and then reduce the dimensionality of the encoding vector through a linear neural network to obtain the peptide modification encoding vector. This method can meet the encoding of most modifications, thereby improving the accuracy of the secondary spectrum prediction.
[0054] (4) Perform one-hot encoding on the charge, and copy the encoding of the charge according to the sequence length of the peptide to obtain the charge encoding vector.
[0055] In specific implementation, when encoding the charge, first perform one-hot encoding on the charge number, and then copy it in the dimension of the peptide sequence, that is, copy the encoding result of the one-hot encoding of the charge number to the same length as the peptide sequence to obtain the charge encoding vector.
[0056] (5) Multiply the normalized collision energy by a preset energy coefficient and then perform one-hot encoding, and copy the encoding of the normalized collision energy according to the sequence length of the peptide to obtain the normalized fragmentation energy encoding vector.
[0057] In specific implementation, when encoding the normalized fragmentation energy, after multiplying the normalized fragmentation energy by the energy coefficient, first perform one-hot encoding, and then copy it in the dimension of the peptide sequence, that is, copy the encoding result of the normalized fragmentation energy to the same length as the peptide sequence to obtain the normalized fragmentation energy encoding vector.
[0058] (6) Obtain the identifier corresponding to the instrument, perform one-hot encoding on the identifier, and copy the encoding of the identifier according to the sequence length of the peptide to obtain the instrument encoding vector.
[0059] In specific implementation, when encoding the instrument, a unique identifier is set for each instrument, such as: Lumos: 1, QE: 2, and so on. Then, one-hot encoding is performed on the identifier, and then it is replicated in the dimension of the peptide sequence, that is, the encoding result of the instrument is replicated to be the same length as the peptide sequence, obtaining the instrument encoding vector.
[0060] Further, as shown in Figure 5 the charge encoding vector, the normalized fragmentation energy encoding vector, and the instrument encoding vector are concatenated to obtain the peptide experimental information encoding vector.
[0061] In one implementation manner, for the aforementioned step S102, that is, when obtaining the fragment ion intensity of the peptide based on the encoding vector and the spectrum prediction model, the following methods can be adopted, including but not limited to the following steps 1 to 2:
[0062] Step 1: Concatenate the encoding vectors to obtain the feature vector of the peptide information.
[0063] In specific implementation, all encoding vectors (peptide sequence encoding vector, peptide structure feature encoding vector, peptide modification encoding vector, charge encoding vector, normalized fragmentation energy encoding vector, and instrument encoding vector) can obtain their respective feature vectors through an embedding layer, and then all feature vectors are concatenated to obtain the feature vector of the peptide information.
[0064] Step 2: Input the feature vector of the peptide information into the spectrum prediction model to obtain the fragment ion intensity of the peptide.
[0065] In specific implementation, the spectrum prediction model includes: an encoding layer and a decoding layer. The encoding layer can extract information such as the fragmentation rule and fragment ion intensity of the peptide from the feature vector, mine the features required for spectrum prediction from the high-dimensional features, and finally output the predicted fragment ion intensity through the decoding layer.
[0066] Based on this, in the embodiment of the present invention, when inputting the feature vector of the peptide information into the spectrum prediction model to obtain the fragment ion intensity of the peptide, the following methods can be adopted, including but not limited to:
[0067] First, convert the feature vector of the peptide information into a feature matrix through the encoding layer.
[0068] Specifically, the main structure of the encoding layer is the encoder architecture of Transformer, including position encoding, multi-head attention mechanism, and feed-forward network. The aforementioned feature vector is converted into a feature matrix H through the Transformer encoder. L, as shown in formula (4), where LayerNorm is layer normalization, X is the input feature vector, Attention(Q, K, V) represents the attention mechanism, and FFN represents the feed-forward network.
[0069] H L = LayerNorm(X + FFN(LayerNorm(X + Attention(Q, K, V)))) (4)
[0070] Then, the feature matrix is input into the decoding layer to obtain the fragment ion intensities of the peptide.
[0071] Specifically, the main structure of the decoding layer is the Transformer decoder structure, including: multi-head attention mechanism, feed-forward network, and linear layer. The dimension 4 of the last layer of the decoding layer represents 4 types of peptide fragment ions, namely, monovalent b ions (b+), divalent b ions (b++), monovalent y ions (y+), and divalent y ions (y++).
[0072] For ease of understanding, the embodiment of the present invention also provides a flowchart of a method for predicting peptide secondary spectra introduced by protein large language model encoding. Refer to Figure 6 shown, which mainly includes the following S1 to S3:
[0073] S1: Encode the information of the peptide, including:
[0074] Index-encode the amino acid sequence of the peptide, convert the peptide sequence from a string to an encoding matrix, and form a peptide sequence encoding vector;
[0075] Encode the post-translational modification of the peptide, embed the encoding vector in the sequence dimension, and form a peptide modification encoding vector of the post-translational modification;
[0076] Encode the normalized fragmentation energy, charge, and instrument of the peptide, expand them in the sequence dimension, form a charge encoding vector, a normalized fragmentation energy encoding vector, and an instrument encoding vector, and splice them to obtain an encoding vector of the peptide experimental information;
[0077] Input the amino acid sequence of the peptide into the protein large language model, output the feature vector of the peptide structure information, and form a peptide structure feature encoding vector.
[0078] S2: Input the feature vector into the encoding layer. The encoding layer extracts the features required for spectrum prediction from the high-dimensional features, and finally outputs the predicted fragment ion intensities through the decoding layer.
[0079] S3: Convert the output peptide fragment ion intensities into a peptide secondary spectrum.
[0080] Specifically, according to the output of the last layer of the decoding layer in S2, each peptide segment will output a vector with a length of the sequence length minus 1 and a dimension of 4, as shown in Table 1. Taking the mass-to-charge ratio of the peptide fragment ions as the abscissa and the fragment ion intensity as the ordinate, after normalizing the fragment ion intensity, the secondary spectrum of the peptide segment can be obtained, as Figure 2 shown.
[0081] Table 1 Schematic table of output fragment ion intensities
[0082]
[0083] The above method provided by the embodiments of the present invention can obtain the sequence coding vector more efficiently through the index coding method, making the network training speed faster; at the same time, introducing the protein large language model to obtain the structural coding information of the peptide segment can improve the accuracy of predicting the spectrum.
[0084] Furthermore, the embodiments of the present invention use the data shown in Table 2 as sample data, and verify the accuracy of the predicted secondary spectrum of the peptide segment in the embodiments of the present invention by comparing the cosine similarity between the predicted secondary spectrum of the peptide segment by the present invention and Prosit and the experimental secondary spectrum.
[0085] Table 2 Source of sample data for the embodiments
[0086]
[0087] Specifically, it includes the following steps:
[0088] Step 1: Search all mass spectrometry data using MaxQuant to obtain the qualitative and quantitative file.
[0089] The specific operation is as follows: The search type is the standard DDA data mode; set carbamidomethylation as the fixed modification; oxidation and acetylation of the protein N-terminus as the variable modification; the digestion type is Trypsin / P, the maximum number of missed cleavage sites is 2, and the peptide mass tolerance is 20 ppm; both PSM and protein FDR are set to 0.01.
[0090] Step 2: Perform peptide spectrum matching on the original mass spectrometry data through the qualitative file, and extract the secondary spectrum of each peptide segment.
[0091] Step 3: Apply the method provided by the present invention and Prosit to all peptide segments respectively to obtain the predicted secondary spectra.
[0092] Step 4: Calculate the cosine similarity between the secondary spectra predicted by the present invention and Prosit and the experimental spectra respectively, and take the average of the cosine similarities of the predicted spectra of all peptide segments.
[0093] The cosine similarity between the predicted spectrogram and the experimental spectrogram obtained in the present invention is compared with that of Prosit, as shown in Table 3. In all datasets, the present invention obtains a higher cosine similarity, and the predicted similarity of the present invention in all datasets basically exceeds 95%, indicating that the method provided by the present invention has extremely high accuracy.
[0094] Table 3 Comparison of Examples between the Present Invention and Prosit
[0095]
[0096] For the method for predicting the secondary spectrogram of a peptide segment introduced with protein large language model encoding provided in the foregoing embodiments, the embodiments of the present invention further provide a device for predicting the secondary spectrogram of a peptide segment introduced with protein large language model encoding. Refer to Figure 7 the structural schematic diagram of a device for predicting the secondary spectrogram of a peptide segment introduced with protein large language model encoding shown in
[0097] An encoding module 701, configured to obtain a peptide segment of a protein and encode the peptide segment information of the peptide segment by using a preset encoding algorithm to obtain an encoded vector of the peptide segment; wherein, the encoding algorithm includes: encoding based on a protein large language model, index encoding, and one-hot encoding;
[0098] A prediction module 702, configured to obtain the fragment ion intensity of the peptide segment based on the encoded vector and a spectrogram prediction model;
[0099] A secondary spectrogram drawing module 703, configured to convert the fragment ion intensity into a secondary spectrogram of the peptide segment.
[0100] The device for predicting the secondary spectrogram of a peptide segment introduced with protein large language model encoding provided by the embodiments of the present invention encodes the peptide segment information of the peptide segment to obtain an encoded vector of the peptide segment by using encoding algorithms such as encoding based on a protein large language model, index encoding, and one-hot encoding, obtains the fragment ion intensity of the peptide segment according to the encoded vector and a spectrogram prediction model, and finally constructs a secondary spectrogram according to the fragment ion intensity of the peptide segment, thereby improving the accuracy of predicting the secondary spectrogram of the peptide segment.
[0101] In one embodiment, the peptide segment information at least includes: amino acid sequence, post-translational modification, charge, normalized collision energy, and instrument; specifically, the encoding module 701 is configured to: encode the amino acid sequence of the peptide segment according to a preset amino acid index to obtain a peptide segment sequence encoding vector; perform structural encoding on the amino acid sequence based on a protein large language model to obtain a peptide segment structural feature encoding vector; add the post-translational modification to the corresponding peptide segment sequence encoding vector, and perform dimensionality reduction on the peptide segment sequence encoding vector after adding the modification based on a linear neural network to obtain a peptide segment modification encoding vector; perform one-hot encoding on the charge, and replicate the encoding of the charge according to the sequence length of the peptide segment to obtain a charge encoding vector; perform one-hot encoding on the normalized collision energy after multiplying it by a preset energy coefficient, and replicate the encoding of the normalized collision energy according to the sequence length of the peptide segment to obtain a normalized fragmentation energy encoding vector; obtain the identifier corresponding to the instrument, perform one-hot encoding on the identifier, and replicate the encoding of the identifier according to the sequence length of the peptide segment to obtain an instrument encoding vector.
[0102] In one embodiment, specifically, the encoding module 701 is configured to: input the amino acid sequence into the pre-trained model of the protein large language model to obtain a structural feature matrix of the amino acid sequence; perform dimensionality reduction on the structural feature matrix based on a fully connected neural network to obtain a peptide segment structural feature encoding vector.
[0103] In one embodiment, specifically, the prediction module 702 is configured to: obtain the encoding vectors and splice them to obtain a feature vector of the peptide segment information; where the encoding vectors include: peptide segment sequence encoding vector, peptide segment structural feature encoding vector, peptide segment modification encoding vector, charge encoding vector, normalized fragmentation energy encoding vector, and instrument encoding vector; input the feature vector of the peptide segment information into a spectral prediction model to obtain the fragment ion intensity of the peptide segment.
[0104] In one embodiment, the spectral prediction model includes: an encoding layer and a decoding layer; specifically, the prediction module 702 is configured to: convert the feature vector of the peptide segment information into a feature matrix through the encoding layer; input the feature matrix into the decoding layer to obtain the fragment ion intensity of the peptide segment.
[0105] In one embodiment, specifically, the secondary spectrum drawing module 703 is configured to: normalize the fragment ion intensity; use the mass-to-charge ratio of the fragment ion as the abscissa and the normalized fragment ion intensity as the ordinate to construct the secondary spectrum of the peptide segment.
[0106] It should be noted that the device provided in the embodiments of the present invention has the same implementation principle and the same technical effects as those in the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding contents in the foregoing method embodiments. The specific values provided in the embodiments of the present invention are only exemplary and are not limited herein.
[0107] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and when the computer program is run by the processor, it executes the method described in any one of the above embodiments.
[0108] Figure 8 FIG. 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes: a processor 80, a memory 81, a bus 82, and a communication interface 83. The processor 80, the communication interface 83, and the memory 81 are connected through the bus 82; the processor 80 is configured to execute an executable module stored in the memory 81, such as a computer program.
[0109] Among them, the memory 81 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 83 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0110] The bus 82 may be an ISA bus, a PCI bus, an EISA bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 8 only a bidirectional arrow is used in FIG. 6, but it does not mean that there is only one bus or one type of bus.
[0111] Among them, the memory 81 is used to store a program. After receiving an execution instruction, the processor 80 executes the program. The method executed by the device defined by the flow process disclosed in any one of the above embodiments of the present invention can be applied to the processor 80 or implemented by the processor 80.
[0112] The processor 80 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 80 or the instructions in the form of software. The above-mentioned processor 80 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 81, and the processor 80 reads the information in the memory 81 and combines its hardware to complete the steps of the above method.
[0113] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the foregoing method embodiments, which will not be elaborated here.
[0114] If the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0115] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily conceive of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for predicting the secondary spectrum of peptide segments introduced by protein large language model encoding, characterized in that, Including: Obtain the peptide segments of the protein, and encode the peptide segment information of the peptide segments using a preset encoding algorithm to obtain the encoded vector of the peptide segments; wherein, the encoding algorithm includes: encoding based on a protein large language model, index encoding, and one-hot encoding; Based on the encoded vector and the spectrum prediction model, obtain the fragment ion intensity of the peptide segments; Convert the fragment ion intensity into the secondary spectrum of the peptide segments.
2. The method according to claim 1, wherein The peptide segment information at least includes: amino acid sequence, post-translational modification, charge, normalized collision energy, and instrument; Encoding the peptide segment information of the peptide segments using a preset encoding algorithm to obtain the encoded vector of the peptide segments, including: Encoding the amino acid sequence of the peptide segments according to a preset amino acid index to obtain a peptide sequence encoded vector; Performing structural encoding on the amino acid sequence based on a protein large language model to obtain a peptide structure feature encoded vector; Adding the post-translational modification to the corresponding peptide sequence encoded vector, and performing dimensionality reduction on the peptide sequence encoded vector with the added modification based on a linear neural network to obtain a peptide modification encoded vector; Performing one-hot encoding on the charge, and replicating the encoding of the charge according to the sequence length of the peptide segments to obtain a charge encoded vector; Performing one-hot encoding on the normalized collision energy after multiplying it by a preset energy coefficient, and replicating the encoding of the normalized collision energy according to the sequence length of the peptide segments to obtain a normalized fragmentation energy encoded vector; Obtain the identifier corresponding to the instrument, perform one-hot encoding on the identifier, and replicate the encoding of the identifier according to the sequence length of the peptide segments to obtain an instrument encoded vector.
3. The method according to claim 2, wherein Performing structural encoding on the amino acid sequence based on a protein large language model to obtain a peptide structure feature encoded vector, including: Inputting the amino acid sequence into the pre-trained model of the protein large language model to obtain the structural feature matrix of the amino acid sequence; Performing dimensionality reduction on the structural feature matrix based on a fully connected neural network to obtain a peptide structure feature encoded vector.
4. The method according to claim 2, wherein Based on the encoded vector and the spectrum prediction model, obtain the fragment ion intensity of the peptide segments, including: Concatenating the encoded vectors to obtain the feature vector of the peptide segment information; wherein, the encoded vectors include: the peptide sequence encoded vector, the peptide structure feature encoded vector, the peptide modification encoded vector, the charge encoded vector, the normalized fragmentation energy encoded vector, and the instrument encoded vector; Inputting the feature vector of the peptide segment information into the spectrum prediction model to obtain the fragment ion intensity of the peptide segments.
5. The method according to claim 4, characterized in that, The spectrum prediction model includes: an encoding layer and a decoding layer; Inputting the feature vector of the peptide segment information into the spectrum prediction model to obtain the fragment ion intensity of the peptide segments, including: Converting the feature vector of the peptide segment information into a feature matrix through the encoding layer; Inputting the feature matrix into the decoding layer to obtain the fragment ion intensity of the peptide segments.
6. The method according to claim 1, characterized in that, Converting the fragment ion intensity into the secondary spectrum of the peptide segments, including: Normalizing the fragment ion intensity; Construct the secondary spectrum of the peptide segment with the mass-to-charge ratio of the fragment ions as the abscissa and the normalized fragment ion intensity as the ordinate.
7. A peptide secondary spectrum prediction device incorporating protein large language model encoding, characterized in that, Including: An encoding module for obtaining the peptide segments of a protein and encoding the peptide segment information of the peptide segments using a preset encoding algorithm to obtain the encoding vectors of the peptide segments; wherein, the encoding algorithm includes: encoding based on a protein large language model, index encoding, and one-hot encoding; A prediction module for obtaining the fragment ion intensity of the peptide segment based on the encoding vector and a spectrum prediction model; A secondary spectrum drawing module for converting the fragment ion intensity into the secondary spectrum of the peptide segment.
8. The device according to claim 7, wherein The peptide segment information at least includes: amino acid sequence, post-translational modification, charge, normalized collision energy, and instrument; specifically, the encoding module is used for: Encoding the amino acid sequence of the peptide segment according to a preset amino acid index to obtain a peptide segment sequence encoding vector; Performing structural encoding on the amino acid sequence based on a protein large language model to obtain a peptide segment structure feature encoding vector; Adding the post-translational modification to the corresponding peptide segment sequence encoding vector and performing dimensionality reduction on the peptide segment sequence encoding vector with the modification added based on a linear neural network to obtain a peptide segment modification encoding vector; Performing one-hot encoding on the charge and replicating the encoding of the charge according to the sequence length of the peptide segment to obtain a charge encoding vector; Performing one-hot encoding on the normalized collision energy after multiplying it by a preset energy coefficient and replicating the encoding of the normalized collision energy according to the sequence length of the peptide segment to obtain a normalized fragmentation energy encoding vector; Obtaining the identifier corresponding to the instrument, performing one-hot encoding on the identifier, and replicating the encoding of the identifier according to the sequence length of the peptide segment to obtain an instrument encoding vector.
9. An electronic device, characterized in that, Including a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is run by the processor, it executes the steps of the method according to any one of claims 1 to 6 above.