Proteomics complex post-translational modification large-scale rapid identification method

CN121122385APending Publication Date: 2025-12-12BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511275221.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-12-12

Smart Images

  • Figure CN121122385A_ABST
    Figure CN121122385A_ABST
Patent Text Reader

Abstract

The invention provides a proteomics complex post-translational modification large-scale rapid identification method which comprises the following steps: preprocessing original protein data of a sample to obtain a primary signal peak cluster and a candidate secondary signal peak set corresponding to each primary signal peak; determining a first-level signal peak and a second-level signal peak pair based on the first-level signal peak cluster and the candidate second-level signal peak set; based on a pre-trained deep learning model and the peak pair of the first-level signal peak and the second-level signal peak, obtaining a matching score of the peak pair of the first-level signal peak and the second-level signal peak; and constructing a target pseudo secondary spectrogram based on the matching score, and performing open search on the target pseudo secondary spectrogram to obtain an identification result containing post-translation modification. According to the method, deconvolution can be carried out on the complex secondary spectrogram under methods including independent collection, the target pseudo secondary spectrogram which can be directly used for open type search is generated, and then accurate identification and deep analysis of accidental modification signals are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biotechnology and proteomics, and particularly relates to a method for rapid identification of complex post-translational modifications at a proteomics scale. BACKGROUND

[0002] The research on post-translational modifications (PTMs) of proteins is an important part of deep proteomics. PTMs add or remove specific chemical groups to proteins in a covalent manner, thereby specifically regulating the activity, folding, stability, cellular localization and interaction with other biological macromolecules of proteins in time and space. These modifications greatly enrich the functions and diversity of proteins and are an important mechanism for regulating cell functions and maintaining cell homeostasis. Abnormal regulation of PTMs is closely related to the occurrence and development of various human diseases, including various cancers such as breast cancer and lung cancer, cardiovascular diseases, metabolic diseases and other pathological processes.

[0003] Due to the high complexity of secondary spectra under the existing proteomics non-dependent acquisition mode, previous analysis mainly adopts methods based on spectral library search. However, due to the limitation of search space, the prediction accuracy of modified peptides and other reasons, these methods have certain limitations to some extent, which may eventually lead to a large number of secondary spectrum peaks being unable to be identified, and some peptide or protein signals in the biological sample being missed. This unexplored information is called "dark matter", including a large amount of post-translational modifications and the like. The "dark matter" may not be high in content in the sample, but has significant biological effects. Therefore, a method is needed to help explore the "dark matter" in the data. SUMMARY

[0004] Therefore, the present application aims to provide a method for rapid identification of complex post-translational modifications at a proteomics scale, which can deconvolve complex secondary spectra under methods including non-dependent acquisition, generate target pseudo-secondary spectra that can be directly used for open search, and further realize accurate identification and deep analysis of unexpected modification signals.

[0005] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows: In a first aspect, the present application provides a method for rapid identification of large-scale complex post-translational modifications in proteomics, comprising: preprocessing raw protein data of a sample to obtain primary signal peak clusters and a candidate secondary signal peak set corresponding to each primary signal peak; determining primary signal peak and secondary signal peak peak pairs based on the primary signal peak clusters and the candidate secondary signal peak set; obtaining a matching score of the primary signal peak and secondary signal peak peak pairs based on a pre-trained deep learning model and the primary signal peak and secondary signal peak peak pairs; constructing a target pseudo-secondary spectrum based on the matching score, and performing open search on the target pseudo-secondary spectrum to obtain an identification result containing post-translational modifications.

[0006] Optionally, the preprocessing of the raw protein data of the sample to obtain the primary signal peak clusters and the candidate secondary signal peak set corresponding to each primary signal peak comprises: analyzing the primary spectrum and the secondary spectrum in the raw protein data using a two-dimensional feature detection algorithm to extract the primary signal peak clusters and the secondary signal peak clusters, respectively; wherein the primary signal peak clusters are complete isotope cluster ion flows of peptide segments in the retention time dimension, and the secondary signal peak clusters are ion flows of single fragment ions; and filtering the secondary signal peak clusters based on a preset matching rule to obtain the candidate secondary signal peak set.

[0007] Optionally, the filtering of the secondary signal peak clusters based on the preset matching rule to obtain the candidate secondary signal peak set comprises: for each primary signal peak, traversing the secondary signal peak clusters within the corresponding fragmentation window to calculate the retention time difference between the secondary signal peak and the primary signal peak; and determining the secondary signal peaks with a retention time difference greater than a filtering threshold as the candidate secondary signal peak set corresponding to the primary signal peak; wherein the filtering threshold is a preset percentage of the retention time of the primary signal peak.

[0008] Optionally, the obtaining of the matching score of the primary signal peak and secondary signal peak peak pairs based on the pre-trained deep learning model and the primary signal peak and secondary signal peak peak pairs comprises: calculating a feature vector of the primary signal peak and secondary signal peak peak pairs; wherein the feature vector at least includes: primary signal single-isotope peak cluster ion flow intensity, secondary signal peak cluster ion flow intensity, and a calculated feature vector; the calculation of the calculated feature vector includes: primary and secondary signal retention time difference, primary and secondary signal overlap ratio, primary single-isotope peak cluster and secondary signal cosine similarity, primary first-isotope peak cluster and secondary signal cosine similarity, and primary second-isotope peak cluster and secondary signal cosine similarity; and inputting the feature vector of the primary signal peak and secondary signal peak peak pairs into the pre-trained deep learning model to obtain the matching score of the primary signal peak and secondary signal peak peak pairs.

[0009] Optionally, the deep learning model is constructed based on a Transformer deep learning architecture, adopts a fixed window size to process ion current intensity curves of the primary signal peak and the secondary signal peak, and processes the ion current intensity curves of the primary signal peak and the secondary signal peak as two independent input paths.

[0010] Optionally, the feature vector of the primary signal peak and the secondary signal peak pair is input into a pre-trained deep learning model to obtain a matching score of the primary signal peak and the secondary signal peak pair, including: inputting the primary signal single-isotope peak cluster ion current intensity and the secondary signal peak cluster ion current intensity into an encoder of a Transformer network respectively to obtain encoding results; inputting the encoding results and the calculated feature vector after splicing into a fully connected network, outputting a matching probability of the primary signal peak and the secondary signal peak pair through a Sigmoid activation function, and optimizing the matching probability using a cross-entropy loss function to obtain the matching score of the primary signal peak and the secondary signal peak pair; wherein the matching probability is in the interval [0, 1].

[0011] Optionally, a target pseudo-secondary spectrum is constructed based on the matching score, including: for each primary signal peak, selecting a preset number of secondary signal peaks based on the matching score to obtain the target pseudo-secondary spectrum.

[0012] In a second aspect, the present application provides a device for large-scale rapid identification of proteomics complex post-translational modifications, including: a preprocessing module for preprocessing original protein data of a sample to obtain a primary signal peak cluster and a candidate secondary signal peak set corresponding to each primary signal peak; a combination module for determining a primary signal peak and a secondary signal peak pair based on the primary signal peak cluster and the candidate secondary signal peak set; a matching module for obtaining a matching score of the primary signal peak and the secondary signal peak pair based on a pre-trained deep learning model and the primary signal peak and the secondary signal peak pair; a spectrum determination module for constructing a target pseudo-secondary spectrum based on the matching score and performing open search on the target pseudo-secondary spectrum to obtain an identification result containing a post-translational modification.

[0013] In a third aspect, the present application provides an electronic device including a processor and a memory, the memory storing computer executable instructions executable by the processor, and the processor executing the computer executable instructions to implement the steps of the method of any one of the first aspect.

[0014] In a fourth aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing a computer program, and the computer program being executed by a processor to perform the steps of the method of any one of the first aspect.

[0015] The present application has the following beneficial effects: The proteinomics complex post-translational modification large-scale rapid identification method provided by the present application comprises the following steps: firstly, the original protein data of a sample is preprocessed to obtain primary signal peak clusters and a candidate secondary signal peak set corresponding to each primary signal peak; then, the primary signal peak and the secondary signal peak peak pair are determined based on the primary signal peak cluster and the candidate secondary signal peak set; next, the matching score of the primary signal peak and the secondary signal peak peak pair is obtained based on the pre-trained deep learning model and the primary signal peak and the secondary signal peak peak pair; finally, the target pseudo secondary spectrum is constructed based on the matching score, and the open search is performed on the target pseudo secondary spectrum to obtain the identification result containing the post-translational modification. In the method, the original protein data is first processed to obtain the primary signal peak and the secondary signal peak peak pair, then the matching score of the primary signal peak and the secondary signal peak peak pair is matched based on the pre-trained deep learning model, and for each primary signal peak, the target pseudo secondary spectrum with higher specificity is constructed according to the matching score, the parent ion mass of the target pseudo secondary spectrum is more accurate, and the fragment ion peak correlation is stronger, so that the open search can be directly performed on the target pseudo secondary spectrum, and the accurate identification and deep analysis of unexpected modification signals are realized.

[0016] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and achieved by the structure particularly pointed out in the description, claims and drawings.

[0017] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are shown as follows. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings without creative labor based on these drawings.

[0019] Figure 1 A flowchart of a proteinomics complex post-translational modification large-scale rapid identification method provided by an embodiment of the present application; Figure 2 A flowchart of a deep learning algorithm provided by an embodiment of the present application; Figure 3 A pseudo secondary spectrum identification result diagram provided by an embodiment of the present application; Figure 4A structural schematic diagram of a proteomics complex post-translational modification large-scale rapid identification device provided by an embodiment of the present application is provided. Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application is provided. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in detail with reference to the drawings. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of the present application.

[0021] At present, due to the high complexity of secondary spectra under the existing proteomics non-dependent acquisition mode, the previous analysis mainly adopts methods such as spectrum library search. However, due to the limitation of search space, the prediction accuracy of modified peptides and other reasons, these methods have certain limitations to some extent, and finally a large number of secondary spectrum peaks may not be identified, and some peptide segments or protein signals in the biological sample are missed. The unexplored information is called "dark matter", including a large number of post-translational modifications and the like. These "dark matter" may not be high in content in the sample, but have significant biological effects. Therefore, a method is needed to help explore the "dark matter" in the data.

[0022] Based on this, the proteomics complex post-translational modification large-scale rapid identification method provided by the embodiment of the present application can deconvolute the complex secondary spectrum under the method including non-dependent acquisition, generate a target pseudo-secondary spectrum that can be directly used for open search, and then realize accurate identification and deep analysis of unexpected modification signals.

[0023] In order to facilitate the understanding of the present embodiment, first, a proteomics complex post-translational modification large-scale rapid identification method disclosed by the present embodiment will be described in detail. The method can be executed by an electronic device, such as a computer, a tablet computer and the like. Referring to the flowchart of a proteomics complex post-translational modification large-scale rapid identification method shown in Figure 1 The method mainly includes the following steps S101 to S104: Step S101: The original protein data of the sample is preprocessed to obtain primary signal peak clusters and a candidate secondary signal peak set corresponding to each primary signal peak.

[0024] In an embodiment, the sample to be analyzed is analyzed by a mass spectrometer to obtain original protein data, i.e., an original RAW file. The original RAW file contains first-order spectrum and second-order spectrum information. Due to differences in the setting of mass-to-charge ratio range and isolation window size, the entire original RAW file can contain dozens to hundreds of different fragmentation windows, wherein each second-order spectrum corresponds to a specific fragmentation window, and the system stores the second-order spectrum information in a classified manner according to the same fragmentation center.

[0025] In an embodiment of the present application, the original protein data can be preprocessed to extract the first-order signal peak cluster and the second-order signal peak cluster, and for each first-order signal peak, the second-order signal peak cluster is coarsely scored in the complete chromatographic interval to be matched to screen out a set containing multiple candidate second-order signal peaks to be matched.

[0026] Step S102: determining a first-order signal peak and a second-order signal peak peak pair based on the first-order signal peak cluster and the candidate second-order signal peak set.

[0027] In an embodiment, for each first-order signal peak, each second-order signal peak in the candidate second-order signal peak set is traversed to form a first-order signal peak and a second-order signal peak peak pair.

[0028] Step S103: obtaining a matching score of the first-order signal peak and the second-order signal peak peak pair based on a pre-trained deep learning model and the first-order signal peak and the second-order signal peak peak pair.

[0029] In an embodiment, the first-order signal peak and the second-order signal peak peak pair are input into a pre-trained deep learning model network to score the matching between the first-order signal peak and the second-order signal peak, and the scoring result (i.e., the matching score) 0xc0+1xc1 is output. Wherein, the element c0 represents the probability of matching failure, and the element c1 represents the probability of matching success.

[0030] Step S104: constructing a target pseudo second-order spectrum based on the matching score, and performing an open search on the target pseudo second-order spectrum to obtain an identification result containing a post-translational modification.

[0031] In an embodiment, for each first-order signal peak, a preset number of second-order signal peaks are selected based on the matching score to obtain a target pseudo second-order spectrum. Specifically, the matching scores are sorted, and for each first-order signal peak, the second-order signal peaks with a matching score greater than 0.5 are selected to form a pseudo second-order spectrum (i.e., a target pseudo second-order spectrum), which is cleaner, the corresponding parent ion mass is more accurate, and the fragment ion peak correlation is stronger. Further, the target pseudo second-order spectrum can be directly used for open search to discover and explore the "dark matter" in the mass spectrometry data.

[0032] The proteinomics complex post-translational modification large-scale rapid identification method provided by the embodiment of the application first processes original protein data to obtain a primary signal peak and a secondary signal peak peak pair, then matches and scores the primary signal peak and the secondary signal peak peak pair based on a pre-trained deep learning model, and for each primary signal peak, a target pseudo-secondary spectrum with higher specificity is constructed according to the matching score, the parent ion mass of the target pseudo-secondary spectrum is more accurate, and the fragment ion peak correlation is stronger, and then the target pseudo-secondary spectrum can be directly used for open search, so that accurate identification and in-depth analysis of unexpected modification signals are realized.

[0033] In an implementation, for the foregoing step S101, that is, when the original protein data of the sample is preprocessed to obtain a primary signal peak cluster and a candidate secondary signal peak set corresponding to each primary signal peak, the following methods can be used, but are not limited to the following methods: First, a two-dimensional feature detection algorithm is used to analyze the primary spectrum and the secondary spectrum in the original protein data, and a primary signal peak cluster and a secondary signal peak cluster are extracted, respectively; wherein the primary signal peak cluster is an isotopic cluster ion flow of a peptide segment in the retention time dimension, and the secondary signal peak cluster is an ion flow of a single fragment ion.

[0034] In specific implementation, a two-dimensional feature detection algorithm is used to analyze the primary spectrum and the secondary spectrum, and a primary signal peak cluster and a secondary signal peak cluster are extracted, respectively. The primary signal peak cluster is defined as an isotopic cluster ion flow of a peptide segment in the retention time dimension, and through accurate determination of signal boundaries and verification of continuity, a reliable candidate primary signal peak cluster set is finally screened out. The secondary signal peak cluster corresponds to the extraction ion flow of a single fragment ion, and a fixed-width sliding window is used to scan and analyze the secondary spectrum under the same fragmentation center along the retention time dimension, and through strict control of key parameters such as spectrum continuity, mass deviation (determined based on mass-to-charge ratio), and peak cluster shape, a high-quality secondary signal peak cluster set is finally determined.

[0035] Then, the secondary signal peak cluster is screened based on a preset matching rule to obtain a candidate secondary signal peak set.

[0036] In specific implementation, for each primary signal peak, the secondary signal peak cluster is traversed in the corresponding fragmentation window, the retention time difference between the secondary signal peak and the primary signal peak is calculated, and the secondary signal peak with a retention time difference greater than a screening threshold is determined as the candidate secondary signal peak set corresponding to the primary signal peak; wherein the screening threshold is a preset percentage of the retention time of the primary signal peak.

[0037] Specifically, for each primary signal peak, all secondary signal peaks in the corresponding fragmentation window are preliminarily screened, the set of secondary signal peaks to be matched is traversed, the retention time difference between the secondary signal peak and the primary signal peak is calculated, and the retention time length of the primary signal peak is 20% as a screening threshold. If the retention time difference is greater than the screening threshold, the secondary signal peak meets the screening condition, and the set of candidate secondary signal peaks to be matched is formed. The coarse matching strategy adopted in the embodiment of the present application effectively reduces the calculation range of subsequent accurate matching, improves the overall data processing efficiency, and finally obtains each primary signal peak and the corresponding set of candidate secondary signal peaks.

[0038] In an embodiment, for the foregoing step S103, i.e., when the matching score of the primary signal peak and the secondary signal peak peak pair is obtained based on the pre-trained deep learning model and the primary signal peak and the secondary signal peak peak pair, the following methods can be used, but are not limited to: First, the feature vector of the primary signal peak and the secondary signal peak peak pair is calculated.

[0039] The feature vector at least includes: primary signal single-isotope peak cluster ion flow intensity, secondary signal peak cluster ion flow intensity, and calculation feature vector; the calculation feature vector includes: primary and secondary signal retention time difference, primary and secondary signal overlap ratio, primary single-isotope peak cluster and secondary signal cosine similarity, primary first-isotope peak cluster and secondary signal cosine similarity, and primary second-isotope peak cluster and secondary signal cosine similarity.

[0040] Specifically, the primary signal single-isotope peak cluster ion flow intensity and the secondary signal peak cluster ion flow intensity are normalized by the maximum value of the extracted original intensity, and the calculation formula is as follows:

[0041] Wherein, is the primary signal single-isotope peak cluster ion flow intensity or the secondary signal peak cluster ion flow intensity, is the original intensity of the primary signal single-isotope peak cluster ion flow or the secondary signal peak cluster ion flow, is the maximum value of the original intensity of the primary signal single-isotope peak cluster ion flow or the secondary signal peak cluster ion flow.

[0042] The primary and secondary signal retention time difference represents the difference in retention time between the primary signal ion flow and the secondary signal ion flow, and the formula is as follows:

[0043] Wherein, is the primary and secondary signal retention time difference, is the retention time of the highest point of the primary signal ion flow intensity, the retention time of the highest point of the intensity of the secondary signal ion current, the length of the retention time of the primary signal ion current.

[0044] The primary and secondary signal overlap ratio represents the overlap ratio of the parent ion XIC and the fragment ion XIC in the retention time, and the formula is as follows:

[0045] wherein, is the primary and secondary signal overlap ratio, is the intensity of the i th parent ion XIC, is the intensity of the i th fragment ion XIC.

[0046] The primary monoisotope peak cluster and secondary signal cosine similarity, the primary first isotope peak cluster and secondary signal cosine similarity, and the primary second isotope peak cluster and secondary signal cosine similarity, and the formula is as follows:

[0047] wherein, is the cosine similarity, is the ion current intensity curve of the primary signal, is the ion current intensity curve of the secondary signal.

[0048] Then, the feature vectors of the primary signal peak and the secondary signal peak peak pair are input into the pre-trained deep learning model to obtain the matching score of the primary signal peak and the secondary signal peak peak pair.

[0049] In an embodiment, the deep learning model is used to calculate the matching score between the primary signal and the secondary signal according to the feature vectors thereof. The model is constructed based on the Transformer deep learning architecture, adopts a fixed window size of 32x1 dimension to process the ion current intensity curves of the primary signal peak and the secondary signal peak, and processes the ion current intensity curves of the primary signal peak and the secondary signal peak as two independent input paths. Among them, the ion current intensity curve of the primary signal is represented as , and the ion current intensity curve of the secondary signal is represented as , and the remaining 5 hand-calculation feature vectors are .

[0050] In specific implementation, when the eigenvectors of the primary signal peak and the secondary signal peak peak pair are input into the pre-trained deep learning model to obtain the matching score of the primary signal peak and the secondary signal peak peak pair, the following modes can be adopted, but are not limited to: first, the primary signal single isotope peak cluster ion flow intensity and the secondary signal peak cluster ion flow intensity are input into the encoder of the Transformer network respectively to obtain the encoding result; then, the encoding result and the calculated eigenvector are spliced and input into the full connection network, the matching probability of the primary signal peak and the secondary signal peak peak pair is output through the Sigmoid activation function, and the matching probability is optimized using the cross entropy loss function to obtain the matching score of the primary signal peak and the secondary signal peak peak pair; wherein the matching probability is in the interval [0, 1].

[0051] Specifically, referring to FIG. 6, Figure 2 First, the ion flow intensity curves of the primary signal and the secondary signal are embedded to map them to a 16-dimensional feature space. Specifically, a linear transformation is used to achieve this, and the embedding operation for the primary signal ion flow intensity curve can be represented as:

[0052] The same operation is performed on the ion flow intensity curve of the secondary signal. The embedded primary signal features and secondary signal features (i.e., the primary signal single isotope peak cluster ion flow intensity and the secondary signal peak cluster ion flow intensity) are input into the encoder part of two Transformer networks with the same structure for processing. The Transformer encoder is stacked with multiple identical layers, each containing a Multi-Head Self-Attention (MHSA) and a Feed-Forward Network (FFN) sublayer. Here, the processing process of the Transformer encoder is illustrated using the primary signal features as an example.

[0053]

[0054] Similarly, the same operation is performed on the secondary signal features After processing by the two Transformer encoder layers, the Transformer- encoded features of the primary signal and the secondary signal are obtained and , both with a dimension of 32x16. The output results of the Transformer and are processed by the full connection network respectively to obtain two 16-dimensional representation vectors and :

[0055] Computing the feature vector After processing through the fully connected network, a 16-dimensional feature vector is obtained The output results of the parent ion and fragment ion path And After flattening and connecting with After processing through the fully connected network, the final feature representation is obtained Y Finally, the matching probability of the two two-dimensional images is output through the Sigmoid activation function, the interval is [0, 1], and the cross entropy loss function is used for optimization, 0 represents the least matching, and 1 represents the most matching.

[0056] Through the selection of the above feature vector, the matching between the primary signal peak and the secondary signal peak can be accurately and objectively reflected, and then the deep learning network constructed by a large amount of data is trained, so that a fine scoring model capable of accurately evaluating the matching performance can be obtained.

[0057] In the embodiment of the application, the public data set is used as the source of the model training set and the test set. The selection indicators of the training set are as follows: mainly according to the DIA-NN report result, the peptide segment signal identified under the condition that the false discovery rate is less than 0.01 is selected. The primary signal peak cluster and the secondary signal peak cluster corresponding to the peptide segment signal are extracted as positive examples, and the primary signal peak cluster and the secondary signal peak cluster belonging to other peptide segments with similar retention time are combined to form negative examples. 30000 positive examples and 30000 negative examples are randomly selected, 75% of which are used for model training, and the remaining 25% are used for testing.

[0058] Further, in the embodiment of the application, the complex secondary spectrum is convolved to generate a pseudo-secondary spectrum with more accurate parent ion mass and more specific fragment ions. The subsequent can be directly used for open search to discover and explore the "dark matter" in the mass spectrum data. In order to verify this conclusion, 20 peptide segments are designed and synthesized in the embodiment of the application to simulate the actual situation of unexpected modification (10 containing 1 modification and 10 containing 2 modifications), which are added to the 293T cell lysate as true positive targets (see Table 1). The 20 peptide segments involve 13 specific mass modifications in total. After non-dependent acquisition of standard samples using a new mass spectrometer Orbitrap Astra, three repeated data are obtained.

[0059] Table 1 Name of synthesized peptide segment

[0060] The same processing is performed on the three repeated data, the primary and secondary signal peak clusters are extracted, the primary and secondary signal matching scores are obtained based on the deep learning model, and finally the more "clean" pseudo-secondary spectrum is generated.

[0061] Table 2 Identification results of synthetic peptide segments

[0062] Using pFind3 to search the exported pseudo-secondary spectrum, it is found that these modified peptide segments can be recalled and the fragment ion matching conditions of these peptide segments in three repeated data are basically consistent and the effect is good, which shows that the method provided in the embodiment of the application can stably export pseudo-secondary spectrum and has good robustness. There is one peptide segment that is not recalled in repeat 1, which is No. 13 peptide segment. This peptide segment successfully constructs a pseudo-secondary spectrum, but the spectrum quality is not high (see Figure 3

[0063] As shown in Table 2, one of the non-dependent acquisition search library software with the highest download rate such as DIA-NN is used to test the synthetic modified peptide segments. If no variable modification is set, DIA-NN cannot report 20 peptide segments, so it is necessary to set specific modifications in DIA-NN. A total of 14 different parameter settings are used, each parameter set contains 1-2 specific modifications contained in the tag peptide in order to maximize the sensitivity of DIA-NN. In this case, DIA-NN cannot correctly recall all the modified peptide segments. When the false discovery rate (FDR) of the parent ion is set to 5%, DIA-NN can only recall 7, and none of the peptide segments containing two modifications can be recalled. Even if the FDR is set to 100%, DIA-NN can only provide a search result with an FDR of 50%, and in this case, 10 peptide segments are reported (see Table 2).

[0064] By comparison with the classic tool DIA-Umpire for constructing pseudo-secondary spectrum, the pseudo-secondary spectrum exported by the embodiment of the application contains more useful information such as modified peptide segments. Here, pFind3 is also used to search the pseudo-secondary spectrum exported by DIA-Umpire, which will miss some modified peptide segments. The missed No. 12 and No. 16 peptide segments, DIA-Umpire did not export their pseudo-secondary spectrum in three repeats, so they cannot be identified in three repeats. The missed No. 08, No. 11, No. 13, No. 19 and No. 20 peptide segments did not export pseudo-secondary spectrum or export pseudo-secondary spectrum with low quality in some repeats, so they cannot be accurately identified.

[0065] In summary, the existing software DIA-NN, DIA-Umpire, etc. does not perform well when analyzing mass spectrum data. The method provided in the embodiment of the application adds a Transformer deep learning model when matching primary and secondary signals, finds the high-dimensional correlation between primary and secondary signals, exports pseudo-secondary spectrum very stably, can effectively retain modification information, and has high reliability and sensitivity. ​

[0066] The method provided by the embodiment of the present application firstly acquires the primary and secondary signal peak clusters in the original data, encodes the primary and secondary signal peak clusters, and obtains high-dimensional representation vectors of the primary and secondary signals; then, based on the signal intensity vector and the signal matching model, matching scores of the primary and secondary signals are obtained; and finally, for each candidate primary signal, a pseudo-secondary spectrum with higher specificity can be constructed according to the matching scores, the parent ion mass of the pseudo-secondary spectrum is more accurate, and the fragment ion peak correlation is stronger. The present application can perform deconvolution on complex secondary spectra under mass spectrum acquisition methods including non-dependent acquisition, generate pseudo-secondary spectra which can be directly used for open search, and finally realize accurate identification and deep analysis of unexpected modification signals.

[0067] For the proteinomics complex post-translational modification large-scale rapid identification method provided by the foregoing embodiment, the embodiment of the present application further provides a proteinomics complex post-translational modification large-scale rapid identification device, which refers to a proteinomics complex post-translational modification large-scale rapid identification device as shown in Figure 4 The proteinomics complex post-translational modification large-scale rapid identification device mainly includes the following parts: The preprocessing module 401 is configured to preprocess the original protein data of the sample to obtain the primary signal peak cluster and the candidate secondary signal peak set corresponding to each primary signal peak.

[0068] The combination module 402 is configured to determine the primary signal peak and secondary signal peak peak pair based on the primary signal peak cluster and the candidate secondary signal peak set.

[0069] The matching module 403 is configured to obtain the matching score of the primary signal peak and secondary signal peak peak pair based on the pre-trained deep learning model and the primary signal peak and secondary signal peak peak pair.

[0070] The spectrum determination module 404 is configured to construct a target pseudo-secondary spectrum based on the matching score, and perform open search on the target pseudo-secondary spectrum to obtain an identification result containing a post-translational modification.

[0071] The proteinomics complex post-translational modification large-scale rapid identification device provided by the foregoing embodiment of the present application firstly processes the original protein data to obtain the primary signal peak and secondary signal peak peak pair, then performs matching scoring on the primary signal peak and secondary signal peak peak pair based on the pre-trained deep learning model, and for each primary signal peak, constructs a target pseudo-secondary spectrum with higher specificity according to the matching score, the parent ion mass of the target pseudo-secondary spectrum is more accurate, and the fragment ion peak correlation is stronger, so that the target pseudo-secondary spectrum can be directly subjected to open search, and accurate identification and deep analysis of unexpected modification signals are realized.

[0072] In an implementation, the preprocessing module 401 is specifically configured to: analyze the primary spectrum and the secondary spectrum in the original protein data by using a two-dimensional feature detection algorithm, and extract primary signal peak clusters and secondary signal peak clusters respectively; the primary signal peak cluster is an isotopic cluster ion flow of a peptide segment in a retention time dimension, and the secondary signal peak cluster is an ion flow of a single fragment ion; and filter the secondary signal peak clusters based on a preset matching rule to obtain a candidate secondary signal peak set.

[0073] In an implementation, the preprocessing module 401 is specifically configured to: for each primary signal peak, traverse the secondary signal peak clusters in a corresponding fragmentation window, and calculate a retention time difference between the secondary signal peak and the primary signal peak; and determine the secondary signal peak with a retention time difference greater than a filtering threshold as a candidate secondary signal peak set corresponding to the primary signal peak; the filtering threshold is a preset percentage of the retention time of the primary signal peak.

[0074] In an implementation, the matching module 403 is specifically configured to: calculate a feature vector of a primary signal peak and a secondary signal peak peak pair; the feature vector at least includes: a primary signal single-isotope peak cluster ion flow intensity, a secondary signal peak cluster ion flow intensity, and a calculation feature vector; the calculation feature vector includes: a primary and secondary signal retention time difference, a primary and secondary signal overlap ratio, a primary single-isotope peak cluster and secondary signal cosine similarity, a primary first-isotope peak cluster and secondary signal cosine similarity, and a primary second-isotope peak cluster and secondary signal cosine similarity; and input the feature vector of the primary signal peak and the secondary signal peak peak pair into a pre-trained deep learning model to obtain a matching score of the primary signal peak and the secondary signal peak peak pair.

[0075] In an implementation, the deep learning model is constructed based on a Transformer deep learning architecture, processes ion flow intensity curves of the primary signal peak and the secondary signal peak by using a fixed window size, and processes the ion flow intensity curves of the primary signal peak and the secondary signal peak as two independent input paths.

[0076] In an implementation, the matching module 403 is specifically configured to: input the primary signal single-isotope peak cluster ion flow intensity and the secondary signal peak cluster ion flow intensity into an encoder of a Transformer network respectively to obtain an encoding result; input the encoding result and the calculation feature vector after splicing into a fully connected network, output a matching probability of the primary signal peak and the secondary signal peak peak pair by using a Sigmoid activation function, and optimize the matching probability by using a cross-entropy loss function to obtain a matching score of the primary signal peak and the secondary signal peak peak pair; the matching probability is in an interval of [0, 1].

[0077] In an implementation, the spectrum determination module 404 is specifically configured to: for each primary signal peak, select a preset number of secondary signal peaks based on the matching scores to obtain a target pseudo-secondary spectrum.

[0078] It should be noted that the device provided by the embodiments of the present application has the same implementation principle and technical effects as the above-mentioned method embodiments. For brevity, the part of the device embodiment not mentioned in the above-mentioned method embodiments can be referred to the corresponding content in the above-mentioned method embodiments. The specific numerical values provided in the embodiments of the present application are only exemplary and are not limited herein.

[0079] The embodiments of the present application also provide an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program performs the method according to any one of the above embodiments when executed by the processor.

[0080] Figure 5 The structure schematic diagram of the electronic device provided by the embodiments of the present application is shown in the figure, and the electronic device 100 includes a processor 50, a memory 51, a bus 52 and a communication interface 53, the processor 50, the communication interface 53 and the memory 51 are connected through the bus 52; the processor 50 is used to execute the executable modules stored in the memory 51, such as computer programs.

[0081] The memory 51 may contain a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 53 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0082] The bus 52 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0083] The memory 51 is used to store programs, and the processor 50 executes the programs after receiving execution instructions. The method executed by the device defined by the flow process disclosed in any one of the above embodiments of the present application can be applied to the processor 50 or realized by the processor 50.

[0084] The processor 50 can be an integrated circuit chip with signal processing capability. In implementation, each step of the above method can be completed by integrated logic circuit of hardware in the processor 50 or by instructions in the form of software. The processor 50 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 51, and the processor 50 reads the information in the memory 51, and combines the hardware to complete the steps of the above method.

[0085] The computer program product of the readable storage medium provided by the embodiments of the present application comprises a computer readable storage medium storing program codes, and the program codes comprise instructions for executing the method described in the foregoing method embodiments. For specific implementation, reference can be made to the foregoing method embodiments, which will not be described here.

[0086] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0087] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacements to some of the technical features. The modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A rapid, large-scale identification method for complex post-translational modifications in proteomics, characterized in that, include: The raw protein data of the sample is preprocessed to obtain a primary signal peak cluster and a set of candidate secondary signal peaks corresponding to each primary signal peak; Based on the primary signal peak cluster and the candidate secondary signal peak set, the primary signal peak and the secondary signal peak pair are determined; Based on the pre-trained deep learning model and the peak pairs of first-level and second-level signal peaks, the matching scores of the peak pairs of first-level and second-level signal peaks are obtained; Based on the matching score, a target pseudo-secondary spectrogram is constructed, and an open search is performed on the pseudo-secondary spectrogram to obtain an identification result that includes post-translational modifications.

2. The method according to claim 1, characterized in that, The raw protein data of the sample is preprocessed to obtain a primary signal peak cluster and a set of candidate secondary signal peaks corresponding to each primary signal peak, including: A two-dimensional feature detection algorithm was used to analyze the primary and secondary spectra of the raw protein data, and primary and secondary signal peak clusters were extracted respectively. The primary signal peak clusters are the ion currents of complete isotopic clusters of peptides in the retention time dimension, and the secondary signal peak clusters are the ion currents of single fragment ions. The secondary signal peak clusters are filtered based on preset matching rules to obtain a set of candidate secondary signal peaks.

3. The method according to claim 2, characterized in that, The secondary signal peak cluster is filtered based on preset matching rules to obtain a candidate set of secondary signal peaks, including: For each primary signal peak, the secondary signal peak cluster is traversed within the corresponding fragmentation window, and the retention time difference between the secondary signal peak and the primary signal peak is calculated. Secondary signal peaks whose retention time difference is greater than a screening threshold are identified as candidate secondary signal peak sets corresponding to the primary signal peak; wherein, the screening threshold is a preset percentage of the retention time of the primary signal peak.

4. The method according to claim 1, characterized in that, Based on a pre-trained deep learning model and peak pairs of primary and secondary signal peaks, matching scores for the peak pairs of primary and secondary signal peaks are obtained, including: Calculate the feature vectors of the primary signal peak and the secondary signal peak pair; wherein, the feature vectors include at least: the ion current intensity of the primary signal single isotope peak cluster, the ion current intensity of the secondary signal peak cluster, and the calculated feature vectors; the calculated feature vectors include: the retention time difference between the primary and secondary signals, the overlap ratio between the primary and secondary signals, the cosine similarity between the primary single isotope peak cluster and the secondary signal, the cosine similarity between the primary first isotope peak cluster and the secondary signal, and the cosine similarity between the primary second isotope peak cluster and the secondary signal; The feature vectors of the first-level signal peak and the second-level signal peak pair are input into a pre-trained deep learning model to obtain the matching score of the first-level signal peak and the second-level signal peak pair.

5. The method according to claim 4, characterized in that, The deep learning model is built on the Transformer deep learning architecture. It uses a fixed window size to process the ion current intensity curves of the first-level signal peak and the second-level signal peak, and processes the ion current intensity curves of the first-level signal peak and the second-level signal peak as two independent input paths.

6. The method according to claim 5, characterized in that, The feature vectors of the first-level signal peak and the second-level signal peak pair are input into a pre-trained deep learning model to obtain the matching score of the first-level signal peak and the second-level signal peak pair, including: The intensity of the ion current of the first-level signal single isotope peak cluster and the intensity of the ion current of the second-level signal peak cluster are respectively input into the encoder of the Transformer network to obtain the encoding result; The encoding result and the calculated feature vector are concatenated and then input into a fully connected network. The matching probability of the first-level signal peak and the second-level signal peak is output through the Sigmoid activation function. The matching probability is optimized using the cross-entropy loss function to obtain the matching score of the first-level signal peak and the second-level signal peak; wherein the matching probability is in the interval [0,1].

7. The method according to claim 1, characterized in that, Constructing a target pseudo-secondary spectrogram based on the matching score includes: For each primary signal peak, a preset number of secondary signal peaks are selected based on the matching score to obtain the target pseudo-secondary spectrum.

8. A device for large-scale rapid identification of complex post-translational modifications in proteomics, characterized in that, include: The preprocessing module is used to preprocess the raw protein data of the sample to obtain the primary signal peak cluster and the set of candidate secondary signal peaks corresponding to each primary signal peak; The combination module is used to determine the peak pairs of primary signal peaks and secondary signal peaks based on the primary signal peak cluster and the candidate secondary signal peak set; The matching module is used to obtain the matching score of the first-level signal peak and the second-level signal peak pair based on the pre-trained deep learning model and the first-level signal peak and the second-level signal peak pair; The spectrum determination module is used to construct a target pseudo-secondary spectrum based on the matching score, and to perform an open search on the target pseudo-secondary spectrum to obtain protein identification results containing post-translational modifications.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method described in any one of claims 1 to 7.