Deep mass spectrometry prediction method, system, device and medium based on data fine-tuning
By constructing a semi-supervised learning task, using the difference between the target peptide and the bait peptide to fine-tune the deep mass spectrometry model, the generalization problem of the deep mass spectrometry model in experimental mass spectrometry data is solved, and the number of peptide identification and recognition capabilities are improved.
Patent Information
- Application Number
- CN202310739711.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-06-20
AI Technical Summary
The existing deep mass spectrometry model has generalization problems when facing experimentally specific mass spectrometry data or specific enzymatic peptides, resulting in a degradation in predictive performance.
Using a data fine-tuning method, by constructing a semi-supervised learning task, using the difference between the target peptide and the bait peptide, semi-supervised fine-tuning of the deep mass spectrometry model is performed, and the model is dynamically adjusted to adapt to the new mass spectrometry experimental data.
The identification ability and identification number of peptides have been improved, and the generalization ability of deep mass spectrometry models in mass spectrometry identification is improved.
Smart Images

Figure CN116741280B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of proteomics, and in particular to a method, system, device and medium for deep mass spectrometry prediction based on data fine-tuning. Background Art
[0002] Protein identification is crucial in proteomics analysis. Identifying a large number of proteins with high accuracy from biological samples is crucial for drug discovery and other applications. In bottom-up proteomics analysis, liquid chromatography-tandem mass spectrometry (LC-MS / MS) is typically used to analyze peptides derived from proteins. The matching process between mass spectra and peptides is crucial for peptide identification and protein inference. The standard approach is database searching. Mass spectrometry database search software, such as Sequest, MaxQuant, and MASCOT, matches peptide sequences to MS / MS spectra based on the similarity between experimental and theoretical mass spectra. To avoid incorrect peptide spectrum matches (PSMs), a threshold is applied to filter out PSMs. A decoy database is used to estimate the distribution of incorrect peptide-to-MS matches, and the threshold is adjusted to the user's desired false discovery rate (FDR). Currently, most database search software do not consider fragment ion intensities. However, incorporating fragment ion intensity information can enhance the separation between true and random matches and improve the accuracy of peptide-to-MS match identification. Commonly used mass spectrometry database search software models theoretical ion intensities from a probabilistic model perspective, but lacks a good screening function for false positives in identification (i.e., incorrectly screening out non-matching peptides into the final results). However, due to the lack of better mass spectrometry intensity modeling methods, traditional database search software still adopts a simple modeling approach.
[0003] In recent years, numerous efforts have been made both domestically and internationally to predict fragment ion intensities, particularly using deep learning for MS / MS mass spectrometry prediction. These efforts have significantly improved the accuracy of mass spectrometry ion intensity prediction, thereby enhancing peptide identification capabilities. Specifically, deep mass spectrometry generation models based on deep learning often employ recurrent neural networks (RNNs) as their underlying architecture, formalizing the task of modeling mass spectrometry ion intensities as a sequence-to-sequence learning task, thereby generating the corresponding MS / MS mass spectrometry ion intensities from the peptide's amino acid sequence. All of these efforts employ a train-prediction paradigm, using paired peptides and mass spectra from standard datasets (such as the ProteomeTools project) to train the prediction model. By screening existing publicly available mass spectrometry data (for example, the Andromeda score for peptide-to-mass spectra matches must be above 100), supervised data pairs for training deep mass spectrometry models are obtained: peptide amino acid sequences and their ion intensity sequences. Once trained, the deep mass spectrometry model is used to generate theoretical spectra from experimental data without considering the specificity of the sample data. Although this method significantly increases the number of reliably identified tryptic and non-tryptic peptides in data-dependent acquisition data, it still suffers from generalization issues due to the lack of training data when faced with experiment-specific mass spectrometry data or peptides digested by specific enzymes, resulting in reduced prediction performance. Summary of the Invention
[0004] The purpose of the present invention is to provide a deep mass spectrometry prediction method, system, equipment and medium based on data fine-tuning, which can effectively alleviate the generalization problem of mass spectrometry prediction, ensure the identification ability of peptide segments, and expand the practical application prospects of deep mass spectrometry models in mass spectrometry identification.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A deep mass spectrometry prediction method based on data fine-tuning, comprising:
[0007] The experimental mass spectrometry peptide data were preprocessed. Based on the differences between target peptides and decoy peptides, multiple sets of semi-supervised data were constructed. Each set of semi-supervised data contained the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra. Each set of semi-supervised data was divided into non-overlapping semi-supervised training data and semi-supervised prediction data.
[0008] The semi-supervised training data in each set of semi-supervised data is used to perform semi-supervised fine-tuning on the deep mass spectrometry model to obtain multiple fine-tuned deep mass spectrometry models;
[0009] The semi-supervised prediction data in each set of semi-supervised data is input into the corresponding fine-tuned deep mass spectrometry model, and the predicted mass spectra of the peptides in the corresponding semi-supervised prediction data output by all fine-tuned deep mass spectrometry models are obtained as the integrated prediction results.
[0010] A deep mass spectrometry prediction system based on data fine-tuning, comprising:
[0011] The data preprocessing unit is used to preprocess the experimental mass spectrometry peptide data. Based on the differences between target peptides and decoy peptides, multiple sets of semi-supervised data are constructed. Each set of semi-supervised data contains the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra. Each set of semi-supervised data is divided into non-overlapping semi-supervised training data and semi-supervised prediction data.
[0012] A semi-supervised fine-tuning unit is used to perform semi-supervised fine-tuning on the deep mass spectrometry model using the semi-supervised training data in each set of semi-supervised data to obtain multiple fine-tuned deep mass spectrometry models;
[0013] The prediction result integration unit is used to input the semi-supervised prediction data in each set of semi-supervised data into the corresponding fine-tuned deep mass spectrometry model, and obtain the predicted mass spectra of the peptides in the corresponding semi-supervised prediction data output by all fine-tuned deep mass spectrometry models as the integrated prediction results.
[0014] A processing device comprising: one or more processors; a memory for storing one or more programs;
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0016] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.
[0017] It can be seen from the technical solution provided by the present invention that a data-based fine-tuning solution is introduced, and the prior knowledge in mass spectrometry identification: the difference between target peptides and bait peptides is used to construct a semi-supervised learning task. On new mass spectrometry experimental data, the deep mass spectrometry model can be dynamically adjusted to generalize it to the current mass spectrometry data, thereby bridging the gap between the deep mass spectrometry model and the current experimental mass spectrometry data, thereby ensuring the recognition ability of peptides and increasing the number of peptide identifications. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A flowchart of a deep mass spectrometry prediction method based on data fine-tuning provided in an embodiment of the present invention;
[0020] Figure 2 A schematic diagram of a deep mass spectrometry prediction method based on data fine-tuning provided by an embodiment of the present invention;
[0021] Figure 3 Schematic diagram of peptide identification results before and after fine-tuning under various enzymatic hydrolysis conditions provided in an embodiment of the present invention;
[0022] Figure 4 A schematic diagram comparing the different methods for discovering specific variant antibodies provided in the embodiments of the present invention;
[0023] Figure 5 A schematic diagram of a deep mass spectrometry prediction system based on data fine-tuning provided by an embodiment of the present invention;
[0024] Figure 6 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] First, the following terms may be used in this article:
[0027] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.
[0028] The following describes in detail a method, system, device, and medium for deep mass spectrometry prediction based on data fine-tuning provided by the present invention. Any information not described in detail in the examples herein is prior art known to those skilled in the art. For any conditions not specified in the examples herein, the conditions were based on those conventional in the art or manufacturer's recommendations.
[0029] Example 1
[0030] The embodiment of the present invention provides a deep mass spectrometry prediction method based on data fine-tuning, which expands the existing Train-Prediction paradigm into the Train-Fine-tune-Predict (TFP) paradigm. Figure 1 As shown, it mainly includes the following steps:
[0031] Step 1: Preprocess the experimental mass spectrometry peptide data and construct multiple sets of semi-supervised data based on the differences between target peptides and decoy peptides.
[0032] In an embodiment of the present invention, each set of semi-supervised data includes the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra (the matching results of all peptides and mass spectra). Each set of semi-supervised data is divided into semi-supervised training data and semi-supervised prediction data according to a set ratio, and the semi-supervised training data and the semi-supervised prediction data do not overlap with each other. At the same time, the collection of semi-supervised prediction data in different sets of semi-supervised data is the matching results of all peptides and mass spectra, and the semi-supervised prediction data in different sets of semi-supervised data do not overlap with each other.
[0033] The preferred implementation of this step is as follows:
[0034] (1) Perform peptide matching on the experimental mass spectrum and mark the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra.
[0035] (2) multiple groups of semi-supervised data are set, each group of semi-supervised data contains the matching results of all target peptides and mass spectra, and the matching results of all decoy peptides and mass spectra; in each group of semi-supervised data, the matching results of peptides and mass spectra are randomly and evenly divided into two parts for cross-verification (K-fold), wherein the matching results of peptides and mass spectra include: the matching results of target peptides and mass spectra, and the matching results of all decoy peptides and mass spectra; one part of the matching results of peptides and mass spectra constitutes a group of semi-supervised prediction data for subsequent prediction of the corresponding fine-tuning model, and the other part constitutes semi-supervised training data for semi-supervised fine-tuning of the subsequent model; in each group of semi-supervised data, the semi-supervised training data and the semi-supervised prediction data do not overlap with each other to avoid data leakage in machine learning, and at the same time, the semi-supervised prediction data of different groups do not overlap with each other, and the semi-supervised prediction data of all groups constitute the full set of matching results of peptides and mass spectra.
[0036] In the embodiment of the present invention, the set ratio can be set by the user based on actual conditions or experience, and the present invention does not limit the numerical value of the ratio. For example, the set ratio can be 1:N-1, that is, the semi-supervised prediction data accounts for 1 part and the semi-supervised training data accounts for N-1 parts, where N is the number of divided data parts, which is an integer greater than or equal to 2.
[0037] (3) Each set of semi-supervised data is vectorized separately; wherein, the peptide segments are one-hot encoded to obtain corresponding encoding vectors, and the peptide segments include: target peptide segments and bait peptide segments; and the intensity information of the mass spectrum is extracted to form the corresponding vector.
[0038] Step 2: Use the semi-supervised training data in each set of semi-supervised data to perform semi-supervised fine-tuning on the deep mass spectrometry model to obtain multiple fine-tuned deep mass spectrometry models.
[0039] In an embodiment of the present invention, the semi-supervised training data in each set of semi-supervised data is used to iteratively semi-supervised fine-tune the deep mass spectrometry model to obtain a fine-tuned deep mass spectrometry model, and finally the same number of fine-tuned deep mass spectrometry models as the number of semi-supervised data groups are obtained.
[0040] In an embodiment of the present invention, a preferred implementation method of semi-supervised fine-tuning of a deep mass spectrometry model iteratively using semi-supervised training data in a set of semi-supervised data is as follows: the mass spectrum in the matching result of the peptide segment and the experimental mass spectrum is used as the actual mass spectrum, the mass spectrum of each peptide segment is predicted using the deep mass spectrometry model, and the similarity score between the predicted mass spectrum and the corresponding actual mass spectrum is calculated, and a number of matching results of peptide segments and mass spectra are screened out from the semi-supervised prediction data according to the similarity score, and pseudo-label values are respectively assigned to the matching results of all screened peptide segments and mass spectra, and the deep mass spectrometry model is optimized using the pseudo-label values; wherein the peptide segments include: target peptide segments and bait peptide segments.
[0041] In an embodiment of the present invention, the method of predicting the mass spectrum of each peptide segment using a deep mass spectrometry model and calculating the similarity score between the predicted mass spectrum and the corresponding actual mass spectrum includes: for each peptide segment, calculating the distance between the predicted mass spectrum and the corresponding actual mass spectrum, and determining the similarity score using the calculated distance.
[0042] In an embodiment of the present invention, screening out a plurality of peptide-mass spectrum matching results from the semi-supervised training data based on the similarity score includes: estimating a false discovery rate based on the similarity score, and selecting a plurality of peptide-mass spectrum matching results whose false discovery rate estimation result exceeds a set threshold.
[0043] In an embodiment of the present invention, the matching results of all screened peptides and mass spectra are respectively assigned pseudo label values, and the deep mass spectrometry model is optimized using the pseudo label values, including: assigning a first pseudo label value to the matching result of the screened target peptide and mass spectrum, and assigning a second pseudo label value to the matching result of the screened bait peptide and mass spectrum; optimizing the deep mass spectrometry model using the first pseudo label value and the second pseudo label value, so that the similarity score of the predicted mass spectrum of the screened target peptide and the corresponding actual mass spectrum of the deep mass spectrometry model is close to the first pseudo label value, and the similarity score of the predicted mass spectrum of the screened bait peptide and the corresponding actual mass spectrum of the deep mass spectrometry model is close to the second pseudo label value.
[0044] Step 3: Input the semi-supervised prediction data in each set of semi-supervised data into the corresponding fine-tuned deep mass spectrometry model, and obtain the predicted mass spectra of the peptides in the corresponding semi-supervised prediction data output by all fine-tuned deep mass spectrometry models as the integrated prediction results.
[0045] The above-mentioned solution provided by the embodiment of the present invention can be applied to existing protein identification software platforms to improve the reliability of deep mass spectrometry prediction and increase the number of protein identifications. It can also be used in medical applications, for example, for the identification of specific antibodies. Of course, the present invention does not limit the specific application directions.
[0046] The above-mentioned solution provided by the embodiment of the present invention introduces a data-based fine-tuning solution and uses prior knowledge in mass spectrometry identification: the difference between target peptides and decoy peptides, to construct a semi-supervised learning task. On new mass spectrometry experimental data, the existing deep mass spectrometry model can be dynamically adjusted to better generalize it to the current mass spectrometry data, thereby bridging the gap between the deep mass spectrometry model and the current experimental mass spectrometry data, thereby ensuring the recognition ability of peptides and increasing the number of peptide identifications.
[0047] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail below with reference to specific embodiments.
[0048] 1. Mass spectrometry data preprocessing.
[0049] 1. Use a general mass spectrometry database search software (such as MaxQuant) to perform peptide segment matching (hereinafter referred to as PSM) on the original mass spectrum, and mark the matching results of the target peptide and the bait peptide, and mark the matching results of the target peptide and the mass spectrum, as well as the matching results of the bait peptide and the mass spectrum.
[0050] 2. Set up multiple groups of semi-supervised data, each group of semi-supervised data contains the matching results of all target peptides and mass spectra, and the matching results of all bait peptides and mass spectra. In order to prevent the data leakage of the matching results of peptides and mass spectra, the present invention performs random data division on each group of semi-supervised data. In the current practice, the matching results of peptides and mass spectra are randomly and evenly divided into two non-overlapping parts, where the matching results of peptides and mass spectra include: the matching results of all target peptides and mass spectra, and the matching results of all bait peptides and mass spectra. The two non-overlapping parts constitute semi-supervised training data and semi-supervised prediction data, wherein the semi-supervised training data can constitute a data set (semi-supervised training data) for fine-tuning a deep mass spectrometry model in the next step. The matching results of any peptide and mass spectra are not shared between the semi-supervised prediction data of different groups of semi-supervised data. As an example, two groups of semi-supervised data can be set, but the number of groups set can be extended to any number.
[0051] 3. Vectorize the matching results between peptides and mass spectra in each set of semi-supervised data. For peptides, one-hot encoding is used to encode the amino acid sequence. For mass spectra, the corresponding intensity information is extracted from the corresponding peptide according to its possible fragmentation mass-to-charge ratio (m / z) to form a vector, considering the b / y ions in the fragmentation, as well as the 1-3 valence fragmentation charges; where m is the mass, z is the charge, and b and y are adjectives used to modify the ions. When the peptide fragment is fragmented, different situations will occur: when the charge appears at the beginning of the peptide sequence, it is a b ion, otherwise it is a y ion. The matching results between peptides and mass spectra here refer to the matching results between the target peptides and mass spectra contained in the semi-supervised data of the group to which they belong, as well as the matching results between the decoy peptides and mass spectra.
[0052] The preprocessing part constructs multiple sets of semi-supervised data, performs random data division, and then processes them into vectorized data that can be used in neural networks.
[0053] 2. Semi-supervised fine-tuning.
[0054] The above preprocessing generates multiple sets of semi-supervised data. The semi-supervised training data in each set is used to perform semi-supervised fine-tuning on the deep mass spectrometry model, resulting in the corresponding fine-tuned deep mass spectrometry model. Since the semi-supervised fine-tuning process is the same, the following example uses the semi-supervised training data from one set of semi-supervised data as an example.
[0055] In the embodiment of the present invention, there is no restriction on the deep mass spectrometry model, and any existing deep mass spectrometry model can be used, which is updated through pseudo-label iteration in semi-supervised learning. In the i-th round of iteration, the similarity score is calculated based on the predicted mass spectrum output by the current deep mass spectrometry prediction model and the corresponding mass spectrum in the semi-supervised training data. Subsequently, the false discovery rate (FDR) is estimated based on the similarity score calculation, and the matching results of high-quality peptides and mass spectra are screened out. Among the high-quality peptide matching results obtained by screening, the matching results of the target peptide and the mass spectrum are assigned a pseudo-label of 1, and the matching results of the bait peptide and the mass spectrum are marked as -1. Subsequently, the model is fine-tuned based on the matching results of the screened peptides and the mass spectrum.
[0056] 1. Similarity score calculation.
[0057] The mass spectrometry model outputs a peptide mass spectrum prediction, which can be used to calculate similarity between the predicted mass spectrum and the actual mass spectrum (i.e., the corresponding mass spectrum in the semi-supervised data). This similarity score can be formalized as the distance between mass spectrum vectors.
[0058] For example, two possible similarity indicators can be considered: Pearson Correlation Coefficient (PCC) and Spectral Angle (SA). The calculation methods are:
[0059]
[0060]
[0061] Among them, p is the vector of predicted mass spectrum, is the vector of actual mass spectrum, represents the mean regularization of the vector, |.|2 represents the L2 regularization of the vector, and T is the transpose symbol.
[0062] 2. False discovery rate estimation and screening of high-quality peptide and mass spectrometry matching results.
[0063] After calculating the similarity score, the false discovery rate estimation indicator commonly used in mass spectrometry identification is used to select high-quality peptide and mass spectrum matching results from the semi-supervised training data, including the matching results of the target peptide and the mass spectrum and the matching results of the bait peptide and the mass spectrum.
[0064] For example, 1% FDR can be selected as the screening threshold to screen out the matching results between peptide segments and mass spectra with an FDR lower than 1%.
[0065] 3. Model fine-tuning.
[0066] The matching result between the target peptide and the mass spectrum after screening is assigned a first pseudo-label value (for example, a value of 1), and the matching result between the decoy peptide and the mass spectrum is assigned a second pseudo-label value (for example, a value of -1). In the i-th iteration, by optimizing the similarity index between the predicted mass spectrum of the target peptide and the actual mass spectrum, the similarity index between the predicted mass spectrum of the decoy peptide and the actual mass spectrum is closer to 1, while the similarity index between the predicted mass spectrum of the decoy peptide and the actual mass spectrum is closer to -1, so that the predicted mass spectrum of the target peptide is more similar to the experimental mass spectrum, while the predicted mass spectrum of the decoy peptide is more different from the experimental mass spectrum.
[0067] For example, the minimum mean square error can be used as the objective function, and batch data gradient descent can be used as the optimization algorithm for the deep mass spectrometry model.
[0068] 3. Fine-tune model integration.
[0069] After obtaining multiple fine-tuned deep mass spectrometry models using multiple sets of semi-supervised data, cross-validation is performed to ensure that the fine-tuned deep mass spectrometry models perform well on the semi-supervised prediction data relative to the semi-supervised training data. Each set of semi-supervised prediction data is then integrated to obtain the predicted mass spectra of the peptides in all the semi-supervised prediction data. Furthermore, similarity calculations can be performed with the actual mass spectra of the peptides in the semi-supervised prediction data to evaluate the performance of the fine-tuned deep mass spectrometry models.
[0070] Figure 2 An example of the present invention is shown, in which two groups of semi-supervised data are divided and trained to obtain two fine-tuned deep mass spectrometry models. However, it should be noted that the present invention does not limit the specific number of fine-tuned models, and users can expand to 3 or more fine-tuned models based on actual conditions or experience.
[0071] The above-mentioned solution provided by the embodiments of the present invention establishes a paradigm for addressing the generalization problem of deep mass spectrometry models. Although the introduction of deep learning into mass spectrometry identification has greatly improved its performance, deep mass spectrometry models still face generalization issues in their predictions when faced with diverse experimental mass spectrometry data across laboratories, species, and mass spectrometers. From a paradigm perspective, this method proposes a semi-supervised fine-tuning step, which can effectively alleviate the generalization problem of mass spectrometry prediction and expand the practical application prospects of deep mass spectrometry models in mass spectrometry identification.
[0072] In order to verify the conclusion, the present invention uses a variety of mass spectrometry data. The present invention was verified on the Bekker et al. dataset with a variety of lytic enzyme mass spectrometry data, in which the deep mass spectrometry model used Prosit2019. Its training data is mostly concentrated in Trypsin hydrolysis. Under other enzymatic hydrolysis conditions (such as Chymo protease), there is a generalization problem. The present invention can effectively increase the number of peptide identifications, such as Figure 3 As shown, Figure 3 The titles of the four parts all refer to different proteases used to cleave proteins. The peptide sequences produced by different proteases are quite different. The variable of protease is used to demonstrate how the method of the present invention can alleviate the fact that the deep mass spectrometry model has poor generalization.
[0073] Furthermore, the present invention can also be applied to the discovery of specific immune antibodies. By introducing a fine-tuning step, almost twice as many specific variant antibodies can be discovered, such as Figure 4 As shown, Figure 4 Prosit in the above is the aforementioned model Prosit2019, and Fine-tuned Prosit is the model after semi-supervised fine-tuning using the solution of the present invention.
[0074] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.
[0075] Example 2
[0076] The present invention also provides a deep mass spectrum prediction system based on data fine-tuning, which is mainly used to implement the method provided in the above embodiment, such as Figure 5 As shown, the system mainly includes:
[0077] The data preprocessing unit is used to preprocess the experimental mass spectrometry peptide data. Based on the differences between target peptides and decoy peptides, multiple sets of semi-supervised data are constructed. Each set of semi-supervised data contains the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra. Each set of semi-supervised data is divided into non-overlapping semi-supervised training data and semi-supervised prediction data.
[0078] A semi-supervised fine-tuning unit is used to perform semi-supervised fine-tuning on the deep mass spectrometry model using the semi-supervised training data in each set of semi-supervised data to obtain multiple fine-tuned deep mass spectrometry models;
[0079] The prediction result integration unit is used to input the semi-supervised prediction data in each set of semi-supervised data into the corresponding fine-tuned deep mass spectrometry model, and obtain the predicted mass spectra of the peptides in the corresponding semi-supervised prediction data output by all fine-tuned deep mass spectrometry models as the integrated prediction results.
[0080] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0081] Example 3
[0082] The present invention also provides a processing device, such as Figure 6 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.
[0083] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0084] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:
[0085] The input device can be a touch screen, image acquisition device, physical button or mouse;
[0086] The output device may be a display terminal;
[0087] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0088] Example 4
[0089] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0090] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0091] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A deep mass spectrometry prediction method based on data fine-tuning, characterized in that: include: The experimental mass spectrometry peptide data were preprocessed. Based on the differences between target peptides and decoy peptides, multiple sets of semi-supervised data were constructed. Each set of semi-supervised data contained the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra. Each set of semi-supervised data was divided into non-overlapping semi-supervised training data and semi-supervised prediction data. The semi-supervised training data in each set of semi-supervised data is used to perform semi-supervised fine-tuning on the deep mass spectrometry model to obtain multiple fine-tuned deep mass spectrometry models; The semi-supervised prediction data in each set of semi-supervised data is input into the corresponding fine-tuned deep mass spectrometry model, and the predicted mass spectra of the peptides in the corresponding semi-supervised prediction data output by all fine-tuned deep mass spectrometry models are obtained as the integrated prediction results; The method of using the semi-supervised training data in each group of semi-supervised data to semi-supervise fine-tune the deep mass spectrometry model separately to obtain multiple fine-tuned deep mass spectrometry models includes: using the semi-supervised training data in a group of semi-supervised data to iteratively semi-supervised fine-tune the deep mass spectrometry model separately to obtain a fine-tuned deep mass spectrometry model, and finally obtaining the same number of fine-tuned deep mass spectrometry models as the number of semi-supervised data groups; wherein, the method of using the semi-supervised training data in a group of semi-supervised data to iteratively fine-tune the deep mass spectrometry model separately includes: using the mass spectrum in the matching result of the peptide segment and the mass spectrum as the actual mass spectrum, using the deep mass spectrometry model to predict the mass spectrum of each peptide segment, and calculating the similarity score between the predicted mass spectrum and the corresponding actual mass spectrum, screening out a number of peptide segment and mass spectrum matching results from the semi-supervised training data according to the similarity score, assigning pseudo-label values to all the screened peptide segment and mass spectrum matching results, and optimizing the deep mass spectrometry model using the pseudo-label values; wherein the peptide segment includes: target peptide segment and decoy peptide segment; The method of assigning pseudo label values to the matching results of all the screened peptides and mass spectra, and optimizing the deep mass spectrometry model using the pseudo label values includes: assigning a first pseudo label value to the matching result of the screened target peptides and mass spectra, and assigning a second pseudo label value to the matching result of the screened bait peptides and mass spectra; optimizing the deep mass spectrometry model using the first pseudo label value and the second pseudo label value, so that the similarity score of the predicted mass spectrum of the screened target peptides and the corresponding actual mass spectrum of the deep mass spectrometry model is close to the first pseudo label value, and the similarity score of the predicted mass spectrum of the screened bait peptides and the corresponding actual mass spectrum of the deep mass spectrometry model is close to the second pseudo label value.
2. The method for deep mass spectrometry prediction based on data fine-tuning according to claim 1, characterized in that: The experimental mass spectrometry peptide data is preprocessed to construct multiple sets of semi-supervised data based on the differences between target peptides and decoy peptides, including: Perform peptide matching on the experimental mass spectrum and mark the matching results of all target peptides and mass spectra, as well as all matching results of all decoy peptides and mass spectra; Multiple groups of semi-supervised data are set. In each group of semi-supervised data, the matching results of all peptides and mass spectra are randomly and evenly divided into two parts, one part constitutes a group of semi-supervised prediction data, and the other part constitutes semi-supervised training data. In each group of semi-supervised data, the semi-supervised training data and the semi-supervised prediction data do not overlap with each other; the collection of semi-supervised prediction data in different groups of semi-supervised data is the matching results of all peptides and mass spectra, and the semi-supervised prediction data in different groups of semi-supervised data do not overlap with each other; the matching results of all peptides and mass spectra include the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra.
3. A deep mass spectrometry prediction method based on data fine-tuning according to claim 1 or 2, characterized in that: The process of preprocessing experimental mass spectrometry peptide data also includes: vectorizing each set of semi-supervised data separately; performing one-hot encoding on the peptides to obtain corresponding encoding vectors, wherein the peptides include: target peptides and bait peptides; and extracting intensity information from the mass spectrum to form a corresponding vector.
4. The method for deep mass spectrometry prediction based on data fine-tuning according to claim 1, characterized in that: The method of predicting the mass spectrum of each peptide segment using the deep mass spectrometry model and calculating the similarity score between the predicted mass spectrum and the corresponding actual mass spectrum includes: for each peptide segment, calculating the distance between the predicted mass spectrum and the corresponding actual mass spectrum, and determining the similarity score using the calculated distance.
5. The method for deep mass spectrometry prediction based on data fine-tuning according to claim 1, characterized in that: The step of screening out a plurality of peptide segment and mass spectrum matching results from the semi-supervised training data according to the similarity score includes estimating a false discovery rate according to the similarity score and selecting a plurality of peptide segment and mass spectrum matching results whose false discovery rate estimation results exceed a set threshold.
6. A deep mass spectrometry prediction system based on data fine-tuning, characterized in that: The method for implementing any one of claims 1 to 5 comprises: The data preprocessing unit is used to preprocess the experimental mass spectrometry peptide data. Based on the differences between target peptides and decoy peptides, multiple sets of semi-supervised data are constructed. Each set of semi-supervised data contains the matching results of all target peptides and mass spectra, as well as the matching results of all decoy peptides and mass spectra. Each set of semi-supervised data is divided into non-overlapping semi-supervised training data and semi-supervised prediction data. A semi-supervised fine-tuning unit is used to perform semi-supervised fine-tuning on the deep mass spectrometry model using the semi-supervised training data in each set of semi-supervised data to obtain multiple fine-tuned deep mass spectrometry models; The prediction result integration unit is used to input the semi-supervised prediction data in each set of semi-supervised data into the corresponding fine-tuned deep mass spectrometry model, and obtain the predicted mass spectra of the peptides in the corresponding semi-supervised prediction data output by all fine-tuned deep mass spectrometry models as the integrated prediction results.
7. A processing device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
8. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.