Mass spectrum spectrum analysis method and device and intelligent terminal

By using a spectral analysis model based on the GoogleNet network structure, screening mass-to-charge ratio intensity values ​​and performing feature compression, the problems of high computational cost and excessively long analysis time in proteomics mass spectrometry are solved, achieving rapid and accurate mass spectrometry spectrum resolution, and improving detection efficiency and clinical application potential.

CN115206422BActive Publication Date: 2026-01-06CHI BIOTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210784151.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2026-01-06
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

Existing proteomics mass spectrometry technology is computationally expensive and takes too long to analyze, which limits its efficiency in scientific research and clinical applications.

Method used

A spectral analysis model based on the GoogleNet network structure is adopted. By screening the mass-to-charge ratio intensity value, spectral features are determined, and feature compression is performed using convolutional layers, mean-sampling layers, and fully connected layers to quickly and accurately resolve protein peptide sequences.

Benefits of technology

It enables rapid and accurate mass spectrometry spectrum interpretation, reduces computational costs, shortens analysis time, and improves detection efficiency, which is beneficial for scientific research output and clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115206422B_ABST
    Figure CN115206422B_ABST
Patent Text Reader

Abstract

The application relates to the field of protein deconvolution analysis, in particular to a mass spectrum analysis method and device and an intelligent terminal. The mass spectrum analysis method comprises the following steps: obtaining a target spectrum, wherein the target spectrum is used for reflecting mass spectrum data of a protein peptide segment; determining a spectrum feature of the target spectrum based on a mass-to-charge ratio of the target spectrum; inputting the spectrum feature of the target spectrum into a preset deconvolution model to obtain an analysis category of the target spectrum, wherein the analysis category is used for reflecting a peptide sequence corresponding to the target spectrum. The mass-to-charge ratio of the target spectrum is used to obtain the spectrum feature, and the spectrum feature is input into the deconvolution model, and the analysis category is obtained through the deconvolution model. The application has the characteristics that the peptide sequence of the protein corresponding to the target spectrum can be effectively, accurately and quickly analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of protein spectral analysis, and in particular to a method, apparatus and intelligent terminal for mass spectrometry spectrum interpretation. Background Technology

[0002] The human genome contains all the information about human development and evolution, revealing the truth about genes encoding proteins. In the post-genomic era, biological research has gradually shifted from genomics to proteomics. Proteins are involved in almost every aspect of cellular function, and protein characterization has become an important part of modern biology, inspiring a new discipline: proteomics. Mass spectrometry is currently the primary technique for protein research, possessing characteristics such as high precision and speed. Advances in mass spectrometry have provided unprecedented speed, sensitivity, and accuracy for protein identification, offering a significant advantage for the research and development of proteomics.

[0003] Currently, most widely used proteomics mass spectrometry techniques utilize computer software algorithms for mass spectrum analysis, but these techniques generally suffer from high computational costs and excessively long analysis times. Summary of the Invention

[0004] The purpose of this application is to provide a method for mass spectrometry spectrum analysis that features short analysis time.

[0005] The above-mentioned objective of this application is achieved through the following technical solution:

[0006] A method for interpreting mass spectrometry spectra, comprising:

[0007] Obtain a target spectrum, wherein the target spectrum is used to reflect the mass spectrometry data of a protein peptide;

[0008] Based on the mass-to-charge ratio of the target spectrum, the spectral characteristics of the target spectrum are determined;

[0009] The spectral features of the target spectrum are input into a preset spectral analysis model to obtain the analysis category of the target spectrum, wherein the analysis category is used to reflect the peptide sequence corresponding to the target spectrum.

[0010] By adopting the above technical solution, the spectral features are obtained based on the mass-to-charge ratio of the target spectrum, and the spectral features are input into the spectral analysis model. The analytical classification is obtained through the spectral analysis model, thereby effectively, accurately and quickly analyzing the peptide sequence of the protein corresponding to the target spectrum.

[0011] Optionally, determining the spectral features of the target spectrum based on the mass-to-charge ratio of the target spectrum includes:

[0012] Based on the intensity value of the mass-to-charge ratio, all the mass-to-charge ratios of the target spectrum are filtered to obtain a number of characteristic mass-to-charge ratios;

[0013] Based on all the mass-to-charge ratios of the target spectrum, the spectral features of the target spectrum are obtained.

[0014] By adopting the above technical solution, the target spectra corresponding to different protein peptide segments have significant differences in mass-to-charge ratio. By using a specified number of mass-to-charge ratios to obtain spectral features, the spectral features are more representative of the target spectra, thereby improving the accuracy of detection and analysis.

[0015] Optionally, the output of the spectral decomposition model includes a feature matrix obtained by compressing the spectral features, and the analytical category corresponding to the feature matrix; the length of the feature matrix is ​​7.

[0016] By adopting the above technical solution, the feature loss after spectral feature compression is reduced, and the utilization rate of protein spectra is improved.

[0017] Optionally, the spectral decomposition model includes four layers for compressing the spectral features to obtain the average value downsampling layer of the feature matrix;

[0018] The process of filtering all mass-to-charge ratios of the target spectrum based on the intensity value of the mass-to-charge ratio yields a number of characteristic mass-to-charge ratios, including:

[0019] Based on the intensity value of the mass-to-charge ratio, all the mass-to-charge ratios of the target spectrum are filtered to obtain 112 characteristic mass-to-charge ratios.

[0020] By adopting the above technical solutions, the spectral interpretation model is more widely applicable to the target spectrum, thereby improving the utilization rate of mass spectrometry data.

[0021] Optionally, the spectral resolution model is trained using the following method:

[0022] Obtain a model training dataset, wherein the model training dataset includes a training spectrogram, spectral features corresponding to the training spectrogram, and peptide labels corresponding to the training spectrogram;

[0023] The spectral features of the training spectra are input into the spectral decomposition model to obtain the training category of the training spectra;

[0024] The spectral decomposition model is trained based on the peptide labels and training categories of the training spectra.

[0025] By adopting the above technical solution, the pre-labeled peptide tags are used as the real results, and the training categories output by the spectral analysis model are used as the actual results. The spectral analysis model is trained using peptide tags and training categories, making the analysis results of the spectral analysis model more accurate and closer to the real values.

[0026] Optionally, obtaining the model training dataset includes:

[0027] Obtain a raw mass spectrometry dataset and a resolution dataset, wherein the raw mass spectrometry dataset contains raw spectra, and the resolution dataset contains resolution results corresponding to the raw spectra;

[0028] The original spectra in the original mass spectrometry dataset are filtered to obtain filtered spectra;

[0029] Remove the filtered spectra from the original mass spectrometry dataset to obtain the mass spectrometry training dataset;

[0030] Obtain the spectral features of the mass spectrometry training dataset;

[0031] Based on the spectral dataset, all the original spectra in the mass spectrometry training dataset are labeled to obtain peptide labels for the mass spectrometry training dataset.

[0032] By adopting the above technical solution, the training data of the spectral interpretation model is screened, and some data with less effective information and insufficient credibility are eliminated in advance, thereby improving the performance of the trained model and making the analysis results of the trained spectral interpretation model more accurate.

[0033] Optionally, the step of filtering the original spectra in the original mass spectrometry dataset to obtain filtered spectra includes:

[0034] Based on the spectral decomposition results of the original spectra, the original spectra whose spectral decomposition results are reversed are selected to obtain the filtered spectra;

[0035] And / or, based on the spectral interpretation results of the original spectra, filter the original spectra whose spectral interpretation results are empty to obtain filtered spectra;

[0036] And / or, based on the spectral interpretation results of the original spectra, filter the original spectra whose scores are less than a score threshold to obtain filtered spectra;

[0037] And / or, based on the number of mass-to-charge ratios in the original spectrum, filter the original images whose mass-to-charge ratios are less than a feature threshold to obtain a filtered spectrum, wherein the feature threshold is less than or equal to the number of mass-to-charge ratios contained in the spectral features;

[0038] And / or, based on the parent ion valence of the original spectra, the original spectra in the original mass spectrometry dataset are grouped to obtain valence classification groups;

[0039] The valence classification groups whose number of spectra is less than the threshold number of valence spectra are used to obtain filtered spectra;

[0040] And / or, based on the peptide sequences in the resolution results, the original spectra in the original mass spectrometry dataset are grouped to obtain peptide classification groups;

[0041] The peptide classification groups with fewer than the number of peptide maps are selected to obtain filtered maps.

[0042] By employing the above technical solutions, original spectra with negative database results, empty spectra, or scores below the scoring threshold are considered to have low reliability. These original spectra are excluded as filter spectra, which optimizes model training. If the mass-to-charge ratio (MCR) count of an original spectrum is less than the feature threshold, it cannot extract or filter out a sufficient number of MCRs as spectral features for subsequent calculations; therefore, it needs to be excluded as a filter spectrum. If the number of original spectra in the valence classification group is less than the valence spectrum count threshold, it indicates insufficient data in that group; similarly, if the number of original spectra in the peptide classification group is less than the peptide spectrum count threshold, it indicates insufficient data in that group. When the data volume in either the valence or peptide classification group is insufficient, the model may struggle to converge during training. Therefore, filtering this data is necessary to ensure rapid model convergence.

[0043] The second main objective of this invention is to provide a mass spectrometry spectrum analysis device.

[0044] A mass spectrometry spectrum analysis device, comprising:

[0045] A spectrum acquisition module is used to acquire a target spectrum, wherein the target spectrum corresponds to a peptide segment of a protein;

[0046] The feature extraction module is used to determine the spectral features of the target spectrum based on the mass-to-charge ratio of the target spectrum;

[0047] The model parsing module is used to input the spectral features of the target spectrum into the preset spectral interpretation model to obtain the parsing category of the target spectrum, wherein the parsing category is used to reflect the peptide sequence corresponding to the target spectrum.

[0048] The third main objective of this invention is to propose a smart terminal.

[0049] A smart terminal includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as a mass spectrometry spectrum analysis method as described above.

[0050] The fourth main objective of this invention is to provide a computer-readable storage medium.

[0051] A computer-readable storage medium storing a computer program capable of being loaded by a processor and executed as in any of the above-described technical solutions for mass spectrometry spectrum analysis. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the mass spectrometry spectrum analysis method of this application.

[0053] Figure 2 This is a schematic diagram of step S2 of the mass spectrometry spectrum analysis method of this application.

[0054] Figure 3 This is a schematic diagram of the spectral resolution model of this application.

[0055] Figure 4 This is a schematic diagram illustrating how the average value downsampling layer of this application compresses spectral features.

[0056] Figure 5 This is a flowchart illustrating the training method of the spectral analysis model in the mass spectrometry analysis method of this application.

[0057] Figure 6 This is a schematic diagram of step E1 of the training method for the spectral interpretation model in the mass spectrometry spectrum analysis method of this application.

[0058] Figure 7 This is a schematic diagram of the optimization process of the spectral resolution model in this application.

[0059] Figure 8 This is a schematic diagram of the selection process for the training spectra in this application.

[0060] Figure 9 This is a schematic diagram of the mass spectrometry spectrum analysis device of this application.

[0061] In the diagram, 1 is the spectrum acquisition module; 2 is the feature extraction module; and 3 is the model parsing module. Detailed Implementation

[0062] Currently, proteomics mass spectrometry technology is mainly used for analysis through specialized software and deep learning algorithms.

[0063] Among specialized software, pFind, Mascot, and Maxquant are widely used for protein spectrometry interpretation. However, they generally suffer from high computational costs and excessively long analysis times. For example, pFind, using a 24-core server, requires an average of about 8 hours to interpret 10GB of spectra. Furthermore, if the protein peptide sequences to be analyzed are complex, it will take even longer. This is equivalent to a single protein mass spectrometry experiment taking at least 8 hours to produce results, highlighting the high computational cost and excessively long analysis time associated with mass spectrometry analysis technology.

[0064] In research on identifying protein peptide sequences using deep learning, some techniques combine CNN and LSTM algorithms to process spectral data. This approach uses CNN to extract image features from the spectrum and LSTM to predict the peptide sequences corresponding to those features. It can sequentially predict multiple amino acids in the protein, predicting the next amino acid from the previous one, achieving an accuracy of up to 75%. However, if any preceding amino acid is predicted incorrectly, the prediction error rate for subsequent amino acids increases significantly.

[0065] The problems with the two methods mentioned above not only greatly limit the efficiency of scientific research output, but also may cause doctors to miss the best treatment time for patients in clinical applications. Therefore, the clinical application of proteomics mass spectrometry technology is also limited.

[0066] Based on the aforementioned technical problems, this application proposes a mass spectrometry spectrum analysis method with short analysis time and high accuracy.

[0067] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0068] Furthermore, the labels for each step in this embodiment are for illustrative purposes only and do not represent a limitation on the order of movement of each step. In practical applications, the order of movement of each step can be adjusted or performed simultaneously as needed, and such adjustments or substitutions are all within the protection scope of this invention.

[0069] The following is in conjunction with the instruction manual appendix. Figures 1-9 The embodiments of this application will be described in further detail.

[0070] This application provides a method for mass spectrometry spectrum analysis, and the main process of the mass spectrometry spectrum analysis method is described below.

[0071] Reference Figure 1 S1. Obtain the target mass spectrometry dataset.

[0072] The target mass spectrometry dataset contains a set of proteomic mass spectrometry data, with multiple target spectra. Each target spectra represents a protein spectrum to be interpreted, and each spectrum represents a peptide segment within the protein. The target mass spectrometry dataset contains mass spectrometry data for this protein peptide segment, reflecting the basic properties of the peptide corresponding to the target spectrum. This data includes the parent ion valence, mass-to-charge ratio (M / C ratio), and intensity value corresponding to the M / C ratio. Specifically, one target spectrum corresponds to one parent ion valence, one target spectrum corresponds to multiple M / C ratios, and one M / C ratio corresponds to one intensity value.

[0073] Specifically, the mass-to-charge ratio, also known as m / z, is the data obtained by dividing mass by charge. The intensity value, also known as intensity, indicates the ion concentration of the parent ion corresponding to the mass-to-charge ratio; a higher intensity value indicates a higher ion concentration. For protein spectra with the same peptide sequence, positions with higher intensities of mass-to-charge ratios are generally close together.

[0074] It is understood that the above mass spectrometry data are all known information in the target spectrum, while the protein peptide corresponding to the target spectrum is unknown. The purpose of the mass spectrometry spectrum interpretation method provided in this application is to quickly and accurately analyze the peptide sequence of the protein peptide corresponding to the target spectrum based on the known information in the target spectrum.

[0075] In one embodiment, the target spectra in the target mass spectrometry dataset are stored in an MGF file.

[0076] S2. Based on the mass-to-charge ratio of the target spectrum, determine the spectral characteristics of the target spectrum.

[0077] Each target spectrum has a spectral feature that reflects the characteristics of the corresponding target spectrum and represents it. In this embodiment, the spectral feature consists of multiple mass-to-charge ratios. Since target spectra corresponding to different peptides have significant differences in their respective mass-to-charge ratios, using a specified number of mass-to-charge ratios to obtain the spectral feature makes the spectral feature highly representative of the target spectrum.

[0078] Reference Figure 1 and Figure 2 Specifically, step S2 includes:

[0079] S21. Based on the intensity value of the mass-to-charge ratio, filter all mass-to-charge ratios of the target spectrum to obtain a number of characteristic mass-to-charge ratios.

[0080] The characteristic mass-to-charge ratio (MTBR) is the MTBR selected from all MTBRs in the target spectrum that best represents the target spectrum. The number of characteristics is a system preset value. In this embodiment, all MTBR intensity values ​​in the target spectrum are read, and the number of MTBRs with the highest intensity values, i.e., the number of MTBRs with the highest ion concentration, are retained as the characteristic MTBRs of the target spectrum.

[0081] S22. Based on the mass-to-charge ratio of all features of the target spectrum, obtain the spectral features of the target spectrum.

[0082] Among them, the spectral features are a one-dimensional feature array consisting of the mass-to-charge ratio of the feature number of features.

[0083] S3. Input the spectral features of the target spectrum into the preset spectral analysis model to obtain the analytical category of the target spectrum.

[0084] The spectral analysis model is a mathematical model trained on the spectral features of a protein spectrum and the corresponding peptide sequences. The resolution category reflects the peptide sequence corresponding to the target spectrum. The spectral analysis model obtains the resolution category by analyzing spectral features composed of multiple mass-to-charge ratios.

[0085] In this embodiment, the spectral analysis model is a mathematical model based on the GoogleNet network structure, modified according to the characteristics of protein spectra. GoogleNet is a deep learning architecture that, compared to other deep learning architectures such as AlexNet and VGG, does not require increasing the network depth (number of layers) to achieve better training results. It reduces negative impacts such as overfitting, vanishing gradients, and exploding gradients, and can utilize computational resources more efficiently, extracting more features with the same amount of computation, thereby improving training results.

[0086] Reference Figure 3 and Figure 4 Specifically, the structure of the spectral analysis model includes convolutional layers, average downsampling layers, Inception layers, and fully connected layers.

[0087] Convolutional layers are used to extract features from the input or the output of the previous neural network layer. In this embodiment, the spectral features input to the spectral model are actually feature arrays composed of multiple mass-to-charge ratios arranged in a matrix, and the data processed by the convolutional layer is also a one-dimensional array.

[0088] The calculation process of the average downsampling layer is as follows: starting from the first value of the input array, the average of three consecutive numbers is taken, and the step size is 2 each time. Therefore, the length of the output array after calculation is reduced by half compared to the original array length, thus achieving the effect of feature compression.

[0089] The Inception layer extracts features through four different computations and then fuses them. The four different computations refer to the different sizes of the convolutional kernels involved in the process. This means that it can concatenate the results of multiple convolutions of different sizes and finally combine the results into an array.

[0090] The purpose of a fully connected layer is to allow the data to learn continuously, calculate the error, and use the error feedback to modify the parameters and continue learning until convergence is achieved. In this embodiment, the fully connected layer is a hidden layer composed of 1*2048 neurons, and the Dropout value of the fully connected layer is set to 0.8, meaning that only 80% of the neurons in the hidden layer are activated during each training process.

[0091] The training data for the spectral analysis model includes the spectral features of a protein spectrum and the corresponding peptide sequences. In practical applications, after inputting the feature array of the spectral features into the model, the convolutional layer performs multiple convolution operations on the feature array based on a preset kernel structure and stride value. The mean-downsampling layer reduces the length of the feature array. Finally, the output of the spectral analysis model is a feature matrix and the probability of this feature matrix belonging to each peptide sequence. The feature matrix can be understood as a compressed spectral feature that represents the characteristics of a target spectrum. By comparing the probabilities of the feature matrix relative to each peptide sequence, the peptide sequence with the highest probability can be identified as the most likely peptide sequence to which the feature matrix belongs, thus analyzing the peptide sequence to which the spectral features belong.

[0092] Specifically, in this embodiment, the feature compression length is 7, meaning that the spectral features are compressed to a length of 7 after passing through multiple average downsampling layers, in order to achieve a high resolution accuracy.

[0093] The specific analysis of the aforementioned beneficial effects is as follows: For mathematical models built using GoogleNet, the length of the output feature matrix is ​​typically 3, 5, 7, or 9. If the length of the compressed feature matrix is ​​set to 3, the feature matrix will struggle to represent the features of a single spectrum, resulting in significant feature loss and poor model accuracy. If the length of the compressed feature matrix is ​​5, there will also be significant feature loss, as the feature matrix cannot represent the features of a single spectrum, leading to poor model accuracy.

[0094] If the length of the compressed feature matrix is ​​set to 5, compared to setting the length to 3, it only reduces some feature loss. Validation within the same dataset yields good results, but validation across different datasets is less effective, resulting in lower model accuracy. Here, a dataset refers to a batch of extracted proteomics mass spectrometry data. For example, the spectral data generated after mass spectrometry of proteins extracted from a batch of 97H cells constitutes a dataset. This dataset is typically split in an 8:2 ratio, with 80% used for training and 20% used for prediction and validation.

[0095] The smaller the length of the compressed feature matrix, the higher the utilization rate of the protein spectrum. If the length of the feature matrix is ​​set to 9, the utilization rate of the protein spectrum will be low, resulting in poor performance in practical applications. Setting the length of the compressed feature matrix to 7 can balance the ability to represent the protein spectrum with the utilization rate of the protein spectrum, achieving a better model performance by balancing the two.

[0096] Furthermore, in this embodiment, the length of the spectral feature is set to 112, meaning the spectral feature contains 112 mass-to-charge ratios, and the number of average downsampling layers is set to 4. During the calculation process of the spectral decomposition model, the original spectral feature of length 112 will undergo feature compression through 4 average downsampling layers, that is, the feature array of the spectral feature will be halved four times, finally resulting in a feature matrix of length 7.

[0097] For protein spectra belonging to different protein peptide segments, the amount of mass-to-charge ratio (M / C ratio) data will vary. After filtering or data cleaning, the actual effective M / C ratio data will be even less. If the length of the spectral features is too large, such as setting the length of the spectral features to 224 and the number of average downsampling layers to 5, the amount of M / C ratio data required for calculation will be too large. This makes it difficult to obtain good prediction results for many protein spectra through the spectral interpretation model, thus narrowing the model's applicability.

[0098] In this embodiment, the length of the spectral features is set to 112, that is, the matrix length of the input content of the spectral decoupling model is 112. The decoupling model is designed with 4 layers of average value downsampling layers to achieve higher accuracy, wider applicability, and faster analysis speed in spectral decoupling analysis.

[0099] Understandably, in other embodiments, if in practical applications only the spectral analysis model is used to analyze a single dataset, without considering the impact of different datasets on validation performance, the length of the feature matrix can be reduced. For example, using a feature matrix of length 5 and employing 4 average downsampling layers, the spectral features would need to include 80 mass-to-charge ratios. Similarly, if in practical applications only peptide sequences with a small / large mass-to-charge ratio dataset are analyzed, fewer / more average downsampling layers can be used. For instance, if a peptide sequence with a large mass-to-charge ratio dataset is required, using 5 average downsampling layers and setting the feature matrix length to 7 would require the spectral features to include 224 mass-to-charge ratios.

[0100] The implementation principle of the mass spectrometry spectrum analysis method provided in this application is as follows: spectral features are obtained based on the mass-to-charge ratio of the target spectrum, and these features are input into the spectral analysis model. The analysis and classification are then obtained through the spectral analysis model. This mass spectrometry spectrum analysis method boasts high accuracy and short analysis time, enabling real-time spectral analysis at a transparent computational cost, i.e., protein spectroscopy. Figure 1 Once generated, it can be analyzed and the analysis results can be obtained quickly, which greatly shortens the analysis time. This not only reduces the limitation on the output efficiency of scientific research projects, but also facilitates rapid clinical testing and is beneficial to the clinical application of proteomics mass spectrometry technology.

[0101] This application provides a training method for a spectral analysis model, and the main process of the training method is described below.

[0102] Reference Figure 5 E1. Obtain the model training dataset.

[0103] The model training dataset contains multiple training spectra, mass spectrometry data corresponding to each training spectra, and peptide labels annotated on each training spectra.

[0104] A training spectrum refers to a protein spectrum, corresponding to the target spectrum, and represents a peptide segment within the protein. Mass spectrometry data reflects the basic properties of the peptide segment corresponding to the training spectrum, including the parent ion valence, mass-to-charge ratio, and intensity value relative to the mass-to-charge ratio. Peptide tags reflect the peptide sequence of the peptide segment corresponding to the training spectrum.

[0105] Reference Figure 5 and Figure 6 Specifically, step E1 includes:

[0106] E11. Obtain the raw mass spectrometry dataset and the decomposition dataset.

[0107] The mass spectrometry raw dataset contains the raw spectrum, and the spectral resolution dataset contains the spectral resolution results corresponding to the raw spectrum. In this embodiment, the spectral resolution dataset is obtained by parsing the raw spectrum using dedicated computer software.

[0108] In this embodiment, all the original images in the mass spectrometry raw dataset will be generated into a dictionary type. This dictionary type contains values ​​such as mass-to-charge ratio, peptide sequence, spectral resolution score, and post-translational modification.

[0109] E12. Filter the original spectra in the original mass spectrometry dataset to obtain the filtered spectra.

[0110] In order to improve the model training effect, it is necessary to filter and clean the data in the original mass spectrometry dataset. Filtered spectra refer to the original spectra with less effective information that need to be filtered out.

[0111] E13. Remove the filtered spectra from the original mass spectrometry dataset to obtain the mass spectrometry training dataset.

[0112] E14. Obtain the spectral features of all training spectra in the mass spectrometry training dataset.

[0113] After removing all filtered spectra from the mass spectrometry training dataset, all remaining original spectra in the mass spectrometry training dataset will be used as training spectra for subsequent model training.

[0114] E15. Based on the spectral decomposition dataset, all the original spectra in the mass spectrometry training dataset are labeled to obtain the peptide labels of the mass spectrometry training dataset.

[0115] Reference Figure 5 and Figure 7 In this embodiment, based on multiple known peptide sequences, each sequence is defined as a category, such as using category numbers 0, 1, 2, etc., and each peptide tag corresponds to a peptide category.

[0116] It is understandable that the training spectrum is equivalent to the protein spectrum before mass spectrometry analysis. Mass spectrometry data represents the known basic attributes of mass spectrometry data, while peptide tags are the actual analytical results obtained after mass spectrometry analysis of the training spectrum. In this embodiment, the specific method for generating peptide tags is as follows: the training spectrum is analyzed by mass spectrometry using dedicated computer software such as pFind. Then, based on the spectral interpretation results, the peptide sequences corresponding to the training spectrum are determined, and peptide tags for the training spectrum are generated according to the peptide classification corresponding to the peptide sequences.

[0117] E2. Input the spectral features of the training spectrum into the spectral decomposition model to obtain the training category of the training spectrum.

[0118] In this process, after the feature array of the spectral features is input into the spectral decomposition model, the convolutional layer performs multiple convolution operations on the feature array based on a preset convolutional kernel structure and stride value. The four-layer mean downsampling layer reduces the length of the feature array. The final output of the spectral decomposition model is a feature matrix of length 7, along with the probability of this feature matrix belonging to each peptide category. The peptide category with the highest probability is selected as the training category of the training spectrum.

[0119] E3. Train the spectral decomposition model based on peptide labels and training categories from the training spectra.

[0120] In this process, the model parameters of the spectral decomposition model are adjusted based on the peptide labels of the training spectra and the analysis results of the training spectra until the peptide labels of the training spectra and the analysis results of the training spectra are within the allowable range, thus obtaining the trained spectral decomposition model.

[0121] Reference Figure 8 To improve model training performance, in step E12 of this embodiment, in-depth data filtering and cleaning are required. In this embodiment, step E12 includes:

[0122] E121. Based on the spectral interpretation results of the original spectrum, filter the original spectra whose spectral interpretation results are reversed to obtain the filtered spectrum.

[0123] Specifically, the spectral results of all original spectra are extracted from the spectral dataset, and the original spectra whose spectral results are reversed are used as the filtered spectra.

[0124] E122. Based on the spectral interpretation results of the original spectra, filter out the original spectra with empty spectral interpretation results to obtain the filtered spectra.

[0125] Specifically, the spectral results of all original spectra are extracted from the spectral dataset, and the original spectra with empty spectral results are used as the filter spectra.

[0126] E123. Based on the spectral interpretation results of the original spectra, filter the original spectra whose scores are less than the score threshold to obtain the filtered spectra.

[0127] Specifically, the scores of the spectral results of all raw spectra are extracted from the spectral dataset, and raw spectra with scores below a threshold are used as filter spectra. The score of the spectral results is also called Raw_Score; the higher the score, the higher the reliability of the spectral results; conversely, the lower the score, the lower the reliability of the spectral results.

[0128] The scoring threshold is a preset value of the system. If the scoring threshold is too high, the number of usable training spectra after screening will be too small. If the scoring threshold is too high, a lot of unreliable spectral results will also be used in model training. In order to improve the utilization rate of training data while taking into account the effectiveness of training data, the scoring threshold is preferably 10 in this embodiment.

[0129] Based on steps E121-E123 above, the original spectra with the following results are considered to have low confidence: the original spectra are reversed, the original spectra are empty, or the original spectra score is less than the score threshold. The purpose of steps E121-E123 is to exclude these original spectra as filter spectra in order to optimize the training of the model.

[0130] E124. Based on the number of mass-to-charge ratios in the original spectrum, filter the original images whose mass-to-charge ratios are less than the feature threshold to obtain the filtered spectrum.

[0131] Specifically, the number of mass-to-charge ratios (MMRs) in all original spectra is obtained, and original images with MMRs less than a feature threshold are used as filter spectra. The feature threshold is greater than or equal to the number of MMRs in the spectral features. If the number of MMRs in an original spectrum is less than the feature threshold, then this original spectrum cannot extract or filter out a sufficient number of MMRs as spectral features for subsequent calculations; therefore, this part of the original spectrum needs to be excluded as a filter spectrum.

[0132] In this embodiment, the feature threshold is preferably 120.

[0133] E125. Based on the parent ion valence of the original spectra, the original spectra in the mass spectrometry raw dataset are grouped to obtain valence classification groups. Then, valence classification groups with fewer than the threshold number of valence spectra are selected to obtain filtered spectra.

[0134] Specifically, based on the parent ion valence of the original spectra, all original spectra are divided into multiple valence classification groups. Then, the number of original spectra in each valence classification group is counted. If the number of original spectra in a valence classification group is less than the valence spectrum count threshold, it indicates that the amount of data in this valence classification group is too small. In this embodiment, the valence spectrum count threshold is preferably 1000.

[0135] In the training process of deep learning models, theoretically, the more data belonging to the same parent ion valence, the better. This is because the protein spectra corresponding to different parent ion valences need to be discussed separately. Therefore, it is necessary to ensure that the number of protein spectra corresponding to each parent ion valence is sufficient. Otherwise, the model will have difficulty converging during training. Thus, it is necessary to remove valence classification groups with insufficient data, that is, to use all the original images contained in this part of the valence classification group as filter spectra.

[0136] In one embodiment, step E125 further includes:

[0137] The difference in the number of original spectra between different valence classification groups is calculated. This difference is compared with a difference threshold. If the difference is greater than the threshold, all original spectra from the valence classification group with fewer original spectra are used as filter spectra. The difference threshold is the number of original spectra from the valence classification group with fewer original spectra.

[0138] Understandably, when a dataset contains multiple valence classification groups, and although the number of original spectra for each valence classification group is greater than 1000, the difference in the number of original spectra between different valence classification groups is too large. In such cases, it is necessary to remove valence classification groups with significantly fewer original spectra. For example, if a dataset contains parent ion valences of +2, +3, +4, +5, and +6, where the average number of original spectra for +2 and +3 exceeds 100,000, while the number of original spectra for +2, +3, +4, +5, and +6 is approximately 20,000, then the original spectra for the valence classification groups of +2, +3, +4, +5, and +6 will be used as filter spectra. If a dataset contains a large number of original spectra for parent ion valences of +4, +5, and +6, it can be used as a case study in the training of the spectral interpretation model.

[0139] E126. Based on the peptide sequences in the spectral decomposition results, the original spectra in the original mass spectrometry dataset are grouped to obtain peptide classification groups. Then, peptide classification groups with fewer than the peptide spectrum number threshold are selected to obtain filtered spectra.

[0140] Specifically, based on the peptide classification of the original spectra, all original spectra are divided into multiple peptide classification groups. Then, the number of original spectra in each peptide classification group is counted. If the number of original spectra in a peptide classification group is less than the peptide spectrum count threshold, it indicates that the amount of data in that peptide classification group is too small. In this embodiment, the peptide spectrum count threshold is 120.

[0141] In the training process of deep learning models, theoretically, the more data belonging to the same peptide category, the better. This is because the protein spectra corresponding to different peptide categories need to be discussed separately. Therefore, it is necessary to ensure that the number of protein spectra corresponding to each peptide category is sufficient. Otherwise, the model will have difficulty converging during training. Thus, it is necessary to remove peptide categories with insufficient data, that is, to use all the original images contained in this part of the peptide category as filter spectra.

[0142] By using steps E125-E126, we can ensure that a high amount of data is maintained for both peptide sequence classification and parent ion valence classification, so that the spectral model can converge quickly during training and the model performance is better.

[0143] The following is an example of the training process of the spectral model using a set of 97H cell whole proteome mass spectrometry data.

[0144] The 97H cell whole proteome mass spectrometry data contained a total of 2,108,428 raw spectra. Based on the interpretation results of all raw spectra, the number of raw spectra with interpretation results of reverse library was 570,722, the number of raw spectra with interpretation results of empty was 82,335, and the number of raw spectra with interpretation score of less than 10 was 1,151,436. After filtering out these raw spectra as filter spectra, 303,935 raw spectra remained.

[0145] The remaining 303,935 original spectra have parent ion valences of +2, +3, +4, +5, and +6. These 303,935 original spectra are divided into five valence groups. The number of original spectra in valence groups +2 and +3 averages over 100,000, while the total number of original spectra in valence groups +4, +5, and +6 is less than 100,000, averaging only around 30,000-40,000. Therefore, the original spectra in valence groups +4, +5, and +6 are filtered out.

[0146] Based on the peptide classification of the remaining original spectra, all original spectra are grouped to obtain multiple peptide classification groups. Then, peptide classification groups with fewer than 120 spectra are selected, and the original spectra in these peptide classification groups are filtered out.

[0147] After the above filtering operations, the final raw spectrum obtained is the training spectrum, including:

[0148] 6356 training spectra with a parent ion valence of +2, including 44 peptide classifications;

[0149] 7856 training spectra with a parent ion valence of +3, including 55 peptide classifications.

[0150] Based on the top 112 mass-to-charge ratios with the highest intensity values ​​in the training spectrum, the spectral features of the training spectrum are obtained; based on the peptide classification to which the training spectrum belongs, the peptide labels of the training spectrum are obtained.

[0151] For example, one of the training spectra has a parent ion valence of +2. The spectral features of this training spectrum include a feature array of 112 mass-to-charge ratios, specifically: [169.13239, 198.08661, 199.07051, 216.09682, 217.1366, ... 1746.78259]. The peptide sequence to which this training spectrum belongs is PVSSAASVYAGAGGSGSR, the peptide classification is 17, and the peptide tag is 17.

[0152] The spectral features of the training spectra are input into the spectral decomposition model to obtain the training category of the training spectra. Then, based on the peptide labels and training categories of the training spectra, the spectral decomposition model can be trained.

[0153] This application also provides a mass spectrometry spectrum analysis apparatus, corresponding to the above-described mass spectrometry spectrum analysis method.

[0154] Reference Figure 9 The mass spectrometry spectrum analysis device includes:

[0155] Spectrum acquisition module 1 is used to acquire target spectra, wherein the target spectra correspond to the peptide segments of a protein;

[0156] Feature extraction module 2 is used to determine the spectral features of the target spectrum based on the mass-to-charge ratio of the target spectrum;

[0157] Model parsing module 3 is used to input the spectral features of the target spectrum into a preset spectral analysis model to obtain the parsing category of the target spectrum, wherein the parsing category is used to reflect the peptide sequence corresponding to the target spectrum.

[0158] Feature extraction module 2 includes:

[0159] The spectrum filtering submodule is used to filter all mass-to-charge ratios of the target spectrum based on the intensity value of the mass-to-charge ratio, and obtain a number of characteristic mass-to-charge ratios.

[0160] The feature combination submodule is used to obtain the spectral features of the target spectrum based on the mass-to-charge ratio of all features of the target spectrum.

[0161] The mass spectrometry spectrum analysis device provided in this embodiment can realize the steps of the aforementioned embodiment due to the functions of its modules and the logical connections between them. Therefore, it can achieve the same technical effect as the aforementioned method. For the principle analysis, please refer to the relevant description of the steps of the aforementioned mass spectrometry spectrum analysis method, which will not be repeated here.

[0162] This application also provides a smart terminal.

[0163] A smart terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The memory stores training data, algorithm formulas, and filtering mechanisms from a training model. The processor provides computational and control capabilities, and when executing the computer program, it implements a mass spectrometry analysis method.

[0164] The smart terminal provided in this embodiment can achieve the same technical effect as the aforementioned method because the computer program in its memory will implement the various steps of the aforementioned method after running on the processor. The principle analysis can be found in the relevant description of the steps of the aforementioned method, and will not be repeated here.

[0165] This application also provides a computer-readable storage medium.

[0166] A computer-readable storage medium includes a memory and a processor. The memory stores a computer program that can be loaded by the processor and executed as described above for mass spectrometry spectrum analysis. When the computer program is executed by the processor, it implements the mass spectrometry spectrum analysis method.

[0167] The readable storage medium provided in this embodiment can achieve the same technical effect as the aforementioned method after the computer program therein is loaded and run on the processor. For the principle analysis, please refer to the relevant description of the steps of the aforementioned method, which will not be repeated here.

[0168] The computer-readable storage medium includes, for example, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0169] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method of mass spectral profile resolution, characterized by, The method comprises: S1, obtaining a target spectrum, wherein the target spectrum is used to reflect mass spectrometry data of a protein peptide; S2, determining a spectrum feature of the target spectrum based on a mass-to-charge ratio of the target spectrum; step S2 comprises: screening all mass-to-charge ratios of the target spectrum based on intensity values of the mass-to-charge ratios to obtain 112 characteristic mass-to-charge ratios; and obtaining the spectrum feature of the target spectrum based on all the characteristic mass-to-charge ratios of the target spectrum; S3, inputting the spectrum feature of the target spectrum into a preset deconvolution model to obtain an analysis category of the target spectrum, wherein the analysis category is used to reflect a peptide sequence corresponding to the target spectrum; the output of the deconvolution model comprises a feature matrix obtained by compressing the spectrum feature and the analysis category corresponding to the feature matrix, and the length of the feature matrix is 7; the deconvolution model comprises four average value down-sampling layers for compressing the spectrum feature to obtain the feature matrix; The deconvolution model is a mathematical model modified based on the network structure of GoogleNet according to the characteristics of protein spectrum, and the deconvolution model is trained in the following manner: obtaining a model training data set, wherein the model training data set comprises training spectra, spectrum features corresponding to the training spectra, and peptide labels corresponding to the training spectra; inputting the spectrum features of the training spectra into the deconvolution model to obtain training categories of the training spectra; training the deconvolution model based on the peptide labels of the training spectra and the training categories of the training spectra.

2. The mass spectral profile resolution method of claim 1, wherein, The method for obtaining the model training data set comprises: obtaining a mass spectrometry original data set and a deconvolution data set, wherein the mass spectrometry original data set comprises original spectra, and the deconvolution data set comprises deconvolution results corresponding to the original spectra; screening the original spectra in the mass spectrometry original data set to obtain filtered spectra; removing the filtered spectra in the mass spectrometry original data set to obtain a mass spectrometry training data set; obtaining spectrum features of the mass spectrometry training data set; annotating all original spectra in the mass spectrometry training data set based on the deconvolution data set to obtain peptide labels of the mass spectrometry training data set.

3. The mass spectral profile resolution method of claim 2, wherein, The method for screening the original spectra in the mass spectrometry original data set to obtain filtered spectra comprises: screening the original spectra whose deconvolution results are anti-libraries based on the deconvolution results of the original spectra to obtain filtered spectra; screening the original spectra whose deconvolution results are empty based on the deconvolution results of the original spectra to obtain filtered spectra; screening the original spectra whose scores of the deconvolution results are less than a score threshold based on the deconvolution results of the original spectra to obtain filtered spectra; screening the original spectra whose numbers of mass-to-charge ratios are less than a feature threshold based on the numbers of mass-to-charge ratios of the original spectra to obtain filtered spectra, wherein the feature threshold is less than or equal to the number of mass-to-charge ratios contained in the spectrum feature. grouping the original spectra in the mass spectrum original data set based on the parent ion valence of the original spectra, to obtain a valence classification group; filtering the valence classification group with a spectrum number less than a valence graph number threshold, to obtain a filtered spectrum. grouping the original spectra in the mass spectrum original data set based on the peptide sequence in the deconvolution result, to obtain a peptide classification group; filtering the peptide classification group with a spectrum number less than a peptide graph number threshold, to obtain a filtered spectrum.

4. A mass spectrum analyzing apparatus characterized by comprising: A mass spectrum spectrum analysis device for implementing a mass spectrum spectrum analysis method according to any one of claims 1 to 3, the mass spectrum spectrum analysis device comprising: a spectrum acquisition module (1) configured to acquire a target spectrum, wherein the target spectrum corresponds to a peptide of a protein; a feature extraction module (2) configured to determine a spectrum feature of the target spectrum based on a mass-to-charge ratio of the target spectrum; a model analysis module (3) configured to input the spectrum feature of the target spectrum into a preset deconvolution model to obtain an analysis category of the target spectrum, wherein the analysis category is used to reflect a peptide sequence corresponding to the target spectrum.

5. A smart terminal, characterized by A computer program stored on a memory and capable of being loaded and executed by a processor to implement a mass spectrum spectrum analysis method according to any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, A computer program stored on a memory and capable of being loaded and executed by a processor to implement a mass spectrum spectrum analysis method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Protein secondary mass spectrum identification method of marker loci based on candidate peptide fragment discrimination

    CN103245714A

  • Deep learning-based protein mass spectrum data analysis method and system

    CN113362899A