A method for judging the reliability of mass spectrometry identification of tumor neoantigen peptides in cells and its application
The reliability of screening tumor neoantigen peptides in cancer cell lines through mass spectrometry detection and machine learning models has solved the problem of frequent false negative results in existing technologies and achieved efficient tumor neoantigen identification.
Patent Information
- Application Number
- CN202510897732.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Existing technologies make it difficult to quickly and accurately identify tumor neoantigen peptides in cancer cell lines. Conventional methods lead to frequent false-negative results, resulting in waste of experimental resources and inefficiency.
Mass spectrometry detection combined with machine learning methods was used to construct a stochastic gradient descent classifier model. By training with historical data and setting the demarcation threshold, the reliability of mass spectrometry identification of tumor neoantigen peptides in the samples to be tested was predicted, and potential samples were screened for subsequent identification.
It improves the success rate of mass spectrometry identification of tumor neoantigens, reduces the waste of experimental steps and cell samples, and improves experimental efficiency and accuracy.
Smart Images

Figure CN120412757B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of protein detection, and specifically relates to a method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells and its application. Background Art
[0002] Tumor neoantigens are short peptides presented on the surface of tumor cells by the MHC class I complex. These peptides often harbor tumor-specific mutations. Because they are not expressed in normal cells, they are ideal targets for tumor immunotherapy and tumor vaccines. However, the low abundance of tumor neoantigens and the complex origins of their mutations make their rapid and accurate identification difficult using conventional techniques.
[0003] With the advent of cancer immunotherapy, tumor neoantigens and their role in cancer treatment have garnered increasing attention and recognition. A growing number of researchers are interested in identifying tumor neoantigens using mass spectrometry. The samples used can be broadly categorized into two types: tumor tissue and cancer cell lines. Because clinically derived tumor tissues are subject to numerous constraints, cancer cell lines, which can be propagated indefinitely, are more popular. Most researchers transduce their target protein of interest into cancer cell lines and then use mass spectrometry to identify peptides derived from the target protein that are presented on the cell surface by the cancer cells. However, these experiments often fail, with mass spectrometry failing to identify any peptides derived from the target protein. However, further investigation reveals that this is often a false negative. To address this issue, we have developed a mass spectrometry-based method that can rapidly determine whether a cancer cell line can yield satisfactory results in mass spectrometry-based tumor neoantigen detection.
[0004] Existing methods often rely on techniques such as RT-PCR (to identify mRNA expression), Western blotting (to identify protein expression), and flow cytometry (to identify cell surface MHC I expression) to measure MHC I or target protein expression in cancer cell lines. However, these techniques all have varying degrees of limitations. RT-PCR can only identify mRNA expression, but mRNA expression does not necessarily equate to protein expression. Western blotting can identify protein expression in cells, but the use of secondary antibodies in Western blotting can distort the final signal, leading to inaccurate expression levels. Flow cytometry can accurately identify cell surface MHC I expression, but it cannot accurately determine whether peptides of the target protein are presented on the cell surface. Furthermore, because both secondary antibodies and fluorescent small molecules amplify the signal, this can distort the final signal, leading to inaccurate expression levels.
[0005] Therefore, providing a method to quickly determine whether the identification of tumor neoantigen peptides in cancer cell lines is successful or not has important application value. Summary of the Invention
[0006] To address the shortcomings of existing technologies, the present invention aims to provide a method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells and its application. The method of the present invention can rapidly determine the success or failure of tumor neoantigen peptide identification in cancer cell lines, thereby resolving the issue of false negatives in mass spectrometry identification of tumor neoantigens in cancer cell lines.
[0007] In order to achieve the purpose of the invention, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells, the method comprising:
[0009] Obtain historical data, including statistical results of the expression of MHC I proteins and target proteins in cells detected by mass spectrometry, as well as the mass spectrometry identification results of tumor neoantigen peptides in the corresponding cells; divide the historical data into positive and negative groups;
[0010] A model is constructed using historical data from the positive and negative groups as a training set, with the feature vectors in the training set being the statistical expression results of the MHCI protein and the target protein. The model is constructed using a stochastic gradient descent classifier with a logarithmic loss function, and ten-fold cross-validation and L1 regularization are used to avoid overfitting. The positive and negative groups are scored based on the model, and the threshold for separating the positive and negative groups is calculated based on the score values.
[0011] Reliability prediction: Use the trained model to score the test samples, compare the score value with the demarcation threshold, and judge the reliability of the test cell sample for the tumor neoantigen peptide mass spectrometry identification experiment.
[0012] The present invention uses conventional mass spectrometry experiments to identify the statistical results of protein expression in cells, and constructs a judgment model based on this. Based on this, a judgment is made on whether the mass spectrometry identification experiment of tumor neoantigen peptides is feasible. Based on the judgment results, it is decided whether the sample should undergo mass spectrometry identification of tumor neoantigen peptides.
[0013] The method of the present invention is advantageous in that conventional mass spectrometry experiments (expression level detection) have a short cycle and require fewer cells, whereas the latter experiment (mass spectrometry identification of tumor neoantigen peptides) requires a long cycle, multiple experimental steps, a complex experimental process, and a large number of cells. Therefore, the present invention uses this experimental approach for preliminary determination or prescreening, and then, based on the results of the preliminary determination, determines whether subsequent mass spectrometry identification of tumor neoantigen peptides is necessary. This effectively reduces experimental risk, improves experimental efficiency, and avoids waste of cell samples.
[0014] Preferably, the pre-treatment steps of the mass spectrometry detection include: lysing cells, extracting proteins, and sequentially performing reduction, alkylation and enzyme cleavage to obtain peptide fragments.
[0015] Preferably, the method for calculating the statistical results of the expression levels of the MHC I protein and the target protein comprises: searching the library using mass spectrometry data analysis software (such as Novor, pFind, Proteome Discoverer, and ProteoScape, etc.) to obtain the number, coverage, and confidence score of the secondary spectra corresponding to the MHC I protein and the target protein; sorting the other proteins obtained by the search according to the corresponding number of secondary spectra and taking the top 10%, calculating the average value of the corresponding number, coverage, and confidence score of the secondary spectra, and normalizing the number, coverage, and confidence score of the secondary spectra of the MHC I protein and the target protein between different cell samples using the above average value, and using the normalized value as the expression statistical result.
[0016] Preferably, the normalized calculation formula is as follows:
[0017] ;
[0018] Where, X i is the number, coverage, and confidence score of the secondary spectra of MHC I protein and target protein; μ is the average number, coverage, and confidence score of the secondary spectra of other proteins in the top 10% of the secondary spectra except MHC I protein and target protein; X norm is the number, coverage, and confidence score of the normalized secondary spectra of MHC I protein and target protein.
[0019] The normalized value is used as a vector, which includes the number of secondary spectra, coverage and confidence score. This vector is the input value of the stochastic gradient descent classifier.
[0020] Preferably, the positive group is an experimental group that has undergone a tumor neoantigen peptide mass spectrometry identification experiment and identified the target protein peptide.
[0021] Preferably, the negative group is an experimental group that has undergone a tumor neoantigen peptide mass spectrometry identification experiment and in which no target protein peptide was identified.
[0022] In the present invention, the positive group and the negative group are divided based on historical data. Both have undergone tumor neoantigen peptide extraction, then mass spectrometry detection, and finally a de novo sequencing algorithm to parse the amino acid sequence of the peptide. The peptide list is then compared to see if there are peptides from the target protein. If there are peptides from the target protein, it is judged as positive, otherwise it is negative.
[0023] Preferably, the method for calculating the demarcation threshold is: among all possible values in the training set, by continuously adjusting the threshold, calculating the true positive rate and false positive rate of samples under different thresholds, and taking the threshold when the difference between the true positive rate and the false positive rate is the largest as the demarcation threshold.
[0024] Preferably, the step of scoring the sample to be tested includes: detecting the expression statistics of MHC I protein and target protein in the cells of the sample to be tested based on mass spectrometry experiments, wherein the expression statistics include: the number of secondary spectra, coverage and confidence score, and inputting the expression statistics into the model for calculation and scoring.
[0025] Preferably, the judgment criteria are: if the score value is lower than the demarcation threshold, it is recorded as negative, and it is not recommended to continue the mass spectrometry identification experiment of tumor neoantigen peptides; otherwise, it is recorded as positive, and it is recommended that the sample continue to undergo the mass spectrometry identification experiment of tumor neoantigen peptides.
[0026] In the present invention, if the score value is lower than the demarcation threshold, it is recorded as negative, indicating that if the sample is subjected to a mass spectrometry identification experiment of tumor neoantigen peptides, there is a high possibility that the peptides derived from the target protein cannot be identified. Therefore, it is not recommended to continue the mass spectrometry identification experiment of tumor neoantigen peptides for the sample; otherwise, it is recorded as positive, indicating that if the sample is subjected to a mass spectrometry identification experiment of tumor neoantigen peptides, there is a high possibility that the peptides derived from the target protein can be identified. Therefore, it is recommended that the sample continue to undergo a mass spectrometry identification experiment of tumor neoantigen peptides.
[0027] In the present invention, a prediction model is constructed using a training set based on historical detection data, and the threshold for dividing negative and positive is determined. The score output by the model is used as the basis for judging the quality of the sample, that is, the probability of whether the peptide segment derived from the target protein in the cell can be effectively identified by the mass spectrometer after being presented to the cell surface.
[0028] In a second aspect, the present invention provides the application of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells described in the first aspect in screening tumor neoantigens.
[0029] The method of the present invention uses the relative abundance of proteins measured by mass spectrometry to more accurately reflect the expression level of proteins, and thus can more accurately measure the success rate of mass spectrometry identification of tumor neoantigens.
[0030] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells as described in the first aspect are implemented.
[0031] In a fourth aspect, the present invention provides a computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, and when the processor executes the computer program, the steps of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells as described in the first aspect are implemented.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] This invention is the first to use mass spectrometry to measure the success rate of tumor neoantigen identification and uses machine learning to build a predictive model to determine the quality of tumor neoantigen samples. Compared with existing methods (RT-PCR, Western blotting, or flow cytometry, etc.), the relative abundance of proteins measured by mass spectrometry can more accurately reflect the expression level of proteins, and thus can more accurately measure the success rate of tumor neoantigen identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the experimental process.
[0035] Figure 2 Schematic diagram of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells.
[0036] Figure 3 It is the classification result of positive and negative data using the trained model. DETAILED DESCRIPTION
[0037] The technical solution of the present invention is further described below by way of specific embodiments. It should be understood by those skilled in the art that the embodiments are merely to help understand the present invention and should not be regarded as specific limitations of the present invention.
[0038] If no specific techniques or conditions are specified in the examples, the experiments were carried out according to the techniques or conditions described in the literature in the field or according to the product instructions. If no manufacturer is specified for the reagents or instruments used, they are all conventional products that can be purchased through regular channels.
[0039] The experimental flow diagram of the present invention is as follows Figure 1 As shown, in the experiment, the samples were subjected to cell lysis, protein extraction, reduction, alkylation and enzyme cleavage in sequence to obtain peptides, which were then detected.
[0040] The schematic diagram of the method for judging the reliability of mass spectrometry identification of tumor neoantigen peptides in cells of the present invention is as follows: Figure 2 As shown, the method includes sample pretreatment, mass spectrometry detection, building a prediction model based on the mass spectrometry detection results, and substituting the detection results of the sample to be tested into the prediction model for result judgment.
[0041] Example 1
[0042] This example provides a method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells.
[0043] 1. Sample pretreatment
[0044] There are cells #1 and #2, both of which express MHC I and a target protein; 2×10 6 The cells were washed three times with PBS buffer. The supernatant was discarded, lysis buffer was added to the cells, and the cells were incubated with rotation at 4°C for 2 hours. The supernatant was collected by centrifugation, which was the cell lysate. Four volumes of acetone (previously stored at -20°C) were added to the cell lysate, mixed thoroughly, and incubated at -20°C for at least 1 hour. The cells were then centrifuged at 15,000 rpm at 4°C for 15 minutes. The supernatant was discarded, the tube was opened, and air-dried for 5 minutes. The protein pellet was dissolved in 10 μL of 8 M urea (100 mM Tris pH 8.5). Sonication was performed for 10 minutes to aid protein solubilization. 0.2 μL of 500 mM TCEP (final concentration 5 mM) was added and the mixture was incubated at room temperature for 20 minutes. 0.2 μL of 500 mM IAA (final concentration 10 mM) was added and the mixture was incubated at room temperature for 15 minutes. 30 μL of 100 mM Tris pH 8.5 was added to dilute the 8 M urea to 2 M. Add trypsin at a 1:(20-100) enzyme-to-protein ratio and digest at 37°C in the dark for 12-16 h. Terminate the digestion reaction by adding 2.2 μL of 90% formic acid to a final concentration of 5%.
[0045] 2. Mass spectrometry detection
[0046] The peptides obtained by enzymatic digestion were loaded onto a reverse-phase chromatography column packed with C18 packing and then washed with an aqueous solution containing 5% acetonitrile and 0.1% formic acid to remove salts from the sample. The column loaded with the peptide sample was then connected to a high-performance liquid chromatograph (HPLC). The peptides were eluted one by one from the column using a 1-hour gradient elution program and then detected by a mass spectrometer.
[0047] The mass spectrometer spray voltage was set to 3.5 kV, the HPLC flow rate was set to 200 nL / min, and the mass spectrometer was set in data-dependent mode. Primary mass spectrometry scans were performed in an orbitrap with a resolution of 60,000. The primary mass spectrometry scan range was 400–2000 m / z. Each primary mass spectrometry scan was followed by eight data-dependent secondary mass spectrometry scans. Peptide fragmentation was performed in a linear ion trap using high-energy collisional dissociation (HCD) mode with a collision energy of 35%. The dynamic exclusion time was set to 30 s.
[0048] 3. Model Scoring
[0049] 3.1 Mass spectrometry data analysis
[0050] Mass spectrometry data were analyzed using Novor software. Novor software was used to search the database to obtain the number of secondary spectra, coverage, and confidence scores corresponding to the MHC I protein and the target protein. The remaining proteins found were ranked by the number of corresponding secondary spectra, and the top 10% were selected. The average number of secondary spectra, coverage, and confidence scores for these proteins were calculated and used to normalize the corresponding values for the MHC I protein and the target protein, which were used to calculate the expression level statistics.
[0051] 3.2 Collecting Historical Data
[0052] Obtain statistical results of the expression levels of MHC I proteins and target proteins in cells detected by mass spectrometry experiments, as well as the mass spectrometry identification results of tumor neoantigen peptides in the corresponding cells; divide historical data into positive and negative groups.
[0053] 4. Build the model
[0054] The historical data of the positive and negative groups are used as a training set, and the feature vectors in the training set are the statistical results of the expression levels of the MHC I protein and the target protein; a stochastic gradient descent classifier is used to build a model, and the loss function is a logarithmic loss function. Ten-fold cross-validation and L1 regularization are used to avoid overfitting; the positive and negative groups are scored based on the model, and the threshold for separating the positive and negative groups is calculated based on the score value;
[0055] By adjusting the threshold, the true positive rate and false positive rate of samples under different thresholds are calculated, and the threshold at which the difference between the true positive rate and the false positive rate is the largest is used as the demarcation threshold.
[0056] In this embodiment, the threshold value for dividing negative and positive is determined to be 17.75 through calculation.
[0057] 5. Reliability prediction
[0058] The trained model is used to score the test samples, and the score is compared with the cut-off threshold to determine the reliability of mass spectrometry identification of tumor neoantigen peptides in cells.
[0059] According to the experimental results, the score of cell 1 is 109.53, which is higher than the threshold of 17.75, so it is identified as positive, that is, the success rate of tumor neoantigen detection by mass spectrometry is high. The score of cell 2 is -29.46, which is lower than the threshold of 17.75, so it is identified as negative, that is, the success rate of tumor neoantigen detection by mass spectrometry is low. (For example, Figure 3 ).
[0060] Both cell 1 and cell 2 underwent mass spectrometry identification of tumor neoantigens. The identification results showed that multiple peptides derived from the target protein were identified in cell 1; while no peptides derived from the target protein were identified in cell 2.
[0061] In summary, the present invention provides a method for judging the reliability of mass spectrometry identification of tumor neoantigen peptides in cells, based on the statistics of historical data and the construction of a mathematical model, scoring the positive group and the negative group based on the model, and calculating the demarcation threshold of the positive group and the negative group based on the score value; and based on the comparison of the score value of the sample to be tested with the demarcation threshold value, the reliability of continuing the mass spectrometry identification experiment can be quickly judged. The experimental method of the present invention is used to make a preliminary judgment or pre-screening, and then based on the results of the preliminary judgment, it is judged whether it is necessary to conduct a subsequent mass spectrometry identification experiment of the tumor neoantigen peptide segment, which can effectively control the experimental risk, improve the experimental efficiency, and avoid the waste of cell samples.
[0062] The applicant declares that the above is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention fall within the scope of protection and disclosure of the present invention.
Claims
1. A method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells, characterized in that: The method comprises: Obtain historical data, obtain statistical results of the expression of MHC I protein and target protein in cells based on mass spectrometry experiments, and mass spectrometry identification results of tumor neoantigen peptides of corresponding cells; divide the historical data into a positive group and a negative group based on whether the tumor neoantigen mass spectrometry identification experiment can identify peptides from the target protein; the calculation method of the statistical results of the expression of MHC I protein and target protein includes: searching the library through mass spectrometry data analysis software to obtain the number, coverage and confidence score of secondary spectra corresponding to MHC I protein and target protein; sorting other proteins obtained from the search according to the number of corresponding secondary spectra and taking the top 10%, calculating the average value of the corresponding number of secondary spectra, coverage and confidence score, normalizing the number, coverage and confidence score of secondary spectra of MHC I protein and target protein between different cell samples with the above average value, and using the normalized value as the expression statistical result; The normalized calculation formula is as follows: ; Where, is the number, coverage, and confidence score of the secondary spectra of MHC I protein and target protein; It is the average of the number of secondary spectra, coverage, and confidence scores of proteins other than MHC I proteins and target proteins whose secondary spectra are ranked in the top 10%; is the number, coverage, and confidence score of the normalized secondary spectra of MHC I protein and target protein; A model is constructed, using historical data of the positive and negative groups as a training set, wherein the feature vectors in the training set are statistical results of the expression levels of the MHC I protein and the target protein identified in the mass spectrometry experiment; a stochastic gradient descent classifier is used to construct the model, a logarithmic loss function is used as the loss function, and ten-fold cross-validation and L1 regularization are used to avoid overfitting; the positive and negative groups are scored based on the model, and a demarcation threshold between the positive and negative groups is calculated based on the score; the demarcation threshold is calculated by continuously adjusting the threshold among all possible values in the training set, calculating the true positive rate and false positive rate of samples under different thresholds, and the threshold at which the difference between the true positive rate and the false positive rate is the largest is used as the demarcation threshold; Reliability prediction uses the trained model to score the test sample. The step of scoring the test sample includes: detecting the expression statistics of MHC I protein and target protein in the cells of the test sample based on mass spectrometry experiment, the expression statistics including the number of secondary spectra, coverage and confidence score, inputting the expression statistics into the model for calculation and scoring; comparing the score value with the demarcation threshold to determine the reliability of the tumor neoantigen peptide mass spectrometry identification experiment of the test cell sample.
2. The method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells according to claim 1, characterized in that: The pre-treatment steps of the mass spectrometry experiment detection include: cell lysis, protein extraction, and sequential reduction, alkylation and enzyme cleavage to obtain peptide segments.
3. The method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells according to claim 1, characterized in that: The positive group is an experimental group in which peptides derived from the target protein can be identified in the tumor neoantigen peptide mass spectrometry identification experiment; The negative group is an experimental group in which no peptides from the target protein were identified in the tumor neoantigen peptide mass spectrometry identification experiment.
4. The method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells according to claim 1, characterized in that: The score value is compared with the cut-off threshold to determine the reliability of the mass spectrometry identification experiment for tumor neoantigen peptides on the cell sample to be tested: if the score value is lower than the cut-off threshold, it is recorded as negative, and it is not recommended to continue the mass spectrometry identification experiment for tumor neoantigen peptides on the sample; otherwise, it is recorded as positive, and it is recommended to continue the mass spectrometry identification experiment for tumor neoantigen peptides on the sample.
5. Use of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells according to any one of claims 1 to 4 in screening tumor neoantigens.
6. A computer-readable storage medium, characterized in that The storage medium stores a computer program, wherein when the computer program is executed by the processor, the steps of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells according to any one of claims 1 to 4 are implemented.
7. A computer device comprising a memory and a processor, wherein a computer program capable of being run on the processor is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method for determining the reliability of mass spectrometry identification of tumor neoantigen peptides in cells according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Model for predicting immunogenicity and application
CN115497613A