A method for assessing the risk of lung cancer recurrence

Mass spectrometry technology detects the proteome in the postoperative tissues of lung cancer patients, combines with the TNM staging system, and builds a logistic regression model, solving the problem of inaccurate risk assessment of lung cancer recurrence in the existing technology, and achieving accurate risk stratification and prognosis prediction for early-stage lung cancer patients.

CN113936734BActive Publication Date: 2025-07-11PROTEINT (TIANJIN) BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111488158.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-07-11
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

The prior art has high recurrence rates after early treatment of lung cancer, especially in patients with stage IB to stage IIIA. It is difficult for existing genetic testing methods to accurately evaluate the risk of recurrence, resulting in the inability to effectively reduce the postoperative recurrence rate.

Method used

Mass spectrometry technology detects the proteome in the postoperative tissues of lung cancer patients, constructs a protein expression matrix, combines with the TNM staging system, and uses logistic regression algorithm to build a risk assessment model to achieve accurate risk stratification of lung cancer patients.

Benefits of technology

It improves the accuracy and detection rate of lung cancer recurrence risk assessment, can identify high-risk patients earlier, help clinicians to conduct timely intervention, and improve patients' prognosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113936734B_ABST
    Figure CN113936734B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for assessing the risk of lung cancer recurrence, including data preprocessing, data quality control, data re-quality control, sample data correction, constructing a protein data matrix, constructing a recurrence model through the protein data matrix and clinical information, screening and testing to select the best prediction model, and finally determining the positive judgment criteria and result interpretation. Compared with the prior art, the method of the present invention has a high detection rate, which is the basis for postoperative risk stratification of early-stage lung cancer patients. Existing research data show that the present invention can have a good prognostic prediction effect on patients with stage I lung cancer. Compared with using only NCCN stratification or traditional TNM staging alone, the combined use of the risk stratification of this model can more accurately identify the prognostic risk, help clinicians recognize the tumor status earlier, perform clinical intervention earlier, and bring better prognosis for patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of a lung cancer prediction method, and specifically, relates to a method for assessing the recurrence risk of lung cancer. Background Art

[0002] Lung cancer has a high incidence and mortality rate and is the number one tumor endangering human life and health. In terms of treatment, the importance of early screening, early diagnosis, and early treatment of lung cancer for the prognosis of patients has always been recognized in the field. However, the current early treatment of lung cancer mainly based on local surgery cannot achieve radical cure for all patients, and the proportion of patients with recurrence and metastasis after surgery still cannot be ignored. At the same time, patients who enter the advanced treatment mode will live with the tumor. Therefore, moving the tumor treatment node forward is crucial for patients to achieve high-quality long-term survival.

[0003] According to the survey data, the postoperative recurrence rate of lung cancer patients from stage IB to stage IIIA exceeds 50%. Even for patients in the extremely early stage of stage IB, the postoperative recurrence rate is still about 32%, and a considerable part of them are distant recurrences. Although we can achieve local radical cure, it is difficult to control distant metastasis. In fact, local recurrence and distant metastasis are all important indicators for measuring the final prognosis of patients. This cruel data indicates that we must reduce the postoperative recurrence rate of the above patients through precise diagnosis and postoperative adjuvant treatment to achieve the purpose of radical resection.

[0004] Currently on the market, there are products for assessing the recurrence risk of stage I-IIA non-squamous NSCLC patients based on multi-gene expression patterns through RNA detection of tumor tissues. The difference in the differentiated population of low, medium, and high risks is smaller than the results of the protein-level model in the present invention. There is currently no such product for constructing a recurrence risk model for stage I lung cancer through mass spectrometry technology. Similar technologies include constructing disease risk models through gene detection technology. Technical methods such as gene detection and transcriptional detection are all links in the central dogma. Summary of the Invention

[0005] Object of the Invention: To solve the problems of the existing technology, the present invention provides a method for assessing the recurrence risk of lung cancer, which can reduce the defects of indirect prediction through genes and transcription.

[0006] Technical Solution: To achieve the above object of the invention, the present invention adopts the following technical solution:

[0007] A method for assessing the recurrence risk of lung cancer, comprising the following steps:

[0008] (1) Data preprocessing; importing the reference standard data file and the data file generated by mass spectrometry into software for data extraction;

[0009] (2) Data quality control: Determine whether the detection error is within the range, whether the sample injection volume meets the standard, and whether there is blood contamination in the sample;

[0010] (3) Data re-quality control: For the data that has passed the quality control in step (2), perform data quality control on the ion information intensity data obtained by the software again;

[0011] (4) Sample data correction: For the data that has passed the quality control in step (3), correct the passed ion intensity and construct the corrected sample data;

[0012] (5) Protein data matrix construction: Obtain sample ion information through the above quality control and correction, construct protein data, and construct a protein expression matrix through recurrence and non-recurrence grouping information;

[0013] (6) Model construction: According to the existing data, randomly select some sample data for classification model training and construction of various algorithms. After successfully constructing the model, verify the model through the remaining data, evaluate the model effect, and select the best prediction model;

[0014] (7) Positive judgment criterion: Through the constructed best prediction model, select the best discrimination threshold under the best ROC condition, map the model values of the recurrence sample data between -1 and 0, and map the non-recurrence sample data values between 0 and 1; After detecting and analyzing the test samples that meet the quality control standards, if the test value finally falls between -1 and 0, it is judged as positive, otherwise it is judged as negative.

[0015] Preferably:

[0016] In step (1), import the reference standard data file and the mass spectrometry-generated data file into the Spectronaut software, select the internally preset FASTA file of Swissprot_Homo, select the preset database LungCancer_lib database, and use the preset method named Umbrella in the settingscheama to start data extraction.

[0017] In step (2), determine whether the liquid phase detection error is within the range based on the sample iRT data; determine whether the mass spectrometry detection error is within the range based on the MS1 / MS2 MassAccuracy data; determine whether the sample injection volume meets the standard based on the total TIC intensity, protein, and peptide identification numbers; determine whether there is blood contamination in the sample based on the protein and peptide identification numbers.

[0018] In step (3), the ion information intensity data obtained by Spectronaut software is subjected to data quality control again: using the data quality control module, filter out the results that do not meet the F.FrgLossType type, remove those with F.MassAccuracyPPM >= 10, or F.MassAccuracyPPM <= -10, and ion intensity >= 1500. Ions that meet the above conditions pass the quality control standard, otherwise the ion information of this sample is removed.

[0019] In step (4), for the data that has passed the quality control in step (3), the ion intensity that has passed is corrected by the data correction method of the in-house algorithm Umbrella software to construct the corrected sample data.

[0020] In step (5), the sample ion information is obtained through quality control and correction. Peptide segments with more than 3 ions under one peptide are retained, and the median intensity of the top 3 ions is used to replace the peptide segment intensity value. For more than >= 1 peptide segment information under one protein, the median intensity of the top 3 peptide segments is used as the final intensity information representing this protein of this sample; the protein data constructed by this method is used to construct a protein expression matrix through the recurrence and non-recurrence grouping information.

[0021] In step (6), according to the existing data, 70% of the sample data is randomly selected. The recurrence and non-recurrence are divided into two groups in a ratio of 7:3 according to the proportion. The 70% sample data is used for the training and construction of the classification models of each algorithm. After successfully constructing the model, the remaining 30% of the data is used for model verification to evaluate the model effect and screen the best prediction model.

[0022] In step (6), the best prediction model is the classification model constructed by the logistics regression algorithm.

[0023] Preferably, the method provided by the present invention is a method for evaluating the recurrence risk of stage I lung cancer.

[0024] Although genetic testing has been widely used in clinical practice, there is still a huge gap from people's expectations. The reason may be that the relationship between genes and diseases remains indirect. We know that proteins in cells are the basis and functional executors of life. Only after post-translational modification of proteins can their activities be changed to perform functions. Genes need to be transcribed into mRNA, translated into proteins, then undergo post-translational modification, and through protein-protein interactions to form complexes, before they have biological activities to regulate and execute specific functions. Therefore, compared with proteins, the relationship between genes and diseases is indirect, and the relationship between proteins and diseases is more direct. In the field of lung cancer, recently, the first detection product based on proteomics for predicting the risk stratification of postoperative recurrence and the benefit of postoperative adjuvant chemotherapy in patients has emerged. By detecting and analyzing the proteome in the postoperative tissues of early lung cancer patients, it realizes the precise risk stratification of patients. There is research evidence that for lung cancer patients evaluated as having a low recurrence risk by TNM (The TNM staging system is the most commonly used tumor staging system internationally. In the TNM staging system: 1. T (the "T" is the first letter of the English word "Tumor" for tumor) refers to the situation of the primary tumor. As the tumor volume increases and the extent of adjacent tissue involvement increases, they are represented by T1 to T4 in sequence; 2. N (the "N" is the first letter of the English word "Node" for lymph node) refers to the involvement of regional lymph nodes. When the lymph nodes are not involved, it is represented by N0. As the degree and extent of lymph node involvement increase, they are represented by N1 to N3 in sequence; 3. M (the "M" is the first letter of the English word "metastasis" for metastasis) refers to distant metastasis (usually hematogenous metastasis). Those without distant metastasis are represented by M0, and those with distant metastasis are represented by M1. On this basis, specific stages are delineated by the combination of the three TNM indicators.). After stratification by Pulmonary Clear Spectrum, about one-third of the patients are still defined as high-risk. In other words, it is extremely necessary to combine the TNM staging of lung cancer with molecular biology testing to form the tumor TNMB (TNMB is based on the TNM classification and adds some other biological methods for combined classification. The "B" refers to Biology) staging.

[0025] Beneficial effects: Compared with the prior art, the method of the present invention has a high detection rate, which is the basis for postoperative risk stratification of early lung cancer patients. Existing research data show that the present invention can have a good prognostic prediction effect on stage I lung cancer patients. Compared with using only NCCN (National Comprehensive Cancer Network in the United States) stratification or traditional TNM staging alone, the combined use of this model for risk stratification can more accurately identify prognostic risks, help clinicians recognize the tumor status earlier, perform clinical intervention earlier, and bring better prognosis for patients. Description of the Drawings

[0026] Figure 1 This is the flowchart of the method for assessing the risk of lung cancer recurrence in the present invention.

[0027] Figure 2 This is the quality control module in the method for assessing the risk of lung cancer recurrence in the present invention.

[0028] Figure 3 This is the protein matrix construction module in the method for assessing the risk of lung cancer recurrence in the present invention.

[0029] Figure 4 This is the model construction module in the method for assessing the risk of lung cancer recurrence in the present invention.

[0030] Figure 5 This is the peak map information of the mass spectrometry RAW file data in the embodiment of the present invention.

[0031] Figure 6 This is the iRT data judgment in the embodiment of the present invention.

[0032] Figure 7 This is the MS1 MassAccuracy data judgment in the embodiment of the present invention.

[0033] Figure 8 This is the MS2 MassAccuracy data judgment in the embodiment of the present invention.

[0034] Figure 9 This is to perform data quality control again on the ion information intensity data obtained by the Spectronaut software in the embodiment of the present invention.

[0035] Figure 10 This is the schematic diagram of the data format after the sample data is corrected in the embodiment of the present invention.

[0036] Figure 11 This is the schematic diagram of the protein expression matrix in the embodiment of the present invention.

[0037] Figure 12 This is the ROC curve of the recurrence risk model constructed for 350 cases of stage I lung cancer samples.

[0038] Figure 13 This is the ROC curve of the recurrence risk model constructed for 150 cases of stage I lung cancer samples.

[0039] Figure 14 This is the survival curve of the model constructed by the method of the present invention, combined with the DFS (disease-free survival) data. Detailed implementation manners

[0040] The following provides a comprehensive description of the solution of the present invention. The described implementation cases are the most preferred implementation manners in the present invention, but the present invention is not limited to the following embodiments.

[0041] Example

[0042] 1. Data preprocessing: Import 10 pre-set reference standard product.raw files into the Spectronaut software, and then import the.raw files generated by high-performance liquid chromatography-mass spectrometry into the Spectronaut software; Select the FASTA file of Swissprot_Homo pre-set internally, select the pre-set database LungCancer_lib database, use the pre-set method named Umbrella in settingscheama, and start data extraction; After the mass spectrometry RAW file is opened with a specific software, the data is as Figure 5 shown, which is the peak map information.

[0043] 2. Data quality control: Perform data quality analysis according to the QCspreadlist; Judge whether the liquid phase detection error meets the range according to the sample iRT data; Judge whether the mass spectrometry detection error meets the range according to the MS1 / MS2 MassAccuracy data; Judge whether the sample injection volume meets the standard according to the total TIC intensity, protein, and peptide identification number; Judge whether there is blood contamination in the sample according to the protein and peptide identification number; The iRT data judgment is as Figure 6 shown; The MS1MassAccuracy data judgment is as Figure 7 shown, and the MS2MassAccuracy data judgment is as Figure 8 shown.

[0044] 3. Perform data quality control on the ion information intensity data obtained by the Spectronaut software again: Use the data quality control module to filter out the results that do not meet the F.FrgLossType type, remove the ions with F.MassAccuracyPPM >= 10 or F.MassAccuracyPPM <= -10 and ion intensity >= 1500. The ions passing through the above conditions meet the quality control standard, otherwise the ion information of this sample is removed. The data format after the previous 1 and 2 treatments is filtered as shown below, as Figure 9 shown.

[0045] 4. Sample data correction: Use the data correction method of the self-developed algorithm Umbrella software to correct the passed 3 ion intensities and construct the corrected sample data. The data format after correction is shown as Figure 10 shown.

[0046] 5. Protein data matrix construction: Obtain sample ion information through the above quality control. Retain the peptide segments with more than 3 ions under a peptide, and replace the peptide segment intensity value with the median intensity of the top 3 ions. For a protein with more than >= 1 peptide segment information, use the median intensity of the top 3 peptide segments as the final intensity information representing the protein of this sample. The protein data constructed by this method is used to construct a protein expression matrix through the recurrence and non-recurrence grouping information. The schematic of the protein expression matrix is as shown in Figure 11 shown.

[0047] 6. Model construction: Based on the existing data (protein expression matrix of 500 cases of stage I lung cancer samples with recurrence and non-recurrence), randomly select 70% of the sample data (recurrence and non-recurrence are divided into two groups according to a ratio of 7:3), and use 70% of the sample data to train and construct classification models for each algorithm. After successfully constructing the model (the AUC value of the ROC curve > 0.94), use the remaining 30% of the data to verify the model and evaluate the model effect (the AUC value of the ROC curve > 0.90). After testing each algorithm, the classification model constructed by the logistics regression algorithm has the best effect and is the final algorithm analysis model adopted.

[0048] 7. Positive judgment criterion: Through the model constructed in 6, under the condition of the best AUC value, map the model values of the recurrence sample data to -1 - 0, and map the non-recurrence sample data values to between 0 - 1. After detecting and analyzing the test samples that meet the quality control standards, if the test value finally falls into -1 - 0, it is judged as positive, otherwise it is judged as negative.

[0049]

Interpretation of test results

[0050] The method for analyzing and determining the test results is as follows:

[0051] 1. The test result of the negative control product (NC) should be negative. If a positive result is detected, there may be problems such as contamination.

[0052] 2. The test result of the positive control product (PC) should be recurrence, corresponding to positive. If a positive result is not detected, it indicates that the performance of the kit is not ideal or there is an error in the operation process, and the test result of this time is invalid.

[0053] Analyze and build a model for the stage I lung cancer mass spectrometry data according to the above method (steps 1 - 7 in the specific implementation method). Among them, a recurrence risk model (350 cases) is constructed, and the area under the ROC curve value (AUC) of the model is 0.94, as shown in Figure 12 shown, and the model has excellent discrimination effect and the model construction is successful.

[0054] Verify the model through the remaining 150 cases of stage I lung cancer samples. The area under the ROC curve value (AUC) of the model is 0.90, as shown in Figure 13As shown, the verification model has excellent discrimination effect and the verification model is successful.

[0055] The risk of first recurrence of lung cancer is evaluated through this model, and the survival curve analysis is carried out in combination with DFS (disease-free survival) data. As Figure 14 shown, for the model prediction, the upper curve is of low risk and the lower curve is of high risk. There are significant differences in DFS between the low risk and high risk predicted by the prediction classification, and the discrimination effect is obvious. It has significant benefit value for guiding clinical treatment in a timely manner for whether there is a first recurrence of lung cancer.

Claims

1. A method for assessing the risk of lung cancer recurrence, characterized in that, The following steps are involved: (1) Data preprocessing; Import the reference standard data file and the data file generated by mass spectrometry into the software for data extraction; (2) Data quality control: determine whether the test error is within the range, whether the sample injection volume meets the standard, and whether the sample is contaminated with blood; (3) Data quality control again: For the data that passed the quality control in step (2), the ion information intensity data obtained by Spectronaut software was subjected to data quality control again: using the data quality control module, the F.FrgLossType type was filtered out and the ions with F.MassAccuracyPPM>=10, or F.MassAccuracyPPM<=-10, or ion intensity>=1500 were removed. The ions that can pass the above screening pass the quality control standard, and the ion information that cannot pass is removed; (4) Sample data correction: for the data that passed the quality control in step (3), the ion intensity that passed was corrected to construct the corrected sample data; (5) Construction of protein data matrix: The sample ion information is obtained through quality control and correction, and the peptide segments with more than 3 ions under one peptide are retained. The median intensity of the top three ions is used to replace the peptide intensity value. For a protein with >= 1 peptide information, the median intensity of the peptide segment is used as the final intensity information representing the sample and protein. The protein data constructed in this way is used to construct a protein expression matrix through recurrence and non-recurrence grouping information. (6) Model construction: Based on the protein expression matrix constructed based on the recurrence and non-recurrence grouping information, some sample data are randomly selected to train and construct the classification model of each algorithm. After the model is successfully constructed, the remaining data are used to verify the model, evaluate the model effect, and select the best prediction model; (7) Positive judgment criteria: By constructing the best prediction model and the best ROC condition, the best discrimination threshold is selected to map the recurrence sample data model value between -1 and 0, and the non-recurrence sample data value is mapped between 0 and 1; after testing and analyzing the test samples that meet the quality control standards, if the test value finally falls between -1 and 0, it is judged as positive, otherwise it is judged as negative.

2. The lung cancer recurrence risk assessment method according to claim 1, wherein In step (1), the reference standard data file and the data file generated by mass spectrometry were imported into the Spectronaut software, the internal preset Swissprot_Homo FASTA file was selected, the preset database LungCancer_lib database was selected, and the preset method named Umbrella was used in the settings to start data extraction.

3. The lung cancer recurrence risk assessment method according to claim 1, wherein, In step (2), determine whether the liquid phase detection error is within the range based on the sample iRT data; determine whether the mass spectrometry detection error is within the range based on the MS1 / MS2 MassAccuracy data; determine whether the sample injection volume meets the standard based on the total TIC intensity, protein, and peptide identification numbers; and determine whether the sample is contaminated with blood based on the protein and peptide identification numbers.

4. The lung cancer recurrence risk assessment method according to claim 1, wherein In step (4), for the data qualified in step (3), the Umbrella software data correction method is used to correct the passed ion intensity, and the corrected sample data is constructed.

5. The lung cancer recurrence risk assessment method according to claim 1, wherein In step (6), according to the protein expression matrix constructed based on the recurrence and non-recurrence grouping information, 70% of the sample data is randomly selected. The recurrence and non-recurrence are divided into two groups according to the ratio of 7:

3. The 70% of the sample data is used to train and construct the classification models of each algorithm. After the model is successfully constructed, the remaining 30% of the data is used for model verification to evaluate the model effect and screen the best prediction model.

6. The lung cancer recurrence risk assessment method according to claim 1, wherein In step (6), the best prediction model is the classification model constructed by the logistics regression algorithm.

7. The lung cancer recurrence risk assessment method according to claim 1, characterized in that, The lung cancer is stage I lung cancer.

Citation Information

Patent Citations

  • Construction and application evaluation of molecular model for predicting postoperative early recurrence risk of liver cancer

    CN110577998A

  • Lung cancer multi-omics detection system

    CN113160883A