Kit for diagnosing non-small cell lung cancer as well as preparation method and application thereof
The combination of 5 biomarkers was screened through proteomics, and combined with machine learning methods, an optimization model for the diagnosis of non-small cell lung cancer was developed, solving the problem of insufficient diagnostic sensitivity and specificity in the prior art, and achieving efficient and accurate diagnostic effects.
Patent Information
- Application Number
- CN202510318392.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art early diagnosis methods for non-small cell lung cancer are insufficient in sensitivity and specificity, and are costly and technically difficult, resulting in many patients being in the advanced stage of the disease at the time of diagnosis.
Eight candidate biomolecular markers were screened through proteomic analysis, and a combination of 5 biomarkers (FERMT2, PRPF19, SART3, CAMP, LDHB) was further screened. Optimized models were developed in combination with machine learning methods to predict patients' risk of disease and achieve convenient and efficient diagnosis of non-small cell lung cancer.
It improves the diagnostic accuracy of non-small cell lung cancer, achieves a prediction accuracy of 95%, provides a golden window for intervention for high-risk populations, and reduces the financial burden of low-risk patients.
Smart Images

Figure CN120161201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine, and particularly to a kit for diagnosing non-small cell lung cancer, a preparation method and uses thereof. Specifically, it includes a kit for distinguishing healthy people from non-small cell lung cancer patients by simultaneously detecting differential expressions of FERMT2, PRPF19, SART3, CAMP, and LDHB proteins and uses thereof. Background Art
[0002] Non-small cell lung cancer (NSCLC) is the most common type of lung cancer and one of the most common and highly lethal cancers worldwide. NSCLC usually has no obvious symptoms in the early stage. As the tumor grows, symptoms such as cough, hemoptysis, dyspnea, and chest pain may occur. Early diagnosis is of great significance for the survival and prognosis of patients. However, current early diagnosis methods for NSCLC mainly rely on imaging examinations and tissue biopsies. These detection methods have problems such as insufficient sensitivity and specificity, as well as relatively high costs and technical difficulties, resulting in many patients being in the advanced stage of the disease when diagnosed with NSCLC. At the same time, the detection of protein biomarker levels has shown great potential in the early diagnosis of cancer. In particular, this is a non-invasive method for detecting protein expression through blood tests, which can provide a simple and reliable approach for the early diagnosis of NSCLC.
[0003] With the development of molecular biology, methods based on detecting the expression levels of specific proteins have gradually become effective means for early diagnosis. Traditional protein biomarkers for NSCLC mainly include cytokeratin fragment (CYFRA21-1), carcinoembryonic antigen (CEA), neuron-specific enolase (NSE), squamous cell carcinoma antigen (SCC-Ag), etc. However, traditional biomarkers have problems of limited specificity and sensitivity and are clinically difficult to play a role in specifically diagnosing NSCLC. On the other hand, with the progress of proteomics technology, a new technical platform has been provided for finding new and effective protein biomarkers. Proteomics takes the proteome of cells, tissues, and organisms as the research object, and then can analyze the changes and differences in protein content at the global level, effectively identifying markers with relatively high specificity and sensitivity that can be applied to the diagnosis of NSCLC. In summary, it is necessary to utilize proteomics technology to develop new markers for diagnosing NSCLC and related detection kits. Summary of the Invention
[0004] The present invention provides a biomolecular marker, a system and an application for diagnosing non-small cell lung cancer, belonging to the field of biomedicine. By proteomic analysis of relevant samples of patients with non-small cell lung cancer, people with benign lung diseases and healthy people, the present invention finds 8 biomarkers that can predict whether the test subject has non-small cell lung cancer; combines these new markers, and further screens a combination containing 5 biomarkers with the most accurate prediction results; then, based on this biomarker combination, uses machine learning methods to develop and optimize the model formula, calculates the risk of the patient's disease, and then realizes the convenient and efficient diagnosis of non-small cell lung cancer, grasps the golden window period of intervention for high-risk groups, reduces the economic burden for low-risk patients, and meets the clinical needs. The technical roadmap of the present invention is shown in Figure 1 。
[0005] To achieve the above object, the technical solutions adopted by the present invention are as follows:
[0006] On the one hand, the present invention provides the use of a biomolecular marker for preparing a reagent for judging whether a patient has non-small cell lung cancer. The biomolecular marker composition comprises one or more of the proteins FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1, and the proteins FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1 respectively contain the sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7 and SEQ ID NO:8 in the sequence listing.
[0007] Since the detection specificity and sensitivity of the molecular markers currently used for detecting non-small cell lung cancer are relatively low, the present invention aims to develop a batch of new biomolecular markers. In this example, through high-throughput proteomic analysis, all proteomic information of non-small cell lung cancer, benign lung diseases and healthy people was obtained. The proteomes of non-small cell lung cancer (n = 26), benign lung diseases (n = 4) and healthy people (n = 20) were obtained by ultra-high performance liquid chromatography ion trap mass spectrometry data-dependent acquisition and data-independent acquisition techniques, and quantitative analysis was performed on 30 protein molecules with the largest differences among non-small cell lung cancer, benign lung diseases and healthy people, and 8 candidate protein molecule markers (FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1) were preliminarily screened.
[0008] Furthermore, the biomolecular marker comprises one or more of FERMT2, PRPF19, SART3, CAMP, and LDHB proteins, and the FERMT2, PRPF19, SART3, CAMP, and LDHB respectively contain the sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5 in the sequence listing.
[0009] Furthermore, the biomolecular marker comprises FERMT2, PRPF19, SART3, CAMP, and LDHB proteins, and the FERMT2, PRPF19, SART3, CAMP, and LDHB respectively contain the sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5 in the sequence listing.
[0010] In order to further obtain a biomarker with better detection effect, the present invention respectively performs ROC curve analysis on the above 8 biomarkers and their combinations, and finds that the combination of FERMT2, PRPF19, SART3, CAMP, and LDHB has the highest sensitivity, accuracy, and AUC value. Therefore, this combination is preferentially selected to predict whether the detection object has non-small cell lung cancer.
[0011] On the other hand, the present invention provides a biomolecular marker composition, which comprises two or more of FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24, and AEBP1 proteins, and the FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24, and AEBP1 proteins respectively contain the sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, and SEQ ID NO:8 in the sequence listing.
[0012] On the other hand, the present invention provides a kit for detecting non-small cell lung cancer, and the kit comprises reagents for detecting the above-mentioned 5 biomolecular markers (FERMT2, PRPF19, SART3, CAMP, and LDHB).
[0013] The reagents include but are not limited to antibodies, peptides, nucleic acid aptamers, and their derivatives.
[0014] On the other hand, the present invention provides a system for diagnosing non-small cell lung cancer, which system comprises a data analysis module; the data analysis module is used for analyzing the detection values of a biomarker combination in a test sample, and the biomarker combination comprises FERMT2, PRPF19, SART3, CAMP, and LDHB proteins, and the FERMT2, PRPF19, SART3, CAMP, and LDHB respectively contain the sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5 in the sequence listing.
[0015] Machine learning adopts different statistical probabilities and data processing techniques, and uses the advantages of computers to perform reinforcement learning from data sets to detect patterns that are difficult to identify in data information. The methods of machine learning are often used in cancer diagnosis and detection, which can help improve researchers' understanding of the occurrence and development processes of cancer. By means of machine learning, the accuracy and specificity of different biomarkers in predicting the occurrence and development of non-small cell lung cancer can be evaluated. By combining the determination of multiple indicators and patient characteristics and testing with different models, the accuracy of cancer detection and diagnostic prediction can be greatly improved. Therefore, based on the levels of 5 protein biomarker molecules, the present invention performs machine learning, constructs 5 evaluation models (Logistic Rsgression model, Random Forest model, XGBoost model, SVM model, and KNN model) for evaluating the risk of non-small cell lung cancer, and validates different models in samples of non-small cell lung cancer patients and healthy individuals to evaluate the accuracy and sensitivity of different models in predicting non-small cell lung cancer, and selects the optimal model. The results show that the prediction result of the constructed Random Forest model is the most accurate; and a linear regression equation is used to fit the diagnostic result of the Random Forest, and an approximate formula is output, which formula can predict the probability that the test subject has non-small cell lung cancer. Specifically, the formula: p = 0.445 + 0.007×FERMT2 + 0.122×PRPF19 + 0.097×SART3 + 0.107×CAMP + 0.018×LDHB, where p is the probability that the test subject has non-small cell lung cancer, and FERMT2, PRPF19, SART3, CAMP, and LDHB respectively represent the corresponding protein concentrations. If p > 0.5, the test result is considered positive, otherwise it is negative. Based on this, the present invention constructs the system.
[0016] Furthermore, the test sample of the system is the serum, plasma, whole blood, secretion, and tissue sample of the test subject; the system is used to detect the presence, relative abundance, or concentration of biomarkers in the sample. In some specific embodiments, an ELISA kit is used to detect the concentration of biomarkers in the sample.
[0017] The presence, absence, or high or low content of the biomarker is a relative concept. For example, the content of the specific biomarker is compared relative to that of non-small cell lung cancer patients or healthy individuals as a reference. In some cases, the concentrations of the above five biomarkers in patients with non-small cell lung cancer are higher than those in healthy individuals, and this increase is statistically significant, such as a significant or highly significant elevation. When a benign lung disease develops into non-small cell lung cancer, the content of these specific biomarkers can be used as a comparison standard for all patients. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the content of the biomarker changes, this change may be a relative increase or a relative decrease, and the difference in this relative increase or relative decrease is significantly different, and of course, it can also be highly significantly different. Therefore, no matter what method is used for detection, a pre-specified value (cut-off value) can be used as a standard. If the value is higher than this value, it is considered that the content has changed, and such a result can be used for prediction or diagnosis.
[0018] In addition, the biomarker described in the present invention can be obtained by detecting the content of the biomarker in a sample by any known method, including but not limited to liquid chromatography, gas chromatography, mass spectrometry, LC-MS, GC-MS, CC-MS, LC-MS-MS, NMR, immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, etc. In other words, as long as a method can accurately detect the content of the biomarker in a sample, it can be used for diagnosing non-small cell lung cancer. As long as the content in the sample is accurately detected, the detected value can be used in the above system for predicting whether a patient has non-small cell lung cancer, and the probability of detecting non-small cell lung cancer can be calculated. It can be understood that the detection is for the molecular biomarker in an individual's sample, and then the detected value is substituted into the equation fitted by the random forest model and the linear regression equation to obtain the disease probability. Such prediction or diagnosis occurs at a certain time, and such detection can be continuous detection, and the progress of the disease can be inferred from the change in the content of certain substances.
[0019] Further, the data analysis module includes an operation equation, which calculates the probability that the detected object has non-small cell lung cancer through an operation method. In some specific embodiments, the calculation formula is: p = 0.445 + 0.007×FERMT2 + 0.122×PRPF19 + 0.097×SART3 + 0.107×CAMP + 0.018×LDHB, where p is the probability that the detected object has non-small cell lung cancer, and FERMT2, PRPF19, SART3, CAMP, and LDHB respectively represent the corresponding protein concentrations. If p > 0.5, the test result is considered positive; otherwise, it is negative.
[0020] Further, the operation equation is derived from a random forest model trained and constructed with the detection values of biomarkers of known samples.
[0021] For the data of the same batch of lung cancer patients, the present invention adopts the same machine learning methods (including the imputation algorithm GAIN and five-fold cross-validation), and verifies the effect of the combination of five biomarkers, namely FERMT2, PRPF19, SART3, CAMP, and LDHB, in predicting non-small cell lung cancer in multiple models. However, in two or more runs, it is found that there are certain differences in the area under the receiver operating characteristic curve (AUC) obtained for the combination of the same 5 biomarkers (AUC≈0.960 and AUC≈0.953 are obtained before and after, etc.). For this difference, the reasons can be scientifically explained from the following aspects:
[0022] 1. Randomness in model training: Algorithms such as random forest and XGBoost often use random sampling (such as random feature selection and random sample selection) to construct each tree during model training. These random processes will result in different specific structures of the model, thus causing slight fluctuations in evaluation metrics (such as AUC). Cross-Validation will also introduce a shuffling operation during the fold division (if shuffle = True is set), or reallocate samples to each fold during different runs, thus bringing subtle differences in the results.
[0023] 2. Internal random initialization of the imputation method (GAIN): The GAIN model is based on a generative adversarial network (GAN), and its training process contains multiple random factors such as random initialization of parameters and random noise input. Even for the same missing data, after multiple imputation runs, the generated imputation values may also have slight differences, thus affecting the subsequent machine learning model.
[0024] 3. Floating - point arithmetic and parallelization: When performing large - scale floating - point operations on a CPU or GPU, thread scheduling and acceleration instructions may cause the operation order to be not exactly the same, thus introducing cumulative rounding errors. This often manifests as occasional differences in the second or third decimal place of metrics such as AUC between different runs. These minor fluctuations at the numerical level are acceptable in statistics and engineering applications and do not affect the scientific judgment of the overall performance of the model.
[0025] Therefore, a slight difference in the AUC value before and after in a machine - learning model is normal and predictable in the field of machine learning and does not affect the scientific judgment of the overall performance of the model. As long as this difference remains within a small range (such as around 0.01 - 0.02, or less than a certain acceptable threshold), it indicates that the algorithm is stable and has good reproducibility.
[0026] Similarly, when evaluating different biomarker combinations, there are also similar fluctuations at the numerical level. On the one hand, this "random difference" reflects the robustness and stability of the algorithm in multiple repeated experiments; on the other hand, if strictly consistent results are required, it can be controlled by fixing all random number seeds, unifying hardware and parallelization settings, etc. By comparing and analyzing the AUC of five biomarker combinations and single features within a unified framework, the present invention proves that this combination can achieve a high and stable AUC in multiple rounds of experiments, highlighting the value and reliability of the biomarker combination in the auxiliary diagnosis of lung cancer patients.
[0027] Furthermore, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0028] The beneficial effects of the present invention include:
[0029] 1. Using proteomics technology, 8 biomolecular markers and a combination containing 5 markers have been screened for the prediction or diagnosis of non - small - cell lung cancer;
[0030] 2. The present invention discovers that in practical applications, the combination containing 5 markers used in conjunction with a random forest model has the highest prediction accuracy, up to 95%.
[0031] 3. The present invention can provide a new method for the diagnosis of clinical non - small - cell lung cancer, a new idea for the identification of proteomics - based biomarkers, a new means for the preoperative evaluation and drug selection of non - small - cell lung cancer, and a new tool for optimizing the treatment plan of non - small - cell lung cancer. Description of the Drawings
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 : Research ideas of the present invention;
[0034] Figure 2 : Analysis of differential protein expression between non-small cell lung cancer and lung benign diseases or healthy populations;
[0035] Figure 3 : KEGG pathway enrichment analysis of differentially expressed proteins between non-small cell lung cancer and lung benign diseases or healthy populations, where the left figure is non-small cell lung cancer VS lung benign diseases, and the right figure is non-small cell lung cancer VS healthy populations;
[0036] Figure 4 : Abundances of 8 differentially expressed protein molecules, where HC (healthy control): FERMT2: 50.67; PRPF19: 0.17; SART3: 3257.72; CAMP: 227.53; LDHB: 39.45; SEC31A: 4437.96; SLC25A24: 1558.09; AEBP1: 0.23; LC (lung cancer patients): FERMT2: 127.95; PRPF19: 0.41; SART3: 5621.23; CAMP: 1415.66; LDHB: 56.82; SEC31A: 2304.53; SLC25A24: 918.00; AEBP1: 0.14;
[0037] Figure 5 : ROC curve analysis based on a single protein molecular marker;
[0038] Figure 6 : ROC curve analysis based on multiple protein molecular markers, where feature represents the protein marker, and the AUC values shown in the figure are the maximum AUC values obtained under the same number of markers;
[0039] Figure 7 : Parameters of the non-small cell lung cancer diagnosis model (Logistic Rsgression model, Random Forest model, and XGBoost model) based on 5 protein markers;
[0040] Figure 8: Parameters of the non-small cell lung cancer diagnosis model (SVM model and KNN model) based on 5 protein markers;
[0041] Figure 9 : ROC curves of 5 models for diagnosing non-small cell lung cancer based on 5 protein markers;
[0042] Figure 10 : Feature importance (SHAP) of the random forest model for diagnosing non-small cell lung cancer based on 5 protein markers;
[0043] Figure 11 : Comparison of the importance of individual protein marker molecules in the random forest model for diagnosing non-small cell lung cancer based on 5 protein markers. Detailed implementation mode
[0044] The present invention will be further elaborated in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. The embodiments are only used to explain the present invention and are not used to limit the scope of the present invention; all other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0045] Unless otherwise specified, the test methods used in the following examples are all conventional methods; the materials, reagents, etc. used, unless otherwise specified, are reagents and materials that can be obtained from commercial channels.
[0046] Example 1: Obtaining candidate molecular markers for diagnosing non-small cell lung cancer
[0047] Based on the data-dependent acquisition (DDA) technology of ultrahigh performance liquid chromatography orbitrap mass spectrometry, precursor ions with high ion intensity are selected and preferentially fragmented to obtain more accurate and abundant mass spectrometry data. Instead, the data-independent acquisition (DIA) technology divides the entire full-scan range of the mass spectrometry into several windows, and selects, fragments, and detects all ions in each window at high speed and cyclically, so as to obtain all fragment information of all ions in the sample without omission and difference. Using these two mainstream data acquisition methods (DDA and DIA) simultaneously can perform high-throughput and high-precision quantification on the proteomes of non-small cell lung cancer, benign lung diseases, and healthy people, providing the possibility for the discovery of molecular markers for non-small cell lung cancer. Therefore, in this embodiment, non-targeted high-throughput proteomics technology and high-resolution mass spectrometry platform DDA and DIA technologies are used to analyze the differences at the protein level between non-small cell lung cancer patients, benign lung diseases, and healthy people, in order to preliminarily obtain molecular markers that can be used to diagnose non-small cell lung cancer. The specific operations are as follows:
[0048] 1.1 Sample processing and proteome sequencing: First, proteins are extracted from the serum or plasma samples of non-small cell lung cancer patients, benign lung disease patients, and healthy control groups. The specific operations are as follows: Use RIPA lysis buffer to extract the proteins in the sample, and add protease and phosphatase inhibitors to prevent degradation; after protein extraction, reduce with dithiothreitol (DTT), and then perform alkylation treatment with iodoacetamide; and use the Thermo Scientific TM BCA Protein Assay Kit of Thermo Fisher Scientific to measure the protein concentration to ensure that the extracted protein quality is qualified. The digestion step uses Trypsin for enzymatic digestion, and the resulting peptide segments are stored at -80 °C. Proteomics analysis uses the Q Exactive TM Plus mass spectrometer of Thermo Fisher Scientific, combined with Easy nLC TMThe 1200 liquid chromatography system uses two techniques, data-dependent acquisition (DDA) and data-independent acquisition (DIA), for high-resolution quantitative proteomics analysis. In the DDA mode, precursor ions with high ion intensity are preferentially selected for fragmentation, while in the DIA mode, the mass spectrometry information of all ions in the sample is comprehensively captured by dividing the mass spectrometry scanning range into multiple windows. Data analysis uses MaxQuant software, and protein identification and quantitative analysis are performed through the UniProt database. The number of proteins identified in different populations by proteomics technology is shown in Table 1.
[0049] 1.2 Data analysis: Protein biomarker molecules with a differential expression fold change greater than 1.5, a significant p-value less than 0.05, and specific functions were screened out in non-small cell lung cancer (n = 26) compared with patients with benign lung diseases (n = 4) or healthy individuals (n = 20) for the next step of research ( Figure 2 ); The biological functions of the above differential proteins were analyzed by KEGG enrichment pathway analysis, and the roles of the differential proteins in inducing non-small cell lung cancer were identified ( Figure 3 ).
[0050] Table 1 Number of proteins identified by proteomics technology in different populations
[0051] Group Sample size Number of identified proteins Non-small cell lung cancer 26 1,978 Benign lung diseases 4 1,176 Healthy population 20 1,842
[0052] As can be seen from Table 1, most (>1,000) proteins in different background samples were identified, indicating that the proteome sequencing technology used in this example has a comprehensive coverage and high sensitivity, and the measured results lay a foundation for subsequent analysis. Then, from the results of differential expression analysis of proteins ( Figure 2 and Table 2), it can be seen that compared with the difference in protein expression between patients with non-small cell lung cancer and patients with benign lung diseases (non-small cell lung cancer VS benign lung disease group), more proteins were up-regulated or down-regulated between patients with non-small cell lung cancer and healthy individuals (non-small cell lung cancer VS healthy individual group). Specifically, compared with healthy individuals, the expression levels of 236 proteins were increased and 189 were down-regulated in patients with non-small cell lung cancer; at the same time, compared with healthy individuals, 152 proteins showed significant changes in expression levels in both patients with non-small cell lung cancer and patients with benign lung diseases, including 63 up-regulated and 89 down-regulated. By Figure 3It can be seen that the proteins with differential expression in the non-small cell lung cancer vs. pulmonary benign disease group are mainly related to the complement pathway (complement and coagulation cascades). This may be because the immune response and inflammatory response of patients with non-small cell lung cancer (NSCLC) are more active. As an important part of the immune system, the complement system will be activated in the tumor microenvironment, promoting inflammation, tumor growth and metastasis. On the other hand, the proteins with differential expression in the non-small cell lung cancer vs. healthy population group are mainly related to phagosome and myosin (complement and coagulation cascades). This may be because in non-small cell lung cancer, the phagosome function of tumor cells is involved in immune escape and intracellular signal transduction. At the same time, the changes in myosin may be closely related to the migration, invasiveness of tumor cells and cell movement in the tumor microenvironment.
[0053] In summary, through proteomics technology, a series of proteins with differential expression in different background samples were screened in this example, providing a theoretical basis for the development of new molecular markers for the diagnosis of non-small cell lung cancer.
[0054] Table 2 Protein differential expression analysis
[0055]
[0056] Example 2: Screening of molecular markers for the diagnosis of non-small cell lung cancer
[0057] According to the analysis results of the proteome sequencing data in Example 1, 30 protein molecules with the greatest differences in non-small cell lung cancer, pulmonary benign diseases and healthy populations were selected and targeted verified in non-small cell lung cancer patients (n = 26) and healthy populations (n = 20). That is, plasma samples of non-small cell lung cancer patients and healthy populations were collected respectively, and the expression levels of 30 protein molecules in the plasma samples were measured and compared by ELISA method to screen out 8 specific protein marker molecules with significant differences in different populations (Table 3 and Figure 4), wherein the steps for measuring protein abundance by ELISA are as follows: First, select a commercial ELISA kit (produced by CUSABIO, ABBEXA, ELK Biosciences, specifically, abx520130 Human FERMT2 ELISA Kit; abx382476 Human PRPF19 ELISA Kit; abx383030 Human SART3 ELISA kit; abx545001 Human SLC25A24 ELISA Kit; abx383089 Human SEC31A ELISA Kit; CSB-EL004476HU Human Antimicrobial Peptide (CAMP) ELISA Kit; CSB-EL012841HU Human L-Lactate Dehydrogenase B Chain (LDHB) ELISA Kit; EK7769 Human AEBP1 ELISA Kit) and prepare reagents such as pre-coated plates or self-coated microplates, standards, detection antibodies, enzyme conjugates, and substrate solutions according to the instructions; After taking serum or plasma samples and centrifuging to remove impurities, dilute appropriately according to the kit requirements; Add the standards and the samples to be tested into the plate wells respectively, seal and incubate at 37°C for 12 hours to allow the target protein to fully bind to the coated antibody; Subsequently, wash the plate 3 - 5 times with PBS or TBS buffer containing Tween-20 to remove unbound substances; Then add biotin- or enzyme-labeled detection antibodies and incubate at 37°C to allow the detection antibodies to bind to the target protein in the sample and wash the plate again; Next, add horseradish peroxidase (HRP) streptavidin or corresponding detection reagents to bind the enzyme, wash the plate, add TMB substrate for color development, and incubate in the dark at room temperature or 37°C until the color development is sufficient, and then terminate the reaction with an acidic solution (such as 2M H2SO4); Finally, read the optical density value at 450 nm (or 630 nm) on an enzyme-linked immunosorbent assay (ELISA) reader, draw a standard curve based on the standards, calculate the protein concentration in the sample, and perform statistical analysis in combination with the experimental data to evaluate the association between protein abundance and diseases such as non-small cell lung cancer. The results show that the concentrations of 8 proteins in non-small cell lung cancer patients are all significantly increased compared with those in healthy individuals (HC), and this result is consistent with the proteomic data. Therefore, the following 8 protein molecular markers are used for further research.
[0058] Table 3 Protein Biomarkers for the Diagnosis of Non-Small Cell Lung Cancer
[0059]
[0060]
[0061] Table 4 ROC Curve Analysis of Single Protein Biomarkers for the Diagnosis of Non-Small Cell Lung Cancer
[0062] Protein molecule AUC value Specificity Sensitivity FERMT2 0.574 0.56 0.67 PRPF19 0.648 0.63 0.71 SART3 0.741 0.80 0.67 CAMP 0.686 0.75 0.65 LDHB 0.630 0.70 0.60 SEC31A 0.345 0.45 0.32 SLC25A24 0.409 0.50 0.38 AEBP1 0.489 0.52 0.46
[0063] Next, in this embodiment, ROC curve analysis was performed on the above-mentioned 8 selected protein molecules and their combinations. For the AUC value of a single molecular marker, logistic regression was used as a single-feature model and appropriate parameters were set to ensure the stability and rationality of the model. Specifically, the regularization method used L2 regularization (penalty='l2') to prevent overfitting of the model; the optimization algorithm selected solver='lbfgs', which is suitable for small and medium-sized datasets and can effectively optimize the log-likelihood function; the maximum number of iterations (max_iter = 1000) was set to 1,000 to ensure convergence and avoid incomplete fitting caused by insufficient iterations; in addition, we set random_state = 42 to fix the random seed and ensure the reproducibility of the experimental results. For the composition, the Logistic Rsgression model was used as the benchmark classification model, but multiple features were combined on the input data and standardized to optimize the generalization ability of the model. In terms of parameters, L2 regularization (penalty='l2') was still used to reduce the risk of overfitting of the model; the optimization algorithm selected solver='liblinear', which is suitable for high-dimensional sparse data and can efficiently optimize the logistic regression model; considering that the combination of multiple features may increase the computational complexity, the maximum number of iterations (max_iter = 2000) was set to 2,000 to ensure that the complex model can fully converge; in addition, the reciprocal C = 1.0 of the regularization strength remained the default value, and if the model effect needs to be optimized, it can be adjusted appropriately according to the experimental situation; similarly, random_state = 42 was set to fix the random seed to ensure the reproducibility of the experimental results. These parameter settings ensure that the logistic regression model can effectively measure the AUC in both single-feature and multi-feature cases, and improve the stability and prediction performance of the model. The results are shown in Table 4 and Figures 5 - 6 . It should be understood that since there are hundreds of combinations that can be obtained by arranging and matching the 8 proteins, Figure 6 only shows the ROC curves of the best combinations (i.e., the combinations with the highest AUC values) when the number of detected protein molecular markers is the same, Figure 6 and the protein types in the best combinations shown in
[0064] Table 5 Best protein marker combinations composed of different numbers of proteins
[0065]
[0066] From Figure 5 it can be seen that among the single markers, the AUC values are arranged in descending order as: SART3 > CAMP > PRPF19; from Figure 6It can be seen that the AUC values of the biomarker combinations are generally higher than those of single biomarkers, indicating that in order to improve the accuracy of diagnosing non-small cell lung cancer, it is advisable to preferentially detect the concentrations of multiple protein molecules simultaneously. Among them, the combination of FERMT2, PRPF19, SART3, CAMP, and LDHB has the highest AUC value, so this combination is preferentially selected as the biomarker combination for diagnosing non-small cell lung cancer. Secondly, it is the combination of FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, and SLC25A24 (7 features). It should be noted that when the number of biomarker types in the combination is <5, as the number of biomarkers increases, the AUC value generally increases. However, when the number of biomarker types >5, the AUC value decreases as the number of biomarkers increases, indicating that the higher the number of biomarkers in the combination, the higher the AUC value is not necessarily the case. This may be because there are cross, overlap, or conflict phenomena among different molecular biomarkers.
[0067] Example 3: Screening of the evaluation model for diagnosing non-small cell lung cancer
[0068] Based on the exploration results of Example 2, machine learning (including the imputation algorithm GAIN and five-fold cross-validation) was performed on the expression levels of the above 5 proteins in the data of lung cancer patients (n = 141) with the levels of the above 5 protein biomarker molecules (i.e., FERMT2, PRPF19, SART3, CAMP, and LDHB proteins) in non-small cell lung cancer patients as the basis. Five evaluation models (Logistic Rsgression model, Random Forest model, XGBoost model, SVM model, and KNN model) for diagnosing non-small cell lung cancer were constructed respectively. The parameter settings of the models are shown in Table 6, and the performance of the constructed models was evaluated (see Figures 7 - 9 ).
[0069] Table 6 Selection of hyperparameters of machine learning models
[0070]
[0071]
[0072] It can be seen from the results that the accuracy, sensitivity, and AUC value of the Random Forest model are higher than those of the other four models, indicating that the Random Forest regression model has the best diagnostic effect. Then, linear regression was used to fit the diagnostic results of the random forest, and the approximate formula was output: p = 0.445 + 0.007×FERMT2 + 0.122×PRPF19 + 0.097×SART3 + 0.107×CAMP + 0.018×LDHB, where p is the probability of the test subject having non-small cell lung cancer, and FERMT2, PRPF19, SART3, CAMP, and LDHB represent the corresponding protein concentrations respectively.
[0073] Next, this embodiment also explored the importance of each component in the above-mentioned random forest model. The TreeExplainer in the shap model was used for evaluation. First, the random forest model was trained using the standardized dataset, and the hyperparameters of the model were set as n_estimators = 100, max_depth = 10, min_samples_split = 2, min_samples_leaf = 1, max_features = "sqrt", criterion = "gini", bootstrap = True. After training, the SHAP values were calculated using TreeExplainer to quantify the importance of each feature to the model decision-making, and a bar chart of feature importance was drawn to visually display the contribution of different features in the random forest classification task. The specific results are shown in Figures 10 - 11 . The results show that PRPF19 has the highest importance weight, followed by FERMT2, CAMP, and SART3, while the relative contribution of LDHB in the model is lower. This indicates that PRPF19 and FERMT2 may be key biomarkers in predicting non-small cell lung cancer. This result provides a theoretical basis for subsequent biological verification and model optimization; further, although the evaluation results of different evaluation methods are inconsistent, no matter which evaluation method is used, the indicators of these two proteins, PRPF19 and FERMT2, occupy a dominant position in the random forest model, which is inconsistent with the evaluation effect of a single biomarker ( Figure 5 ), which may be because of the different types of models.
[0074] To further verify the accuracy of the diagnostic results of the constructed model, the above five models were also used to diagnose new clinically known non-small cell lung cancer patient samples (n = 50) and healthy population samples (n = 50). If p > 0.5, the test result is considered positive; otherwise, it is negative. The specific diagnostic results are shown in Table 7.
[0075] Table 7 Diagnostic results of different models
[0076]
[0077] As can be seen from Table 7, in practical applications, the diagnostic results of the Random Forest model have the highest accuracy, followed by the XGBoost model. This result is consistent with the ROC curve analysis result. Moreover, when processing this batch of samples, the AUC value of the Random Forest model is also as high as 0.93, and the specificity and sensitivity are 0.90 and 0.85 respectively, further indicating that the Random Forest model is most suitable for constructing a model for diagnosing non-small cell lung cancer based on the combination of FERMT2, PRPF19, SART3, CAMP, and LDHB. In other words, in order to maximize the accuracy of diagnosing non-small cell lung cancer, a random forest model constructed from the combination of FERMT2, PRPF19, SART3, CAMP, and LDHB markers is preferably selected.
[0078] Example 4: A method for diagnosing non-small cell lung cancer
[0079] Combined with the exploration results of Examples 1 to 3, this example provides a method for diagnosing whether a person has non-small cell lung cancer. The steps of the method are as follows:
[0080] 4.1 Collect samples
[0081] The samples include serum, plasma, whole blood, secretions, and tissue samples. In this example, plasma is preferably used.
[0082] 4.2 Detect the concentrations of FERMT2, PRPF19, SART3, CAMP, and LDHB proteins in the samples
[0083] In this example, the ELISA method is used to detect the concentrations of the above 5 proteins. The specific experimental steps are the same as those described in Example 2. It should be understood that in addition to the ELISA method, other methods can also be used to detect the protein concentrations, such as mass spectrometry. In other words, no matter which method is used to detect the proteins, as long as the measured protein concentrations are accurate, the detection results are applicable to the method provided in this example.
[0084] 4.3 Substitute the measured concentrations of FERMT2, PRPF19, SART3, CAMP, and LDHB proteins into p = 0.445 + 0.007×FERMT2 + 0.122×PRPF19 + 0.097×SART3 + 0.107×CAMP + 0.018×LDHB, where p is the probability that the test subject has non-small cell lung cancer. If p > 0.5, it is considered that the test subject may have non-small cell cancer; otherwise, it is considered that the test subject does not have non-small cell cancer.
[0085] The foregoing has shown and described the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes that fall within the meaning and scope of the equivalent elements of the claims in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A use of a biomolecular marker for preparing a reagent for determining whether a patient has non-small cell lung cancer, characterized in that: The biomolecule marker composition comprises one or more of FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1 proteins, and the FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1 proteins respectively contain the sequences shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7 and SEQ ID NO: 8 in the sequence list.
2. The use according to claim 1, characterized in that The biomolecular markers include one or more of FERMT2, PRPF19, SART3, CAMP, and LDHB proteins, and the FERMT2, PRPF19, SART3, CAMP, and LDHB contain the sequences shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, and SEQ ID NO: 5 in the sequence list, respectively.
3. The use according to claim 2, characterized in that The biomolecular markers include FERMT2, PRPF19, SART3, CAMP, and LDHB proteins, and the FERMT2, PRPF19, SART3, CAMP, and LDHB contain the sequences shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, and SEQ ID NO: 5 in the sequence table, respectively.
4. A biomolecule marker composition, characterized in that: The molecular marker composition comprises two or more of FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1 proteins, and the FERMT2, PRPF19, SART3, CAMP, LDHB, SEC31A, SLC25A24 and AEBP1 proteins respectively contain the sequences shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7 and SEQ ID NO: 8 in the sequence list.
5. A kit for detecting non-small cell lung cancer, characterized in that: The kit comprises a reagent for detecting the biomolecule marker for the use according to any one of claims 1 to 3.
6. A system for diagnosing non-small cell lung cancer, characterized in that: The system comprises a data analysis module; the data analysis module is used to analyze the detection value of the biomolecular marker composition in the test sample, the biomolecular marker composition comprises FERMT2, PRPF19, SART3, CAMP, LDHB proteins, and the FERMT2, PRPF19, SART3, CAMP and LDHB respectively contain the sequences shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4 and SEQ ID NO: 5 in the sequence list.
7. The system according to claim 6, characterized in that The detection samples of the system are serum, plasma, whole blood, secretions and tissue samples of the detection object; the system is used to detect the presence or relative abundance or concentration of biomarkers in the samples.
8. The system according to claim 6, characterized in that The data analysis module includes an operation equation, which calculates the probability that the detected object suffers from non-small cell lung cancer through an operation method.
9. The system according to claim 8, characterized in that The operation equation is derived from a random forest model constructed by training the detection values of biomarkers of known samples.
10. The system according to claim 6, characterized in that The system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values of the biomarkers; the data input interface is used to input the detection values of the biomarkers, and the data output interface is used to output the prediction results.