Method for predicting EGFR gene mutation in non-small cell lung cancer based on endogenous formaldehyde and serum tumor marker combined kit

By analyzing the serum tumor markers and endogenous formaldehyde concentrations in NSCLC patients, using the triple method of CEA+SSCC-Ag+FA, combined with IBM SPSS27.0 statistical software to draw the ROC curve, successfully predicting EGFR gene mutations in NSCLC, solving the problem of low prediction accuracy in the existing technology and improving prediction accuracy.

CN119985755AInactive Publication Date: 2025-05-13BAYANNUR CITY HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131230.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively predict EGFR gene mutations in non-small cell lung cancer (NSCLC), and second-generation sequencing is expensive and the results are not easily repeated.

Method used

By analyzing the serum tumor markers and endogenous formaldehyde concentrations in NSCLC patients, the ROC curve was drawn with IBM SPSS27.0 statistical software to predict EGFR gene mutations.

Benefits of technology

Accurate prediction of EGFR gene mutations was achieved. The AUC of the triple method reached 0.77 and the P value was 0.001, which improved the prediction accuracy and provided a basis for EGFR gene detection in patients with lung adenocarcinoma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119985755A_ABST
    Figure CN119985755A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of lung cancer, in particular to a method for predicting EGFR gene mutation in non-small cell lung cancer based on an endogenous formaldehyde and serum tumor marker combined kit. Comprising the following steps: 1, detecting the concentration of endogenous formaldehyde FA in a serum sample; 2, detecting the concentration of tumor markers in the serum specimen by using a chemiluminescence method, wherein the tumor markers comprise squamous cell carcinoma antigen SCC-Ag and carcino-embryonic antigen CEA; and 3, inputting the concentrations of endogenous formaldehyde FA, a serum tumor marker CEA and a squamous epithelial cell carcinoma antigen SCC-Ag by adopting IBM SPSS 27.0 statistical software, drawing an ROC curve for diagnosing the EGFR mutation state, and predicting EGFR gene mutation according to the ROC curve. According to the invention, the prediction of the accuracy of the next-generation sequencing result by serology is realized through a CEA + SSCC-Ag + FA triple method, tests show that the AUC of the CEA + SSCC-Ag + FA triple method reaches 0.77, the P value is 0.001, the prediction accuracy is the highest, and a basis is further provided for EGFR gene detection of lung adenocarcinoma patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lung cancer, and in particular to a method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers. Background Art

[0002] With the advent of the era of precision medicine, lung cancer patients have taken individualized targeted drug therapy with genes as therapeutic targets to a new level. Tumor markers are specific molecules produced and released by cancer cells in the human body and can be detected in the patient's body. The detection of tumor markers is a simple, fast and inexpensive method that can be quickly and accurately detected by the laboratories of major hospitals, including CEA, CA125, CA199, NSE and CYFRA21-1. Studies have found that CEA serum levels are significantly increased in NSCLC, while CYFRA21-1 expression levels are low. Other studies have shown that CA199 levels are significantly correlated with EGFR mutations.

[0003] In recent years, the results of precision medicine using next-generation sequencing have become an important reference indicator for targeted drug use. On July 10, 2018, the "Expert Consensus on the Application of Next-Generation Sequencing Technology in Precision Medicine Diagnosis of Tumors" systematically elaborated on the quality requirements of NGS technology, clinical tumor-related NGS testing content, sample processing, sequencing process, data management, informatics analysis, result report interpretation and consultation, and discussed various issues in the clinical application of NGS.

[0004] The body itself can produce endogenous formaldehyde. Endogenous formaldehyde can be produced by the metabolism of DNA, RNA, reducing sugars, aldehydes, lipids, etc., or by oxidative stress and enzymatic reactions, such as the deamination of primary amines catalyzed by semicarbazide sensitive amine oxidase (SSAO). Long-term exposure to formaldehyde may cause RNA cross-linking and DNA cross-linking, leading to carcinogenesis in different organs. And the results of preliminary laboratory studies have shown that the expression of endogenous formaldehyde is increased in lung cancer and breast cancer tissues. The laboratory can use high-performance liquid chromatography (HPLC) to measure the concentration of endogenous formaldehyde in serum, urine, and tissues. If the concentration of endogenous formaldehyde related to the EGFR mutation status can be determined. However, endogenous formaldehyde has not been reported as a standard for predicting the accuracy of mutation sites. Therefore, the researchers speculated whether endogenous formaldehyde can also be used as an indicator, combined with serum tumor markers, to analyze the correlation with mutation sites (such as EGFR), in order to obtain a two- or three-way method to predict the accuracy of mutation site detection.

[0005] In short, it is urgent for the second-generation sequencing to seek new methods and new angles for quality control to ensure safety and effectiveness. Relevant management and supervision should be multi-pronged, from source to process to clinical result analysis, so that the output data is more accurate and reliable, providing an additional guarantee for the selection of targeted drugs. Summary of the invention

[0006] In order to comprehensively solve the above problems, the present invention aims to predict the accuracy of mutation site detection by analyzing serum tumor markers and endogenous formaldehyde concentrations in NSCLC patients to obtain a triple method, and further provide a basis for EGFR gene testing in patients with lung adenocarcinoma.

[0007] In order to achieve the above object, the present invention provides a method for predicting EGFR gene mutation in non-small cell lung cancer based on endogenous formaldehyde and serum tumor marker combined kit, comprising:

[0008] Step 1: Detect endogenous formaldehyde in serum samples;

[0009] Step 2: Detect tumor markers in serum samples using chemiluminescence;

[0010] Step 3: Using IBM SPSS27.0 statistical software, input the concentrations of endogenous formaldehyde FA, serum tumor marker CEA, and squamous cell carcinoma antigen SCC-Ag, draw the ROC curve for diagnosing EGFR mutation status, and predict EGFR gene mutation based on the curve.

[0011] Preferably, step 1 comprises:

[0012] Step 1.1: Centrifuge the serum sample at 4°C and take the supernatant;

[0013] Step 1.2: Add 10% trichloroacetic acid to the supernatant, centrifuge at 4°C, and take the supernatant;

[0014] Step 1.3: After the supernatant of step 1.2, 2,4-dinitrophenylhydrazine (1 g / L) and acetonitrile are mixed evenly, the mixture is placed in a water bath;

[0015] Step 1.4: Centrifuge at 4°C for 10 min, 13,000 r / min, and take the supernatant for HPLC analysis.

[0016] Preferably, the centrifuge speed in step 1 is 13000 rpm / min and the time is 10 min.

[0017] Preferably, the centrifugation time in step 1.2 is 30 min, and the centrifuge speed is 13000 r / min.

[0018] Preferably, in step 1.3, the supernatant is 0.4 ml, 2,4-dinitrophenylhydrazine (1 g / L) is 0.1 ml, and acetonitrile is 0.5 ml.

[0019] Preferably, in step 1.3, the mixture is placed in a 60° C. water bath for 30 min.

[0020] Preferably, in step 3, the value of the area under the ROC curve AUC is obtained based on IBM SPSS27.0 statistical software, thereby realizing the prediction of EGFR gene mutation.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] The current second-generation sequencing is expensive and the results are not easy to be repeated, while endogenous formaldehyde and tumor serum markers are easy to detect in the laboratory. The CEA+SSCC-Ag+FA triple method of the present invention realizes the prediction of the accuracy of the second-generation sequencing results by serology. Through experiments, the AUC of the CEA+SSCC-Ag+FA triple method of the present invention reached 0.77, the P value was 0.001, and the prediction accuracy was the highest, thereby realizing the prediction of EGFR gene mutations. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0024] In the attached picture:

[0025] Figure 1 This is the ROC curve diagram of the prediction of EGFR mutation by serum CEA, SCC-Ag and FA in the present invention. DETAILED DESCRIPTION

[0026] The following combination Figure 1 The preferred embodiments of the present invention are described. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0027] Embodiment 1:

[0028] 1. Research subjects

[0029] A total of 210 patients with NSCLC pathologically confirmed to be diagnosed from 2019 to 2022 were collected. According to the inclusion and exclusion criteria, 186 patients with lung adenocarcinoma who had both serum tumor markers and genetic testing data of pathological tissue specimens or blood using the NGS method were selected as research subjects. The basic clinical information of the research subjects and clinical data such as the results of various auxiliary examinations including serum tumor markers and genetic testing reports were collected. The serum endogenous formaldehyde concentration was detected by HPLC, and the data were comprehensively analyzed and studied.

[0030] 1.1 Inclusion criteria

[0031] (1) Inpatients admitted between 2019 and 2022, with the patient's informed consent and approval by the Medical Ethics Committee;

[0032] (2) The pathological diagnosis of the surgical specimen is lung adenocarcinoma, and it is reviewed and verified by two or more pathologists. The diagnosis and staging standards of lung cancer are unified according to the latest guidelines;

[0033] 1.2 Exclusion criteria

[0034] (1) Combined with other malignant tumors;

[0035] (2) In the absence of complete clinical case data, only the surgery to obtain pathological tissue for a definitive diagnosis was used for patients with multiple hospitalizations;

[0036] (3) NGS genetic testing failed and no genetic results were obtained.

[0037] 2. Detection Method

[0038] 2.1 Genetic testing

[0039] Paraffin tissue sections of lung cancer patients were collected in the pathology department, and the sections were reviewed and verified by more than two pathologists. Through the steps of sample nucleic acid extraction, library amplification, hybridization capture, and machine sequencing, NGS technology was used to sequence genes closely related to tumors, including EGFR, ALK, ROS1, etc. The detection items included detection of point mutations (SNV), small fragment insertion and deletion (INDEL), copy number variation (CNV) and fusion (FUSION) and other mutation forms, each gene mutation type, mutation site and mutation abundance.

[0040] Based on the completion of the above preparations, the present invention provides a method for predicting EGFR gene mutation in non-small cell lung cancer based on endogenous formaldehyde and serum tumor marker combined kit, comprising:

[0041] Step 1: Detect the concentration of endogenous formaldehyde FA in the serum sample; including:

[0042] Step 1.1: Centrifuge the serum sample at 4°C for 10 min at a speed of 13,000 rpm / min, and collect the supernatant;

[0043] Step 1.2: Add 10% trichloroacetic acid to the supernatant, centrifuge at 4°C for 30 min at a speed of 13000 r / min, and take the supernatant;

[0044] Step 1.3: Take 0.4 ml of the supernatant from step 1.2, 0.1 ml of 2,4-dinitrophenylhydrazine (1 g / L) and 0.5 ml of acetonitrile, mix well, and incubate in a water bath at 60°C for 30 min;

[0045] Step 1.4: After water bath, centrifuge at 4°C for 10 min at a speed of 13,000 r / min, and take the supernatant for HPLC analysis.

[0046] Step 2: Detect tumor markers in serum samples using chemiluminescence; tumor markers include cytokeratin 19 fragment (CYFRA21-1), squamous cell carcinoma antigen (SCC-Ag), carcinoembryonic antigen (CEA), carbohydrate antigen 125 (CA125), carbohydrate antigen 199 (CA199) and carbohydrate antigen 153 (CA153).

[0047] Step 3: Using IBM SPSS27.0 statistical software, input the concentrations of endogenous formaldehyde FA, serum tumor marker CEA and squamous cell carcinoma antigen SCC-Ag, draw the ROC curve for diagnosing EGFR mutation status, predict EGFR gene mutation based on the ROC curve, and obtain the value of the area under the ROC curve AUC based on IBM SPSS27.0 statistical software. The area under the ROC curve is between 0.1 and 1. As a numerical value, it can intuitively evaluate the quality of the classifier, and the larger the value, the better.

[0048] As shown in the following experiment, the ROC area of ​​the CEA+SSCC-Ag+FA triple method of the present invention is the largest, so the CEA+SSCC-Ag+FA triple method of the present invention improves the predictive value and further realizes the prediction of EGFR gene mutation.

[0049] Specifically, IBM SPSS27.0 statistical software was used for analysis. Quantitative data that followed normal distribution were expressed as "x±s", and t-test was used for inter-group comparison; those that did not follow normal distribution were expressed as "P50 (P25, P75)", and non-parametric rank sum test was used for inter-group comparison. Qualitative data were expressed as "frequency (constituent ratio)", and chi-square test or Fisher's exact probability method was used for inter-group comparison. Samples were sorted according to the prediction results of the learner, and the samples were predicted one by one in this order as positive examples. The values ​​of two important quantities (TPR and FPR) were calculated each time, and they were plotted as horizontal and vertical coordinates respectively. Spearman correlation coefficient was used to analyze the correlation between each indicator and EGFR mutation status; the receiver operating characteristic curve (ROC curve) between each indicator and EGFR genotype was drawn, and the variables that may be factors affecting EGFR gene mutation were included in binary logistic regression analysis to evaluate the predictive value of each indicator for EGFR gene mutation. P < 0.05 was considered statistically significant.

[0050] Result analysis:

[0051] 1. Diagnostic evaluation of endogenous formaldehyde, serum tumor marker CEA, and SCC-Ag in predicting EGFR mutation status in patients with lung adenocarcinoma

[0052] EGFR mutation was grouped information, and the ROC curves of endogenous formaldehyde, serum tumor marker CEA, and SCC-Ag levels for diagnosing EGFR mutation status were analyzed and drawn. The larger the predictive value AUC, the more valuable it is, and above 0.7 has predictive value.

[0053] The area under the curve (AUC) corresponding to CEA was 0.646, and its 95% confidence interval (CI) was 0.562-0.729;

[0054] The area under the curve (AUC) corresponding to SCC-Ag was 0.294, and its 95% confidence interval (CI) was 0.574-0.838;

[0055] The area under the curve (AUC) of FA was 0.546, and its 95% confidence interval (CI) was 0.449-0.642.

[0056] The maximum AUC of FA+CEA+SCC-Ag combined detection was 0.770, with a 95% CI of 0.637-0.903, and the diagnostic value for EGFR mutation was the highest, as shown in Table 1. Figure 1 .

[0057] Table 1 Analysis of serum CEA and CA153 for the diagnosis of EGFR mutation

[0058]

[0059]

[0060] 2. Binary Logistic Regression Analysis for Predicting EGFR Gene Mutation in Patients with Lung Adenocarcinoma

[0061] Statistical results showed that CES, SCC-Ag and endogenous formaldehyde showed differences in EGFR gene mutation or not. The indicators with statistically significant differences in univariate analysis and the indicators with more significant significance in current related studies were used as independent variables, and the EGFR mutation status of patients was used as the dependent variable. Binary Logistic regression analysis was performed to analyze the influencing factors affecting the EGFR mutation status. The results showed that CEA had a good positive predictive value for EGFR mutation. CEA (P=0.024, OR=1.002, 95%CI: 1.000-1.004) was a risk factor for predicting the EGFR mutation status of lung adenocarcinoma patients, and had an important predictive value for the EGFR mutation status of lung adenocarcinoma patients. However, the predictive value of SCC-Ag (P=0.861, OR=1.004, 95%CI: 0.956-1.055) and FA (P=0.075, OR=1.029, 95%CI: 0.997-1.062) still needs to be further verified by experiments (see Table 2).

[0062] Table 2 Binary Logistic regression analysis of predictive factors for EGFR mutation

[0063]

[0064] 3. Relationship between general information and EGFR mutation in patients with lung adenocarcinoma

[0065] Among the 186 cases in this study, the youngest was 32 years old and the oldest was 90 years old, with an average age of 64.33±10.14 years old. There were 57 cases (30.65%) younger than 60 years old and 129 cases (69.35%) older than 60 years old. Among them, 76 cases were detected with EGFR gene mutation, and the EGFR gene mutation rate was 40.86%. According to EGFR mutation, they were divided into positive and negative groups. In the EGFR positive group, 49 cases were older than 60 years old and 27 cases were younger than or equal to 60 years old; in the negative group, 80 cases were older than 60 years old and 30 cases were younger than or equal to 60 years old. The average age of the EGFR positive and negative groups was 62.51±10.33 years old and 65.58±9.87 years old, respectively. There was no statistically significant difference in age between the two groups (P=0.259). There were 24 males and 52 females in the EGFR-positive group; there were 76 males and 34 females in the EGFR-negative group. The difference between the two groups was statistically significant (P=0.020). The number of gene mutations in female patients was higher than that in males. There were 53 smokers and 23 non-smokers in the EGFR-positive group, and 58 smokers and 52 non-smokers in the EGFR-negative group. The difference between the two groups in terms of smoking or not was statistically significant (P=0.020).

[0066] In the EGFR-positive group, 6 patients had distant organ and / or lymph node metastasis and 70 patients had no metastasis. In the EGFR-negative group, 9 patients had distant organ and / or lymph node metastasis and 101 patients had no metastasis. There was no significant difference between the two groups (P>0.05). See Table 3 for details.

[0067] Table 3 Relationship between EGFR mutation and clinical characteristics of patients with lung adenocarcinoma

[0068]

[0069]

[0070] 4. Analysis of differences between endogenous formaldehyde / tumor markers and EGFR positive and negative groups

[0071] Compared with the EGFR positive and negative groups, the difference in serum tumor marker CEA levels between the two groups was statistically significant (P = 0.021, respectively), and the other tumor markers were not significantly different between the two groups (P > 0.05), as shown in Table 4. Patients with high CEA levels are more likely to have EGFR mutations, as shown in Table 5.

[0072] Table 4 Differential analysis of serum tumor markers, endogenous formaldehyde and EGFR mutation in patients with lung adenocarcinoma

[0073]

[0074]

[0075] Table 5 Correlation analysis between clinical characteristics and EGFR gene expression in patients with lung adenocarcinoma

[0076]

[0077] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers, characterized in that: include Step 1: Detect the concentration of endogenous formaldehyde FA in serum samples; Step 2: Detect the concentration of tumor markers in serum samples using chemiluminescence method. Tumor markers include squamous cell carcinoma antigen SCC-Ag and carcinoembryonic antigen CEA; Step 3: Using IBM SPSS27.0 statistical software, input the concentrations of endogenous formaldehyde FA, serum tumor marker CEA, and squamous cell carcinoma antigen SCC-Ag, draw the ROC curve for diagnosing EGFR mutation status, and predict EGFR gene mutation based on the ROC curve.

2. The method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers according to claim 1, characterized in that: Step 1 includes: Step 1.1: Centrifuge the serum sample at 4°C and take the supernatant; Step 1.2: Add 10% trichloroacetic acid to the supernatant, centrifuge at 4°C, and take the supernatant; Step 1.3: After the supernatant of step 1.2, 2,4-dinitrophenylhydrazine (1 g / L) and acetonitrile are mixed evenly, the mixture is placed in a water bath; Step 1.4: After water bath, centrifuge at 4°C for 10 min and take the supernatant for HPLC analysis.

3. The method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers according to claim 2, characterized in that: The centrifuge speed in step 1.1 was 13000 rpm / min and the time was 10 min.

4. The method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers according to claim 3, characterized in that: The centrifugation time in step 1.2 is 30 min, and the centrifuge speed is 13000 r / min.

5. The method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers according to claim 4, characterized in that: In step 1.3, the supernatant is 0.4 ml, 2,4-dinitrophenylhydrazine (1 g / L) is 0.1 ml, and acetonitrile is 0.5 ml.

6. The method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers according to claim 5, characterized in that: In step 1.3, place in a 60°C water bath for 30 min.

7. The method for predicting EGFR gene mutation in non-small cell lung cancer based on a combined kit of endogenous formaldehyde and serum tumor markers according to claim 6, characterized in that: In step 3, the value of the area under the ROC curve (AUC) is obtained based on IBM SPSS 27.0 statistical software to predict EGFR gene mutations.