Early gastric cancer prognostic difference gene and recurrence prediction model

By screening and constructing a predictive model for early gastric cancer recurrence based on differentially expressed genes, the problem of the lack of accurate predictive models in existing technologies has been solved, enabling efficient risk assessment of early gastric cancer recurrence and the formulation of individualized treatment plans.

CN114941031BActive Publication Date: 2026-04-14PEKING UNION MEDICAL COLLEGE HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current technologies lack accurate models for predicting early gastric cancer recurrence, making it difficult to effectively guide individualized patient follow-up plans in clinical practice.

Method used

By screening differentially expressed genes (mcDEGs) that show monotonic changes in expression during the process of gastritis → low-grade intraepithelial neoplasia → high-grade intraepithelial neoplasia → early gastric cancer, and constructing a recurrence prediction model using multivariate Cox regression and decision tree algorithms, combined with real-time quantitative polymerase chain reaction to detect gene expression, an early gastric cancer recurrence risk assessment system was established.

Benefits of technology

It achieves high sensitivity and high specificity in predicting early gastric cancer recurrence, and can guide clinicians to adjust the frequency of patient follow-up, thereby improving treatment efficiency and patient comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114941031B_ABST
    Figure CN114941031B_ABST
Patent Text Reader

Abstract

The application relates to the establishment of an early gastric cancer recurrence prediction model. By using two batches of gene chip transcriptome data GSE130823 and GSE55696, 25 potential genes related to early gastric cancer recurrence are screened out, and an early gastric cancer recurrence prediction model based on eight genes AREG, LOC100507520, MMD, CH3L1, FOS, CCL20, CXCR2 and BATF3 is established. The model has excellent sensitivity, that is, all the patients predicted to not relapse do not relapse, and the frequency of reexamination and follow-up of the patients can be adjusted according to the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biological diagnostics, specifically to differentially expressed genes in the prognosis of early gastric cancer (EGC), and to the establishment of a recurrence prediction model. Background Technology

[0002] Gastric cancer is one of the most common tumors that significantly impact human health. Numerous studies have shown that the progression of gastric cancer follows a clear, multi-stage, progressive process: from initial inflammation and atrophy, to precancerous lesions (including LGIN and HGIN), then to early-stage gastric cancer, and further to advanced gastric cancer (AGC). Early-stage gastric cancer refers to gastric cancer with or without lymph node metastasis, confined to the gastric mucosa or submucosa. Because patients with early-stage gastric cancer generally have a longer overall survival, another prognostic indicator of clinical concern is tumor recurrence. The 5-year recurrence rate for EGC patients after ESD treatment is between 3% and 9%. The assessment and prediction of recurrence risk directly determines the subsequent follow-up plan for different patients. Therefore, an efficient and accurate recurrence prediction model can effectively guide clinicians in developing individualized patient follow-up plans and has significant clinical value.

[0003] Risk factors associated with recurrence in patients with early gastric cancer (EGC) mainly include tumor lesion characteristics (such as lesion size, pathological classification, and depth of tumor invasion) and endoscopic and surgical procedures (such as intraoperative bleeding and completeness of lesion resection). Patients with tumors larger than 20 mm are more likely to experience recurrence. Patients with poorly differentiated tumors have a higher risk of recurrence compared to those with highly differentiated tumors. Patients with shorter surgical or operative times, less intraoperative bleeding, and complete lesion resection have a relatively lower probability of recurrence. In addition, advanced age and Helicobacter pylori (Hp) infection are independent risk factors for metachronous recurrence in EGC patients. Studies on early gastric cancer recurrence are mostly focused on clinicopathological factors, with limited research at the genetic level, and a lack of accurate tumor recurrence prediction models.

[0004] Accordingly, this study used whole transcriptome data from EGC specimens to screen for genes with significantly different expression changes that exhibited monotonically increasing or decreasing expression during tumor evolution (gastritis → LGIN → HGIN → EGC)—namely, mcDEGs. Using recurrence as the outcome, three tumor recurrence prediction models were constructed using three methods: cluster analysis, risk scoring based on multivariate Cox regression, and decision tree analysis. Patient samples were prospectively collected to detect the expression of the corresponding genes, validating the predictive efficacy of the models. Furthermore, the application value of the models in clinical follow-up and personalized treatment was explored. Summary of the Invention

[0005] This study first screened differentially expressed genes (mcDEGs) exhibiting monotonically changing expression in the progression from gastritis / control tissues → low-grade intraepithelial neoplasia (LGIN) → high-grade intraepithelial neoplasia (HGIN) → EGC from two sets of gene chip transcriptome data (GSE130823 and GSE55696). Potentially tumor recurrence-associated mcDEGs were then identified using both t-tests and univariate Cox regression analysis. Next, using stage I / II patients from the external dataset GSE62254 (which includes prognostic data) as the training set, the identified mcDEGs were used as training variables to predict tumor recurrence, thus constructing a recurrence prediction model based on a decision tree algorithm. Furthermore, a prospective study was conducted to collect 16 HGIN or EGC patients as a validation set (4 relapsed and 12 non-relapsed). The expression levels of the corresponding mcDEGs were detected by quantitative real-time polymerase chain reaction (qRT-PCR) and used as a test set to input into the model to test the predictive efficacy (sensitivity, specificity, etc.) of the model.

[0006] This invention provides an early gastric cancer recurrence prediction model: Genes exhibiting monotonically changing differentially expressed (mcDEGs) expression in the gastritis / control tissue → low-grade intraepithelial neoplasia (LGIN) → high-grade intraepithelial neoplasia (HGIN) → EGC stages are screened from gene chip transcriptome data. Potentially tumor recurrence-associated mcDEGs are further screened using both t-tests and univariate Cox regression analysis. Using stage I / II patients from the external dataset GSE62254 (containing prognostic data) as the training set, the selected mcDEGs are used as training variables, and the predicted outcome is tumor recurrence. A recurrence prediction model based on a decision tree algorithm is constructed. After screening and pruning based on factors such as parameter importance and collinearity, the final gene is selected as the prediction indicator.

[0007] A gene combination for assessing the risk of recurrence in early gastric cancer, the gene combination comprising AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3.

[0008] The above gene combination was used in the preparation of a kit to assess the risk of early gastric cancer recurrence.

[0009] The reagents for detecting changes in the expression of AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3 genes are used in the preparation of a kit for predicting the risk of gastric cancer recurrence, where the gastric cancer is early-stage gastric cancer, and the reagents are PCR detection reagents.

[0010] The present invention also provides a kit for predicting the risk of early gastric cancer recurrence: characterized in that it includes reagents for detecting changes in the expression of AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3.

[0011] The present invention also provides an apparatus, system and / or model for determining early gastric cancer recurrence, which includes assessments of AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3.

[0012] The present invention also provides a gene that can be used to independently predict the risk of early gastric cancer recurrence: any of the following genes can be selected: FOS, AREG, SNCA, MMD, CHI3L1, KCNMB4, CHN1, BATF3, LOC100507520.

[0013] A kit for predicting the risk of early gastric cancer recurrence includes reagents for detecting changes in the expression of one or more genes: FOS, AREG, SNCA, MMD, CHI3L1, KCNMB4, CHN1, BATF3, LOC100507520, and AP1G1.

[0014] The present invention also provides an apparatus, system and / or model for determining early gastric cancer recurrence, which includes assessments of FOS, AREG, SNCA, MMD, CHI3L1, KCNMB4, CHN1, BATF3, LOC100507520 and AP1G1.

[0015] The gene and combination model of the present invention has excellent sensitivity, that is, all patients predicted to be non-relapsed did not relapse. The clinical significance is that the frequency of follow-up examinations for these patients can be adjusted accordingly. Attached Figure Description

[0016] Figure 1. Research flowchart: mcDEGs: differentially expressed genes with monotonic changes; HGIN: high-grade intraepithelial neoplasia; EGC: early gastric cancer.

[0017] Figure 2. Screening of differentially expressed genes with monotonic changes in the GSE130823 and GSE55696 datasets.

[0018] Figure 3. Forest plot of univariate Cox regression analysis of differentially expressed genes with monotonic changes.

[0019] Figure 4. Classification results of the external dataset based on the decision tree prediction model (tree diagram)

[0020] The first row of numbers in each ellipse represents the recurrence status, with 0 indicating no recurrence and 1 indicating recurrence; the second row of numbers is the Gini coefficient; and the third row of numbers is the percentage of patients in that category out of the total.

[0021] Figure 5 ROC curves predicted by a decision tree model from an external dataset

[0022] Figure 6 The validation set uses the ROC curves predicted by the decision tree model. Detailed Implementation

[0023] Example 1: The Process of This Study

[0024] This study first screened differentially expressed genes (mcDEGs) exhibiting monotonically changing expression in the progression from gastritis / control tissues → low-grade intraepithelial neoplasia (LGIN) → high-grade intraepithelial neoplasia (HGIN) → EGC from two sets of gene chip transcriptome data (GSE130823 and GSE55696). Potentially tumor recurrence-associated mcDEGs were then identified using both t-tests and univariate Cox regression analysis. Next, using stage I / II patients from the external dataset GSE62254 (which includes prognostic data) as the training set, the identified mcDEGs were used as training variables to predict tumor recurrence, thus constructing a recurrence prediction model based on a decision tree algorithm. Furthermore, a prospective validation set of 16 HGIN or EGC patients (4 relapsed, 12 non-relapsed) was collected. The expression levels of corresponding mcDEGs were detected by quantitative real-time polymerase chain reaction (qRT-PCR) and used as the test set input into the model to test the model's predictive efficacy (sensitivity, specificity, etc.). The flowchart of this study is shown below. Figure 1 As shown.

[0025] Example 2: Basic clinical information of patients included in the study

[0026] The first batch of gene chip specimens included 94 samples. The study subjects were patients diagnosed with LGIN, HGIN, or EGC at the Department of Gastroenterology, Peking Union Medical College Hospital, between 2011 and 2015. The test results were stored in the Gene Expression Comprehensive Database, accession number GSE130823. The second batch of gene chip specimens included 77 samples. The study subjects were patients diagnosed with LGIN, HGIN, EGC, and gastritis at the Department of Gastroenterology, Peking Union Medical College Hospital, between March 2010 and May 2013, accession number GSE55696. The third batch of validation set included 16 patients and 32 samples. These patients were those who visited the Department of Gastroenterology, Peking Union Medical College Hospital, between January 2018 and June 2019 and had regular follow-up visits. They had clear biopsy results within the past year. The acquisition and pathological interpretation of biopsy specimens were the same as the first two batches of samples. A total of 16 patients (6 with HGIN and 10 with EGC) and 32 samples were ultimately included. Four patients experienced recurrence, while 12 remained relapse-free. Recurrence criteria were based on the "Consensus Opinion on Screening and Endoscopic Diagnosis and Treatment of Early Gastric Cancer in China" (Changsha Edition, 2014). Specimens were tested using a LightCycler 480 qRT-PCR instrument (Roche, Switzerland).

[0027] Example 3: Training and Validation of a Risk Score Prediction Model Based on Multifactor Cox Regression Analysis

[0028] In the initial screening of mcDEGs, univariate Cox regression analysis was used to identify mcDEGs significantly associated with relapse as the outcome, based on data from stage I / II patients in the external dataset GSE62254. Further correlation tests were performed on each mcDEG using the R language corrplot package (0.88), and the glmnet package (4.1.1) was used to include relapse-related mcDEGs in LASSO regression analysis, excluding unnecessary or multicollinear genes. Multivariate Cox regression analysis was then performed on the remaining genes to determine whether they significantly affected relapse, and a formula was constructed to calculate the patient's risk score. The risk score calculation formula used in this study is as follows: Risk score = ∑(XJ * coefJ), where XJ is the gene expression level of the mcDEGs included in the multivariate Cox regression analysis after normalization, and coefJ is the coefficient of the corresponding gene in the multivariate Cox regression analysis. Based on the constructed formula, a risk score was calculated for each patient, and the optimal cut-off value was determined using X-tile software. Patients with risk scores above the cut-off value were assigned to the high-risk recurrence group, and those with risk scores below the cut-off value were assigned to the low-risk recurrence group. Finally, the Log-rank test and Kaplan-Meier survival analysis were used to compare whether there was a significant difference in recurrence outcomes between the two groups, and the model grouping was compared with the actual recurrence situation to calculate the accuracy of the model's predictions.

[0029] After training the risk scoring formula using an external dataset and defining the critical value, a validation set was constructed. The expression of the corresponding mcDEGs was input, and the risk scores of each patient were calculated. The survival curves of patients in the high-risk relapse group and the low-risk relapse group were compared to see if there were significant differences. The sensitivity and specificity of the model were also calculated.

[0030] Example 4: Training and Validation of a Recurrence Prediction Model Based on Cluster Analysis

[0031] To obtain the most accurate mcDEGs combination, a traversal approach was adopted, exhaustively performing cluster analysis on all permutations and combinations of mcDEGs, calculating the classification accuracy, and selecting the mcDEGs combination with the highest accuracy as the input parameter for the final cluster analysis model. After obtaining the normalized gene expression values ​​of confirmatory samples, a single confirmatory sample was selected and mixed with the previous training samples to construct a new cluster. Using the gene expression value of the mcDEGs combination with the highest accuracy as a parameter, unsupervised cluster analysis was performed to obtain the classification of the patient corresponding to that sample in the model. All confirmatory samples were examined one by one to observe whether their classification by the cluster model matched the actual relapse situation, and the final sensitivity and specificity of the model were calculated.

[0032] The formulas for calculating accuracy, sensitivity, and specificity are as follows:

[0033] Accuracy = (TP + TN) / (TP + FP + TN + FN)

[0034] Sensitivity = TP / (TP + FN)

[0035] Specificity = TN / (FP+TN).

[0036] TP, TN, FP, and FN refer to the true positive rate, true negative rate, false positive rate, and false negative rate, respectively.

[0037] Example 5: Training and Validation of a Recurrence Prediction Model Based on Decision Trees and Random Forests

[0038] The study first employed the R language rpart package (4.1.15) for decision tree analysis. The input variable parameter was the expression level of mcDEGs obtained from the initial screening. Data from stage I / II patients in the GSE62254 external dataset was used as the training set, and data from 16 confirmatory sample patients were used as the validation set. The corresponding decision tree model was constructed, and the confusion matrix of the validation set was output. The ROC curve was plotted using the ROC language ROCR package (1.0.11), the area under the ROC curve was calculated, and the sensitivity and specificity of the model were determined.

[0039] This study used the t-test for comparisons between continuous variables and Fisher's exact test or chi-square analysis for comparisons between non-continuous variables. The significance criterion was a Z3 value < 0.05. Multiple comparisons were corrected using the False Discovery Rate (FDR) method.

[0040] Example 6: Results of Differential Gene Screening

[0041] Based on the GSE130823 dataset, 75 genes with significantly increased expression and monotonically increasing expression were identified in the diseased tissues, while 4 genes with monotonically decreasing expression were identified. Based on the GSE55696 dataset, 40 genes with significantly increased expression and monotonically decreasing expression were identified in the diseased tissues, while 4 genes with monotonically decreasing expression were identified. The intersection of the differentially expressed genes from the two datasets yielded 32 genes; the union yielded 91 genes. Figure 2 As shown.

[0042] After merging the two sets of genes, clinical data of stage I / II patients from the external dataset GSE62254, which includes prognostic data, were selected. With relapse as the outcome, univariate Cox regression analysis was performed on each mcDEG. The results showed that 21 genes were independent influencing factors associated with patient relapse. The forest plotted was as follows: Figure 3 As shown.

[0043] To minimize the possibility of omissions in gene screening, patient information from the same external dataset was selected. Patients were divided into relapse and non-relapse groups based on their relapse status. A row-wise t-test was performed on each of the 91 genes obtained earlier. The results showed that 22 genes showed significant differences in expression between the two groups, of which 18 genes were consistent with those obtained from univariate Cox regression screening.

[0044] Based on the combined results of univariate Cox regression analysis and t-test, a total of 25 mcDEGs that were significantly associated with early gastric cancer recurrence were selected (see Table 1). Therefore, these genes were used as input genes for subsequent model training and validation.

[0045] Table 1. mcDEGs that showed significant p-values ​​in univariate Cox regression analysis or t-tests on external datasets.

[0046]

[0047] GO and KEGG analyses were performed on these 25 mcDEGs. GO analysis showed that the main enriched functions were related to immune regulation, involving a wide range of innate immune functions. Among them, genes such as CTSC, SNCA, PLA2G7, and S100A8 play effects in multiple physiological functions. KEGG analysis showed that four genes—CXCR2, TNFSF15, IL13RA2, and CCL20—are located in cytokine interaction pathways. Additionally, S100A8, CCL20, and FOS are all involved in the IL-17 signaling pathway.

[0048] Example 7: Expression of mcDEGs in confirmatory samples

[0049] After completing the mcDEGs screening, this study prospectively collected lesion tissue specimens and adjacent mucosal specimens from 16 patients with HGIN or EGC (4 with recurrence and 12 without recurrence), and used qRT-PCR technology to detect the expression of 25 mcDEGs.

[0050] Example 8: Training and validation results of the prediction model built based on unsupervised clustering analysis algorithm

[0051] Using the expression levels of the selected mcDEGs as input variables, and the expression levels and relapse rates of corresponding genes in stage I / II patients from the external dataset GSE62254 as the training set, unsupervised clustering analysis based on the Ward.D algorithm was performed to divide patients into two groups. Log-rank tests were then conducted on the two groups, and Kaplan-Meier survival curves were plotted. In cases where there was a significant difference in survival between the two groups, the model accuracy was calculated as the proportion of true positive patients (classified as high-risk but actually experiencing relapse) to true negative patients (classified as low-risk but actually not experiencing relapse) relative to the total sample.

[0052] Since the prediction results of cluster analysis models depend on the number and combination of genes selected, an exhaustive approach was adopted to perform cluster analysis on all permutations and combinations of the 25 mcDEGs to compare the accuracy of the models. A total of 33,554,431 permutations and combinations were found.

[0053] Cluster analysis, survival analysis, and accuracy calculation were performed on 33,554,431 gene combinations. The results showed that the model accuracy gradually increased as the number of genes increased from a single gene. However, once the number of genes reached a certain level, the model accuracy not only failed to improve with further increases in gene count but also showed a downward trend. Comparing the accuracy of the 33,554,431 gene combinations, the three gene combinations with the highest accuracy were selected: two combinations of 10 genes each and one combination of 11 genes, both with an accuracy of 77.17% (see Table 2).

[0054] Table 2. Gene combinations included in the three optimal clustering analysis models.

[0055]

[0056] Example 9: Training and Validation Results of Prediction Model Based on Decision Tree

[0057] Decision tree analysis was performed using the RPART package in R. The expression values ​​of 25 differentially expressed genes with monotonic changes selected from stage I / II patients in the GSE62254 dataset were used as prediction variables, and relapse outcome was used as the classification result for model training. To avoid overfitting, the parameter `minsplit` was set to 10, and the remaining parameters used their default values. The importance scores of the 25 genes in the decision tree analysis are shown in Table 3. The top five most important genes were MMP12, AREG, CCL20, CHI3L1, and FOS.

[0058] Table 3. Importance scores of 325 monotonic differentially expressed genes in the decision tree model.

[0059]

[0060] After screening and pruning based on factors such as parameter importance and collinearity, the model ultimately retained only eight genes: AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3, as the final predictive indicators. The resulting classification dendrogram is shown below. Figure 4 As shown.

[0061] The ROC curve was plotted based on the trained model, and the results are as follows: Figure 5 As shown, the calculated AUC is 0.895, indicating that the model training effect is good.

[0062] The model was validated using gene expression and relapse rates of 16 confirmatory patients as the validation set. The predicted relapse rates were compared with the actual relapse rates. Results showed that among 12 non-relapsed patients, 5 were predicted as relapsed and 7 were predicted as non-relapsed, with a misclassification rate of 41.7%. All 4 relapsed patients were correctly predicted, with a misclassification rate of 0% (as shown in Table 4). The overall sensitivity of the model was 100%, the specificity was 58.3%, and the AUC value was 0.792 (as shown in Table 4). Figure 6 (As shown).

[0063] This study used transcriptome data from two batches of gene chips, GSE130823 and GSE55696, to screen 25 genes potentially associated with early gastric cancer recurrence. A recurrence prediction model based on eight genes—AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3—was established. Sixteen patients were included as a validation set, and the expression levels of these genes were detected using qRT-PCR. The tumor recurrence prediction model trained using machine learning algorithms showed good predictive efficacy in both the training and validation sets. The prediction model trained using a decision tree algorithm achieved a sensitivity of 100%, a specificity of 58.3%, and an area under the curve (AUC) of 0.792 in the validation set. This suggests that these differentially expressed genes, whose expression shows monotonic changes at different stages of gastric cancer evolution, can predict the risk of tumor recurrence in patients. Predictive models built on machine learning algorithms can uncover the complex potential relationship between genes and tumor recurrence, possessing excellent predictive efficacy. They can provide guidance for clinicians in developing individualized follow-up plans for EGC patients. In future explorations, decision tree mapping of the expression levels of eight genes—AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3—can be used to preliminarily predict the recurrence probability of patients. Because this model has excellent sensitivity—that is, all patients predicted not to relapse did not relapse—the clinical significance is that the frequency of follow-up visits for these patients can be adjusted accordingly.

[0064] Table 4. Validation results and confusion matrix of the prediction model based on decision tree

[0065]

[0066] A near 100% sensitivity means that after treatment, genetic testing can identify all high-risk patients for recurrence in early-stage gastric cancer. For these patients, clinicians can increase the frequency of follow-up visits and monitor for recurrence promptly using methods such as endoscopy. If recurrence occurs, it ensures timely diagnosis and treatment, improving patient survival. On the other hand, 100% sensitivity also means that patients classified as low-risk for recurrence by the model have an extremely low probability of subsequent recurrence. Clinically, HGIN or EGC patients typically undergo endoscopic follow-up every six months to one year after ESD treatment. For patients classified as low-risk for recurrence by the predictive model, extending the follow-up period can be considered, reducing the number of endoscopic re-examinations, improving patient comfort during treatment and follow-up, and reducing treatment costs, thus alleviating the financial burden on patients.

[0067] Example 10: Establishment of Independent Predictive Indicators for Early Gastric Cancer Prognosis

[0068] Further analysis of these three gene combinations revealed a high degree of gene overlap. Gene combination A and combination B differed by only one gene, while gene combination C differed from combination B by only one additional gene. This suggests a complex and non-linear relationship among these genes, indicating that they may have a stronger correlation with early gastric cancer recurrence. Therefore, further verification is needed to determine whether these highly overlapping genes can serve as independent predictors of early gastric cancer prognosis.

[0069] FOS, AREG, SNCA, MMD, KCNMB4, CHN1, BATF3, LOC100507520, MICALL2, CTSC, and AP1G1 are differentially expressed genes included in multiple prediction models. SNCA belongs to the synuclein family and is often associated with neurological diseases such as Parkinson's and Alzheimer's. MMD is a monocyte tomacrophage differentiation-associated gene. CHN1 encodes a GTPase activator protein, primarily involved in neurotransmission. BATF3 encodes an AP-1 family transcription factor involved in regulating dendritic cell differentiation within the immune system. LOC100507520 is a non-coding RNA with limited research reports. MICALL2 encodes a cytoskeleton regulatory protein. CTSC, or cathepsin C, has been shown in multiple studies to promote the progression and metastasis of tumors such as breast cancer and liver cancer. AP1G1 is a gamma-adaptin protein belonging to the large subunit family of adaptor complexes. However, there is still a lack of research on the correlation between the above-mentioned genes and the prognosis of early gastric cancer.

[0070] Immunohistochemistry was used to detect the expression of FOS, AREG, SNCA, MMD, CHI3L1, KCNMB4, CHN1, BATF3, LOC100507520, and AP1G1 in 40 patients with recurrent EGC. Quantitative real-time polymerase chain reaction (qRT-PCR) was used to detect the mRNA expression of each gene in cancer tissues. Chi-square test or Fisher's test was used to analyze the correlation between gene expression and clinicopathological factors. Univariate analysis assessed the association between clinicopathological factors, including FOS, AREG, SNCA, MMD, KCNMB4, CHN1, BATF3, LOC100507520, MICALL2, CTSC, and AP1G1, and EGC recurrence. Multivariate analysis identified independent risk factors for EGC recurrence.

[0071] Trizol was used for mRNA extraction, and the StepOnePlus real-time PCR system and SYBR Green method were used for cDNA synthesis and quantitative PCR. GAPDH was used as an internal control for 2−ΔΔ CT. The mean mRNA level in adjacent tumor tissue was set to 1.0, and other mRNA levels were normalized to this baseline. The primer sequences for GAPDH and individual genes are as follows:

[0072] GAPDH upstream primer: 5'- tggagaatgagaggtgggatg -3';

[0073] GAPDH downstream primer: 5'- gagcttcacgttcttgtatctg -3';

[0074] FOS upstream primer: 5'-actctcatagtttcttccctaag-3';

[0075] FOS downstream primer: 5'- ttccactgagggcttgggc -3';

[0076] AREG upstream primer: 5'-cacatcttttacgcttgtcaa-3';

[0077] AREG downstream primer: 5'-caggatgagtggctgtccc-3';

[0078] SNCA upstream primer: 5'- tgtattcatgaaaggac -3';

[0079] SNCA downstream primer: 5'- ttcaggttcgtagtcttga -3';

[0080] MMD upstream primer: 5'-atgtgtgatagaatggttatctatt-3';

[0081] MMD downstream primer: 5'-gaacacagcctttatact-3';

[0082] CHI3L1 upstream primer: 5'-gttgatgataagttcacgggt-3';

[0083] CHI3L1 downstream primer: 5'-tgtaataatatttaattgtgc-3';

[0084] KCNMB4 upstream primer: 5'- ctcggcttgtttctcatcatct-3';

[0085] KCNMB4 downstream primer: 5'- ttgggtaagagaacttgcgc -3';

[0086] CHN1 upstream primer: 5'- agtattatggaagagag-3';

[0087] CHN1 downstream primer: 5'- agccatcttgacatcttcaat -3';

[0088] BATF3 upstream primer: 5'- tcctgcagaggagcgtcg-3';

[0089] BATF3 downstream primer: 5'- ttcatcggggcaagcagccg -3';

[0090] upstream primer for LOC100507520: 5'-tgagaactccgagatgcattag-3';

[0091] LOC100507520 downstream primer: 5'- gctagttgagatgtcgatagtgc -3';

[0092] AP1G1 upstream primer: 5'- ttacagacaaacgcattggctatt-3';

[0093] AP1G1 downstream primer: 5'-agctatgaatgatatattagcac-3'.

[0094] Table 5. Comparison of mcDEGs expression differences between lesion tissues and control tissues in patients with early gastric cancer recurrence.

[0095]

[0096] Studies have shown that the mRNA levels of FOS, AREG, SNCA, MMD, CHI3L1, KCNMB4, CHN1, BATF3, LOC100507520, and AP1G1 in early gastric cancer tissues are significantly higher than those in adjacent normal tissues, and they are highly expressed only in cancerous tissues. The expression of these genes is significantly associated with early gastric cancer recurrence (P=0.002), and these genes can be identified as independent biomarkers for the diagnosis of early gastric cancer (P=0.001). sequence list <110> Peking Union Medical College Hospital, Chinese Academy of Medical Sciences <120> Early gastric cancer prognostic differential genes and recurrence prediction models <160> twenty two <170> SIPOSequenceListing 1.0 <210> 1 <211> twenty one <212> DNA <213> Artificial sequence <400> 1 tggagaatga gaggtgggat g 21 <210> 2 <211> twenty two <212> DNA <213> Artificial sequence <400> 2 gagcttcacg ttcttgtatc tg 22 <210> 3 <211> twenty three <212> DNA <213> Artificial sequence <400> 3 actctcatag tttcttccct aag 23 <210> 4 <211> 19 <212> DNA <213> Artificial sequence <400> 4 ttccactgag ggcttgggc 19 <210> 5 <211> twenty one <212> DNA <213> Artificial sequence <400> 5 cacatctttt acgcttgtca a 21 <210> 6 <211> 19 <212> DNA <213> Artificial sequence <400> 6 caggatgagt ggctgtccc 19 <210> 7 <211> 17 <212> DNA <213> Artificial sequence <400> 7 tgtattcatg aaaggac 17 <210> 8 <211> 19 <212> DNA <213> Artificial sequence <400> 8 ttcaggttcg tagtcttga 19 <210> 9 <211> 25 <212> DNA <213> Artificial sequence <400> 9 atgtgtgata gaatggttatctatt 25 <210> 10 <211> 18 <212> DNA <213> Artificial sequence <400> 10 gaacacagcc tttatact 18 <210> 11 <211> twenty one <212> DNA <213> Artificial sequence <400> 11 gttgatgata agttcacggg t 21 <210> 12 <211> twenty one <212> DNA <213> Artificial sequence <400> 12 tgtaataata tttaattgtg c 21 <210> 13 <211> twenty two <212> DNA <213> Artificial sequence <400> 13 ctcggcttgt ttctcatcat ct 22 <210> 14 <211> 20 <212> DNA <213> Artificial sequence <400> 14 ttgggtaaga gaacttgcgc 20 <210> 15 <211> 17 <212> DNA <213> Artificial sequence <400> 15 agtattatgg aagagag 17 <210> 16 <211> twenty one <212> DNA <213> Artificial sequence <400> 16 agccatcttg acatcttcaa t 21 <210> 17 <211> 18 <212> DNA <213> Artificial sequence <400> 17 tcctgcagag gagcgtcg 18 <210> 18 <211> 20 <212> DNA <213> Artificial sequence <400> 18 ttcatcgggg caagcagccg 20 <210> 19 <211> twenty two <212> DNA <213> Artificial sequence <400> 19 tgagaactcc gagatgcatt ag 22 <210> 20 <211> twenty three <212> DNA <213> Artificial sequence <400> 20 gctagttgag atgtcgatag tgc 23 <210> twenty one <211> twenty four <212> DNA <213> Artificial sequence <400> twenty one ttacagacaa acgcattggc tatt 24 <210> twenty two <211> twenty three <212> DNA <213> Artificial sequence <400> twenty two agctatgaat gatatattag cac 23

Claims

1. The use of reagents for detecting changes in the expression of AREG, LOC100507520, MMD, CHI3L1, FOS, CCL20, CXCR2, and BATF3 in the preparation of a kit for assessing the risk of recurrence in early gastric cancer: characterized in that... This reagent includes a combination of primers, the primer sequences of which are shown below: FOS upstream primer: 5'-actctcatagtttcttccctaag-3'; FOS downstream primer: 5'-ttccactgagggcttgggc-3'; AREG upstream primer: 5'-cacatcttttacgcttgtcaa-3'; AREG downstream primer: 5'-caggatgagtggctgtccc-3'; MMD upstream primer: 5'-atgtgtgatagaatggttatctatt-3'; MMD downstream primer: 5'-gaacacagcctttatact-3'; CHI3L1 upstream primer: 5'-gttgatgataagttcacgggt-3'; CHI3L1 downstream primer: 5'-tgtaataatatttaattgtgc-3'; BATF3 upstream primer: 5'-tcctgcagaggagcgtcg-3'; BATF3 downstream primer: 5'-ttcatcggggcaagcagccg-3'; LOC100507520 upstream primer: 5'-tgagaactccgagatgcattag-3'; LOC100507520 downstream primer: 5'-gctagttgagatgtcgatagtgc-3'; This reagent also contains an internal reference primer: GAPDH upstream primer: 5'-tggagaatgagaggtgggatg-3'; GAPDH downstream primer: 5'-gagcttcacgttcttgtatctg-3'.

Citation Information

Patent Citations

  • Markers for cancer detection

    US20120039811A1