Reagent, kit, chip and system for detecting SNP site of CLDN18 gene and application thereof
By using a kit and system to detect SNP sites in the CLDN18 gene, and utilizing the rs6804932 site as a genetic marker, combined with genome-wide association analysis and eQTL co-localization analysis, a precise lung cancer risk assessment model was constructed. This model solved the problem of lung cancer risk stratification in non-smokers and achieved high predictive efficacy and biological mechanism support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WEST CHINA HOSPITAL SICHUAN UNIV
- Filing Date
- 2026-06-08
- Publication Date
- 2026-07-03
AI Technical Summary
Current technologies lack precise lung cancer risk stratification strategies for non-smokers, existing biomarkers have insufficient predictive efficacy and lack biological mechanism support, making it difficult to develop transferable testing products or clinical application systems.
We developed reagents, kits, chips, and systems for detecting SNP sites in the CLDN18 gene. Using the rs6804932 site as a genetic marker, we screened for lung cancer risk in East Asian and non-smoking populations using PCR amplification, gene chip, and nucleic acid sequencing technologies. We then constructed a lung cancer risk assessment model by combining genome-wide association analysis, fine mapping, and eQTL co-localization analysis.
After considering age, sex, and smoking status, the model incorporating the rs6804932 locus achieved an AUC of 0.712, demonstrating good risk prediction performance. The rs6804932 locus carrying the A allele is associated with a lower risk of lung cancer and has promising application prospects.
Smart Images

Figure CN122326754A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of tumor genetic testing technology, specifically involving reagents, kits, chips, systems, and their applications for detecting SNP sites in the CLDN18 gene. Background Technology
[0002] Lung cancer is one of the leading causes of cancer-related morbidity and mortality worldwide. Although smoking is considered the most significant risk factor for lung cancer, numerous epidemiological studies in recent years have shown a continuous increase in the incidence of lung cancer among non-smokers, particularly in Asian populations. This suggests that, in addition to environmental exposure, genetic factors play a crucial role in the development and progression of lung cancer in non-smokers. Currently, early lung cancer screening mainly relies on imaging techniques such as low-dose CT (LDCT), but this method suffers from problems such as over-screening, high false-positive rates, and high costs. Furthermore, a precise risk stratification strategy for non-smokers is lacking. Therefore, developing risk prediction methods based on genetic biomarkers is of great importance for achieving early screening and precise prevention of lung cancer. In recent years, genome-wide association studies (GWAS) have identified several susceptibility loci associated with lung cancer, but most studies focus on smoking-related lung cancer, and the explanatory power of reported loci in non-smokers is limited. In addition, existing biomarkers generally suffer from the following problems: 1) Insufficient predictive effectiveness (low AUC); 2) Lack of functional validation and biological mechanism support; 3) No transferable testing products or clinical application system has yet been formed.
[0003] The CLDN18 (Claudin 18) gene encodes a tight junction protein, an important molecule for maintaining epithelial barrier function, and is specifically expressed in lung tissue. Previous studies have suggested that CLDN18 may play an important role in the development of lung cancer, but the role of its genetic variations in predicting lung cancer risk remains unclear.
[0004] Therefore, it is urgent to screen and validate genetic markers with high predictive power, especially SNP loci with protective or risk indication effects in non-smokers in East Asian populations, in order to establish a more accurate lung cancer risk assessment system. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides reagents, kits, chips, systems, and their applications for detecting SNP sites in the CLDN18 gene.
[0006] This invention provides the application of a detection reagent for the CLDN18 gene SNP site in the preparation of kits, chips, or systems for screening lung cancer, wherein the SNP site is rs6804932.
[0007] Furthermore, when the allele at the SNP locus is A, the risk of developing lung cancer is low; when the allele at the SNP locus is G, the risk of developing lung cancer is high.
[0008] Furthermore, the detection reagent includes at least one of PCR amplification reagents, gene chip reagents, and nucleic acid sequencing reagents.
[0009] Furthermore, the detection reagent includes primers with nucleotide sequences as shown in SEQ ID NO. 1-4.
[0010] Furthermore, the screening subjects were selected from East Asian populations and / or non-smokers.
[0011] Furthermore, the screening samples include blood, lymph, cerebrospinal fluid, and lung tissue samples.
[0012] The present invention provides a kit or chip for screening lung cancer, which includes a detection reagent for the CLDN18 gene SNP site, wherein the SNP site is rs6804932.
[0013] Furthermore, when the allele at the SNP locus is A, the risk of developing lung cancer is low; when the allele at the SNP locus is G, the risk of developing lung cancer is high.
[0014] The present invention provides a system for screening lung cancer, the system comprising: a computing device for predicting lung cancer susceptibility based on sequencing results of a CLDN18 gene SNP site, wherein the SNP site is rs6804932.
[0015] Furthermore, the system includes the following modules: The test result collection module is configured to collect sequencing results of SNP sites in the CLDN18 gene; The calculation module is configured to calculate the risk of lung cancer based on the sequencing results collected from the detection results module.
[0016] Furthermore, the SNP site of the CLDN18 gene is rs6804932. In this invention, the reference genome for CLDN18 is GRCh38, and the coordinate position of rs6804932 is chromosome 3, Pos138032158.
[0017] This invention, based on genome-wide association analysis (GWAS), systematically identified single nucleotide polymorphism (SNP) sites associated with lung cancer pathogenesis. Combining fine mapping, transcriptome analysis, and eQTL co-localization analysis, it was the first to confirm that rs6804932 in the CLDN18 gene is a potential protective variant against lung cancer. The results showed that carrying the A allele at the rs6804932 site was significantly associated with a lower risk of lung cancer. Further analysis of the lung cancer risk assessment model revealed that, after considering age, sex, and smoking status, including the rs6804932 site improved the model's AUC to 0.712, demonstrating good risk prediction performance. These results indicate that the CLDN18 rs6804932 site can serve as an important genetic biomarker for predicting lung cancer risk.
[0018] Therefore, the rs6804932 site of the CLDN18 gene has promising applications as a target in the detection of genetic susceptibility to lung cancer, risk assessment, and the preparation of related testing products.
[0019] Obviously, based on the above description of the present invention, and according to common technical knowledge and conventional methods in the field, various other modifications, substitutions or alterations can be made without departing from the basic technical concept of the present invention.
[0020] The following detailed embodiments further illustrate the above-described content of the present invention. However, this should not be construed as limiting the scope of the present invention to the following examples. All technologies implemented based on the above-described content of the present invention fall within the scope of the present invention. Attached Figure Description
[0021] Figure 1 This is a quality control chart for the LCCWCH cohort sample level; where, Figure 1 'a' is a graph showing the distribution of missing samples. Figure 1 b is the distribution of the heterozygous / homozygous ratio (Het / Hom plot). Figure 1 c is the distribution of insertions plot. Figure 1 d is the distribution of the insertion / deletion ratio (Int / Del). Figure 1 e is a distribution diagram of the number of SNPs. Figure 1 f is the distribution of Ti / Tv (transformation / transversion ratio); the dashed line represents the quality control threshold determined based on the mean of all samples ± 3 standard deviations.
[0022] Figure 2This is a regional association map obtained from GWAS analysis; where, Figure 2 a is the signal result diagram of the 54-56 Mb region on chromosome 2. Figure 2 b is the signal result diagram of the locus in the 132-134.5 Mb region of chromosome 2. Figure 2 c is the signal result diagram of the locus in the 24.5-26.5 Mb region of chromosome 3. Figure 2 Image d shows the signal results for the locus in the 73.8-74.6 Mb region of chromosome 4. Figure 2 e is the signal result diagram of the 0.8-1.8 Mb region on chromosome 5. Figure 2 f is the signal result diagram of the 53-56 Mb region site on chromosome 10.
[0023] Figure 3 This is a regional association map obtained from GWAS analysis; where, Figure 3 a is the signal result diagram of the locus in the 103.4-104.4 Mb region of chromosome 10. Figure 3 b is the signal result diagram of the locus in the 67-68.5 Mb region of chromosome 11. Figure 3 c is the signal result diagram of the locus in the 41.8-42.6 Mb region of chromosome 15. Figure 3 d represents the signal results of the locus in the 45.4-46.2 Mb region of chromosome 19.
[0024] Figure 4 This is a Manhattan plot obtained from a GWAS analysis. The dashed line represents the genome-wide significance threshold; sites exceeding this threshold are considered significantly associated with lung cancer risk at the genome-wide level.
[0025] Figure 5 This is a graph showing the results of a stratified genome-wide association analysis. The left graph is the stratified Manhattan plot; the right graph is the corresponding QQ plot, used to assess the distribution of association signals and potential systematic bias. The upward curve at the tail of the QQ plot indicates the presence of genuine association signals, with more significant deviations in the non-smoker group. The genome expansion factor λ is close to 1 and less than 1.05, suggesting that the results are less affected by population stratification or systematic bias.
[0026] Figure 6 This figure shows the comparison of effect sizes and correlation coefficients between the GWAS results of the West China Hospital Lung Cancer Cohort (LCCWCH) and the results of the Independent Repeated Validation Cohort Japan Lung Cancer Cohort (BBJ).
[0027] Figure 7 This is a Manhattan diagram of the TWAS (Transcriptome-Wide Association Study).
[0028] Figure 8 This is a graph showing the relationship between lung cancer risk sites on chromosome 3 and CLDN18 gene expression; among them, Figure 8 a is a Manhattan plot of the eQTL region, showing the association between genetic variation in this region and CLDN18 expression levels. Figure 8 b is a GWAS regional association map, showing the association signals between genetic variations in the same region and the risk of lung cancer. Figure 8 c is the co-localization analysis plot of GWAS and eQTL, suggesting evidence of co-localization between lung cancer risk signals and CLDN18 expression regulation signals in this region. Specifically, rs6804932 showed a significant association in both GWAS and eQTL analyses. Figure 8 d is a box plot of CLDN18 expression levels for different rs6804932 genotypes, used to assess the allele-specific expression effect at this locus.
[0029] Figure 9 This is a diagram showing the results of an integrated genetic analysis from GWAS significant loci to the functional mechanisms of target genes. Figure 9 a represents the co-location analysis results of GWAS and CLDN18 eQTL. Figure 9 b is the allele dose-response result of rs6804932 on CLDN18 expression. Figure 9 c represents the co-location analysis results of rs9830151 and CLDN18 eQTL. Figure 9 d represents the regulatory effect of rs9830151 on TP63 expression in lung tissue. Beta represents the regression coefficient in the linear regression model, which reflects the magnitude and direction of the effect of this allele on the expression level of the target gene.
[0030] Figure 10 This is a graph showing the results of an independent TWAS analysis based on GTEx v8 lung tissue data.
[0031] Figure 11 This image shows the expression characteristics, clinical prognostic value, and paired sample validation results of the CLDN18 gene in lung adenocarcinoma (LUAD); among them, Figure 11 a represents the integrated expression spectrum of TCGA+GTEx. Figure 11 b represents the Kaplan-Meier curve of overall survival for TCGA-LUAD. The significance of the difference between the two survival curves is P = 0.022, the hazard ratio (HR) is 0.96, and CI represents the confidence interval. Figure 11 c represents the results of paired sample expression analysis.
[0032] Figure 12 This is a graph showing the results of multiplex immunofluorescence analysis; in which, Figure 12 Image a shows multiplex immunofluorescence micrographs of tumor lung tissue and normal lung tissue. Figure 12b is a graph showing the quantitative statistical results of multiplex immunofluorescence analysis. Figure 12 c represents the correlation analysis results between CLDN18 protein expression and patient survival. The log-rank P-value of the difference between the two survival curves is 0.0331, the hazard ratio (HR) is 0.5535, and CI represents the confidence interval.
[0033] Figure 13 The figure shows the effect of CLDN18 overexpression on lung cancer; among them, Figure 13 Figure a shows tumor images in the control group and the CLDN18 overexpression group in a subcutaneous tumor model established using LLC cells. Figure 13 b is a graph showing the statistical results of tumor weight in an LLC cell-based subcutaneous tumor model; Figure 13 c shows tumor images of the control group and the CLDN18 overexpression group in an orthotopic lung cancer tumor model established using LLC cells. Figure 13 d is a graph showing the statistical results of tumor weight in the LLC in situ lung cancer model.
[0034] Figure 14 This is a diagram showing the results of the functional annotation; where, Figure 14 a is a map showing the distribution of variant sites in the CLDN18 gene region. Figure 14 Figure b shows the results of the luciferase reporter gene experiment in A549 cells. Figure 14 c shows the results of the luciferase reporter gene experiment in H1299 cells.
[0035] Figure 15 The figure shows the results of the Sanger method for verifying the rs6804932 mutation site in CLDN18.
[0036] Figure 16 ROC curves for a multivariate lung cancer risk assessment model constructed based on age, sex, smoking status, and the CLDN18 rs6804932 locus. Detailed Implementation
[0037] Unless otherwise specified, all reagents and materials used in the following examples and experimental cases are commercially available.
[0038] Example 1: Composition and Usage of the PCR Detection and Sequencing Kit of the Present Invention I. Reagent Kit The kit in this example includes amplification reagents for amplifying the rs6804932 mutation site of the CLDN18 gene, as well as reagents for Sanger sequencing.
[0039] 1. Amplification reagents The PCR amplification reagents are used to amplify a DNA sequence containing the SNP site, and their composition is shown in Table 1.
[0040] Table 1 PCR Amplification Reagents The PCR mixture in Table 1 includes components required for conventional PCR such as Taq enzyme, dNTPs, and magnesium ions; primer pair information is shown in Table 2.
[0041] Table 2 Primers used for gene amplification 2. Sequencing reagents The reagent comprises the components shown in Table 3.
[0042] Table 3. Reagents for Genome Variation Typing Detection (including purification reagents) The F primer is the sequencing amplification primer, and its nucleotide sequence is: ATGTGGATAGAAGGAAATGAA (SEQ ID NO. 3, forward primer); AGGACTTGCTGTGTTCCATTA (SEQ ID NO. 4, reverse primer).
[0043] II. Instructions for using the reagent kit 1. DNA extraction Peripheral venous blood samples were collected from healthy individuals and lung cancer patients. The samples were placed in blood collection tubes containing EDTA anticoagulant, gently inverted to mix, and then stored briefly at 4°C before genomic DNA extraction. A commercially available peripheral blood genomic DNA extraction kit was used, following the manufacturer's instructions. The specific steps were as follows: An appropriate amount of whole blood sample was taken, and erythrocyte lysis buffer was added and thoroughly mixed. After incubation at room temperature, the sample was centrifuged, and the supernatant was discarded to remove erythrocytes. Cell lysis buffer and proteinase K were then added to fully lyse leukocytes and release genomic DNA. After incubation at 56°C for a certain period, binding buffer and anhydrous ethanol were added to allow the DNA to bind to a silica gel adsorption column. The mixture was transferred to the adsorption column and centrifuged. The filtrate was discarded, and the column was washed sequentially to remove proteins, salt ions, and other impurities. Finally, the adsorption column was placed in a new centrifuge tube, and an appropriate amount of nuclease-free water or elution buffer was added. After standing at room temperature, the column was centrifuged to obtain purified genomic DNA. The concentration and purity of the extracted DNA were detected using a UV spectrophotometer or a quantitative fluorescence spectrometer, and DNA integrity was assessed by agarose gel electrophoresis. DNA samples that meet the quality requirements are used for subsequent PCR amplification, Sanger sequencing, or genotyping analysis.
[0044] 2. PCR amplification and detection DNA fragments containing the detected mutation sites were amplified by PCR. The PCR amplification system for mutation sites is shown in Table 4.
[0045] Table 4 PCR reaction conditions are shown in Table 5: Table 5 The PCR product was detected by agarose gel electrophoresis. The rs6804932 mutation site is located on the CLDN18 gene, and the PCR product is 505 bp in length.
[0046] 3. Sanger sequencing detection (1) Purification of PCR products: The system is shown in Table 6, and the reaction conditions are shown in Table 5.
[0047] Table 6 (2) Sanger sequencing The aforementioned genotyping reagents were used as sequencing amplification reagents for Sanger sequencing of the purified PCR products.
[0048] The aforementioned genotyping reagents were used as sequencing amplification reagents for Sanger sequencing of purified PCR products to detect the genotype of the CLDN18 gene at the rs6804932 locus. The specific method is as follows: Using extracted genomic DNA as a template, the CLDN18 gene fragment containing the rs6804932 locus was amplified using specific primers. After enzymatic purification of the PCR product, sequencing primers, BigDye Mix, 5× sequencing buffer, and ddH2O were added for Sanger sequencing. After the reaction, the sequencing product was purified, and the sequence was read using a capillary electrophoresis sequencer. The obtained sequencing peak diagram was compared with the CLDN18 reference sequence to determine the base type at the rs6804932 locus and to identify the sample genotype. The results are interpreted as follows: If sequencing results show a change from G to A at position 138032158 of the CLDN18 gene, it indicates that the sample carries the A allele at the rs6804932 mutation site detected in this invention. The A allele at the rs6804932 site is associated with a lower risk of lung cancer, and its protective effect shows a dose-dependent trend; that is, the higher the copy number of the A allele, the lower the risk of lung cancer. Specifically, when the genotype at the rs6804932 site is AA, it indicates the lowest risk of lung cancer; when the genotype is AG, it indicates a lower risk of lung cancer than the GG genotype; and when the genotype is GG, it indicates a relatively higher risk of lung cancer.
[0049] The technical solution of the present invention will be further explained through experiments below.
[0050] Experiment Example 1: Multi-omics integrated analysis to screen key candidate genes related to lung cancer 1. Research Cohort This invention, based at West China Hospital of Sichuan University, established the West China Hospital Lung Cancer Cohort (LCCWCH) to systematically analyze the genetic susceptibility to lung cancer in the Chinese population. The cohort included 2,687 lung cancer patients and 8,620 healthy controls. Whole-genome sequencing of peripheral blood samples was performed using the Illumina NovaSeq 6000 platform. After standardized data processing and rigorous quality control, the quality control results are as follows: Figure 1 As shown, 2,564 cases and 8,370 controls were ultimately retained for subsequent analysis. Principal component analysis revealed strong genetic homogeneity in the cohort. This study has been approved by the Ethics Committee of West China Hospital, Sichuan University.
[0051] 2. Genome-wide association study (GWAS) To identify common genetic variants associated with lung cancer risk in the Chinese population, principal component analysis and genome-wide association analysis (GWAS) were performed on 2,564 lung cancer patients and 8,370 controls in the LCCWCH cohort.
[0052] Principal component analysis showed no significant population stratification, which remained consistent after stratification by pathological type, sex, and smoking status. GWAS ultimately identified 11 genes with genome-wide significance (P ≤ 5 × 10⁻⁶). -8 Genetic loci significantly associated with lung cancer risk ( ) Figure 2-4 ). Among them, 4 sites ( TP63 (3q28) TERT (5p15.33) ATP6V1G2-DDX39B (6p21.32) and STN1 (10q24.33) is a previously reported locus, and the other 7 loci ( NCKAP5 (2q21.2) CLDN18 (3q22.3) MTHFD2L (4q13.3) PCDH15 (10q21.1) FAM86C2P (11q13.2) VPS39 (15q15.1) SYMPK (19q13.32) is a newly discovered associated site.
[0053] 3. Further screening of associated sites This experiment identified seven newly discovered lung cancer-associated sites, located in areas involved in microtubule regulation. NCKAP5 (2q21.2) Encoding epithelial tight junction protein CLDN18 (3q22.3), related to folic acid metabolism MTHFD2L (4q13.3), involved in calcium-dependent cell adhesion PCDH15(10q21.1), FAM86C2P (11q13.2), annotated as a long non-coding RNA, is involved in membrane transport. VPS39 (15q15.1) and the encoding of scaffold proteins SYMPK (19q13.32). Further conditional and fine-mapping analyses were conducted in four hotspot regions, identifying 15 potential causal variants. Most of these variants were located in introns or intergenic regions, suggesting their potential involvement in lung cancer development by regulating gene expression. Among them, those located in... CLDN18 rs6804932 was identified as the primary candidate function SNP.
[0054] 4. Stratification analysis based on smoking status In the LCCWCH cohort, this study performed principal component analysis and genome-wide association analysis (GWAS) on 2,564 lung cancer patients and 8,370 controls, and further stratified analysis based on smoking status. The results are as follows: Figure 5 As shown. Stratified analysis of smoking status shows that among smokers, MTHFD2L The (4q13.3) locus was significantly associated with lung cancer risk (P = 2.70 × 10⁻⁶). -12 Nine risk loci with genome-wide significance were identified in the non-smoker population, including five newly discovered loci: NCKAP5 (2q21.2) CLDN18 (3q22.3) MTHFD2L (4q13.3) VPS39 (15q15.1) and SYMPK (19q13.32).
[0055] 5. The verification queue verifies the GWAS signal. This experimental example further validated the results based on previous research, and the results are as follows: Figure 6 As shown. Specifically, the identified SNPs were replicated and validated in an independent East Asian population cohort—BioBank Japan (BBJ, lung cancer: 4,444 cases, healthy individuals: 174,282)—to assess the robustness of the GWAS results. Of the 11 lead SNPs, 9 loci were available in the BBJ data, including CLDN18 (3q22.3) Overall, the direction and magnitude of the effect sizes among the studies were generally consistent (Pearson r = 0.68, P = 0.042), further supporting the robustness of the results of this invention.
[0056] 6. Colocation Analysis A lung tissue-specific eQTL dataset was constructed based on adjacent normal lung tissue from 153 patients in the West China Lung Cancer Cohort (WCLCTC). The Coloc method was used to integrate this eQTL dataset with GWAS results, as shown below. Figure 7 As shown, three genes with strong evidence of colocalization were identified (PPH4>0.80): CLDN18 (P = 3.51 × 10) -8 (Z = -5.51) DCBLD1 (P = 2.27 × 10) -7 (Z = -5.18) and BNIP3P5 (P = 1.87 × 10) -7 Z = -5.21). Among them, rs6804932 (3q22.3) and CLDN18 The eQTL signal was highly colocalized (PPH4 = 0.95), and its lung cancer protective allele A was significantly increased. CLDN18 Expression (β = 0.46, P = 5.74 × 10) -16 ), the result is as follows Figure 8 As shown.
[0057] Validation was further confirmed using GTEx v8 lung tissue eQTL data. CLDN18 and TP63 Colocalization relationships were observed, with CLDN18 expression colocalizing with lung cancer GWAS signaling at rs6804932 (3q22.3) (β = 0.24, P = 3.71 × 10⁻⁶). -25 TP63 expression was found to co-localize with lung cancer GWAS signaling at the rs9830151 (3q28) site (β = 0.12, P = 4.24 × 10⁻⁶). -2 The result is as follows: Figure 9 As shown.
[0058] Independent TWAS analysis based on GTEx v8 lung tissue data further validated TP63 (P = 1.03 × 10) -10 (Z = -6.46) and CLDN18 (P = 1.84 × 10) -7 A significant association was found (Z = -5.22), as shown in the results. Figure 10 As shown.
[0059] In summary, the multi-omics integrated analysis consistently supports CLDN18 as a key candidate gene for genetic susceptibility to lung cancer. Further analysis revealed that the effector allele at the CLDN18 rs6804932 locus is A, with an effect size OR of 0.80, suggesting that individuals carrying the A allele have a lower risk of developing lung cancer; and the protective effect is further enhanced with increasing copy number of the A allele.
[0060] Experimental Example 2: The Relationship Between CLDN18 and Lung Cancer 1. The relationship between CLDN18 expression level and lung cancer To assess the relationship between CLDN18 expression levels and the development and progression of lung cancer, this study integrated and analyzed data from the TCGA lung cancer cohort and GTEx v8 normal lung tissue. Specifically, transcriptome expression data and corresponding clinical information of lung cancer tissues were obtained from the TCGA database, and transcriptome expression data of normal lung tissues were obtained from the GTEx v8 database. To reduce technical bias between different data sources, gene annotation information was first standardized, the CLDN18 expression matrix was extracted, and the expression values were converted into comparable standardized expression levels. Subsequently, a pooled analysis was performed on TCGA tumor samples and GTEx normal lung tissue samples to compare the differences in CLDN18 expression between the two groups. Furthermore, based on clinical follow-up information of TCGA lung cancer patients, patients were grouped according to CLDN18 expression levels to further analyze its relationship with overall survival.
[0061] The results are as follows Figure 11 As shown. Compared with normal GTEx lung tissue, CLDN18 expression was significantly reduced in TCGA lung cancer tissue, and patients with low CLDN18 expression had poorer overall survival. Figure 11 (ab). Further validation using WCLCTC cohort and single-cell transcriptome data revealed that CLDN18 expression was highest in normal lung tissue and gradually decreased with increasing tumor malignancy. These results suggest that decreased CLDN18 expression may be associated with the development and progression of lung cancer and poor prognosis.
[0062] 2. Multiplex immunofluorescence analysis To further verify the relationship between CLDN18 protein expression and lung cancer progression and changes in epithelial integrity, this experiment performed multiplex immunofluorescence analysis on lung cancer tissue microarrays.
[0063] Multiplex immunofluorescence staining was performed on lung cancer tissue microarrays to detect the protein expression and spatial distribution of CLDN18, E-cadherin, and N-cadherin. Briefly, tissue microarray sections were dewaxed, hydrated, and then subjected to antigen retrieval, followed by blocking of non-specific binding sites with blocking buffer. Sections were sequentially incubated with primary antibodies against CLDN18, E-cadherin, and N-cadherin, and then developed with corresponding fluorescently labeled secondary antibodies (Akoya, Cat#NEL861001KT). After each round of staining, antibody elution or heat retrieval was performed according to the experimental protocol before proceeding to the next round of antibody incubation. After all labeling was completed, cell nuclei were counterstained with DAPI, and the slides were mounted with anti-fluorescence quenching mounting medium. Further image analysis software was used to identify tissue regions and cell nuclei, and the fluorescence intensity and percentage of positive cells for each marker were calculated.
[0064] Based on the median CLDN18 expression level score, the samples were divided into high-expression and low-expression groups, and the relationship between CLDN18 expression and overall patient survival was further evaluated. Simultaneously, the association between CLDN18 and E-cadherin and N-cadherin expression levels was further analyzed to assess whether CLDN18 downregulation is associated with epithelial-mesenchymal transition (EMT)-related phenotypic changes.
[0065] The results are as follows Figure 12 As shown, CLDN18 protein expression was lowest in tumor tissues, and its low expression was significantly associated with poorer overall survival in patients. Further analysis revealed that CLDN18 expression levels were positively correlated with the epithelial marker E-cadherin and negatively correlated with the mesenchymal marker N-cadherin. These results suggest that CLDN18 downregulation may promote lung cancer progression by facilitating epithelial-mesenchymal transition (EMT) and disrupting epithelial structural integrity.
[0066] 3. Functional Experiment To further verify the functional role of CLDN18 in lung cancer progression, this experiment constructed a stable CLDN18 overexpression model in LLC lung cancer cells. First, using the mouse-derived CLDN18 reference transcript NM_019815.3 from the NCBI database as a template, specific amplification primers were designed based on its coding sequence (upstream primer: CGGAATTCGCCACCATGGCTCTGCCTCGGAG (SEQ ID NO. 5); downstream primer: CCGCTCGAGTCACACGTAGTCCTTGCGGTC (SEQ ID NO. 6)). The complete coding region sequence of CLDN18 was obtained by PCR amplification. Subsequently, the amplification product was cloned into a lentiviral expression vector. For LLC cells, the GL180 lentiviral overexpression vector carrying the mouse-derived CLDN18 coding sequence was used for subsequent experiments. After construction, the CLDN18 overexpression lentiviral vector was packaged, and the viral supernatant was collected and used to infect LLC cells. After infection, stable overexpression cell lines were selected using appropriate antibiotics. LLC cells infected with the empty vector served as a negative control group. Finally, the mRNA and protein expression levels of CLDN18 were detected by real-time quantitative PCR (qRT-PCR) and Western blot to verify the successful construction of the CLDN18 overexpression model and its expression efficiency. The overexpression sequence is as follows: (SEQ ID NO. 7).
[0067] Mouse subcutaneous tumorigenesis experiment: Control group LLC cells and CLDN18 overexpression group LLC cells were resuspended in a mixture of serum-free culture medium and matrix gel (50% each) and cultured at the same cell number (1×10⁻⁶). 6 The tumor was subcutaneously injected into immunized mice. The long and short diameters of the tumor were measured periodically after injection, and the tumor volume was calculated using a formula. Mice were sacrificed at the experimental endpoint, and the tumor tissue was dissected and weighed.
[0068] Mouse orthotopic lung cancer model: To further simulate the lung tumor microenvironment, control group and CLDN18 overexpression group LLC cells were used to construct an orthotopic lung cancer model. A lung tumor burden model was established by intrapulmonary injection via the pleural cavity. After inoculation, the long and short diameters of the tumor were measured periodically, and the tumor volume was calculated using the formula: Tumor volume = 1 / 2 × long diameter × short diameter. 2 .
[0069] Animal experiment results such as Figure 13 As shown, overexpression of CLDN18 significantly inhibited LLC cell proliferation and markedly reduced tumor growth capacity in both mouse subcutaneous tumor and orthotopic lung cancer models. These results suggest that downregulation of CLDN18 may accelerate lung cancer progression by weakening epithelial barrier function and promoting epithelial-mesenchymal transition (EMT).
[0070] 4. Functional annotations To elucidate the potential regulatory function of the CLDN18 rs6804932 locus, this study performed functional annotation of the genomic region containing this locus and validated its allele-specific regulatory role using a dual-luciferase reporter assay. Based on genome-wide association analysis (GWAS) results, genetic annotation information for the CLDN18 rs6804932 locus and its adjacent regions was extracted. The location of this locus within the 3'UTR of the CLDN18 gene was determined using a genome annotation database, and eQTL analysis results were integrated to assess the relationship between different alleles of rs6804932 and CLDN18 expression levels. Using the rs6804932 genotype as the independent variable and CLDN18 expression level as the dependent variable, this study analyzed whether this locus acts as a cis-eQTL to regulate CLDN18 expression.
[0071] Construction of the dual-luciferase reporter vector: Using human genomic DNA as a template, the CLDN18 3'UTR fragment containing the rs6804932 site was amplified. Fragments carrying the G and A alleles were constructed separately and cloned into the downstream region of the luciferase gene in the dual-luciferase reporter vector to mimic the regulatory role of the 3'UTR on reporter gene expression. The primer sequences were as follows: upstream primer: 3'UTR-F: CGACGCGTTGACTTCTGAGGACTTTGAG (SEQ ID NO. 8); downstream primer: UTR-R: ATAAGAATGCGGCCGCTCAGGAGGTGAGGTTGAG (SEQ ID NO. 9). A549 and H1299 lung cancer cells were cultured in a medium containing 10% fetal bovine serum and 1% penicillin and streptomycin at 37°C and 5% CO2. One day before transfection, cells were seeded into 24-well or 96-well plates to achieve a cell confluence of approximately 70%–80% at transfection. Reporter plasmids carrying the rs6804932-G or rs6804932-A alleles were transfected into A549 and H1299 cells, respectively. Simultaneously, an internal control luciferase plasmid was co-transfected to correct for differences in transfection efficiency. Empty vectors or basic reporter vectors served as negative controls. 24–48 hours after transfection, the culture medium was discarded, and cells were lysed using cell lysis buffer. Following the instructions of the dual-luciferase reporter gene assay kit, the activities of firefly luciferase and Renilla luciferase were measured sequentially. Using Renilla luciferase activity as an internal control, firefly luciferase activity was standardized, and relative luciferase activity was calculated. The relative luciferase activity difference between the rs6804932-A and rs6804932-G constructs was compared to assess the regulatory effect of different alleles of this SNP on reporter gene expression.
[0072] The results are as follows Figure 14 As shown, rs6804932 is located in the 3' untranslated region (3'UTR) of CLDN18 and is significantly correlated with CLDN18 expression level as a cis-eQTL. Further experiments showed that in A549 and H1299 lung cancer cells, reporter vectors carrying the A allele exhibited higher luciferase activity than vectors carrying the G allele, suggesting that the A allele can enhance post-transcriptional regulatory activity in this region or increase CLDN18 expression level.
[0073] In conclusion, rs6804932 may reduce the risk of lung cancer by upregulating CLDN18 expression and maintaining epithelial integrity.
[0074] Experiment 3: Further validation of CLDN18 as a key candidate gene for lung cancer susceptibility. To further verify the application value of CLDN18-related genetic variants in lung cancer susceptibility assessment, this experiment used Sanger sequencing to perform genotyping verification at the CLDN18 rs6804932 locus. The specific detection reagents and methods used were the kit and usage instructions described in Example 1.
[0075] The obtained sequencing peak diagram was compared with the reference genome sequence to determine the base type at the rs6804932 locus and to identify the sample genotype as GG, AG, or AA. Figure 15 As shown.
[0076] To further validate the application value of CLDN18-related genetic variants in assessing susceptibility to lung cancer, this study conducted a risk prediction analysis based on the West China Lung Cancer Validation Cohort (306 lung cancer patients and 1538 healthy controls). This cohort is a separate dataset from the aforementioned West China Hospital Lung Cancer Cohort (LCCWCH). First, genotypic information at the rs6804932 locus was extracted, and a genetic risk score was constructed by combining the effect direction and effect size of this locus. Simultaneously, clinical phenotypes such as age, sex, and smoking status were incorporated into the model, with lung cancer prevalence status as the outcome variable. A multivariate logistic regression model was used to assess the association between the risk score and the risk of lung cancer incidence.
[0077] Subsequently, receiver operating characteristic (ROC) curves were plotted based on the model's predicted probabilities, and the area under the curve (AUC) was calculated to evaluate the model's ability to distinguish between lung cancer patients and healthy controls. The optimal risk score cutoff value was determined using the Youden's Index maximization method, and its calculation formula is as follows: Youden's Index = Sensitivity + Specificity-1.
[0078] Here, Sensitivity represents sensitivity, and Specificity represents specificity. When the Youden index reaches its maximum value, the corresponding risk score is defined as the optimal cutoff value, and the model's sensitivity, specificity, and overall discriminative power are further calculated.
[0079] The results are as follows Figure 16As shown, the lung cancer risk assessment model based on the CLDN18 rs6804932 locus demonstrated good discriminative ability in the West China Lung Cancer Cohort validation set. After combining age, sex, smoking status, and the CLDN18 rs6804932 locus, the model's AUC reached 0.712, indicating its certain predictive ability for lung cancer risk. Compared with the basic model that only included age, sex, and smoking status, the addition of the rs6804932 locus improved both the model's sensitivity and specificity. These results further support CLDN18 as a key candidate gene for lung cancer genetic susceptibility and suggest that its related genetic variations have potential value in lung cancer risk assessment and auxiliary screening.
[0080] As can be seen from the above embodiments and experimental examples, this invention identifies single nucleotide polymorphism (SNP) sites associated with lung cancer incidence based on genome-wide association analysis (GWAS), and, combined with fine mapping, transcriptome analysis, and eQTL co-localization analysis, for the first time identifies rs6804932 in the CLDN18 gene as a potential protective variant against lung cancer. Experimental results show that carrying the A allele at the rs6804932 site is significantly associated with a lower risk of lung cancer. Further risk assessment models show that, after considering age, sex, and smoking status, including the rs6804932 site increases the model's AUC to 0.712, demonstrating good risk prediction performance. Therefore, the rs6804932 site in the CLDN18 gene can serve as an important biomarker for lung cancer risk prediction and auxiliary screening, showing promising application prospects in early lung cancer risk assessment, genetic susceptibility testing, and the development of related diagnostic products.
Claims
1. The application of CLDN18 gene SNP site detection reagents in the preparation of kits, chips, or systems for screening lung cancer, characterized in that: The SNP site is rs6804932.
2. The application of the CLDN18 gene SNP site detection reagent according to claim 1 in the preparation of kits, chips, or systems for screening lung cancer, characterized in that: When the allele at the SNP locus is A, the risk of developing lung cancer is low; when the allele at the SNP locus is G, the risk of developing lung cancer is high.
3. The application of the CLDN18 gene SNP site detection reagent according to claim 1 in the preparation of kits, chips, or systems for screening lung cancer, characterized in that: The detection reagents include at least one of PCR amplification reagents, gene chip reagents, and nucleic acid sequencing reagents.
4. The application of the CLDN18 gene SNP site detection reagent according to claim 3 in the preparation of kits, chips, or systems for screening lung cancer, characterized in that: The detection reagent includes primers with nucleotide sequences as shown in SEQ ID NO. 1-4.
5. The application of the CLDN18 gene SNP site detection reagent according to claim 1 in the preparation of kits, chips, or systems for screening lung cancer, characterized in that: The screening participants were selected from East Asian populations and / or non-smokers.
6. The application of the CLDN18 gene SNP site detection reagent according to claim 1 in the preparation of kits, chips, or systems for screening lung cancer, characterized in that: Screening samples include blood, lymph, cerebrospinal fluid, and lung tissue samples.
7. A kit or chip for screening lung cancer, characterized in that: It includes a detection reagent for the CLDN18 gene SNP site, which is rs6804932.
8. The lung cancer screening kit or chip according to claim 7, characterized in that: When the allele at the SNP locus is A, the risk of developing lung cancer is low; when the allele at the SNP locus is G, the risk of developing lung cancer is high.
9. A system for screening lung cancer, characterized in that, The system includes a computational device for predicting lung cancer susceptibility based on sequencing results of a SNP site in the CLDN18 gene, wherein the SNP site is rs6804932.
10. The system according to claim 9, characterized in that, The system includes the following modules: The test result collection module is configured to collect sequencing results of SNP sites in the CLDN18 gene; The calculation module is configured to calculate the risk of lung cancer based on the sequencing results collected from the detection results module.