A model for colorectal cancer prognosis prediction and application thereof
By constructing a prognostic prediction model for colorectal cancer based on genes such as ARL6IP4, the problem of lacking effective biomarkers in existing technologies has been solved, enabling stable and effective prediction and early intervention for the prognosis of colorectal cancer patients, and reducing the incidence of adverse prognoses.
Patent Information
- Application Number
- CN202411772356.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2044-12-04
AI Technical Summary
The lack of effective biomarkers in current technologies to predict the prognosis of colorectal cancer leads to significant differences in clinical outcomes among colorectal cancer patients at the same TNM stage. There is an urgent need for more effective methods and biomarkers to predict the prognosis of colorectal cancer.
Six genes—ARL6IP4, ATP2A1, CRYAB, ELFN2, MAP2, MAT1A, ORC1, and POU4F1—were used as a combination of biomarkers. The expression levels of these genes in the samples were detected by methods such as nucleic acid sequencing, nucleic acid hybridization, chromatography, mass spectrometry, digital imaging, protein immunoassay, and dye technology. A risk scoring model for predicting the prognosis of colorectal cancer was constructed, and the Cox regression model was used for prediction.
This study provides a stable and effective prognostic prediction model for colorectal cancer, which improves the ability to predict the prognosis of colorectal cancer patients. It can identify high-risk patients and conduct early monitoring and intervention, reduce the incidence of adverse prognoses, and improve patient outcomes.
Smart Images

Figure CN119351557B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of biological medicine, and relates to a disease prognosis prediction model, in particular, to a model for colorectal cancer prognosis prediction and application thereof. BACKGROUND
[0002] Colorectal cancer (CRC) is the most common malignant tumor of the digestive system and the second leading cause of cancer-related deaths worldwide. Despite significant advances in current treatment methods, the prognosis of CRC patients remains dismal. Colorectal cancer is highly heterogeneous and complex, with a variety of differences in clinical presentation and prognosis. According to the American Joint Committee on Cancer (AJCC), the main prognostic indicator for colorectal cancer is still the tumor node metastasis (TNM) staging system. However, there are significant differences in clinical outcomes among colorectal cancer patients with the same TNM stage. Therefore, there is an urgent need to find more effective means and potential biomarkers to predict the prognosis of colorectal cancer.
[0003] Liquid-liquid phase separation (LLPS) is a phenomenon caused by weak interactions between biomolecules, leading to the formation of membraneless organelles, also known as biomolecular condensates. Biomolecular condensates are the basis for the spatiotemporal coordination of biological activities within cells and play a crucial role in many life activities in living organisms, including gene expression regulation, signal transduction, and regulation of cellular metabolism. Disruption of this precise spatiotemporal regulation can lead to devastating pathological consequences, such as tumors. LLPS plays a crucial role in the occurrence and development of tumors. One study reported that SHP2 gene mutations directly confer phase separation ability to promote MAPK activation, thereby promoting tumorigenesis. Another study found that DDX21 phase separation condensates target and activate MCM5, which in turn promotes the activation of epithelial-mesenchymal transition (EMT), which contributes to CRC metastasis. More and more studies have begun to construct prognosis models for different cancer types based on LLPS-related genes (LRGs). However, no risk model related to LRGs has been established in CRC. SUMMARY
[0004] To make up for the shortcomings of the prior art, the purpose of the present application is to provide a risk score model related to LRGs for prognosis prediction of colorectal cancer.
[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions: The first aspect of the present application provides a biomarker combination for prognosis prediction of colorectal cancer.
[0006] Further, the biomarker combination comprises ARL6IP4, ATP2A1, CRYAB, ELFN2, MAP2, MAT1A, ORC1 and POU4F1.
[0007] In the present application, the term "biomarker" means a compound, preferably a gene, which is differentially present (i.e. increased or decreased) in a biological sample from a subject or a group of subjects having a first phenotype (e.g. suffering from a disease) compared to a biological sample from a subject or a group of subjects having a second phenotype (e.g. not suffering from a disease). The term "biomarker" generally refers to the presence / concentration / amount of one gene or the presence / concentration / amount of two or more genes. In the present application, the biomarker comprises ARL6IP4, ATP2A1, CRYAB, ELFN2, MAP2, MAT1A, ORC1 and POU4F1. In the present application, the biomarker comprises the gene and its encoded protein and homologs, mutations, and isoforms thereof. The term encompasses the full-length, unprocessed molecule, as well as any form of the molecule derived by processing in the cell. The term encompasses naturally occurring variants of the biomarker (e.g. splice variants or allelic variants).
[0008] A second aspect of the present application provides the use of a reagent for detecting the expression level of the biomarker combination according to the first aspect of the present application in the manufacture of a product for the prediction of colorectal cancer.
[0009] Further, the reagent comprises a reagent for detecting the expression level of the biomarker combination in a sample by nucleic acid sequencing technology, nucleic acid hybridization technology, chromatography technology, mass spectrometry technology, digital imaging technology, protein immunization technology, dye technology, and / or next generation sequencing technology.
[0010] Further, the reagent comprises a reagent for detecting the mRNA expression level of the biomarker combination and / or a reagent for detecting the protein expression level of the biomarker combination.
[0011] Preferably, the reagent comprises a reagent for detecting the level of cDNA complementary to the mRNA transcribed from the biomarker combination.
[0012] Preferably, the reagent is selected from a primer or a probe.
[0013] Further, the reagent comprises a reagent for detecting the level of polypeptide or protein encoded by the biomarker combination.
[0014] Preferably, the reagent is selected from an antibody, an antibody fragment or an affinity protein.
[0015] Further, the primer according to the present application can be prepared by a chemical synthesis method well known to those skilled in the art, appropriately designed by using a method well known to those skilled in the art with reference to known information, and prepared by a chemical synthesis.
[0016] Further, the probe according to the present application can be prepared by a chemical synthesis, appropriately designed by using a method well known to those skilled in the art with reference to known information, and prepared by a chemical synthesis, or can be prepared by preparing a gene containing a desired nucleic acid sequence from biological material, and amplifying it using a primer designed to amplify the desired nucleic acid sequence.
[0017] Further, the probe hybridized to the nucleic acid sequence of the gene can be DNA, RNA, DNA-RNA chimera, PNA, or other derivatives. The length of the probe is not limited, and any length can be used as long as specific hybridization is achieved, and specific binding to the nucleotide sequence of interest is achieved. The length of the probe can be as short as 25, 20, 15, 13, or 10 bases. Also, the length of the probe can be as long as 60, 80, 100, 150, 300 base pairs or more, or even the entire gene. Since different probe lengths have different effects on hybridization efficiency and signal specificity, the length of the probe is usually at least 14 base pairs, and the length of the probe is usually not more than 30 base pairs, and the length of the nucleotide sequence complementary to the nucleotide sequence of interest is preferably 15 to 25 base pairs. The self-complementary sequence of the probe is preferably less than 4 base pairs to avoid affecting the hybridization efficiency.
[0018] A third aspect of the present application provides a product for prognosis prediction of colorectal cancer.
[0019] Further, the product comprises a reagent for detecting the expression level of the biomarker combination according to the first aspect of the present application.
[0020] Further, the reagent comprises a reagent for detecting the mRNA expression level of the biomarker combination and / or a reagent for detecting the protein expression level of the biomarker combination.
[0021] Preferably, the reagent is selected from a primer, a probe, an antibody, an antibody fragment, or an affinity protein.
[0022] Preferably, the product further comprises a total RNA extraction reagent, a reverse transcription reagent, and / or a next-generation sequencing reagent.
[0023] Further, the detection of the expression level of the biomarker combination according to the present application can employ an assay method known in the art, including but not limited to a method of detecting the amount of RNA transcript of the gene in the biomarker combination or the amount of polypeptide encoded by the gene in the biomarker combination.
[0024] Preferably, the RNA transcripts of the genes can be converted into complementary cDNA by methods known in the art, and the amount of the RNA transcripts can be obtained by determining the amount of the complementary cDNA. The amount of the RNA transcripts of the genes or the complementary cDNA thereof can be normalized to the amount of total RNA or total cDNA in the sample or to the amount of the RNA transcripts of a set of housekeeping genes or the complementary cDNA thereof.
[0025] Preferably, the RNA transcripts can be detected and quantified by methods such as hybridization, amplification or sequencing, including but not limited to methods of hybridizing the RNA transcripts with probes or primers, methods of detecting the amount of the RNA transcripts or the corresponding cDNA products thereof by various quantitative PCR techniques or sequencing techniques based on polymerase chain reaction (PCR). The quantitative PCR techniques include but are not limited to fluorescent quantitative PCR, real-time PCR or semi-quantitative PCR techniques. The sequencing techniques include but are not limited to Sanger sequencing, next-generation sequencing, third-generation sequencing and single-cell sequencing, etc.
[0026] The fourth aspect of the present application provides a risk score model for prognosis prediction of colorectal cancer.
[0027] Further, the risk score model comprises the following biomarker combination: ARL6IP4, ATP2A1, CRYAB, ELFN2, MAP2, MAT1A, ORC1 and POU4F1.
[0028] Preferably, the risk score model calculates the risk score by using the following regression equation: Risk score = n , wherein n is the number of genes used for prognosis prediction, Exp i is the expression level of each gene, Coef i is the regression coefficient of each gene; the risk score is low risk when it is lower than the median value, and the risk score is high risk when it is higher than the median value.
[0029] The fifth aspect of the present application provides a device or system for prognosis prediction of colorectal cancer.
[0030] Further, the device or system comprises: a data acquisition module for acquiring the expression profile data of the biomarker combination of the first aspect of the present application in the sample of the subject to be tested; a prediction module for providing the expression profile data of the biomarker combination obtained by the data acquisition module as input data to the trained prediction model, wherein the prediction model is trained to predict the prognosis of the subject based on the expression profile data of the biomarker combination of the subject; and a prediction result output module for obtaining the output result of the prediction model in the prediction module to obtain the prognosis prediction result of the subject.
[0031] Preferably, the prediction model is a Cox regression model.
[0032] More preferably, the Cox regression model is a LASSO Cox regression model.
[0033] Most preferably, the prediction model is the risk score model of the fourth aspect of the present application.
[0034] The sixth aspect of the present application provides a computer device.
[0035] Further, the computer device comprises a memory and a processor, the memory stores a program, and the processor implements the following method when executing the program: obtaining the biomarker combination expression profile data in the sample of the subject to be tested according to the first aspect of the present application; providing the biomarker combination expression profile data as input data to the trained prediction model; and outputting the prognosis prediction result of the subject to be tested.
[0036] Preferably, the prediction model is a Cox regression model.
[0037] More preferably, the Cox regression model is a LASSO Cox regression model.
[0038] Most preferably, the prediction model is the risk score model of the fourth aspect of the present application.
[0039] Further, the computer device of the present application includes (but is not limited to) any kind of personal computer, server and other terminals that can interact with users through keyboards, touchpads or voice control devices. The computer device herein can also include mobile terminals, which include (but are not limited to) any kind of electronic devices that can interact with users through keyboards, touchpads or voice control devices, such as tablet computers, smart phones, personal digital assistants (PDAs), smart wearable devices and other terminals. The network in which the computer device is located includes (but is not limited to) the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN).
[0040] Further, the memory of the present application includes non-volatile and / or volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The code of the operating system is stored on the memory. For example, the memory also stores codes or instructions, and by running these codes or instructions, the risk score model for prognosis prediction of colorectal cancer provided by the embodiments disclosed herein can be realized. The volatile memory can include random access memory (RAM) or external cache memory.
[0041] Further, the computer device of the present application can include a processor, a memory, an external interface, a display and an input device connected by a system bus. The processor is used to provide computing and control capabilities. The display of the computer device can be a liquid crystal display or an electronic ink display, and the input device can be a touch layer overlaid on the display, or can be a button, trackball or touchpad provided on the housing of the computer device, or can be an external keyboard, touchpad or mouse.
[0042] The processor can include one or more microprocessors, digital processors. The processor can call the program code stored in the memory to perform related functions. The processor, also known as a central processing unit (CPU), can be a very large scale integrated circuit, which is an operation core (Core) and a control core (Control Unit).
[0043] The seventh aspect of the present application provides a computer readable storage medium.
[0044] Further, the computer readable storage medium includes a stored computer program.
[0045] The computer program controls the computer readable storage medium to implement the method of the sixth aspect of the present application when running.
[0046] The eighth aspect of the present application provides any one of the following applications.
[0047] Further, the application includes: 1) the application of the biomarker combination of the first aspect of the present application in establishing a prognosis risk score model of colorectal cancer; 2) the application of the risk score model of the fourth aspect of the present application in predicting the sensitivity of the colorectal cancer patient to the chemotherapeutic drug; wherein when the risk score is low, the colorectal cancer patient is sensitive to the chemotherapeutic drug.
[0048] Preferably, the chemotherapeutic drug includes oxaliplatin, 5-fluorouracil (5-FU), irinotecan, dabrafenib, lapatinib, and ulixertinib.
[0049] The ninth aspect of the present application provides the use of a reagent for detecting the expression level of ARL6IP4 in a sample in the preparation of a product for diagnosing colorectal cancer.
[0050] Further, the product includes a chip, a kit, a test paper or a high-throughput sequencing platform.
[0051] In the present disclosure, the kit comprises a reagent for detecting the expression level of ARL6IP4 in a sample. For example, the kit can be an RT-PCR kit, a DNA chip kit, an ELISA kit, a protein chip kit, a rapid kit, or an MRM (multiple reaction monitoring) kit.
[0052] For example, the diagnostic kit can further comprise elements necessary for reverse transcription polymerase chain reaction. The RT-PCR kit comprises a pair of primers specific to the gene encoding the ARL6IP4 protein. Each primer is a nucleotide having specificity to the nucleic acid sequence of the gene, and can be about 7 to 50 bp in length, more particularly about 10 to 39 bp. In addition, the kit can further comprise primers specific to the nucleic acid sequence of a control gene. In addition, the RT-PCR kit can comprise test tubes or suitable vessels, reaction buffers (different pH and magnesium concentrations), deoxynucleotides (dNTPs), enzymes (e.g., Taq polymerase and reverse transcriptase), deoxyribonuclease inhibitors, ribonuclease inhibitors, DEPC-water, and sterile water.
[0053] In addition, the diagnostic kit of the present disclosure can comprise elements necessary for operating a DNA chip. The DNA chip kit can comprise a substrate bound to a gene or cDNA or an oligonucleotide corresponding to a fragment thereof, and reagents, agents, and enzymes for constructing a fluorescently labeled probe. In addition, the substrate can comprise a control gene or cDNA or an oligonucleotide corresponding to a fragment thereof.
[0054] In some embodiments, the kit of the present disclosure can comprise elements necessary for performing ELISA. The ELISA kit can comprise an antibody specific to a protein. The antibody has high selectivity and affinity to the marker protein, has no cross-reactivity with other proteins, and can be a monoclonal antibody, a polyclonal antibody, or a recombinant antibody. In addition, the ELISA kit can comprise an antibody specific to a control protein. In addition, the ELISA kit can further comprise a reagent capable of detecting the bound antibody, for example, a labeled secondary antibody, a chromophore, an enzyme (e.g., conjugated with an antibody), a substrate thereof, or a substance capable of binding to the antibody.
[0055] In addition, the present application also provides a method for constructing the risk score model for predicting the prognosis of colorectal cancer according to the fourth aspect of the present application, and the method comprises the following steps: 1) candidate gene analysis: screening LLPS-related genes (LRGs) in the Cancer Genome Atlas database; 2) model establishment: using the LRGs obtained in step 1) to determine the genes constituting the prognostic gene label by LASSO Cox analysis, and the genes comprise ARL6IP4, ATP2A1, CRYAB, ELFN2, MAP2, MAT1A, ORC1 and POU4F1.
[0056] Further, the method for constructing further comprises verifying the efficiency of the model.
[0057] Compared with the prior art, the present application has the advantages and beneficial effects that: the present application provides a novel model for predicting the prognosis of colorectal cancer, and the stability and effectiveness of the model are verified by a validation dataset of colorectal cancer patients, the model provides reliable biomarkers for the prognosis evaluation of colorectal cancer patients, improves the ability of predicting the prognosis of colorectal cancer patients, can effectively identify colorectal cancer patients with high-risk prognosis, and can be used for early monitoring and effective intervention in the clinic, so as to reduce the incidence of adverse prognosis and mortality of colorectal cancer patients, and improve the prognosis of colorectal cancer patients, and the model has a wide application prospect in the clinic. BRIEF DESCRIPTION OF DRAWINGS
[0058] Figure 1 A schematic diagram for the present study is shown in Figure 1. Figure 2 Figure 2 shows the LRGs differential expression analysis and functional annotation diagram; wherein, Figure 2A is a Venn diagram showing the intersection of genes identified in the DrLLPS database and TCGA-CRC; Figure 2B is a heatmap showing the expression changes of LRGs in CRC according to the adjusted p-value <0.05, |log2FC|>1; Figure 2C is a PCA analysis diagram of differentially expressed LRGs; Figure 2D is a GO enrichment analysis diagram of differentially expressed LRGs in biological process (BP), cellular component (CC) and molecular function (MF); Figure 2E is a KEGG enrichment analysis diagram of differentially expressed LRGs; Figure 3 Figure 3 shows the construction of a prognosis risk model related to LLPS; wherein, Figure 3A is a single-factor Cox regression analysis of LRGs in TCGA-CRC; Figures 3B and 3C determine 16 prognostic genes through LASSO regression and screening of the optimal parameters of the differentially expressed LRGs related to prognosis; Figure 3D is a multivariate Cox regression analysis of the prognostic genes obtained from the LASSO regression analysis; Figures 3E-3G are PCA diagrams showing the distribution of low-risk and high-risk in all genes, LRGs and risk genes; Figure 4Figures showing the prognostic performance and validation of the risk signature of LRGs; wherein, panels A-C are K-M curves showing the OS of two risk groups in TCGA-CRC in training (A), testing (B) and whole (C) cohorts; panels D-F are the risk score of model genes, survival status (red dots represent death, blue dots represent survival) and gene expression distribution in training (D), testing (E) and whole (F) cohorts; panels G-I are ROC curves used to evaluate the sensitivity and specificity of 1-year, 3-year and 5-year survival rate of risk score in training (G), testing (H) and whole (I) cohorts; panel J is a K-M curve showing the OS of two risk groups in GSE39582; panel K is a K-M curve showing the RFS of two risk groups in GSE39582; panel L is a K-M curve showing the prognosis of tumor stage I-II subgroup; Figure 5 Figures showing the ROC curve analysis, Nomogram construction and risk model of clinical characteristics; wherein, panel A is a prognostic ROC curve of risk score and other clinical factors; panel B is a nomogram developed to predict 1, 3, 5-year survival rate of colorectal cancer patients; panel C is a calibration curve used to evaluate the accuracy of the nomogram; panel D is a C-index curve used to evaluate the consistency index of the nomogram; panel E is a heatmap showing the relationship between the expression levels of 8 genes and various clinicopathological characteristics; Figure 6 Figures showing the correlation between risk signature and immune cell infiltration; wherein, panel A is a difference analysis chart of immune cell infiltration between two risk groups; panel B is a correlation analysis chart of infiltrating immune cell types and risk score; panel C is a distribution bar chart of infiltrating immune cells between high and low risk groups; panel D is a correlation analysis chart of infiltrating immune cell types and 8 LRGs; Figure 7 Figures showing the correlation between risk signature and immune cell infiltration; wherein, panel A is a difference analysis chart of immune cell infiltration between two risk groups; panel B is a correlation analysis chart of infiltrating immune cell types and risk score; panel C is a distribution bar chart of infiltrating immune cells between high and low risk groups; panel D is a correlation analysis chart of infiltrating immune cell types and 8 LRGs; Figure 8 Figures showing the correlation between risk signature and TMB; wherein, panels A, B show the top 15 mutated genes in the high risk group and the low risk group, respectively; panel C is a K-M analysis chart of high TMB group and low TMB group; panel D is a K-M curve under the joint influence of risk score and TMB; Figure 9The graph shows the correlation between ARL6IP4 expression levels and phase separation ability in colorectal cancer. A represents the ARL6IP4 mRNA levels in colorectal cancer and normal tissues from the TCGA database; B compares the ARL6IP4 mRNA levels in 10 pairs of colorectal cancer and normal tissues; C shows a three-dimensional reconstructed image of HCT116 live cells transfected with the GFP-ARL6IP4 plasmid obtained using confocal laser scanning microscopy; D shows the image (top) and quantitative analysis (bottom) of GFP-ARL6IP4 FRAP, with data expressed as Mean ± SE, n = 3 droplets; E shows the fusion of two GFP-ARL6IP4 droplets to form a larger droplet. Figure 10 The graphs show the ROC curves of ARL6IP4, where A is the ROC curve of ARL6IP4 in the database and B is the ROC curve of ARL6IP4 in the clinical samples. Detailed Implementation
[0059] The present invention will be further illustrated below with reference to specific embodiments. These embodiments are for illustrative purposes only and should not be construed as limiting the invention. Those skilled in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the invention. The scope of the invention is defined by the claims and their equivalents. Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods; the reagents, biological materials, etc., used in the following embodiments are commercially available unless otherwise specified.
[0060] Example 1: Establishment of a risk scoring model for predicting the prognosis of colorectal cancer
[0061] I. Experimental Methods
[0062] 1. Data source: RNA sequencing data, corresponding clinicopathological features and mutation data of CRC were retrieved from the TCGA website (https: / / portal.gdc.cancer.gov / ). This website contains 571 tumor samples and 44 adjacent normal samples.
[0063] The GSE39582 dataset, containing 566 CRC samples, was downloaded from the GEO website (https: / / www.ncbi.nlm.nih.gov / geo / ) and used for the validation cohort. The list of LRGs was selected from DrLLPS (http: / / llps.biocuckoo.cn / ), an online database containing 150 scaffold proteins that drive LLPS, 987 regulators that help regulate LLPS, and 8148 potential target proteins that may influence membrane-bound compartment formation. 3611 LRGs identified in Homo sapiens were screened for further research.
[0064] 2. LRGs Differential Expression Analysis and Functional Labeling: Venn diagrams were used to display the intersection genes between TCGA-CRC and LRGs; the limma package was used to screen for differentially expressed genes between colorectal cancer and adjacent normal samples, with |log2FC|>1 and false discovery rate (FDR)<0.05 as filtering criteria; heatmap analysis was generated using ggplot2 and pheatmap packages; principal component analysis (PCA) was used to visualize the distribution including normal and tumor samples; GO and KEGG were used to explore the potential functions and pathways of LRGs.
[0065] 3. Construction and Validation of the LRGs Risk Model: The expression data were log2 transformed for subsequent analysis; the batch effect between TCGA-CRC and GSE39582 was adjusted using the R package "sva". By excluding genes not simultaneously present in the GEO and TCGA databases, we screened 773 differentially expressed LRGs. Patients were randomly assigned to a training cohort and a testing cohort in a 1:1 ratio. We then assessed whether there were statistically significant differences in clinical characteristics (including age, sex, tumor stage, T, N, M) between the two groups. In the training cohort, univariate Cox regression was used to select LRGs with prognostic value. In univariate Cox regression analysis, p < 0.01 was considered statistically significant. To avoid overfitting, we performed LASSO Cox regression using the R package "glmnet", followed by multivariate Cox regression to further screen prognostic LRGs. Finally, we constructed a new prognostic feature using 8 genes. We calculated the risk score for each sample using the following formula: Risk Score = Based on the median risk score calculated from the training cohort, all patients were divided into low-risk and high-risk groups. To further evaluate feasibility, GSE39582 was used as an external validation cohort, with risk score calculation methods consistent with those used in the training group. Furthermore, principal component analysis (PCA) was employed to study the distribution of risk scores. To determine the prognosis of the two patient groups, Kaplan-Meier (KM) analysis was performed using the R packages “survival” and “survminer”. Receiver operating characteristic (ROC) analysis was used to validate the predictive specificity and sensitivity of the risk model, and the area under the curves (AUC) was calculated.
[0066] 4. Correlation of LRGs risk model with clinical factors: To evaluate the prognostic value of the risk score, univariate and multivariate Cox regression analysis were used to screen independent risk factors significantly associated with survival (p < 0.05). ROC curves were used to evaluate the sensitivity and specificity of the risk score and clinical factors by the "timeROC" package. In the process of constructing nomograms, we utilized the "survival" and "rms" packages and included gender, risk score, age, and tumor stage as variables in the analysis. We plotted calibration curves to assess the consistency between the actual probability and the predicted probability of 1-year, 3-year, and 5-year overall survival (OS). The C-index was used to evaluate the predictive ability of the nomogram.
[0067] II. Experimental Results
[0068] 1. LRGs differential expression analysis and functional annotation
[0069] The flowchart of our study is shown in Figure 1 . We intersected the genes in the TCGA-CRC cohort with the LRGs in the DrLLPS database, resulting in 3530 LRGs ( Figure 2 A). Then, in the TCGA-CRC cohort, a total of 3526 genes with expression data were screened for subsequent differential expression analysis between colorectal cancer samples and normal samples. According to the thresholds of FDR < 0.05 and |log2FC| > 1, we identified 828 differentially expressed LRGs, of which 299 were down-regulated and 529 were up-regulated. Heatmaps show the top 50 up-regulated and down-regulated differentially expressed LRGs, respectively ( Figure 2 B). Based on these differentially expressed LRGs, PCA showed that there was a significant difference between normal samples and colorectal cancer samples ( Figure 2 C). We performed GO and KEGG pathway annotation analysis to reveal the potential biological functions of differentially expressed LRGs in CRC. GO analysis showed that these differentially expressed LRGs were related to the process of LLPS, such as membrane-free organelle assembly ( Figure 2 D). KEGG analysis showed that these differentially expressed LRGs were related to tumor-related pathways, such as cell cycle, DNA replication, HIF-1 signaling pathway, platinum resistance, p53 signaling pathway, etc. Figure 2 E).
[0070] 2. Construction and validation of the predictive risk model based on LRGs
[0071] CRC samples from the TCGA database were randomly divided into a training group (50%, n=271 cases) and a validation group (50%, n=271 cases). Furthermore, clinical characteristics were compared between the two groups, and no differences were found. Next, to construct new prognostic features, univariate Cox regression analysis was performed in the training cohort to preliminarily screen genes associated with prognosis. 43 genes were identified as prognostic genes for colorectal cancer, of which 10 were considered protective genes (HR<1), and the remaining 33 were considered risk genes (HR>1). Figure 3 A). Furthermore, we performed LASSO Cox regression to reduce overfitting ( Figure 3 B, C), and then perform multivariate Cox regression analysis (B, C), Figure 3 D). We obtained eight genes associated with overall survival (OS) in colorectal cancer patients (ARL6IP4, ATP2A1, CRYAB, ELFN2, MAP2, MAT1A, ORC1, and POU4F1) to construct prognostic features. Then, a risk score was calculated for each colorectal cancer patient using the following algorithm: Risk score = Coef i and Exp i These represent the coefficients and expression levels associated with each risk gene, respectively. Therefore, we can calculate the risk score for each sample using the above formula. Patients were divided into high-risk and low-risk groups based on the median risk score of the training cohort. We analyzed the differential expression of eight prognostic characteristic genes in the high-risk and low-risk groups. PCA showed the expression of all genes in TCGA-CRC patients (…). Figure 3 E), LRGs ( Figure 3 F) and risk genes ( Figure 3 The distribution of risk genes (G) was observed. Of particular note, colorectal cancer patients were distributed in significantly different directions based on their risk genes, suggesting that risk characteristics may effectively distinguish between high-risk and low-risk groups. KM analysis showed that the overall survival (OS) in the high-risk group was significantly lower. Figure 4 Both risk score and PFS were significantly lower in the high-risk group than in the low-risk group. The distribution of risk score and survival status showed that in the training cohort, mortality increased with increasing risk score. The heatmap illustrates the expression of eight genes in the high-risk and low-risk groups. Figure 4 D). Next, we used the same formula to calculate a risk score for each patient in the validation cohort and the entire cohort, which were used to validate the effectiveness of this risk characteristic. The results were similar to those observed in the training group. Figure 4 B, C, E, F). Furthermore, the ROC curve was used to evaluate the model's predictive ability. In the training cohort, the AUCs for 1 year, 3 years, and 5 years were 0.863, 0.840, and 0.752, respectively. Figure 4G). For the validation queue, the AUCs for 1 year, 3 years, and 5 years were 0.669, 0.676, and 0.648, respectively. Figure 4 H), and the AUCs for the entire queue were 0.748, 0.749, and 0.700, respectively. Figure 4 These results highlight the superior predictive accuracy of our model. Furthermore, we introduced the GEO dataset (GSE39582) as an external validation cohort. The KM curves for GSE39582 show that the prognosis is worse in the high-risk group than in the low-risk group, consistent with the prognostic results observed in the TCGA cohort. Figure 4 JK). Furthermore, to evaluate the effectiveness of the LRGs risk model in predicting the prognosis of colorectal cancer patients, we categorized the entire patient population into different subgroups based on clinical characteristics, including tumor stage I-II or III-IV, male or female, younger (≤65 years) or older (>~65 years), N0 or N1-N2, and M0 or M1. These results affirmed the effectiveness and robustness of the risk model in predicting survival outcomes in colorectal cancer patients.
[0072] 3. Independent prognostic significance and clinical characteristics of the risk model: Univariate Cox regression analysis showed that risk score, age, tumor stage, and T, N, and M stages were important prognostic factors affecting overall survival (OS) in colorectal cancer patients. Furthermore, multivariate analysis showed that risk score was an independent prognostic indicator of OS (HR 1.030, 95% CI 1.019-1.041, p<0.001). In addition, ROC analysis showed that the AUC of risk score was 0.749, which was superior to tumor stage (AUC=0.720) in predicting biomarkers. Figure 5 A). We created a prognostic nomogram that included gender, risk score, age, and tumor stage ( Figure 5 B). The calibration curves show that the nomogram's predictions of the 1-year, 3-year, and 5-year OS probabilities are in very good agreement with the ideal predictions. Figure 5 C). The C-index value of the nomogram was significantly higher than that of other clinical features ( Figure 5 D). These results indicate that nomograms are effective in predicting patient prognosis and can serve as a tool for clinical decision-making. Furthermore, risk scores differ when grouped by tumor stage, T stage, N stage, and M stage. The heatmap shows the correlation between the expression levels of eight LRGs in the TCGA database and clinicopathological features of CRC patients ( Figure 5 E).
[0073] Example 2: Relationship between risk score and immune characteristics and drug sensitivity
[0074] I. Experimental Methods
[0075] 1. Immune Characterization and Tumor Mutation Burden Analysis: The CIBERSORT algorithm was used to assess differences in 22 immune cell types between high-risk and low-risk groups. Furthermore, correlation analysis was used to explore the relationship between immune cell types and characteristic genes. The “estimate” R package was used to calculate matrix, immune, and estimate scores. We then compared the differences between the two risk groups. We used the “maftools” package to plot a waterfall plot to visualize the mutation landscape in our sample. We used the “survival” package to investigate the relationship between TMB, risk score, and patient survival.
[0076] 2. Immunotherapy Response and Drug Sensitivity: The Tumor Immune Dysfunction and Rejection (TIDE) algorithm (http: / / tide.dfci.harvard.edu / ) was used to assess the responsiveness of the two patient groups to immune checkpoint blockade (ICB) therapy. Differences in drug sensitivity between the two groups were analyzed using the "oncoppredict" package in R. Detailed data were obtained from the Cancer Drug Sensitivity Genomics (GDSC) database (https: / / www.cancerrxgene.org / ).
[0077] II. Experimental Results
[0078] 1. Relationship between risk score and immune cell infiltration
[0079] The tumor immune microenvironment plays a crucial role in regulating cancer progression, clinical outcomes, and treatment response, a fact widely recognized. To explore the correlation between immune cell characteristics and risk scores, we first assessed the relative scores of 22 tumor-infiltrating immune cells in each CRC sample using the CIBERSORT algorithm. Compared to the low-risk group, we observed more memory B cells and T regulatory cells (Tregs) infiltration in the high-risk group, but fewer CD8+ T cells, M1 macrophages, and resting dendritic cells. Figure 6 A). Spearman correlation analysis showed that the risk score was negatively correlated with resting dendritic cells, macrophage M1 cells, and CD8+ T cells, but positively correlated with memory B cells and Tregs. Figure 6 B). The composition of each typical immune cell is as follows: Figure 6 As shown in C. Furthermore... Figure 6 The results showed that all eight genes were associated with immune cell infiltration, suggesting that these genes may serve as potential targets for CRC immunotherapy.
[0080] 2. Relationship between risk characteristics, immune properties, and drug sensitivity
[0081] The ESTIMATE algorithm showed that the interstitial score and ESTIMATE score of the high-risk group were higher than those of the low-risk group, while there was no significant difference in immune scores between the two groups. Figure 7 A), indicating that interstitial cell infiltration increases with increasing risk score. Analysis of immune-related functional differences showed that the high-risk group had lower APC co-inhibition scores, cell lysis activity scores, pro-inflammatory scores, MHC class I scores, and T cell co-inhibition scores ( Figure 7 B). Furthermore, the TIDE score in the high-risk group was higher than that in the low-risk group ( Figure 7 CE).
[0082] 3. Correlation between risk characteristics and TMB
[0083] We analyzed the mutation status of CRC in the high-risk and low-risk groups. Maftools analysis revealed the top 15 most frequently mutated genes in both the high-risk and low-risk groups (…). Figure 8 A, B). We found that the mutation incidence was higher in the high-risk group compared to the low-risk group. In particular, the most common mutated gene was APC. Previous studies have shown that high TMB predicts poor prognosis for various types of cancer. In assessing the prognostic value of TMB, we observed no statistically significant difference in survival expectations between the high-TMB and low-TMB groups, but the prognosis tended to be worse in the high-TMB group. Figure 8 C). Next, colorectal cancer patients were divided into four groups based on TMB and risk scores for survival analysis: low TMB + low risk, low TMB + high risk, high TMB + low risk, and high TMB + high risk. The results showed that patients in the low TMB + low risk group had better survival outcomes than the other three groups. Figure 8 D).
[0084] Example 3: Expression level and phase separation ability of ARL6IP4 in colorectal cancer
[0085] I. Experimental Methods
[0086] 1. Clinical Specimens, RNA Extraction, and qRT-PCR: Ten pairs of colorectal cancer tissues and corresponding adjacent normal tissues (ANTs) were collected from patients with colorectal cancer who underwent radical or palliative resection at the Department of Gastrointestinal Surgery, Affiliated Hospital of Xuzhou Medical University, as a cohort, with the consent of all patients. Total RNA was extracted from the cohort using RNA Isolater Total RNA Extraction Reagent (Vazyme, Nanjing, China). 1000 ng of RNA was reverse transcribed using the SweScript RT I First-Strand cDNA Synthesis Kit (Servicebio, Wuhan, China). cDNA analysis was performed using real-time quantitative polymerase chain reaction (qRT-PCR) on a LightCycler 96 system (Roche, Switzerland) using the ChamQ SYBR qPCR Master Mix (Vazyme, Nanjing, China). GAPDH was used as an endogenous control. Primer information: ARL6IP4 F: TGACCAAGGAGGAGTGGGAT, ARL6IP4 R: RCCTCTAGGACCTCGCCATCT.
[0087] 2. Phase separation assay for colorectal cancer cells: HCT116 cells were cultured in glass-bottomed dishes and then transfected with the ARL6IP4-GFP plasmid. GFP-labeled spots were observed using a confocal laser scanning microscope (CLSM, Leica STELLARIS 5, Germany). Photobleaching followed by fluorescence recovery (FRAP) was performed using a full-power 488 nm laser. Time-lapse images were captured after photobleaching to monitor the recovery of fluorescence intensity every second over a total duration of 40 seconds. GraphPadPrism software was used for quantitative analysis and graphical representation of the data.
[0088] II. Experimental Results
[0089] To investigate whether phase segregation exists among the genes that construct the risk characteristics of this study, we selected ARL6IP4, which had the highest coefficient in multivariate Cox regression analysis, for experimental validation. First, the upregulation of ARL6IP4 was confirmed in the TCGA-CRC cohort (…). Figure 9 A). Furthermore, we assessed the expression level of ARL6IP4 in 10 pairs of fresh frozen tissue specimens using quantitative qRT-PCR. The results showed that ARL6IP4 was significantly upregulated in most CRC tissues compared to ANTs (A). Figure 9 B). Subsequently, after ectopic transfection of the ARL6IP4-GFP plasmid, we observed the formation of droplet-like structures by ARL6IP4 in live CRC cells using 3D confocal microscopy. Figure 9C). To assess the fluidity of ARL6IP4 condensates, we performed FRAP experiments. The results showed that fluorescence recovery in ARL6IP4 condensates was rapid, occurring within ~40 seconds after photobleaching Figure 9 D). In addition, by time-lapse imaging, we observed that punctate structures exhibited fusion behavior in live CRC cells Figure 9 E). In summary, our results show that ARL6IP4 is upregulated in CRC and phase separates.
[0090] Further, we performed ROC analysis of ARL6IP4 in colorectal cancer patients using tumor (n=647) and normal tissues (n=51) in the TCGA database, and found that its AUC in the database was 0.839 Figure 10 A); further, we verified the diagnostic value of ARL6IP4 in colorectal cancer patients by performing ROC analysis of tumor (n=40) and normal tissues (n=40) in a clinical cohort, and found that its AUC in the clinic was 0.734 Figure 10 B). Taken together, these results show that ARL6IP4 has good diagnostic performance.
[0091] The above description of the embodiments is only for the purpose of understanding the method of the present application and its core idea. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and modifications can be made to the present application, and these improvements and modifications will also fall within the scope of protection of the claims of the present application.
Claims
1. Use of a reagent for detecting the expression level of ARL6IP4 in a sample in the manufacture of a product for diagnosing colorectal cancer.
Citation Information
Patent Citations
Novel model for prognosis prediction and diagnosis of colorectal cancer and application thereof
CN114594259A
Biomarker and model for prognosis risk prediction of colorectal cancer and application of biomarker and model
CN116030880A
Composition for diagnosis of pancreatic cancer
WO2023177206A1