Marker panels, products, systems and uses thereof for colorectal cancer prognosis
By using the cosine distance calculation of 14 gene markers combinations and the CMS typed characteristic genome expression template, the accuracy of the evaluation of high-risk subtypes in colorectal cancer patients was solved, and more accurate prognostic evaluation and treatment decisions for bowel cancer were achieved.
Patent Information
- Application Number
- CN202410127289.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-01-30
AI Technical Summary
The lack of specific molecular markers in the prior art is used to evaluate high-risk subtypes in patients with colorectal cancer, resulting in a lack of accuracy in adjuvant chemotherapy decisions, especially for stage II/III patients, it is difficult to accurately identify individuals who truly benefit.
A combination of 14 gene markers (IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4) was used to evaluate the prognosis of intestinal cancer. The cosine distance calculation of gene expression data and the genomic expression template of CMS typing was determined to predict patient prognosis.
It improves the identification accuracy of high-risk patients, enhances the clinical prediction ability of bowel cancer prognosis, and supports more accurate treatment decisions.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure HDA0004688716540000011
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine; in particular, to a marker group, product, system and use thereof for evaluating the prognosis of colorectal cancer. Background Art
[0002] Colorectal cancer (CRC) has become the third most common malignant tumor in my country. For patients who have undergone surgery for colorectal cancer, postoperative pathological evaluation is the most important prognostic indicator and the basis for postoperative treatment. Patients with advanced CRC should receive adjuvant therapy after surgery, but the decision on adjuvant therapy for stage II / III patients remains controversial.
[0003] Retrospective studies have found that approximately 10-20% of stage II patients will experience recurrence and metastasis after surgery. Furthermore, high-risk stage II patients may also benefit from adjuvant chemotherapy. Currently, the assessment of whether stage II colorectal cancer patients require adjuvant chemotherapy is primarily based on histological criteria: tumor invasion depth, degree of differentiation, presence of lymphovascular invasion, presence of perineural invasion, total number and number of positive lymph nodes, and pathological findings of resection margins. Molecular studies suggest that patients with MSI-H (microsatellite instability high) or dMMR (mismatch repair deficient) tumors benefit less from 5-FU chemotherapy. For stage III CRC patients, the appropriate duration of chemotherapy remains controversial. Recent results from the IDEA study indicate that the 5-year survival difference between 3 and 6 months of adjuvant chemotherapy in stage III patients is minimal. In low-risk stage III, 4 cycles of XELOX have a significantly better 5-year survival rate than 8 cycles of XELOX. In high-risk stage III, 4 cycles of XELOX only reduce the 5-year survival rate by 1% compared with 8 cycles of XELOX, but with significantly reduced toxicity and side effects. Furthermore, current clinical treatment results suggest that adjuvant chemotherapy is not unique to high-risk stage II / III CRC. The key to these issues lies in the lack of specific molecular markers to define or determine "high risk." Therefore, how to screen specific molecular markers, conduct precise molecular characterization and subtype analysis of stage II / III CRC patients, and identify truly "high-risk" patients who can benefit from adjuvant chemotherapy remains a pressing clinical challenge. Summary of the Invention
[0004] To address these technical issues, the inventors, through extensive and in-depth research and analysis, have discovered and refined a panel of markers that can be used to assess the prognosis of colorectal cancer. This panel enables precise molecular characterization and subtype analysis of colorectal cancer patients, significantly improving the accuracy of identifying truly "high-risk" patients and clinical prognosis.
[0005] On the one hand, the present invention provides a marker panel for evaluating the prognosis of colorectal cancer in a subject, which has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4.
[0006] On the other hand, the present invention provides a product or product group for evaluating the prognosis of colorectal cancer in a subject, which comprises a reagent for detecting each marker of a marker group from a biological sample of the subject, wherein the marker group has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4.
[0007] In another aspect, provided herein is a system for assessing the prognosis of bowel cancer in a subject, comprising a memory and one or more processors;
[0008] Wherein, the memory comprises:
[0009] expression data of each gene of a marker panel from a biological sample of a subject,
[0010] CMS typing feature genome expression template data, and
[0011] one or more processor-executable instructions;
[0012] The marker panel comprises the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4; and
[0013] The one or more processor-executable instructions are configured to:
[0014] (a) obtaining gene expression data of the marker panel from a biological sample of a subject;
[0015] (b) Calculate the cosine distance between the gene expression data in (a) and the CMS typing feature gene expression template;
[0016] (c) determining the CMS type based on the cosine distance calculated in (b); and
[0017] (d) Determine the subject's colorectal cancer prognosis based on the CMS classification obtained in (c).
[0018] In another aspect, a computer-readable medium is provided, comprising:
[0019] Expression level data of each gene of a marker panel from a biological sample of a subject, wherein the marker panel comprises the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4;
[0020] CMS typing feature genome expression template data, and
[0021] Instructions for performing a method comprising:
[0022] (a) obtaining gene expression data of a marker panel from a biological sample of a subject;
[0023] (b) Calculate the cosine distance between the gene expression data in (a) and the CMS typing characteristic gene group expression template;
[0024] (c) determining the CMS type based on the cosine distance calculated in (b); and
[0025] (d) Determine the subject's colorectal cancer prognosis based on the CMS classification obtained in (c).
[0026] In another aspect, the present invention provides an electronic device loaded with the computer-readable medium.
[0027] On the other hand, the present invention provides a reagent for detecting each marker of a marker group from a biological sample of a subject in the preparation of a product or product group for evaluating the prognosis of colorectal cancer in a subject, wherein the marker group has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4.
[0028] On the other hand, the present invention provides a product or product group comprising reagents for detecting each marker of a marker group from a biological sample of a subject for use in evaluating the prognosis of colorectal cancer in a subject, wherein the marker group has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will be further described below with reference to the accompanying drawings, which are merely illustrative of exemplary embodiments of the present invention and are not intended to limit the scope of the present invention.
[0030] Figure 1 is an exemplary flow chart of an overall method for screening prognostic markers for colorectal cancer according to one of the exemplary embodiments herein;
[0031] Figure 2 This is the result of the volcano plot of differentially expressed genes in Example 1 of this article;
[0032] Figure 3 is a correlation coefficient matrix diagram between the CMS40 markers in Example 1 herein;
[0033] Figure 4 This is the result of survival analysis of the CMS molecular typing data obtained using the CMS40 marker group in Example 1 herein;
[0034] Figure 5 This is the result of survival analysis of the CMS molecular typing data obtained using the CMS20 marker group in Example 1 herein;
[0035] Figure 6 This is the result of survival analysis of the CMS molecular typing data obtained using the CMS14 marker group in Example 1 herein;
[0036] Figure 7 This is the result of a three-category recurrence-free survival analysis of the CMS molecular typing data obtained using the CMS40 marker group in Example 3 herein. DETAILED DESCRIPTION
[0037] The meaning of scientific and technological terms in this application is consistent with the general understanding of those skilled in the art, unless otherwise specified. In this application, "one" or its combination with various quantifiers includes both singular and plural meanings, unless otherwise specified. In this application, for the same parameter or variable, when multiple numerical values, numerical ranges, or combinations thereof are given for description, it is equivalent to specifically revealing these numerical values, range end values, and numerical ranges formed by any combination thereof. In this application, any numerical value, whether or not it is accompanied by a modifier such as "about", covers an approximate range that can be understood by those skilled in the art, such as plus or minus 10%, 5%, etc. In this article, each "embodiment" refers equally to and covers the implementation methods of various methods and systems of this application. In this application, one or more technical features in any implementation method can be freely combined with one or more technical features in any one or more other implementation methods, and the implementation methods obtained thereby also belong to the content disclosed in this application.
[0038] The following lists some terms used in the embodiments of the present invention. Within the scope of the present specification and claims, the relevant terms are defined as follows. Other terms not listed here have the commonly used definitions in the art and their meanings are well known to those skilled in the art.
[0039] As used herein, the term "cancer" refers to the presence of cells that have typical characteristics of a cancerous cell, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features known in the art. In one embodiment, "cancer" can be intestinal cancer or liver cancer. In one embodiment, "cancer" can include pre-malignant cancers as well as malignant cancers.
[0040] In one embodiment, as will be appreciated by those skilled in the art, the methods as described herein do not involve steps performed by a doctor / physician. Therefore, before the final diagnosis performed by a physician can be provided to the subject, the results obtained by the methods as described herein need to be combined with clinical data and other clinical manifestations. The final diagnosis of whether a subject suffers from bowel cancer is within the scope of the physician and is not considered to be part of this disclosure. Therefore, the terms "determine", "detect" and "diagnose" as used herein refer to the probability or possibility of identifying a subject suffering from a disease (such as bowel cancer) at any stage of development or determining the subject's susceptibility to developing the disease. In one embodiment, "diagnosis", "determination" and "detection" are performed before symptoms are manifested. In one embodiment, "diagnosis", "determination" and "detection" allow clinicians / physicians (in combination with other clinical manifestations) to confirm bowel cancer in a subject suspected of having bowel cancer.
[0041] As used herein, the term "sample", "sample" or "biological sample" means a sample collected from a subject for detecting the type and amount of colorectal cancer markers therein. The subject sample may be from the circulatory system, that is, from the blood, or may not be from the circulatory system, that is, not from the blood. The subject sample may be any sample suitable for detecting colorectal cancer markers, and its sources include tissue, whole blood, bone marrow, pleural fluid, peritoneal fluid, central spinal fluid, milk, urine, tears, sweat, saliva, organ secretions, and washing fluids of the bronchi, nasal cavity, throat, etc. In some embodiments, the biological sample is selected from: fresh tissue samples, frozen tissue samples or paraffin-embedded tissue samples (such as FFPE samples).
[0042] As used herein, the term "prognosis" refers to an estimate of the likely outcome (with or without treatment) within a certain time period in the future based on the subject's current condition. Generally, the result is in the form of a probability (%) of the outcome, such as the probability of cure (%), recurrence (%), or death (%).
[0043] As used herein, the term "correlation analysis" refers to a statistical method used to investigate whether a relationship exists between two or more random variables. Its primary purpose is to determine whether there is a statistical correlation or dependence between two or more variables and to quantify the extent and form of this correlation. Basic methods of correlation analysis include linear correlation, rank correlation, and distance correlation.
[0044] As used herein, the term "correlation coefficient" is a statistic used in correlation analysis to quantify the closeness of the relationship between variables. It reflects the strength of the linear relationship between two variables. A larger absolute value indicates a stronger linear correlation between the two variables; a value closer to 0 indicates a weaker correlation. Commonly used correlation coefficients include the Pearson correlation coefficient and the Spearman correlation coefficient. The correlation coefficient referred to in this study is the Pearson correlation coefficient.
[0045] As used herein, the term "quantile normalization" is a commonly used sequencing data processing method. Quantile normalization is an important preprocessing step in RNA-seq analysis, which can effectively eliminate systematic errors between different samples and make the samples comparable. The main idea is: first, the expression data of a gene in all samples are merged together, sorted, and the expression corresponding to each sequence position (such as 1%, 5%, 25%, etc. quantiles) is calculated. Then, for each sample, the original expression is replaced by the average expression of the quantile sequence corresponding to the original expression of the gene in the sample. Finally, the above process is repeated for all genes in the sample. In this way, through quantile conversion, the expression distribution of different batches and different samples tends to be consistent, eliminating the influence of different sequencing depths and technical errors between samples. The advantages of quantile normalization are: simple and effective, easy to implement the program, no reference sample is required, it is robust to missing values and outliers, and the relative size relationship of expression between samples remains unchanged.
[0046] As used herein, the term "consensus molecular subtypes" or "CMS" is a gene expression-based molecular typing method for colorectal cancer. CMS typing is determined by calculating the gene expression similarity of each sample with these four patterns, and the method uses cosine similarity and classification model to determine which type the sample belongs to. CMS typing is related to patient prognosis and drug efficacy, and can guide precision treatment. Compared with a single biomarker, CMS typing comprehensively utilizes whole genome expression information to more comprehensively reflect the biological characteristics of the tumor. It is an important tool for precision medicine in colorectal cancer.
[0047] The term "CMS1-inflammatory" is a consensus molecular subtype of colorectal cancer, representing inflammatory colorectal cancer. Its key characteristics are: increased immune cell infiltration, particularly T lymphocytes and macrophages; upregulation of inflammation-related genes, such as inflammatory cytokines IL-6 and IL-8; increased expression of immune checkpoints such as PD-L1; activation of genes associated with natural killer (NK) T cells and TH1 T cells; and often higher microsatellite instability. It is associated with a better prognosis and higher sensitivity to immunotherapy, such as anti-PD-1.
[0048] The term "CMS2-Transit Amplifying" is a consensus molecular subtype of colorectal cancer, representing a highly proliferative, poorly differentiated colorectal cancer subtype. Its main characteristics are: high expression of genes related to intestinal epithelial cell proliferation and transmission, such as cell cycle proteins, proliferation-associated antigens, etc.; downregulation of intestinal epithelial cell differentiation-related genes; abnormal activation of the WNT pathway; mutations in associated tumorigenic driver genes such as APC, TP53, and KRAS; pathological manifestations of poorly differentiated adenocarcinoma; poor prognosis; and sensitivity to standard chemotherapy.
[0049] The term "CMS2-Enterocyte" is a consensus molecular subtype of colorectal cancer, representing a differentiated intestinal epithelial subtype associated with the differentiation and absorptive functions of normal intestinal epithelial cells. Key features include: high expression of genes associated with intestinal epithelial differentiation and absorption, such as alkaline phosphatase and intestinal alkaline phosphatase; upregulation of cell adhesion proteins such as E-cadherin; downregulation of the WNT pathway; driver genes including APC, KRAS, and TP53; pathological diagnosis of well-differentiated adenocarcinoma; favorable prognosis; and sensitivity to standard first-line chemotherapy.
[0050] The term "CMS3-Goblet like" is a consensus molecular subtype of colorectal cancer, representing the goblet cell-like colorectal cancer subtype. Its main characteristics include: high expression of goblet cell differentiation-related genes, such as MUC2 and TFF3; activation of mucus production-related pathways; often activation of the RAS pathway, and KRAS mutations are more common; the pathological type is mostly mucinous adenocarcinoma; less immune infiltration; more occurrence in the right (colon) hemicolon; poor sensitivity to standard first-line chemotherapy; and poor prognosis.
[0051] The term "CMS4-Stem like" is a consensus molecular subtype of colorectal cancer, representing a stem cell-like subtype of colorectal cancer. Its main characteristics include: high expression of stem cell marker genes, such as LGR5 and EPHB2; upregulation of EMT-related genes; activation of the WNT pathway and Notch pathway; high tumor heterogeneity and poor differentiation; pathological types are mostly signet ring cell carcinoma and mucinous carcinoma; tumor stem cells proliferate and are resistant to treatment; prognosis is poor; and the risk of tumor recurrence is high.
[0052] As used in this article, "p-value" represents the probability of the observed data appearing in the hypothesis space. Specifically, the p-value represents the probability of obtaining data equal to or more extreme than the observed result when the null hypothesis of the hypothesis test is true. Generally speaking, if the p-value is very small, for example, less than 0.01, it means that the result is extremely unlikely to be a random event under the null hypothesis, and the null hypothesis is rejected, that is, the result is statistically significant. If the p-value is very large, for example, greater than 0.05, the null hypothesis cannot be rejected, that is, the result is not statistically significant. The smaller the p-value, the higher the statistical significance of the result. Commonly used significance judgment thresholds are 0.05 and 0.01. Therefore, the p-value reflects the probability of observing the current result under the premise that the null hypothesis is true, and is an important basis for judging whether the hypothesis test results are significant. The smaller the p-value, the more significant the result.
[0053] As used in this article, the term "t-test" refers to a statistical method used to test whether two sample means are significantly different. The basic idea behind a t-test is to construct a hypothesis, calculate the observed statistic (t-value), determine the p-value based on the t-distribution, and then use the p-value to determine whether the null hypothesis is true.
[0054] As used herein, the term "survival analysis" means that survival analysis requires the preparation of survival time (time), status (status) and other characteristic data. Among them, the status is generally represented by 0 (no) or 1 (yes) to indicate whether an event (such as death) has occurred.
[0055] As used herein, the terms "subject," "patient," and "subject" are used interchangeably and generally refer to a mammal, such as a bovine, equine, ovine, porcine, canine, feline, rodent, or primate, such as a human or non-human mammal.
[0056] A. Marker panel for assessing the prognosis of colorectal cancer
[0057] On the one hand, the present invention provides a marker panel for evaluating the prognosis of colorectal cancer in a subject, which has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4.
[0058] Herein, the marker panel having the above-mentioned 14 gene markers may also be equivalently referred to as "CMS14" or "CMS14 marker panel."
[0059] As used herein, a "product set" refers to two or more products that can be provided together (e.g., in the same kit or packaging (such as a kit)) or separately (e.g., not in the same kit or packaging (such as a kit)), and these products are to be used in combination and cannot be used individually (e.g., for assessing the prognosis of bowel cancer in a subject).
[0060] In some embodiments, the intestinal cancer is colorectal cancer. In some embodiments, the intestinal cancer is stage II / III colorectal cancer. In some embodiments, the intestinal cancer can be selected from rectal cancer, left colon cancer or right colon cancer.
[0061] In some embodiments, the expression (eg, expression level) of each gene marker in the marker panel can be used to determine the CMS typing of a subject.
[0062] In some embodiments, the comparison of the expression (eg, expression level) of each gene marker in the marker group with the CMS typing characteristic gene set expression template can be used to determine the CMS typing of the subject to further evaluate the subject's colorectal cancer prognosis.
[0063] As used herein, the term "CMS typing characteristic genome expression template" refers to a pre-set classification template, for example, a pre-set classification template data table. The classification template may be a set of labeled genes, which are labeled with different CMS classifications. With respect to a specific gene, if the gene is labeled as a specific category, the gene has a higher expected expression in samples belonging to that category compared to samples that do not belong to that category. In an exemplary embodiment, the classification template data table may include at least two groups (e.g., two columns) of information: probes (e.g., Entrez IDs) and categories (e.g., CMS categories). In another exemplary embodiment, the classification template data table may include three groups (e.g., three columns) of information: probes (e.g., Entrez IDs), categories (e.g., CMS categories), and gene symbols.
[0064] In some embodiments, the comparison of the expression (e.g., expression level) of each gene marker in the marker group with the CMS typing characteristic gene set expression template is expressed as the cosine distance between the expression (e.g., expression level) of each gene marker in the CMS14 marker group and the CMS typing characteristic gene set expression template.
[0065] In this article, the term "cosine distance" is defined as the characteristic distance, which is the default characteristic distance of the nearest template prediction model. The shortest distance d indicates that the sample is closest to the characteristic genome expression template of a certain CMS typing, that is, it is most likely to be this typing. As a measure of the typing p-value for statistical significance test, a random permutation test is used. By randomly extracting characteristic genes (the default value is 1000 times) to generate a random distribution of characteristic distances, the distance between the tested sample and the typing characteristic template is compared with the randomly generated distance distribution and the corrected false discovery rate (FDR) to calculate the p-value of the significance test. The smaller the p-value, the stronger the statistical significance of the shortest cosine characteristic distance, which means that the predicted CMS typing is more reliable (usually the threshold for statistical significance p is p<0.05).
[0066] In some embodiments, the CMS classification may include one or more of the following: CMS1-Inflammatory, CMS2-Transit Amplifying, CMS2-Enterocyte, CMS3-Goblet like, and CMS4-Stem like.
[0067] In one embodiment, the IFIT3 gene has the nucleotide structure shown in ENSG00000119917, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0068] In one embodiment, the CXCL13 gene has the nucleotide structure shown in ENSG00000156234, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0069] In one embodiment, the STAT1 gene has the nucleotide structure shown in ENSG00000115415, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0070] In one embodiment, the CXCL9 gene has the nucleotide structure shown in ENSG00000138755, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0071] In one embodiment, the CA4 gene has a nucleotide structure as shown in ENSG00000167434, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0072] In one embodiment, the AQP8 gene has the nucleotide structure shown in ENSG00000103375, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0073] In one embodiment, the SLC4A4 gene has the nucleotide structure shown in ENSG00000080493, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0074] In one embodiment, the EREG gene has a nucleotide structure as shown in ENSG00000124882, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0075] In one embodiment, the AREG gene has a nucleotide structure as shown in ENSG00000109321, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0076] In one embodiment, the REG4 gene has a nucleotide structure as shown in ENSG00000134193, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0077] In one embodiment, the SPINK4 gene has a nucleotide structure as shown in ENSG00000122711, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0078] In one embodiment, the SFRP2 gene has the nucleotide structure shown in ENSG00000145423, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0079] In one embodiment, the ZEB2 gene has the nucleotide structure shown in ENSG00000169554, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0080] In one embodiment, the SFRP4 gene has the nucleotide structure shown in ENSG00000106483, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0081] Without intending to be bound by a particular theory, it was found that in the CMS14 marker panel, IFIT3, CXCL13, STAT1 and CXCL9 were beneficial for specifically distinguishing CMS1-inflammatory type; CA4, AQP8 and SLC4A4 were beneficial for specifically distinguishing CMS2-intestinal epithelial cell type; EREG and AREG were beneficial for specifically distinguishing CMS2-transient proliferative type; SPINK4 and REG4 were beneficial for specifically distinguishing CMS3-cup type; SFRP2, ZEB2 and SFRP4 were beneficial for specifically distinguishing CMS4-stem type.
[0082] In some embodiments, a marker panel for assessing the prognosis of colorectal cancer in a subject may further include the following six gene markers, based on the CMS14 marker panel: CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1. The resulting marker panel of 20 gene markers may also be equivalently referred to herein as "CMS20" or the "CMS20 marker panel."
[0083] In some embodiments, the marker panel for evaluating the prognosis of colorectal cancer in a subject has the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8, SLC4A4, EREG, AREG, SPINK4, REG4, MUC2, SFRP2, ZEB1, ZEB2 and SFRP4.
[0084] In one embodiment, the CA1 gene has the nucleotide structure shown in ENSG00000133742, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0085] In one embodiment, the CLCA4 gene has the nucleotide structure shown in ENSG00000016602, or a nucleotide sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto.
[0086] In one embodiment, the MS4A12 gene has the nucleotide structure shown in ENSG00000071203, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0087] In one embodiment, the CLDN8 gene has the nucleotide structure shown in ENSG00000156284, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0088] In one embodiment, the MUC2 gene has the nucleotide structure shown in ENSG00000198788, or a nucleotide sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto.
[0089] In one embodiment, the ZEB1 gene has the nucleotide structure shown in ENSG00000148516, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0090] Without intending to be bound by a particular theory, it was found that in the CMS20 marker panel, IFIT3, CXCL13, STAT1 and CXCL9 were beneficial for specifically distinguishing CMS1-inflammatory type; CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8 and SLC4A4 were beneficial for specifically distinguishing CMS2-enterothelial cell type; EREG and AREG were beneficial for specifically distinguishing CMS2-transient proliferative type; SPINK4, REG4 and MUC2 were beneficial for specifically distinguishing CMS3-cup type; SFRP2, ZEB1, ZEB2 and SFRP4 were beneficial for specifically distinguishing CMS4-stem type.
[0091] In some embodiments, the marker panel for assessing the prognosis of colorectal cancer in a subject may further include the following 20 gene markers based on the CMS20 marker panel: CXCL10, AIM2, GBP5, CXCL11, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2, and TWIST1. The marker panel with 40 gene markers thus formed may also be equivalently referred to herein as "CMS40" or "CMS40 marker panel."
[0092] In some embodiments, the marker panel for evaluating the prognosis of colorectal cancer in a subject has the following 40 gene markers: CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, CXCL9, CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, CA2, ASCL2, VAV3, CELP, EREG, RNF43, AREG, SPINK4, REG4, MLPH, TFF3, MUC2, AQP3, COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2 and TWIST1.
[0093] In one embodiment, the CXCL10 gene has the nucleotide structure shown in ENSG00000169245, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0094] In one embodiment, the AIM2 gene has the nucleotide structure shown in ENSG00000163568, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0095] In one embodiment, the GBP5 gene has a nucleotide structure as shown in ENSG00000154451, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0096] In one embodiment, the CXCL11 gene has the nucleotide structure shown in ENSG00000169248, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0097] In one embodiment, the KRT20 gene has the nucleotide structure shown in ENSG00000171431, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0098] In one embodiment, the SLC26A3 gene has the nucleotide structure shown in ENSG00000091138, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0099] In one embodiment, the CA2 gene has a nucleotide structure as shown in ENSG00000104267, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0100] In one embodiment, the ASCL2 gene has the nucleotide structure shown in ENSG00000183734, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0101] In one embodiment, the VAV3 gene has the nucleotide structure shown in ENSG00000134215, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0102] In one embodiment, the CELP gene has the nucleotide structure shown in ENSG00000170827, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0103] In one embodiment, the RNF43 gene has a nucleotide structure as shown in ENSG00000108375, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0104] In one embodiment, the MLPH gene has the nucleotide structure shown in ENSG00000115648, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0105] In one embodiment, the TFF3 gene has a nucleotide structure as shown in ENSG00000160180, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0106] In one embodiment, the AQP3 gene has the nucleotide structure shown in ENSG00000165272, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0107] In one embodiment, the COL3A1 gene has the nucleotide structure shown in ENSG00000168542, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0108] In one embodiment, the SNAI2 gene has a nucleotide structure as shown in ENSG00000019549, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0109] In one embodiment, the CCDC80 gene has a nucleotide structure as shown in ENSG00000091986, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0110] In one embodiment, the AEBP1 gene has the nucleotide structure shown in ENSG00000106624, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0111] In one embodiment, the TIMP2 gene has the nucleotide structure shown in ENSG00000035862, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0112] In one embodiment, the TWIST1 gene has the nucleotide structure shown in ENSG00000122691, or a nucleotide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0113] Without intending to be bound by a particular theory, it was found that among the CMS40 marker panel, CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, and CXCL9 were beneficial for specifically distinguishing CMS1-inflammatory type; CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, and CA2 were beneficial for specifically distinguishing CMS2-intestinal epithelial cell type; and ASCL2 , VAV3, CELP, EREG, RNF43 and AREG are conducive to the specific differentiation of CMS2-transient proliferative type; SPINK4, REG4, MLPH, TFF3, MUC2 and AQP3 are conducive to the specific differentiation of CMS3-cup type; COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2 and TWIST1 are conducive to the specific differentiation of CMS4-stem type.
[0114] B. Products for assessing the prognosis of colorectal cancer
[0115] On the other hand, provided herein is a product or product set for assessing the prognosis of colorectal cancer in a subject, comprising reagents for detecting each marker of a marker set from a biological sample of the subject, wherein the marker set has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4 (i.e., the "CMS14 marker set").
[0116] As used herein, the term "product panel" is intended to refer to a combination of more than one product, for example, two, three, four, or more products. Gene markers in the marker panel described herein may be present separately in different products within the product panel. More than one product within the product panel can be combined to assess a subject's prognosis for colorectal cancer.
[0117] In some embodiments, the marker panel may further include the following six gene markers based on the CMS14 marker panel: CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1 (i.e., forming the "CMS20 marker panel"). In some such embodiments, the marker panel has the following 20 gene markers: CXCL10, AIM2, GBP5, CXCL11, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2, and TWIST1.
[0118] In some embodiments, the marker panel may further include the following 20 gene markers on the basis of the CMS20 marker panel: CXCL10, AIM2, GBP5, CXCL9, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2 and TWIST1 (i.e., forming the "CMS40 marker panel"). In some such embodiments, the marker panel has the following 40 gene markers: CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, CXCL9, CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, CA2, ASCL2, VAV3, CELP, EREG, RNF43, AREG, SPINK4, REG4, MLPH, TFF3, MUC2, AQP3, COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2, and TWIST1.
[0119] In some embodiments, the product can be a reagent product or a kit. In some embodiments, the product can be a combination of products selected from the group consisting of a reagent product or a kit.
[0120] In some embodiments, the reagents for detecting each marker of a marker panel in a biological sample from a subject can be packaged (e.g., encapsulated) in a container contained in the product. In some embodiments, the container is a sealed container (e.g., a sealed container with a lid).
[0121] In some embodiments, the reagent for detecting each marker of a marker panel in a biological sample from a subject is a reagent that facilitates detecting the expression of each gene marker in the marker panel in the biological sample.
[0122] As used herein, "expression" includes the production of mRNA from a gene or a portion of a gene, the production of a protein encoded by the RNA or gene portion, and the appearance of a detectable substance associated with expression. For example, the binding of a cDNA binding ligand (such as an antibody) to a gene or other oligonucleotide, protein, or protein fragment, as well as the chromogenic portion of the binding ligand, are all included within the scope of the term "expression." An increase in half-dot density on an immunoblot, such as a Western blot, is also included within the scope of the term "expression" based on biological molecules.
[0123] In some embodiments, the reagent is a reagent capable of detecting the mRNA level of the marker. Such reagents are well known in the art and include but are not limited to nucleic acid probes that specifically bind to the target sequence, primers that amplify the target sequence, non-specific fluorescent dyes (e.g., SYBR Green I) or combinations thereof.
[0124] In some embodiments, the nucleic acid probe may be a single-labeled nucleic acid probe, such as a radionuclide (such as 32P, 3H, 35S, etc.) labeled probe, a biotin-labeled probe, a horseradish peroxidase-labeled probe, a digoxigenin-labeled probe, or a fluorescent group (such as FITC, FAM, TET, HEX, TAMRA, Cy3, Cy5, etc.) labeled probe; the nucleic acid probe may also be a dual-labeled nucleic acid probe, such as a Taqman probe, a molecular beacon, a displacement probe, a scorpion primer probe, a QUAL probe, a FRET probe, etc.
[0125] In some embodiments, the reagent is a reagent capable of detecting protein levels of the marker. In some embodiments, the reagent for detecting protein levels of the marker includes reagents required for immunological detection; the immunological detection method is selected from ELISA, Elispot, Western blotting, or surface plasmon resonance. Reagents required for immunological detection are well known in the art and include, but are not limited to, antibodies and targeting peptides that specifically bind to at least one of the proteins ZNF33B, PRKX, LEF1, FKBP1A, SERPINB8, and SULT1B1.
[0126] In some embodiments, the reagent carries a detectable label, such as an enzyme (such as horseradish peroxidase, alkaline phosphatase, etc.), a radionuclide (such as 3H, 125I, 35S, 14C, 32P, etc.), a fluorescent dye (such as FITC, TRITC, PE, Texas Red, quantum dots, Cy7, Alexa 750, etc.), an acridinium ester compound, a magnetic bead, colloidal gold or colored glass or plastic (such as polystyrene, polypropylene, latex, etc.) beads, and biotin for binding to avidin (such as streptavidin) modified with the above-mentioned label.
[0127] In some embodiments, the product may further comprise reagents for pre-treating the sample. In some embodiments, the reagents for pre-treating the sample include, but are not limited to, the following reagents: a diluent (e.g., phosphate buffered saline or normal saline) for diluting the sample; an anticoagulant (e.g., heparin) for preventing blood clotting.
[0128] In some embodiments, the product further includes an apparatus (eg, a tool and / or instrument) for detecting gene expression levels in a subject.
[0129] In some embodiments, the product further comprises reagents and / or instruments (eg, tools and / or apparatus) for detecting other disease markers.
[0130] C. System for assessing the prognosis of colorectal cancer
[0131] In another aspect, provided herein is a system for assessing the prognosis of bowel cancer in a subject, comprising a memory and one or more processors;
[0132] Wherein, the memory comprises:
[0133] expression data of each gene of a marker panel from a biological sample of a subject,
[0134] CMS typing feature genome expression template data, and
[0135] one or more processor-executable instructions;
[0136] wherein the marker panel comprises the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4 (i.e., the "CMS14 marker panel"); and
[0137] The one or more processor-executable instructions are configured to:
[0138] (a) obtaining gene expression data of a marker panel CMS14 from a biological sample of a subject;
[0139] (b) Calculate the cosine distance between the gene expression data in (a) and the CMS typing feature gene expression template;
[0140] (c) determining the CMS type based on the cosine distance calculated in (b); and
[0141] (d) Determine the subject's colorectal cancer prognosis based on the CMS classification obtained in (c).
[0142] In some embodiments, the marker panel may further include the following six gene markers based on the CMS14 marker panel: CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1 (i.e., forming a "CMS20 marker panel"). In some such embodiments, the marker panel has the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8, SLC4A4, EREG, AREG, SPINK4, REG4, MUC2, SFRP2, ZEB1, ZEB2, and SFRP4.
[0143] In some embodiments, the marker panel may further include the following 20 gene markers on the basis of the CMS20 marker panel: CXCL10, AIM2, GBP5, CXCL11, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2 and TWIST1 (i.e., forming the "CMS40 marker panel"). In some such embodiments, the marker panel has the following 40 gene markers: CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, CXCL9, CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, CA2, ASCL2, VAV3, CELP, EREG, RNF43, AREG, SPINK4, REG4, MLPH, TFF3, MUC2, AQP3, COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2, and TWIST1.
[0144] In some embodiments, the intestinal cancer is colorectal cancer. In some embodiments, the intestinal cancer is stage II / III colorectal cancer. In some embodiments, the intestinal cancer can be selected from rectal cancer, left colon cancer or right colon cancer.
[0145] In some embodiments, expression data for each gene of a marker panel from a biological sample from a subject is derived from RNA sequencing data of the sample tissue.
[0146] Sequencing can include any suitable sequencing technology known to those skilled in the art.In some embodiments, sequencing includes high-throughput sequencing.
[0147] In some embodiments, the expression data of each gene of the marker panel of the biological sample from the subject is obtained by normalizing the RNA sequencing data of the sample tissue.
[0148] Normalization can include any suitable normalization method known to those skilled in the art. In some embodiments, normalization includes quantile normalization. In some embodiments, quantile normalization includes calculating the log2 value of the original sequenced molecule count for each sample, then sorting the log2 molecule counts for each sample, calculating the arithmetic mean of the log2 molecule counts for all samples corresponding to the order, and replacing the log2 values with the arithmetic mean to form a normalized molecule count matrix.
[0149] In some embodiments, the CMS classification may include one or more of the following: CMS1 - inflammatory, CMS2 - transient proliferative, CMS2 - enterocyte, CMS3 - goblet, and CMS4 - stem.
[0150] In some embodiments, the system includes the following modules:
[0151] Sequencing library construction module: This module is used to construct sequencing libraries from sample RNA;
[0152] Quantitative sequencing module: This module is used to quantify and sequence the sequencing library;
[0153] Data normalization module: This module is used to normalize the quantitative and sequencing results;
[0154] CMS molecular typing module: This module is used to perform molecular typing on data normalization results.
[0155] D. Computer-readable media
[0156] In another aspect, a computer-readable medium is provided, comprising:
[0157] Gene expression data for a marker panel from a biological sample of a subject, wherein the marker panel comprises the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4 (i.e., the "CMS14 marker panel").
[0158] CMS typing feature genome expression template data, and
[0159] Instructions for performing a method comprising:
[0160] (a) obtaining gene expression data of a marker panel from a biological sample of a subject;
[0161] (b) Calculate the cosine distance between the gene expression data in (a) and the CMS typing characteristic gene group expression template;
[0162] (c) determining the CMS type based on the cosine distance calculated in (b); and
[0163] (d) Determine the subject's colorectal cancer prognosis based on the CMS classification obtained in (c).
[0164] In some embodiments, the marker panel may further include the following six gene markers based on the CMS14 marker panel: CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1 (i.e., forming a "CMS20 marker panel"). In some such embodiments, the marker panel has the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8, SLC4A4, EREG, AREG, SPINK4, REG4, MUC2, SFRP2, ZEB1, ZEB2, and SFRP4.
[0165] In some embodiments, the marker panel may further include the following 20 gene markers on the basis of the CMS20 marker panel: CXCL10, AIM2, GBP5, CXCL9, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2 and TWIST1 (i.e., forming the "CMS40 marker panel"). In some such embodiments, the marker panel has the following 40 gene markers: CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, CXCL9, CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, CA2, ASCL2, VAV3, CELP, EREG, RNF43, AREG, SPINK4, REG4, MLPH, TFF3, MUC2, AQP3, COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2, and TWIST1.
[0166] In some embodiments, the intestinal cancer is colorectal cancer. In some embodiments, the intestinal cancer is stage II / III colorectal cancer. In some embodiments, the intestinal cancer can be selected from rectal cancer, left colon cancer or right colon cancer.
[0167] In some embodiments, expression data for each gene of a marker panel from a biological sample from a subject is derived from RNA sequencing data of the sample tissue.
[0168] Sequencing can include any suitable sequencing technology known to those skilled in the art.In some embodiments, sequencing includes high-throughput sequencing.
[0169] In some embodiments, the expression data of each gene of the marker panel of the biological sample from the subject is obtained by normalizing the RNA sequencing data of the sample tissue.
[0170] Normalization can include any suitable normalization method known to those skilled in the art. In some embodiments, normalization includes quantile normalization. In some embodiments, quantile normalization includes calculating the log2 value of the original sequenced molecule count for each sample, then sorting the log2 molecule counts for each sample, calculating the arithmetic mean of the log2 molecule counts for all samples corresponding to the order, and replacing the log2 values with the arithmetic mean to form a normalized molecule count matrix.
[0171] In some embodiments, the CMS classification may include one or more of the following: CMS1 - inflammatory, CMS2 - transient proliferative, CMS2 - enterocyte, CMS3 - goblet, and CMS4 - stem.
[0172] In another aspect, the present invention provides an electronic device loaded with the computer-readable medium.
[0173] E. Application
[0174] On the other hand, the present invention provides the use of a reagent for detecting each marker of a marker panel from a biological sample of a subject in the preparation of a product or product panel for assessing the prognosis of colorectal cancer in a subject, wherein the marker panel has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4 (i.e., the "CMS14 marker panel").
[0175] On the other hand, provided herein is a product or product group comprising reagents for detecting each marker of a marker group from a biological sample of a subject for use in assessing the prognosis of colorectal cancer in a subject, wherein the marker group has the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2 and SFRP4 (i.e., the "CMS14 marker group").
[0176] In some embodiments, the marker panel may further include the following six gene markers based on the CMS14 marker panel: CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1 (i.e., forming a "CMS20 marker panel"). In some such embodiments, the marker panel has the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8, SLC4A4, EREG, AREG, SPINK4, REG4, MUC2, SFRP2, ZEB1, ZEB2, and SFRP4.
[0177] In some embodiments, the marker panel may further include the following 20 gene markers on the basis of the CMS20 marker panel: CXCL10, AIM2, GBP5, CXCL9, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2 and TWIST1 (i.e., forming the "CMS40 marker panel"). In some such embodiments, the marker panel has the following 40 gene markers: CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, CXCL9, CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, CA2, ASCL2, VAV3, CELP, EREG, RNF43, AREG, SPINK4, REG4, MLPH, TFF3, MUC2, AQP3, COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2, and TWIST1.
[0178] After extensive and in-depth research and analysis, the inventors identified a list of genes associated with CMS typing through extensive screening and analysis. They then performed differentially expressed gene analysis on the expression data to identify differentially expressed genes for each CMS type. They then analyzed and selected overexpressed genes for each CMS type, and refined the specific marker panels described herein. The inventors surprisingly discovered that the marker panels described herein can achieve superior CMS typing results with a relatively small number of markers and can be used to accurately assess colorectal cancer prognosis. This provides a low-cost, efficient, and high-quality approach to colorectal cancer prognosis assessment.
[0179] Example
[0180] The present invention will be further described below in conjunction with specific exemplary embodiments. It should be understood that these embodiments are intended only to illustrate the present invention and are not intended to limit the scope of the present invention. Those skilled in the art may make appropriate modifications and variations to the present invention, and these modifications and variations are all within the scope of the present invention.
[0181] Example 1 Screening of prognostic markers for colorectal cancer
[0182] This study screened and analyzed prognostic markers using samples from 1,116 colorectal cancer patients from three hospitals in Shanghai. Figure 1 The general method for screening prognostic markers for colorectal cancer is demonstrated. The specific steps are as follows:
[0183] (I) Determining the gene list associated with CMS typing
[0184] 1) Data Conversion and Normalization: Quantile normalization was used for RNA sequencing data from fresh and FFPE tissues. The specific process is as follows: Due to the different technical characteristics of sample preparation and sequencing between fresh and FFPE tissues, the molecular count distributions of the raw sequencing data differ significantly. To eliminate technical bias caused by sample source (e.g., fresh tissue vs. FFPE), data from samples from different sources need to be normalized for comparability. First, the logarithm (LOG2) of the raw sequencing molecular counts for each sample was calculated to make the value distribution more symmetrical. To account for the presence of zero values, a small value of 0.25 was added to zero values to prevent them from becoming NA or missing values after logarithmization. Then, the Log2 values of each sample were sorted, and the average Log2 value of all samples at the same rank position was calculated. For each gene, the raw Log2 value was replaced by the rank-order average of the Log2 values across samples. This ensures that samples from different sources, processed using the same normalization method, will have the same distribution. This eliminates technical bias in the sample source and allows for more reliable conclusions to be drawn by performing difference analysis between different groups based on normalized data.
[0185] 2) Construction of CMS template: The CMS typing model uses the expression changes (e.g., up-regulation or down-regulation) of characteristic gene groups associated with gene pathway activity, signal transduction, and cell biological activity processes in each CMS typing to determine whether the sample under test has a characteristic pattern associated with a certain CMS typing. The CMS typing in this article is based on the reported 786 genes (Sadanandam, A. (2013) Nat Med. 19 (5): 619-25), and its related information (including updated information of these genes in databases such as NCBI), as well as the related gene groups discovered subsequently, to comprehensively screen and establish a gene sequencing combination to train and establish the typing model. The typing model algorithm uses the nearest template prediction (Nearest Template Predictions), using the expression changes (e.g., up-regulation or down-regulation) of characteristic gene groups related to gene pathway activity, signal transduction, and cell biological activity in each CMS typing as a template, and calculates the cosine distance between the normalized molecular quantity value distribution of the characteristic gene group of each sample and the CMS typing characteristic gene group expression template.
[0186] This cosine distance is defined as the characteristic distance, which is the default characteristic distance of the nearest template prediction model. The shortest distance d indicates that the sample is closest to the characteristic genome expression template of a certain CMS typing, that is, it is most likely to belong to that typing. As a measure of the typing P value for statistical significance test, a random permutation test is used. By randomly extracting characteristic genes (the default value is 1000 times) to generate a random distribution of characteristic distances, the distance between the tested sample and the typing characteristic template is compared with the randomly generated distance distribution and the corrected false discovery rate (FDR) to calculate the P value of the significance test. The smaller the P value, the stronger the statistical significance of the shortest cosine characteristic distance, which means that the predicted CMS typing is more reliable (usually the threshold for statistical significance P is P<0.05).
[0187] Marker genes associated with CMS typing are defined in advance. Without loss of generality, it is assumed that group A genes (nA) are upregulated in CMS type A, but not expressed or downregulated in CMS type B. Similarly, group B genes (nB) are upregulated in CMS type B, but not expressed or downregulated in CMS type A. Group A plus group B genes constitute the characteristic genome templates A and B of the gene expression patterns of CMS type A and type B typing. From the N genes of the tested sample, the normalized values of nA+nB characteristic genes are extracted, and compared with the two templates A and B respectively, and the characteristic cosine distance relative to A or B is calculated. The typing of the template with the closest distance becomes the predicted CMS typing. When calculating the statistical significance of the characteristic distance, nA+nB genes are randomly extracted 1000 times from the N genes to generate the zero distribution of the characteristic distance d. By comparing the characteristic distance of the tested sample with the zero distribution, the calibrated P value of statistical significance is calculated. The red and blue colors in the heatmap indicate upregulated and downregulated gene expression, respectively. All samples were classified into the following five CMS types: CMS1-inflammatory type, CMS2-enterocyte type, CMS2-transient proliferative type, CMS3-cup type, and CMS4-stem type.
[0188] (2) Perform differentially expressed gene analysis on expression data to obtain differentially expressed genes of each CMS type
[0189] A total of 838 samples from a tertiary hospital in Shanghai were used as a training data set to screen a prognostic marker panel. The R package limma was used to perform differentially expressed gene (DEG) analysis using the eBayes algorithm. The results included reading expression data, setting up an expression data matrix, converting data using the voom function, fitting a linear model using lmFit, adjusting variance and p-value using eBayes, obtaining topTable results, drawing MD plots, and drawing volcano plots (e.g., Figure 2 ) and other steps to obtain the differentially expressed genes of each CMS type.
[0190] (III) Selecting overexpressed genes for each CMS type to obtain the marker group CMS40
[0191] To select overexpressed genes for each CMS type, we first selected genes with a fold change greater than 1.3 from the differentially expressed marker panel for each CMS type, with a maximum of 10 genes selected for each CMS type. We then used gene ontology annotation analysis to identify 40 genes representing the five CMS types, which we named the "CMS40" marker panel. As shown in Table 1 , the CMS40 marker panel includes: CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11, CXCL9, CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4, CA2, ASCL2, VAV3, CELP, EREG, RNF43, AREG, SPINK4, REG4, MLPH, TFF3, MUC2, AQP3, COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2, and TWIST1. Among them, CXCL10, AIM2, IFIT3, CXCL13, STAT1, GBP5, CXCL11 and CXCL9 could specifically distinguish the CMS1-inflammatory (CMS1_Inflammatory) type; CA4, CA1, CLCA4, KRT20, MS4A12, AQP8, CLDN8, SLC26A3, SLC4A4 and CA2 could specifically distinguish the CMS2-enterocyte (CMS2_Enterocyte) type; ASCL2, VAV3, CELP, EREG, RNF43 and AREG could specifically distinguish the CMS2-transit-proliferating (CMS2_TransitAmplifying) type; SPINK4, REG4, MLPH, TFF3, MUC2 and AQP3 could specifically distinguish the CMS3-goblet (CMS3_Goblet) type. like) type; COL3A1, SFRP2, ZEB1, SNAI2, ZEB2, SFRP4, CCDC80, AEBP1, TIMP2 and TWIST1 can specifically distinguish CMS4-stem (CMS4_Stem like) type.
[0192] Table 1 List of 40 significantly differentially expressed genes with fold change greater than 1.3
[0193]
[0194] (IV) Further comprehensive verification of the screened marker group and intersection to obtain the marker group CMS20
[0195] We compared the marker combination obtained from the above screening (containing 40 candidate genes) with the empirically verified colorectal cancer molecular typing genes (containing 38 candidate genes, as a positive control marker group), and the resulting intersection contained 20 gene markers. These 20 common genes are gene markers that have been confirmed to be associated with colorectal cancer through both algorithm prediction and empirical verification. They are more accurate and reliable than either algorithm prediction or empirical verification alone. These 20 genes were assigned to a high-confidence colorectal cancer molecular typing marker group and named the "CMS20" marker group. Specifically, the CMS20 marker group includes: IFIT3, CXCL13, STAT1, CXCL9, CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8, SLC4A4, EREG, AREG, SPINK4, REG4, MUC2, SFRP2, ZEB1, ZEB2, and SFRP4. Among them, IFIT3, CXCL13, STAT1 and CXCL9 can specifically distinguish the CMS1-inflammatory (CMS1_Innammatory) type; CA4, CA1, CLCA4, MS4A12, AQP8, CLDN8 and SLC4A4 can specifically distinguish the CMS2-enterothelial cell (CMS2_Entcrocyte) type; EREG and AREG can specifically distinguish the CMS2-transient proliferation (CMS2_Transit Amplifying) type; SPINK4, REG4 and MUC2 can specifically distinguish the CMS3-goblet-like (CMS3_Goblet like) type; SFRP2, ZEB1, ZEB2 and SFRP4 can specifically distinguish the CMS4-stem-like (CMS4_Stem like) type.
[0196] Table 2 CMS20 marker panel
[0197] CMS Type Gene EntrezID Fold change t P-value adj.P.Val CMS1_Inflammatory iFit3 3437 1.933130837 9.657039362 1.90086E-21 1.81612E-20 CMS1_Inflammatory CXCL13 10563 2.529919558 14.07397641 2.38204E-42 3.77791E-40 CMS1_Inflammatory STAT1 6772 2.518828709 13.92543994 1.51279E-41 1.49955E-39 CMS1_Inflammatory CXCL9 4283 2.398874858 13.11158448 2.89751E-37 2.08884E-35 CMS2_Enterocyte CA4 762 2.366377764 17.35527122 1.23907E-61 9.82581E-59 CMS2_Enterocyte CA1 759 2.340091454 17.08073666 6.50317E-60 2.57851E-57 CMS2_Enterocyte CLCA4 22802 2.249182095 16.21269278 1.34659E-54 3.55949E-52 CMS2_Enterocyte MS4A12 54860 2.215540592 15.90077477 9.83181E-53 1.55933E-50 CMS2_Enterocyte AQP8 343 2.197753239 15.6196028 4.47165E-51 5.91003E-49 CMS2_Enterocyte CLDN8 9073 2.118084287 14.70039096 8.33649E-46 8.26355E-44 CMS2_Enterocyte SLC4A4 8671 2.091970501 14.62326613 2.25184E-45 1.98412E-43 CMS2_Transit.amplifying EREG 2069 2.058542252 14.52587039 7.85432E-45 1.03808E-42 CMS2_Transit.amplifying AREG 374 2.004373762 13.90426091 1.96663E-41 1.94942E-39 CMS3_Goblet.like SPINK4 27290 2.530018071 12.07638181 4.11256E-32 3.26126E-29 CMS3_Goblet.like REG4 83998 2.50208751 11.93404367 1.97628E-31 7.83596E-29 CMS3_Goblet.like MUC2 4583 2.110217066 9.715236529 1.11207E-21 2.20468E-19 CMS4_Stem.like SFRP2 6423 1.649504053 9.299433542 4.81886E-20 4.19929E-19 CMS4_Stem.like ZEB1 6935 1.570443115 8.203515337 4.96341E-16 3.0992E-15 CMS4_stem.like ZEB2 9839 1.587685106 8.452358663 6.65993E-17 4.51395E-16 CMS4_Stem.like SFRP4 6424 2.205499829 15.19508982 1.29815E-48 1.71572E-46
[0198] (V) Using correlation analysis to discover alternative relationships between prognostic markers
[0199] Correlation analysis was used to find the substitution relationship between the 40 prognostic markers in CMS40, and the marker group was further refined: the pcarson correlation coefficient between the markers was calculated using the corrplot function in R language to form a correlation coefficient matrix diagram (such as Figure 3 ). According to the correlation coefficient matrix diagram, the following conclusions can be drawn:
[0200] a. The correlation coefficient between CA1 and CA4 is 0.74, indicating that CA1 and CA4 are interchangeable.
[0201] b. The correlation coefficient between CA1 and CA2 is 0.69, indicating that CA1 and CA2 can replace each other;
[0202] c. The correlation coefficient between CA1 and CLCA4 was 0.69, indicating that CA1 and CLCA4 can replace each other;
[0203] d. The correlation coefficient between CA1 and MS4A12 is 0.75, indicating that CA1 and MS4A12 can replace each other;
[0204] e. The correlation coefficient between CA1 and CLDNB is 0.71, indicating that CA1 and CLDNB can replace each other;
[0205] f. The correlation coefficient between CA4 and MS4A12 is 0.69, indicating that CA4 and MS4A12 can replace each other;
[0206] g. The correlation coefficient between CA4 and CLCA4 was 0.67, indicating that CA4 and CLCA4 can replace each other;
[0207] h. The correlation coefficient between CA2 and MS4A12 is 0.70, indicating that CA2 and MS4A12 can replace each other;
[0208] i. The correlation coefficient between CLCA4 and MS4A12 is 0.70, indicating that CLCA4 and MS4A12 can replace each other;
[0209] j. The correlation coefficient between REG4 and SPINK4 is 0.69, indicating that REG4 and SPINK4 can replace each other;
[0210] k. The correlation coefficient between REG4 and MUC2 is 0.67, and REG4 and MUC2 can replace each other;
[0211] The correlation coefficient between SPINK4 and MUC2 was 0.75, indicating that SPINK4 and MUC2 can replace each other.
[0212] The correlation coefficient between m.EREG and AREG is 0.73, and EREG and AREG can replace each other;
[0213] The correlation coefficient among CXCL9, CXCL10, and CXCL11 is at least 0.69, indicating that they can be substituted for each other.
[0214] (VI) Further refine the marker group based on the substitution relationship between markers to obtain the marker group CMS14
[0215] Based on the CMS20 panel, combined with the results of gene correlation and differential analysis, the panel was further refined to obtain a list of 14 genes, named the "CMS14" panel. Specifically, the CMS14 panel includes: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4. Among them, IFIT3, CXCL13, STAT1 and CXCL9 can specifically distinguish the CMS1-inflammatory (CMS1_Inflammatory) type; CA4, AQP8 and SLC4A4 can specifically distinguish the CMS2-enterocyte (CMS2_Enterocyte) type; AREG and EREG can specifically distinguish the CMS2-transit amplifying (CMS2_Transit Amplifying) type; SPINK4 and REG4 can specifically distinguish the CMS3-goblet-like (CMS3_Goblet like) type; SFRP2, ZEB2 and SFRP4 can specifically distinguish the CMS4-stem-like (CMS4_Stem like) type.
[0216] Table 3 CMS14 gene list
[0217]
[0218] (VII) All 1,116 colorectal cancer samples (including 278 samples from two other hospitals) were used to evaluate the CMS molecular typing ability of the marker panel
[0219] Based on data from 1,116 samples, CMS molecular typing was performed using a combination of CMS40, CMS20, and CMS14 markers. Relapse-free survival data, including relapse status and time, was also collected for all 1,116 samples.
[0220] Recurrence-free survival analysis was performed based on the CMS molecular typing results, recurrence status, and time of all samples obtained using the CMS40 marker panel. Figure 4As shown, the horizontal axis is the recurrence-free survival time in months; the vertical axis is the recurrence-free survival rate. The light green curve in the figure represents CMS1_Inflammatory, the yellow curve represents CMS2_Enterocyte, the blue curve represents CMS2_Transit.amplifying, the dark green curve represents CMS3_Goblet.like, and the red curve represents CMS4_Stem.like. The five curves are relatively similar when the survival rate is close to 1. When the survival rate decreases to near 0, the divergence between the curves increases, indicating that there are significant prognostic differences between different groups. The overall difference between the five curves is significant (p<0.0001), which shows that the CMS40 characteristic gene can be used as an important biomarker to accurately predict the survival of colorectal cancer patients.
[0221] Recurrence-free survival analysis was performed based on the CMS molecular typing results, recurrence status, and time of all samples obtained using the CMS20 marker panel. Figure 5 As shown, the horizontal axis is the recurrence-free survival time in months; the vertical axis is the recurrence-free survival rate. The light green curve in the figure represents CMS1_Inflammatory, the yellow curve represents CMS2_Enterocyte, the blue curve represents CMS2_Transit.amplifying, the dark green curve represents CMS3_Goblet.like, and the red curve represents CMS4_Stem.like. The five curves are relatively similar when the survival rate is close to 1. When the survival rate decreases to near 0, the divergence between the curves increases, indicating that there are significant prognostic differences between different groups. The overall difference between the five curves is significant (p<0.017), which shows that the CMS20 characteristic gene can be used as an important biomarker to accurately predict the survival of colorectal cancer patients.
[0222] Disease-free survival analysis was performed based on the CMS molecular typing results, disease status, and time of all samples obtained using the CMS14 marker panel. Figure 6As shown, the horizontal axis is the disease-free survival time in months; the vertical axis is the disease-free survival rate. The red curve in the figure represents CMS1_Inflammatory, the blue curve represents CMS2_Enterocyte, the green curve represents CMS2_Transit.amplifying, the orange curve represents CMS3_Goblet.like, and the purple curve represents CMS4_Stem.like. The five curves are relatively similar when the survival rate is close to 1. When the survival rate decreases to near 0, the divergence between the curves increases, indicating that there are significant prognostic differences between different groups. The overall difference of the five curves is significant (p = 0.0015), which shows that the CMS14 characteristic gene can be used as an important biomarker to accurately predict the survival of colorectal cancer patients.
[0223] Example 2 Molecular typing for evaluating the prognosis of colorectal cancer
[0224] (1) Sequencing library construction
[0225] (1) Sample RNA extraction
[0226] RNA was extracted and purified from FFPE samples using an RNA extraction kit according to the manufacturer's instructions. The extracted RNA was accurately quantified (using a Qubit fluorescence quantifier is recommended) and stored at -70°C.
[0227] (2) Sequencing library construction
[0228] 1) Reverse transcription of RNA samples into cDNA: RNA samples are subjected to reverse transcriptase reaction to synthesize complementary DNA (cDNA).
[0229] 2) Molecular tagging: A unique molecular barcode is added to each cDNA sample. A molecular barcode is a short sequence tag that labels each original template molecule during the amplification process, allowing for accurate subsequent expression calculations.
[0230] 3) Purification: Use magnetic beads or silica column purification to purify the barcoded cDNA product from the reaction system to remove impurities such as proteins.
[0231] 4) First round of PCR reaction: Perform the first round of polymerase chain reaction (PCR) using gene-specific primers to amplify the marker gene to obtain a sufficient amount of template.
[0232] 5) PCR product purification: Purify the first-round PCR product again to prevent impurities from affecting subsequent reactions. Magnetic beads or silica column purification can be used for purification.
[0233] 6) Second round of adapter sequence PCR reaction: A second round of PCR was performed, adding sample multiple indexes and sequencing platform universal adapter sequences for distinguishing samples and binding to the sequencing chip.
[0234] 7) Sequencing library purification: Finally, the final sequencing library is purified again using magnetic bead purification.
[0235] 8) Sequencing Library Quantification: Accurately quantify the library to control the amount of sequencing applied to the machine and ensure that each sample reaches the required sequencing depth. Common quantification methods include fluorescent quantitative PCR and chip electrophoresis.
[0236] (2) Quantitative data acquisition
[0237] The gene expression data of the colorectal cancer prognostic marker combination is obtained, and its detection technology includes but is not limited to real-time fluorescence quantitative qPCR technology, gene chip technology and high-throughput whole transcriptome (or targeted gene transcriptome) sequencing technology. This embodiment is described using targeted gene transcriptome sequencing as an example.
[0238] The process of this technology includes:
[0239] 1. Based on the results of colorectal cancer prognostic marker screening, any combination of CMS14, CMS20, or CMS40 markers was used as sequencing targets.
[0240] 2. Isolate total RNA from tumor samples using an RNA extraction kit. Assess the quality and concentration of the RNA.
[0241] 3. Generate a cDNA library of the target genes through reverse transcription and PCR amplification. A unique molecular barcode is added during the amplification process to calculate expression levels.
[0242] 4. Use a high-throughput sequencing platform such as Illumina to perform single-end or double-end sequencing on each sample.
[0243] 5. Align the reads to the reference genome and obtain the expression matrix of each gene based on the unique molecular tag counts.
[0244] 6. Comprehensively analyze the expression matrix of multiple samples to obtain a combined prognostic model.
[0245] (3) Data normalization: In order to eliminate the technical bias caused by the source of the sample (such as fresh tissue or FFPE), the data of samples from different sources need to be normalized to make them comparable. First, the logarithm of the original sequencing molecular count of each sample (LOG2 value) is calculated to make the value distribution more symmetrical. Taking into account the existence of zero values, a very small number of 0.25 is uniformly added to the zero value to avoid becoming NA or missing values after taking the logarithm. Then, the LOG2 values of each sample are sorted, and the average LOG2 value of all samples at the same ordinal position is calculated. For each gene, the original LOG2 value is replaced by the ordinal average of its LOG2 value in each sample. In this way, samples from different sources are processed by the same normalization method, and the resulting values will obey the same distribution. This eliminates the technical bias of the sample source, and more reliable conclusions can be obtained by performing difference analysis between different groups based on the normalized data.
[0246] (IV) CMS molecular subtype: Based on the normalized expression matrix, the cosine similarity of each sample across the five characteristic expression patterns of CMS typing (CMS1 inflammatory type, CMS2 intestinal epithelial type, CMS2 transient proliferative type, CMS3 goblet cell type, and CMS4 stem cell type) was calculated using the CMS14, CMS20, or CMS40 marker panel. The CMS molecular subtype to which each sample belonged was determined by maximizing the cosine similarity.
[0247] Example 3 Study on the correlation between CMS40 and survival of patients with colon cancer
[0248] (1) Sample source: FFPE samples from 229 patients with stage III colon cancer from a tertiary hospital in Shanghai.
[0249] (2) Clinical sample data: Progression-free survival data of 229 samples were collected.
[0250] (3) Obtaining gene expression data:
[0251] According to the method in the molecular typing system for evaluating the prognosis of colorectal cancer in Example 2, sequencing library construction and quantitative data acquisition were performed to obtain expression data of all 40 characteristic genes in the CMS40 marker group of 229 clinical samples.
[0252] (IV) Data normalization: First, all data were Log2 transformed, and then the expression data were normalized using the quantile normalization method.
[0253] (V) CMS molecular typing: Based on the normalized expression matrix, the CMS40 marker panel was used to calculate the cosine similarity of each sample across the five characteristic expression patterns of CMS typing (CMS1: inflammatory type, CMS2: intestinal epithelial type, CMS2: transient proliferative type, CMS3: goblet cell type, and CMS4: stem cell type). The CMS subtype to which each sample belonged was determined by maximizing the cosine similarity.
[0254] (6) Survival analysis:
[0255] All samples were divided into three groups according to the five types of CMS molecular typing:
[0256] ① Low-risk group: samples of CMS1_Inflammatory and CMS2_Transit.amplifying were grouped together;
[0257] ② Intermediate-risk group: samples of CMS2 enterocyte type CMS2_Enterocyte and CMS3 goblet cell type CMS3_Goblet.like are grouped together;
[0258] ③ High-risk group: Samples of CMS4 stem cell type CMS4_Stem.like are grouped together.
[0259] The results of the progression-free survival analysis of patients were as follows Figure 7 As shown, the horizontal axis represents disease-free survival (DFS) in months, and the vertical axis represents DFS. The red curve represents the low-risk group, comprised of samples from two CMS subtypes: CMS1_Inflammatory and CMS2_Transit.amplifying; the blue curve represents the intermediate-risk group, comprised of samples from two CMS subtypes: CMS2_Enterocyte and CMS3_Goblet.like; and the green curve represents the high-risk group, comprised of samples from one CMS subtype: CMS4_Stem.like. The three curves are relatively similar when the survival rate approaches 1. As the survival rate decreases toward 0, the divergence between the curves increases, indicating significant prognostic differences between the different groups. The difference in survival curves between the low-risk group (CMS1_Inflammatory and CMS2_Transit.amplifying) and the intermediate-risk group (CMS2_Enterocyte and CMS3_Goblet.like) was significant (p=0.0069). At the same time, the difference in survival curves between the low-risk group and the high-risk group (CMS4_Stem.like) was very significant (p<0.0001), indicating that the CMS40 gene marker group can be used as an important biomarker group to accurately predict the survival of colorectal cancer patients.
[0260] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form or substance. It should be pointed out that ordinary technicians in this technical field can make several improvements and supplements without departing from the method of the present invention, and these improvements and supplements should also be regarded as the scope of protection of the present invention. Any equivalent changes, modifications and evolutions made by technicians familiar with this profession without departing from the spirit and scope of the present invention by using the technical content disclosed above are all equivalent embodiments of the present invention; at the same time, any equivalent changes, modifications and evolutions made to the above embodiments based on the essential technology of the present invention are still within the scope of the technical solution of the present invention.
[0261] At the same time, this application uses specific terms to describe the embodiments of this application. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned twice or multiple times in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.
[0262] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of the present application are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
Claims
1. Use of a reagent for detecting each marker of a marker panel in a biological sample from a subject in the preparation of a product or product panel for assessing the prognosis of colorectal cancer in a subject by CMS typing of colorectal cancer, wherein: The marker panel is: a marker panel consisting of the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4; a marker panel consisting of the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, SFRP4, CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1; or A marker panel consisting of the following 40 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, SFRP4, CA1, CLCA4, MS4A12, CLDN8, MUC2, ZEB1, CXCL10, AIM2, GBP5, CXCL11, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2, and TWIST1.
2. A system for assessing the prognosis of colorectal cancer in a subject, comprising a memory and one or more processors; in, The memory comprises: expression data of each gene of a marker panel from a biological sample of a subject, CMS typing feature genome expression template data, and one or more processor-executable instructions; Wherein, the marker group is used for CMS typing of colorectal cancer; Wherein, the marker group is: a marker panel consisting of the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4; a marker panel consisting of the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, SFRP4, CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1; or a marker panel consisting of the following 40 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, SFRP4, CA1, CLCA4, MS4A12, CLDN8, MUC2, ZEB1, CXCL10, AIM2, GBP5, CXCL11, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2, and TWIST1; and The one or more processor-executable instructions are configured to: (a) obtaining gene expression data of the marker panel from a biological sample of a subject; (b) Calculate the cosine distance between the gene expression data in (a) and the CMS typing feature gene expression template; (c) Determine the CMS type based on the cosine distance calculated in (b); and (d) Determine the subject's colorectal cancer prognosis based on the CMS classification obtained in (c).
3. A computer-readable medium comprising: Gene expression data of a marker panel from a biological sample of a subject, the marker panel being used for CMS typing of colorectal cancer; wherein the marker panel is: a marker panel consisting of the following 14 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, and SFRP4; a marker panel consisting of the following 20 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, SFRP4, CA1, CLCA4, MS4A12, CLDN8, MUC2, and ZEB1; or a marker panel consisting of the following 40 gene markers: IFIT3, CXCL13, STAT1, CXCL9, CA4, AQP8, SLC4A4, AREG, EREG, SPINK4, REG4, SFRP2, ZEB2, SFRP4, CA1, CLCA4, MS4A12, CLDN8, MUC2, ZEB1, CXCL10, AIM2, GBP5, CXCL11, KRT20, SLC26A3, CA2, ASCL2, VAV3, CELP, RNF43, MLPH, TFF3, AQP3, COL3A1, SNAI2, CCDC80, AEBP1, TIMP2, and TWIST1; CMS typing feature genome expression template data, and Instructions for performing a method comprising: (a) obtaining gene expression data of a marker panel from a biological sample of a subject; (b) Calculate the cosine distance between the gene expression data in (a) and the CMS typing characteristic gene group expression template; (c) Determine the CMS type based on the cosine distance calculated in (b); and (d) Determine the subject's colorectal cancer prognosis based on the CMS classification obtained in (c).
4. An electronic device loaded with the computer-readable medium according to claim 3.
5. The use of claim 1, the system of claim 2, the computer-readable medium of claim 3, or the electronic device of claim 4, wherein: The intestinal cancer is stage II / III colorectal cancer.
6. The system according to claim 2, wherein Includes the following modules: Sequencing library construction module: This module is used to construct sequencing libraries from sample RNA; Quantitative sequencing module: This module is used to quantify and sequence the sequencing library; Data normalization module: This module is used to normalize the quantitative and sequencing results; CMS molecular typing module: This module is used to perform molecular typing on data normalization results.
7. The use according to claim 1, wherein The reagent is a reagent for detecting the expression of each gene marker in the marker group in the biological sample.
Citation Information
Patent Citations
Prognosis prediction for colorectal cancer
CN101389957A
Prognostic marker and prognostic risk assessment model for metastatic colon adenocarcinoma and application of prognostic marker and prognostic risk assessment model
CN112143809A