A cancer prognosis survival analysis method based on collaborative co-expression network

By building a collaborative co-expression network, prognostic biomarkers in cancer multiomics data were extracted and sample clustering was performed, the problems of biomarker selection and prognostic subtype discovery in cancer prognostic survival analysis were solved, and accurate prediction of prognostic categories of cancer patients and targeted therapy support were achieved.

CN119601086BActive Publication Date: 2025-08-12NORTHEAST FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510034865.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-08-12
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively select biomarkers closely related to cancer prognostic survival and discover prognostic survival subtypes, resulting in limited clinical application of cancer prognostic survival analysis research.

Method used

Using a method based on a collaborative co-expression network, a collaborative co-expression network is constructed by collecting multiomics data, prognostic biomarkers are extracted, and sample clustering is carried out for survival analysis, so as to achieve accurate prediction of the prognostic categories of cancer patients.

Benefits of technology

Accurate prediction of the prognosis category of cancer patients is achieved, and the molecular mechanisms of differences in cancer prognosis survival are revealed for biologists, and the clinical physicians are assisted in targeted and precise treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601086B_ABST
    Figure CN119601086B_ABST
Patent Text Reader

Abstract

The present invention is a method for cancer prognosis survival analysis based on a collaborative co-expression network. The present invention relates to the technical field of cancer tumor prognosis analysis. The present invention collects multi-omics data of cancer based on a public database and pre-processes the data; based on the processed data, a collaborative co-expression network for cancer multi-omics data is constructed; based on the collaborative co-expression network, prognosis biomarkers are extracted; based on the extracted prognosis biomarkers, sample clustering for survival analysis is performed, and prognosis survival analysis is performed. The present invention achieves accurate prediction of the prognosis category to which cancer patients belong, provides statistical support for biologists to reveal the molecular mechanisms that lead to differences in cancer prognosis and survival, and assists clinicians in ultimately achieving targeted precision treatment for prognosis biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cancer tumor prognosis analysis, and is a cancer prognosis survival analysis method based on a collaborative co-expression network. Background Art

[0002] Malignant tumors (commonly known as cancer, hereinafter referred to as such) seriously endanger human survival and health due to their high recurrence rate, rapid growth rate and low cure rate. Surgery supplemented by postoperative radiotherapy and chemotherapy is currently the main method for treating cancer. However, the invasion and metastasis characteristics of cancer cells limit the ability of surgery to cure cancer in one go, and radiotherapy and chemotherapy bring pain to cancer patients due to their side effects. Prognostic survival analysis can determine the risk of postoperative recurrence, metastasis and even death of patients, provide the necessary basis for individualized treatment of cancer patients after surgery, and maximize the survival time of patients. Prognostic survival analysis refers to: constructing an association model between the phenotype or molecular indicators of existing cases and their survival time, and measuring the phenotype or molecular indicators of new patients based on this model to predict their survival time. In fact, cancer recurrence, metastasis and patient death all have corresponding time nodes. In the present invention, the time period from surgery to the occurrence of the above events is collectively referred to as survival time.

[0003] Studies have found no clear correlation between phenotypic indicators (such as personal information and clinical indicators) and survival in cancer patients. Therefore, constructing two-dimensional matrices (rows representing different features, columns corresponding to different samples) based on multi-omics data (such as genomics, transcriptomics, and proteomics) and extracting features (i.e., biomarkers such as DNA, RNA, and proteins) that are closely related to survival is a key step in solving cancer prognosis and survival analysis. This process is known as feature selection in the fields of pattern recognition and machine learning. First, the high-dimensionality and small sample size of the constructed two-dimensional matrices renders classic feature selection methods ineffective in pattern recognition. Limited by the number of cases and the cost of data collection, the number of samples (also known as patients or subjects) often differs by at least two orders of magnitude from the number of features (also known as variables), making traditional feature selection methods inevitably prone to overfitting. Furthermore, using a cancer patient's omics matrix as input and their prognostic survival time as output further complicates feature selection. Depending on the follow-up, the patient's survival time corresponds to the time from surgery to the occurrence of an event (e.g., recurrence, metastasis, death, etc.) or the last follow-up. If a patient is lost to follow-up or is still alive at the end of the study, only the follow-up time (right-censored time) is recorded, not the patient's true survival time.

[0004] To extract features closely related to the prognostic survival time of cancer patients from matrices representing different omics, existing research focuses on the following two aspects: 1. The idea of screening associated features. The correlation between each feature and the prognostic survival time is calculated to determine whether the feature should be selected, or constraints are used to filter out noisy features, retaining only features closely related to prognostic survival. 2. The idea of combining feature clusters. Closely related features are combined into feature clusters to represent prognostic biomarkers. Both of the above approaches have problems: the former emphasizes the correlation between features and survival time, but ignores the relationship between features; the latter focuses on the correlation between features, but ignores the connection between features and survival time.

[0005] Cancer prognostic survival analysis faces two pressing challenges: prognostic biomarker extraction and prognostic class discovery. The former aims to select key features closely associated with prognostic survival for subsequent construction of survival indices; the latter aims to identify potential prognostic subtypes within specific cancers or pan-cancer populations based on different prognostic survival outcomes, thereby predicting patient survival risk categories and assisting in personalized clinical treatment. The two are mutually reinforcing: prognostic biomarker extraction removes redundant features from two-dimensional matrices, while sample clustering within a two-dimensional matrix in a low-dimensional feature space helps improve the stability and reliability of the discovered prognostic classes. Prognostic class discovery divides samples into two or more categories based on prognostic survival time and attempts to extract biomarkers with significant differences between the two or more categories as key prognostic features. Research on prognostic class discovery primarily focuses on univariate sample classification strategies and univariate sample clustering strategies. Research on prognostic biomarker extraction, on the other hand, focuses on multivariate feature selection strategies, univariate feature combination strategies, and correlated feature clustering strategies.

[0006] To date, the two challenges of selecting key biomarkers closely associated with cancer survival and identifying potential prognostic survival subtypes for specific cancers remain largely unresolved, severely limiting the clinical application of cancer prognostic survival analysis. To address these two issues, existing research focuses on either the relationship between features and survival or the relationship between features, focusing on the five aforementioned strategies. An effective approach that organically combines these two focuses and five strategies is urgently needed. Prognostic biomarkers are first extracted, followed by clustering samples based on their corresponding feature subspace matrices or the output of regression models to predict survival risk and even survival for new samples. This aligns with the logical order of cognition. Given the continuous nature of survival, sample clustering strategies are superior to sample classification strategies. More crucial than univariate or multivariate approaches is the ability to simultaneously select prognostic biomarkers that are closely associated with prognostic survival while also focusing on the correlations between these biomarkers. This relationship remains a blind spot in existing research. Furthermore, no research has yet been conducted using cancer multi-omics data to investigate causal relationships between prognostic survival biomarkers. Summary of the Invention

[0007] To address these issues, the present invention proposes a cancer prognosis and survival analysis method based on a collaborative co-expression network. The proposed method will be applied to the prognosis and survival analysis of multi-omics data from various malignant tumors, including malignant brain glioma, breast cancer, primary liver cancer, gastric adenocarcinoma, and esophageal squamous cell carcinoma. Biomarkers closely associated with prognosis and survival will be selected, and a prognosis and survival index will be constructed. This will be used to discover and compare different prognosis categories of samples, infer causal relationships between prognosis biomarkers, and accurately predict the prognosis category of cancer patients. This will provide statistical support for biologists to reveal the molecular mechanisms that lead to differences in cancer prognosis and survival, and for clinicians to ultimately achieve targeted precision therapy based on prognosis biomarkers.

[0008] The present invention provides the following technical solutions:

[0009] A cancer prognosis survival analysis method based on a collaborative co-expression network comprises the following steps:

[0010] Step 1: Collect multi-omics data on cancer based on public databases and preprocess the data;

[0011] Step 2: Based on the processed data, construct a collaborative co-expression network for cancer multi-omics data;

[0012] Step 3: Extract prognostic biomarkers based on the collaborative co-expression network;

[0013] Step 4: Based on the extracted prognostic biomarkers, sample clustering for survival analysis is performed, and prognostic survival analysis is performed.

[0014] Preferably, the step 1 is specifically:

[0015] Step 1.1: For the target tumor, collect high-quality RNA-seq data of the FPKM expression type from the TCGA public database, convert the expression value to log2 TPM+1, and perform normalization to ensure that the overall expression level of the data conforms to the normal distribution;

[0016] Step 1.2: Collect DNA methylation data of target tumor samples from the TCGA database, preferably high-throughput methylation sequencing data, and use Minfi and ChAMPR packages for background correction, normalization, probe filtering, and batch effect correction;

[0017] Step 1.3: For the target tumor, collect high-quality CNV data of the Segment Mean type from the TCGA public database through the Genomic Data Commons and convert the data into the GISTIC2 format. Perform data cleaning methods on the CNV data, including removing outliers, correcting errors, and removing redundant information. Then, standardize the CNV data to eliminate the influence of experimental batch effects and sample processing differences. Z-score standardization is used to standardize the CNV data.

[0018] Step 1.3: Collect FASTQ-formatted ChIP-Seq data from the TCGA database for the target tumor samples. Use FastQC to perform quality control on the raw sequencing data to identify and address low-quality reads, contamination, or other potential issues. Use Trimmomatic or Cutadapt to remove adapter sequences and low-quality reads. Use alignment tools to align the cleaned reads to the reference genome. Use SAMtools to filter out reads with low alignment quality, multiple alignments, and nonspecific alignments. Use Picard Tools or SAMtools to remove PCR duplicates to ensure data accuracy.

[0019] Preferably, the step 2 is specifically as follows:

[0020] Cox regression based on pairwise feature enumeration is used to fit the prognosis survival time, and Pearson correlation based on pairwise feature enumeration is used to measure the correlation between features. The two metrics are unified to label the weights of any two nodes in a fully connected network. The cluster center automatic search clustering algorithm based on density descending is used to delete the edges with lower weights in the network, thereby realizing network construction and feature cluster mining. The cluster center automatic search algorithm based on density descending is used to cluster the weights of all edges in the network in descending density, deleting the edges with lower weight density from the network until disconnected feature clusters appear.

[0021] Step 2.1: Based on the Cox regression, the pairwise feature joint enumeration is performed. The Cox proportional hazard regression model is used to fit the survival time by enumerating the pairwise features, thereby revealing the significant relationship between any pair of features and prognosis survival. The regression coefficient β corresponding to any pair of features is obtained by performing maximum likelihood estimation on the logarithmic partial likelihood function of the hazard function. i ,β j ) T , when it is assumed that the features are uncorrelated, then any component of the regression coefficient obeys the Gaussian distribution, and the Wald test is used to determine whether the regression coefficient component deviates significantly from zero;

[0022] By reordering the samples and their corresponding survival outcomes, the sample size is further expanded. A reordered p-value corresponding to the Wald statistic is generated for each feature. All pairwise features are enumerated to obtain pairwise p-values, which are used to select significant feature pairs closely related to survival from all pairwise feature combinations. To combat overfitting, a penalty constraint is imposed on the log-partial likelihood function corresponding to the hazard function. Considering that the hazard ratio may vary with survival time, a weighted Cox regression model is used instead.

[0023] Step 2.2: Pearson correlation-based joint enumeration of two features. To measure the correlation between two features, the Pearson correlation coefficient r is used to calculate the correlation between any two features, and the t statistic of n samples is constructed based on this. The statistic is expressed as: By using the sample reordering method, the order of sample values of one feature in any pair of features is disrupted, and the above t statistic is calculated repeatedly. The corresponding reordered p-value is expressed as:

[0024]

[0025] Among them, t(i,j) is the statistic corresponding to the correlation coefficient r of features i and j, t b (i, j) represents the statistic generated by randomly reordering the values of any feature, and B is the number of random reorderings;

[0026] Step 2.3: Determine the association relationship of network nodes. The collaborative co-expression network uses individual features as network nodes and the relationship between two features as edges. The weight of the edge between nodes is calculated by integrating the significant relationship between the feature and survival time and the correlation between features. The following metric is constructed to measure the weight of the edge between nodes:

[0027]

[0028] Where p2(i) and p2(j) represent the binary Cox regression reranking p-values obtained by combining the two features i and j, respectively, and p1(i) and p1(j) represent the univariate Cox regression reranking p-values obtained by combining the two features i and j, respectively. represents rounding up; the first term of m(i, j) is the "synergistic" part that measures the contribution of features i and j to survival regression, and the second term is derived from the re-ranked p-value obtained by the joint enumeration of two features based on Pearson correlation, corresponding to the "co-expression" part; m(i, j) is used to measure the association between nodes i and j: the larger its value, the more likely it is that nodes i and j are derived from the same feature cluster, and vice versa. When the value of m(i, j) is less than zero, its value is set to the maximum value of the weight in the network, and the weighted undirected graph constructed is called a synergistic-co-expression network;

[0029] Step 2.4: Based on the feature cluster mining of the collaborative-coexpression network, a density-decreasing cluster center automatic search clustering algorithm is introduced to measure the node association relationship of the weight of the collaborative coexpression network edge. The measurement threshold of the edge in the network is limited by the optimal automatic estimation of the number of categories and cluster boundaries to determine whether there is an associated edge between any two nodes in the network, thereby degenerating the network from a weighted undirected graph to an undirected graph.

[0030] Preferably, the step 3 is specifically:

[0031] Step 3.1: Bottom-up “cross-cluster enumeration” biomarker selection. Based on the “cross-cluster hypothesis,” univariate enumeration is performed within each feature cluster. Univariate variables from c different feature clusters are combined into a multivariate, and their regression coefficients on the Cox proportional hazards function are calculated.

[0032] Step 3.2: Top-down "cross-cluster random" biomarker selection, introducing a sample resampling voting strategy, randomly selecting a single variable within each feature cluster, combining the single variables from c different feature clusters into a multivariate, and calculating its regression coefficient for the Cox survival hazard ratio function; assuming that the features are independent of each other, analogous to the pairwise joint enumeration method based on Cox regression, each set of "cross-cluster" multivariates can obtain corresponding c re-ranked p-values, and in each round, votes are counted for the multivariates with significant re-ranked p-values; after l rounds of voting, the features with the highest votes in each feature cluster are selected to form the significant prognostic feature set;

[0033] Step 3.3: Unsupervised feature selection of joint multivariates: Individual variables are extracted from different feature clusters to form the joint multivariate for regression analysis. For bottom-up "cross-cluster enumeration" biomarker selection, c re-ranked p-values corresponding to each enumerated multivariate are recorded. For top-down "cross-cluster random" biomarker selection, a threshold is set to record multiple sets of c-dimensional p-value vectors corresponding to multiple randomly extracted multivariates. In the c-dimensional feature space, all enumerated or re-ranked p-value vectors that meet the threshold are automatically clustered based on density-decreasing cluster centroid search. The clusters closest to the origin are selected to form the c-dimensional feature set, and the Mahalanobis distance from the scattered points to the origin is calculated. When bottom-up "cross-cluster enumeration" biomarker selection is used, the multidimensional variable closest to the origin is selected as the significant prognostic feature set. If top-down "cross-cluster random" biomarker selection is used, it is still necessary to determine whether the multidimensional variable closest to the origin corresponds to the feature combination with the highest cumulative votes.

[0034] Step 3.4: Construction of prognostic survival analysis indicators. Taking the Cox proportional hazards regression model as an example, since it is difficult to estimate the baseline hazard function or baseline survival function, the linear part of the risk score is used to construct the prognostic survival indicator. The risk score of the sample is a linear combination of the corresponding sample values of the prognostic characteristics, and the combination coefficient is the regression coefficient.

[0035] Preferably, the step 4 is specifically as follows:

[0036] Step 4.1: Sample clustering for prognostic survival analysis: To discover prognostic survival categories, key biomarkers from different omics or prognostic survival analysis indicators built from different omics data are combined as input to the convolutional neural network. The output is regarded as a new feature subspace, and the samples are automatically clustered based on density-decreasing cluster centroids to discover the prognostic category to which the samples belong.

[0037] Step 4.2: Draw Kaplan-Meier survival analysis curves for samples from different categories to observe the changes in survival time within the category and the differences in survival time between categories. Use the log-rank test to measure whether there are significant quantitative differences in survival time between different categories. Draw a risk score survival analysis chart to observe the changes in survival time between samples from different categories.

[0038] Step 4.3: Analysis of significant differences in prognosis between classes: Use statistics based on sample distribution differences to achieve class comparison and evaluate whether there are significant differences in risk scores between different categories. On this basis, try to use training samples to construct different types of classifiers, and use the prognostic survival results of test samples to verify the accuracy and effectiveness of the classification results.

[0039] Preferably, the cross-group assumption is:

[0040] The significant feature set closely related to the prognosis of survival time is not entirely composed of a single significant variable related to the prognosis of survival time, and there is no obvious correlation between the variables in the significant feature set.

[0041] A cancer prognosis survival analysis system based on a collaborative co-expression network, comprising:

[0042] A data acquisition module, which collects multi-omics data of cancer based on public databases and pre-processes the data;

[0043] A network construction module, wherein the network construction module constructs a collaborative co-expression network for cancer multi-omics data based on the processed data;

[0044] An extraction module, which extracts prognostic biomarkers based on a collaborative co-expression network;

[0045] A clustering module, which performs sample clustering for survival analysis based on extracted prognostic biomarkers;

[0046] An analysis module is used to perform prognostic survival analysis based on the results of sample clustering.

[0047] A computer-readable storage medium stores a computer program, which is executed by a processor to implement a cancer prognosis survival analysis method based on a collaborative co-expression network.

[0048] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements a cancer prognosis survival analysis method based on a collaborative co-expression network when executing the computer program.

[0049] The present invention has the following beneficial effects:

[0050] Compared with the prior art, the present invention has the following advantages:

[0051] This invention is applied to cancer prognosis and survival analysis. By enabling the selection of prognostic biomarkers, the establishment of survival indicators, the discovery and comparison of prognostic categories, and the prediction of their respective categories, it provides statistical support for revealing the molecular mechanisms that lead to differences in cancer prognosis and survival, and ultimately for achieving targeted precision therapy for key biomarkers. The innovation of this invention is primarily reflected in the following two aspects:

[0052] (1) The concept of synergy-coexpression is proposed, and the correlation metrics based on survival regression model and coexpression network are combined. The cluster center automatic search clustering technology based on density descending is introduced to construct the synergy-coexpression network.

[0053] (2) Based on the assumption that significant biomarkers are cross-cluster, a bottom-up biomarker enumeration method and a top-down biomarker random selection method are designed to achieve joint multivariate unsupervised biomarker selection;

[0054] The present invention uses mature theories such as Cox regression, Pearson correlation, unsupervised clustering, supervised classification, differential expression analysis, Bayesian network, etc. to provide strong theoretical support for the implementation of the present invention. The technical route is reasonably laid out, the logical relationship between each part is close, and it is advanced layer by layer. First, a metric based on Cox regression and Pearson correlation is constructed to measure the weight of the edge between nodes, and a clustering algorithm is used to degenerate the synergistic-coexpression network from a weighted undirected graph to an undirected graph; then, for the bottom-up or top-down feature selection results, a clustering algorithm is introduced to realize unsupervised feature selection of joint multiple variables and construct a prognostic survival analysis indicator; then, the survival analysis indicators of the constructed prognostic features are clustered to complete the prognostic category discovery and class comparison; finally, the Bayesian network is used to realize the causal relationship inference of the prognostic features, and the survival risk category to which the new sample belongs is predicted. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0056] Figure 1 This is a framework diagram for the research content of cancer multi-omics prognostic survival analysis based on collaboration-co-expression network;

[0057] Figure 2 This is a technical roadmap of the method of the present invention;

[0058] Figure 3 Constructing schematics for collaboration-co-expression networks for cancer multi-omics data;

[0059] Figure 4 Schematic diagram for prognostic biomarker extraction;

[0060] Figure 5 Schematic diagram for prognostic category discovery in cancer multi-omics data;

[0061] Figure 6 This is the result of hierarchical clustering on all pairwise significant miRNA expression levels;

[0062] Figure 7 The results of hierarchical clustering on all pairwise significant miRNA correlation coefficients. DETAILED DESCRIPTION

[0063] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0064] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0065] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0066] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0067] The present invention is described in detail below with reference to specific embodiments. Specific embodiment one:

[0069] according to Figure 1-Figure 7 As shown, the specific optimization technical solution adopted by the present invention to solve the above technical problems is: the present invention relates to a cancer prognosis survival analysis method based on a collaborative co-expression network.

[0070] A cancer prognosis survival analysis method based on a collaborative co-expression network comprises the following steps:

[0071] Step 1: Collect multi-omics data on cancer based on public databases and preprocess the data;

[0072] Step 2: Based on the processed data, construct a collaborative co-expression network for cancer multi-omics data;

[0073] Step 3: Extract prognostic biomarkers based on the collaborative co-expression network;

[0074] Step 4: Based on the extracted prognostic biomarkers, the samples for survival analysis are clustered and prognostic survival analysis is performed. Specific embodiment two:

[0076] The difference between the second embodiment of the present invention and the first embodiment is that:

[0077] The step 1 is specifically as follows:

[0078] Step 1.1: For the target tumor, collect high-quality RNA-seq data of the FPKM expression type from the TCGA public database, convert the expression value to log2 TPM+1, and perform normalization to ensure that the overall expression level of the data conforms to the normal distribution;

[0079] Step 1.2: Collect DNA methylation data of target tumor samples from the TCGA database, preferably high-throughput methylation sequencing data, and use Minfi and ChAMPR packages for background correction, normalization, probe filtering, and batch effect correction;

[0080] Step 1.3: For the target tumor, collect high-quality CNV data of the Segment Mean type from the TCGA public database through the Genomic Data Commons and convert the data into the GISTIC2 format. Perform data cleaning methods on the CNV data, including removing outliers, correcting errors, and removing redundant information. Then, standardize the CNV data to eliminate the influence of experimental batch effects and sample processing differences. Z-score standardization is used to standardize the CNV data.

[0081] Step 1.3: Collect FASTQ-formatted ChIP-Seq data from the TCGA database for the target tumor samples. Use FastQC to perform quality control on the raw sequencing data to identify and address low-quality reads, contamination, or other potential issues. Use Trimmomatic or Cutadapt to remove adapter sequences and low-quality reads. Use alignment tools to align the cleaned reads to the reference genome. Use SAMtools to filter out reads with low alignment quality, multiple alignments, and nonspecific alignments. Use Picard Tools or SAMtools to remove PCR duplicates to ensure data accuracy. Specific embodiment three:

[0083] The only difference between the third embodiment of the present invention and the second embodiment is that:

[0084] The step 2 is specifically as follows:

[0085] Cox regression based on pairwise feature enumeration is used to fit the prognosis survival time, and Pearson correlation based on pairwise feature enumeration is used to measure the correlation between features. The two metrics are unified to label the weights of any two nodes in a fully connected network. The cluster center automatic search clustering algorithm based on density descending is used to delete the edges with lower weights in the network, thereby realizing network construction and feature cluster mining. The cluster center automatic search algorithm based on density descending is used to cluster the weights of all edges in the network in descending density, deleting the edges with lower weight density from the network until disconnected feature clusters appear.

[0086] Step 2.1: Based on the Cox regression, the pairwise feature joint enumeration is performed. The Cox proportional hazard regression model is used to fit the survival time by enumerating the pairwise features, thereby revealing the significant relationship between any pair of features and prognosis survival. The regression coefficient β corresponding to any pair of features is obtained by performing maximum likelihood estimation on the logarithmic partial likelihood function of the hazard function. i ,β j ) T , when it is assumed that the features are uncorrelated, then any component of the regression coefficient obeys the Gaussian distribution, and the Wald test is used to determine whether the regression coefficient component deviates significantly from zero;

[0087] By reordering the samples and their corresponding survival outcomes, the sample size is further expanded. A reordered p-value corresponding to the Wald statistic is generated for each feature. All pairwise features are enumerated to obtain pairwise p-values, which are used to select significant feature pairs closely related to survival from all pairwise feature combinations. To combat overfitting, a penalty constraint is imposed on the log-partial likelihood function corresponding to the hazard function. Considering that the hazard ratio may vary with survival time, a weighted Cox regression model is used instead.

[0088] Step 2.2: Pearson correlation-based joint enumeration of two features. To measure the correlation between two features, the Pearson correlation coefficient r is used to calculate the correlation between any two features, and the t statistic of n samples is constructed based on this. The statistic is expressed as: By using the sample reordering method, the order of sample values of one feature in any pair of features is disrupted, and the above t statistic is calculated repeatedly. The corresponding reordered p-value is expressed as:

[0089]

[0090] Among them, t(i,j) is the statistic corresponding to the correlation coefficient r of features i and j, t b (i, j) represents the statistic generated by randomly reordering the values of any feature, and B is the number of random reorderings;

[0091] Step 2.3: Determine the association relationship of network nodes. The collaborative co-expression network uses individual features as network nodes and the relationship between two features as edges. The weight of the edge between nodes is calculated by integrating the significant relationship between the feature and survival time and the correlation between features. The following metric is constructed to measure the weight of the edge between nodes:

[0092]

[0093] Where p2(i) and p2(j) represent the binary Cox regression reranking p-values obtained by combining the two features i and j, respectively, and p1(i) and p1(j) represent the univariate Cox regression reranking p-values obtained by combining the two features i and j, respectively. represents rounding up; the first term of m(i, j) is the "synergistic" part that measures the contribution of features i and j to survival regression, and the second term is derived from the re-ranked p-value obtained by the joint enumeration of two features based on Pearson correlation, corresponding to the "co-expression" part; m(i, j) is used to measure the association between nodes i and j: the larger its value, the more likely it is that nodes i and j are derived from the same feature cluster, and vice versa. When the value of m(i, j) is less than zero, its value is set to the maximum value of the weight in the network, and the weighted undirected graph constructed is called a synergistic-co-expression network;

[0094] Step 2.4: Based on the feature cluster mining of the collaborative-coexpression network, a density-decreasing cluster center automatic search clustering algorithm is introduced to measure the node association relationship of the weight of the collaborative coexpression network edge. The measurement threshold of the edge in the network is limited by the optimal automatic estimation of the number of categories and cluster boundaries to determine whether there is an associated edge between any two nodes in the network, thereby degenerating the network from a weighted undirected graph to an undirected graph. Specific embodiment four:

[0096] The only difference between the fourth embodiment of the present invention and the third embodiment is that:

[0097] The step 3 is specifically as follows:

[0098] Step 3.1: Bottom-up “cross-cluster enumeration” biomarker selection. Based on the “cross-cluster hypothesis,” univariate enumeration is performed within each feature cluster. Univariate variables from c different feature clusters are combined into a multivariate, and their regression coefficients on the Cox proportional hazards function are calculated.

[0099] Step 3.2: Top-down "cross-cluster random" biomarker selection, introducing a sample resampling voting strategy, randomly selecting a single variable within each feature cluster, combining the single variables from c different feature clusters into a multivariate, and calculating its regression coefficient for the Cox survival hazard ratio function; assuming that the features are independent of each other, analogous to the pairwise joint enumeration method based on Cox regression, each set of "cross-cluster" multivariates can obtain corresponding c re-ranked p-values, and in each round, votes are counted for the multivariates with significant re-ranked p-values; after l rounds of voting, the features with the highest votes in each feature cluster are selected to form the significant prognostic feature set;

[0100] Step 3.3: Unsupervised feature selection of joint multivariates: Individual variables are extracted from different feature clusters to form the joint multivariate for regression analysis. For bottom-up "cross-cluster enumeration" biomarker selection, c re-ranked p-values corresponding to each enumerated multivariate are recorded. For top-down "cross-cluster random" biomarker selection, a threshold is set to record multiple sets of c-dimensional p-value vectors corresponding to multiple randomly extracted multivariates. In the c-dimensional feature space, all enumerated or re-ranked p-value vectors that meet the threshold are automatically clustered based on density-decreasing cluster centroid search. The clusters closest to the origin are selected to form the c-dimensional feature set, and the Mahalanobis distance from the scattered points to the origin is calculated. When bottom-up "cross-cluster enumeration" biomarker selection is used, the multidimensional variable closest to the origin is selected as the significant prognostic feature set. If top-down "cross-cluster random" biomarker selection is used, it is still necessary to determine whether the multidimensional variable closest to the origin corresponds to the feature combination with the highest cumulative votes.

[0101] Step 3.4: Construction of prognostic survival analysis indicators. Taking the Cox proportional hazards regression model as an example, since it is difficult to estimate the baseline hazard function or baseline survival function, the linear part of the risk score is used to construct the prognostic survival indicator. The risk score of the sample is a linear combination of the corresponding sample values of the prognostic characteristics, and the combination coefficient is the regression coefficient. Specific embodiment five:

[0103] The only difference between the fifth embodiment of the present invention and the fourth embodiment is that:

[0104] The step 4 is specifically as follows:

[0105] Step 4.1: Sample clustering for prognostic survival analysis: To discover prognostic survival categories, key biomarkers from different omics or prognostic survival analysis indicators built from different omics data are combined as input to the convolutional neural network. The output is regarded as a new feature subspace, and the samples are automatically clustered based on density-decreasing cluster centroids to discover the prognostic category to which the samples belong.

[0106] Step 4.2: Draw Kaplan-Meier survival analysis curves for samples from different categories to observe the changes in survival time within the category and the differences in survival time between categories. Use the log-rank test to measure whether there are significant quantitative differences in survival time between different categories. Draw a risk score survival analysis chart to observe the changes in survival time between samples from different categories.

[0107] Step 4.3: Analysis of significant differences in prognosis between classes: Use statistics based on sample distribution differences to achieve class comparison and evaluate whether there are significant differences in risk scores between different categories. On this basis, try to use training samples to construct different types of classifiers, and use the prognostic survival results of test samples to verify the accuracy and effectiveness of the classification results. Specific embodiment six:

[0109] The only difference between the sixth embodiment of the present invention and the fifth embodiment is that:

[0110] Synergy represents the correlation between the feature pair composed of two biomarkers and prognostic survival time; co-expression represents the correlation between two biomarkers. Specific embodiment seven:

[0112] The only difference between the seventh embodiment of the present invention and the sixth embodiment is that:

[0113] The cross-group assumptions are as follows:

[0114] The significant feature set closely related to the prognosis of survival time is not entirely composed of a single significant variable related to the prognosis of survival time, and there is no obvious correlation between the variables in the significant feature set. Specific embodiment eight:

[0116] The only difference between the eighth embodiment of the present invention and the seventh embodiment is that:

[0117] The present invention provides a cancer prognosis survival analysis system based on a collaborative co-expression network, the system comprising:

[0118] A data acquisition module, which collects multi-omics data of cancer based on public databases and pre-processes the data;

[0119] A network construction module, wherein the network construction module constructs a collaborative co-expression network for cancer multi-omics data based on the processed data;

[0120] An extraction module, which extracts prognostic biomarkers based on a collaborative co-expression network;

[0121] A clustering module is used to cluster samples for survival analysis based on extracted prognostic biomarkers and perform prognostic survival analysis. Specific embodiment nine:

[0123] The only difference between the ninth embodiment of the present invention and the eighth embodiment is that:

[0124] The present invention provides a computer-readable storage medium having a computer program stored thereon. The program is executed by a processor to implement a cancer prognosis survival analysis method based on a collaborative co-expression network. Specific embodiment ten:

[0126] The only difference between the tenth embodiment of the present invention and the ninth embodiment is that:

[0127] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements a cancer prognosis survival analysis method based on a collaborative co-expression network when executing the computer program. Specific embodiment eleven:

[0129] The only difference between the eleventh embodiment of the present invention and the tenth embodiment is that:

[0130] The present invention will be described in further detail below with reference to the accompanying drawings and specific embodiments.

[0131] Previous research in this paper found that the set of significant features closely related to prognostic survival time is not entirely composed of a single significant variable related to prognostic survival time; that is, there is no obvious correlation between the variables in the significant feature set. (In fact, this hypothesis originated from a series of research works such as the covariate selection method related to tumor prognostic survival proposed previously.) This phenomenon is called . This shows that: the various features in the significant feature set come from different feature clusters in the co-expression network model, rather than the same feature cluster where a single significant feature is located. This is completely contrary to the practice of selecting the cluster where a single significant variable is located from the co-expression network model to achieve feature selection. Take the two features that constitute a significant feature pair (that is, a tuple) as an example: one and only one of the features is closely related to the prognosis survival time; when the other feature is added, the fitting performance of the feature tuple for the prognosis survival time is further enhanced. In fact, the existing co-expression network model is a tree structure, and focusing on the correlation between a pair of features and the prognosis survival time is precisely the key to feature set graphing (that is, networking).

[0132] Based on the above considerations, the cross-group hypothesis is regarded as a collaborative expression of the binary group of prognostic significant features. A prognostic survival analysis method based on the collaborative-co-expression network is proposed to achieve the integration of the five types of strategies in existing research (such as Figure 1 ). Among them, Represents the correlation between the feature pair composed of two biomarkers and prognostic survival time; It represents the correlation between two biomarkers.

[0133] We will construct synergistic co-expression networks for multi-omics data, including RNA-seq, ChIP-seq, DNA methylation, and copy number variation data, and select prognostic biomarkers from these networks. We will also combine prognostic biomarkers from different omics to generate prognostic survival analysis metrics, and construct Bayesian networks to infer causal relationships between different prognostic biomarkers. We will conduct research in the following three areas.

[0134] (1) Research on the construction method of collaborative-co-expression network for cancer multi-omics data

[0135] The two-dimensional matrices corresponding to the multi-omics data are used as input to construct the corresponding synergy-co-expression network. Figure 1As shown in the red box in the figure, it is divided into two parts: "Synergy" and "Co-expression." On the one hand, it focuses on the relationship between biomarkers and survival time: a bottom-up feature combination strategy is used to design a pairwise association metric between biomarkers and survival time, which serves as the "Synergy" part of the constructed network. Unlike the univariate feature combination strategy, the "Synergy" part of the constructed network focuses only on the joint enumeration of pairwise features. On the other hand, it explores the correlation between biomarkers: a correlated feature clustering strategy is used to traverse the correlation between pairwise biomarkers, which serves as the "Co-expression" part of the constructed network. Unlike the correlated feature clustering strategy, the "Co-expression" part of the constructed network focuses on the correlation between pairwise features, rather than clustering feature correlations. The synergy-co-expression network treats features as network nodes and the associations between pairwise features as edges. The key to constructing this network is to determine whether there is an edge that associates any two nodes. The present invention attempts to unify the "synergy" and "co-expression" information of pairwise features into the weight of the edge connecting any two nodes, and uses a density-decreasing cluster center automatic search clustering algorithm to screen out edges with smaller weights, thereby obtaining a synergy-co-expression network for cancer multi-omics data.

[0136] (2) Research on the selection method of prognostic biomarkers by combining multiple variables

[0137] Combined multivariate prognostic biomarker selection Figure 1 As shown in the green box in , it is divided into two parts: "cross-cluster enumeration" and "cross-cluster random". The above two prognostic biomarker selection methods are both based on the cross-cluster hypothesis, that is, the significant features related to survival time come from different feature clusters (or feature clusters, for the convenience of network representation, they are all referred to as feature clusters below). For the constructed collaborative-coexpression network, the present invention will try two feature selection methods, "cross-cluster enumeration" from bottom to top and "cross-cluster random" from top to bottom, based on the idea of "cross-cluster", to obtain a set of key biomarkers closely related to survival time. It should be further explained that when constructing a collaborative-coexpression network, only two features need to be concerned about; the "cross-cluster enumeration" and "cross-cluster random" here are both carried out between different feature clusters in the network. If the constructed collaborative-coexpression network has c feature clusters, the dimension of each round of enumeration or random selection of features is c.

[0138] (3) Research on prognostic category discovery methods for cancer multi-omics data

[0139] Prognostic survival categories were found to be Figure 1As shown in the blue box in the figure, for the prognostic biomarker sets selected from different omics data, the present invention intends to carry out the following research in sequence: construct a prognostic survival analysis indicator by combining prognostic biomarkers from different omics; use a bottom-up sample clustering strategy on the constructed prognostic survival analysis indicator to discover the prognostic survival category of the sample; to verify the reliability of the selected biomarkers and sample categories, use a bottom-up sample classification strategy to determine the multivariate statistical hypothesis test results and classification results of the sample prognostic category; and realize the prediction of the category to which the new sample belongs.

[0140] Research objectives

[0141] In cancer multi-omics analysis, by taking into account the correlation between features and between features and prognostic survival time, selecting key biomarkers closely related to prognostic survival, and discovering possible prognostic survival subtypes of specific cancers, have become two urgent problems to be solved. The present invention will be oriented towards cancer multi-omics data with survival time and survival event occurrence status (such as death or not), and will construct a collaborative-coexpression network based on the cross-cluster hypothesis; it will be divided into "cross-cluster enumeration" and "cross-cluster random" feature selection methods to achieve joint multivariate prognostic biomarker selection; for the submatrix corresponding to the selected prognostic biomarker and its prognostic survival analysis index, an unsupervised clustering method is used to discover different risk categories of sample prognostic survival; the present invention realizes accurate prediction of the prognostic category to which cancer patients belong. In this way, more accurate statistical models and computational tools are provided for biologists to use cancer multi-omics data to study prognostic survival analysis.

[0142] Key issues to be addressed

[0143] In order to achieve the above research objectives, the present invention intends to focus on solving the following three problems.

[0144] (1) Determination of the association relationship between nodes in the collaboration-coexpression network

[0145] The synergy-coexpression network is a small-world network that combines pairwise features. The "cross-cluster hypothesis" states that a pair of features significantly correlated with survival time originates from different clusters, while correlated pairs of features originate from the same cluster. A key challenge is how to unify the metrics of "synergy" and "coexpression" on the edges of the synergy-coexpression network, calculate their weights, and automatically remove low-weight edges to obtain disconnected clusters of features.

[0146] (2) Unsupervised feature selection of joint multivariate

[0147] After establishing a synergistic-coexpression network, a set of prognostic biomarkers needs to be selected from this network for a range of subsequent applications (such as designing prognostic survival analysis indicators). Based on the "cross-cluster hypothesis," two feature selection methods can be used: bottom-up "cross-cluster enumeration" and top-down "cross-cluster randomization." In either approach, key features must be selected from different feature clusters to form the joint multivariate matrix for Cox regression analysis. How to achieve optimal selection of prognostic features in an unsupervised manner is a key issue that needs to be addressed.

[0148] The present invention insists on using the "unsupervised feature selection" expression method, and its viewpoint is as follows: in the process of constructing the collaboration-coexpression network, survival information is indeed used when calculating the weights of the edges between two nodes; however, in the step of realizing the joint multivariate feature selection, only the proposed unsupervised clustering algorithm is used to perform unsupervised clustering on the cross-cluster enumeration or cross-cluster randomly selected feature scores or p-values. The entire clustering process does not involve prognostic survival time or survival status at all. For a detailed description, please see the "Proposed Research Plan" section below.

[0149] (3) Prognostic category discovery for cancer multi-omics data

[0150] Over the past decade, there has been considerable debate over whether to perform univariate sample classification or univariate sample clustering. The conflict centers on how to categorize samples into different survival categories. A key issue that needs to be addressed is how to construct a cross-omics prognostic survival model based on key biomarkers closely associated with prognostic survival time, using data from diverse omics datasets. This model integrates clustering and classification strategies.

[0151] Study plan

[0152] In response to the above research content and research objectives, this paper uses the public multi-omics cancer dataset (TCGA) with prognostic and survival information as the research object, covering the collection and analysis of transcriptomics (RNA-seq), genomics (DNA methylation, CNV), and proteomics (Chip-seq) data. The specific data preprocessing scheme for the multi-omics data of a specific cancer is as follows:

[0153] ① For the target tumor, we plan to collect high-quality RNA-seq data of the FPKM expression type from the TCGA public database and convert the expression values into log2(TPM+1) for normalization. This aims to ensure that the overall expression level of the data conforms to the normal distribution and ensure the accuracy of the experiment.

[0154] ②Collect DNA methylation data of target tumor samples from the TCGA database, preferably using high-throughput methylation sequencing data (such as Illumina 450K or EPIC BeadChip). Use R packages such as Minfi and ChAMP for background correction, normalization, probe filtering, and batch effect correction.

[0155] ③ For target tumors, high-quality CNV data of the SegmentMean type were collected from the TCGA public database through the Genomic Data Commons (GDC). For further analysis, the data were converted to the GISTIC2 format. To ensure data quality, the CNV data were first cleaned using methods such as outlier removal, error correction, and redundant information removal. Subsequently, the CNV data were normalized to eliminate the effects of experimental batch effects and sample processing differences. Z-score normalization was proposed for the CNV data.

[0156] ④ FASTQ-formatted chromatin immunoprecipitation sequencing (ChIP-Seq) data from target tumor samples will be collected from the TCGA database. FastQC will be used to perform quality control on the raw sequencing data to identify and address low-quality reads, contamination, or other potential issues. Tools such as Trimmomatic or Cutadapt will be used to remove adapter sequences and low-quality reads. Alignment tools (such as BWA, Bowtie2, or STAR) will be used to align the cleaned reads to the reference genome. SAMtools will be used to filter out reads with low alignment quality, multiple alignments, and nonspecific alignments. Further experiments will use tools such as Picard Tools or SAMtools to remove PCR duplicates to ensure data accuracy.

[0157] Based on the above data, the present invention is divided into four parts to carry out the prognosis survival analysis research. The overall technical roadmap is as follows: Figure 2 As shown in Figure 2. Solid arrows indicate the sequence of the research plan. For multi-omics data from a specific cancer, the processing plan is as follows: After preprocessing the data from different omics, the synergistic-coexpression network is constructed and prognostic biomarkers are extracted. The prognostic biomarkers from different omics are combined as input to a convolutional neural network, and their output is used for prognostic category discovery. Based on this, causal relationships are inferred for the prognostic biomarker sets from different omics.

[0158] (1) Construction of collaborative-co-expression networks for cancer multi-omics data

[0159] This part serves as the basis for the entire project. The specific research work includes: how to combine pairwise features to achieve the fitting of prognosis survival; how to calculate the correlation between pairwise features; on this basis, consider how to unify the fitting results and correlation metrics with the weight metrics of the edges between network nodes and the key issue of deleting edges with lower weights.

[0160] like Figure 3 As shown in the figure, for a specific cancer omics data set, Cox regression based on pairwise feature enumeration is proposed to fit the prognostic survival time. Pearson correlation based on pairwise feature enumeration is used to measure the correlation between features. The two metrics mentioned above are unified to label the weights of any two nodes in a fully connected network. The proposed cluster center automatic search clustering algorithm based on density descending is used to remove edges with lower weights in the network, thereby achieving network construction and feature cluster mining. Here, the proposed cluster center automatic search algorithm based on density descending is used to cluster the weights of all edges in the network in descending density, removing edges with lower weight density (corresponding to the several low-density Gaussian distributions where the weights are located) from the network until disconnected feature clusters are formed. The details of the proposed clustering algorithm are detailed in the "Research Foundation" section below and the corresponding references.

[0161] Based on the joint enumeration of two features of Cox regression. The research on survival analysis involves a variety of models. The Cox proportional hazard regression model has been integrated into various application statistical software because it is easy to understand and implement, and has been widely used. The present invention adopts the Cox proportional hazard regression model to fit the survival time by enumerating two features, thereby revealing the significant relationship between any pair of features and prognosis of survival. By performing maximum likelihood estimation on the logarithmic partial likelihood function of the hazard function, the regression coefficient β corresponding to any two features can be obtained = (β i ,β j ) T . If it is assumed that the features are uncorrelated, then any component of the regression coefficient obeys a Gaussian distribution. The Wald test can be used to determine whether the regression coefficient component deviates significantly from zero. By reordering the samples and their corresponding survival outcomes to further expand the sample size, a reordered p-value corresponding to the Wald statistic can be generated for each feature. Enumerating all pairwise features to obtain paired p-values can be used to select significant feature pairs that are closely related to survival from all pairwise feature combinations. To deal with overfitting, it is possible to consider adding a penalty constraint to the log-partial likelihood function corresponding to the hazard function; considering that the hazard ratio may change with survival time, a weighted Cox regression model can be used instead.

[0162] Pairwise feature joint enumeration based on Pearson correlation. To measure the correlation between two features, the Pearson correlation coefficient r is used to calculate the correlation between any two features and use it to construct the t statistic for n samples. This statistic can be expressed as: Still using the sample reordering method, we disrupt the order of sample values of any two features and repeat the calculation of the above t statistic. The corresponding reordered p-value can be expressed as:

[0163]

[0164] Among them, t(i,j) is the statistic corresponding to the correlation coefficient r of features i and j, t b (i, j) represents the statistic generated by randomly reordering the values of any feature, and B is the number of random reorderings.

[0165] Determination of the association relationship of network nodes. The collaborative-co-expression network uses a single feature as a network node and the relationship between two features as an edge. The weight of the edge between nodes is calculated by comprehensively analyzing the significant relationship between the features and the survival time and the correlation between the features. On the one hand, the Cox regression model based on the joint enumeration of two features identifies the association relationship between paired network nodes and survival time. The smaller the re-ranking p-value corresponding to any pair of feature nodes, the greater the contribution of these two nodes to the regression of the survival risk proportional function. On the other hand, the Pearson correlation coefficient based on the joint enumeration of two features identifies the correlation between two nodes. The smaller the re-ranking p-value corresponding to any pair of feature nodes, the more correlated the two nodes are. According to the "cross-cluster hypothesis", two features that are significantly correlated with survival time come from different feature clusters; and the more correlated two features are, the more likely they are to come from the same feature cluster. Therefore, the present invention intends to construct the following metric for measuring the weight of the edge between nodes:

[0166]

[0167] Where p2(i) and p2(j) represent the binary Cox regression reranking p-values obtained by combining the two features i and j, respectively, and p1(i) and p1(j) represent the univariate Cox regression reranking p-values obtained by combining the two features i and j, respectively. Represents rounding up. The first term of m(i,j) is the "synergistic" part that measures the contribution of features i and j to survival regression, and the second term is derived from the re-ranked p-value obtained by the joint enumeration of two features based on Pearson correlation, which corresponds to the "co-expression" part. m(i,j) is used to measure the association between nodes i and j: the larger its value, the more likely it is that nodes i and j are derived from the same feature cluster, and vice versa. In addition, when the value of m(i,j) is less than zero, its value is set to the maximum value of the weight in the network. The weighted undirected graph constructed in this way is called .

[0168] Feature cluster mining based on synergy-coexpression networks. This paper introduces a proposed density-decreasing cluster center automatic search clustering algorithm based on the weight of the edges in the synergy-coexpression network, which measures the node association relationship. By automatically estimating the number of categories and cluster boundaries, it defines the edge measurement threshold in the network to determine whether any two nodes in the network have an associated edge. This method degenerates the network from a weighted undirected graph to an undirected graph. Disconnected nodes are considered to originate from different feature clusters, and the synergy-coexpression network is degenerated into several non-overlapping feature clusters.

[0169] (2) Prognostic biomarker extraction

[0170] This part is based on the "cross-cluster hypothesis" and aims to extract a set of significant biomarkers closely related to prognostic survival (i.e., prognostic feature set). Specific research work includes: how to enumerate all cross-cluster feature combinations from the bottom up to select a prognostic biomarker set; how to randomly extract cross-cluster feature combinations from the top down to select a prognostic biomarker set; for any of the above feature selection schemes, consider the key issue of how to automatically select a prognostic feature set from a large number of multidimensional features, that is, the problem of unsupervised feature selection of joint multiple variables; how to construct clinical indicators for survival analysis based on the selected prognostic feature set.

[0171] like Figure 4 As shown, the resulting synergistic-coexpression network clearly identifies distinct feature clusters. Based on the "cross-cluster hypothesis," a bottom-up cross-cluster feature enumeration and a top-down random cross-cluster feature selection were employed on a specific omics matrix. Ignoring the correlations between significant features, an unsupervised feature selection scheme for multivariate measurement was designed based on a density-decreasing cluster centroid automatic search clustering algorithm. On this basis, the selected prognostic features were used to calculate risk scores based on the Cox proportional hazards model, which were then used to construct prognostic survival analysis indicators.

[0172] Bottom-up "cross-cluster enumeration" biomarker selection. Based on the "cross-cluster hypothesis," single variables are enumerated within each feature cluster, and the single variables from c different feature clusters are combined into multivariate variables, and their regression coefficients for the Cox survival hazard ratio function are calculated. This method is essentially a multivariate joint enumeration based on Cox regression. Assuming that the features are independent of each other, analogous to the pairwise feature joint enumeration method based on Cox regression, each set of "cross-cluster" multivariate variables can obtain corresponding c re-ranked p-values, which can be used to select a set of significant prognostic features that are most closely related to survival from all c-dimensional multivariates. Similarly, it is possible to consider adding penalty constraints to the log-likelihood function corresponding to the hazard function or using a weighted Cox regression model to improve the existing model.

[0173] Top-down "cross-cluster random" biomarker selection. If the number of feature clusters is too large, a top-down "cross-cluster random" feature selection method should be considered. A sample resampling voting strategy is introduced, randomly selecting a single variable within each feature cluster. Single variables from c different feature clusters are combined into a multivariate, and their regression coefficients for the Cox survival hazard ratio function are calculated. Assuming that features are independent of each other, analogous to the pairwise joint enumeration method based on Cox regression, each set of "cross-cluster" multivariates will obtain corresponding c reranked p-values. Multivariates with significant reranked p-values are voted on in each round. After l rounds of voting, the features with the highest votes in each feature cluster are selected to form the significant prognostic feature set.

[0174] Unsupervised feature selection of joint multivariates. Regardless of the aforementioned methods, individual variables must be extracted from different feature clusters to form the joint multivariate for regression analysis. For bottom-up "cross-cluster enumeration" biomarker selection, c re-ranked p-values corresponding to each enumerated multivariate are recorded. For top-down "cross-cluster random" biomarker selection, a threshold is set to record multiple sets of c-dimensional p-value vectors corresponding to multiple randomly extracted multivariates. In the c-dimensional feature space, all enumerated or re-ranked p-value vectors that meet the threshold are automatically clustered using a density-decreasing cluster centroid search. The clusters closest to the origin are selected to form the c-dimensional feature set, and the Mahalanobis distances of the scattered points to the origin are calculated. If bottom-up "cross-cluster enumeration" biomarker selection is used, the multidimensional variable closest to the origin is selected as the significant prognostic feature set. If top-down "cross-cluster random" biomarker selection is used, it is still necessary to determine whether the multidimensional variable closest to the origin corresponds to the feature combination with the highest cumulative votes.

[0175] Prognostic survival analysis indicator construction. Taking the Cox proportional hazards regression model as an example, due to the difficulty in estimating the baseline hazard function or baseline survival function, the risk score (risk score) - the linear portion of this risk function - is often used to construct a prognostic survival indicator. The risk score of a sample is a linear combination of the corresponding sample values of the prognostic features, and the combination coefficient is the regression coefficient. If other models are used, it is possible to consider directly using the submatrix element values corresponding to the prognostic feature set as the indicator for sample prognostic survival analysis.

[0176] (3) Prognostic category discovery for cancer multi-omics data

[0177] This part aims to discover different categories of prognostic samples. Specific research work includes: how to cluster samples based on the established prognostic survival analysis indicators; how to verify the reliability of prognostic categories based on sample distribution differences. It is planned to use a cluster center automatic search clustering algorithm based on density descending to cluster samples of prognostic survival analysis indicators, and to use hypothesis testing technology to analyze the degree of expression differences between categories. The schematic diagram of prognostic category discovery is shown in the figure below. Figure 5 shown.

[0178] Sample clustering for prognostic survival analysis. To identify prognostic survival categories, key biomarkers derived from different omics or prognostic survival analysis indicators constructed from different omics data are combined as input to a convolutional neural network. The output is treated as a new feature subspace, and samples are automatically clustered using a density-decreasing cluster centroid search to identify the prognostic category to which the samples belong. Kaplan-Meier survival analysis curves are plotted for samples from different categories to observe changes in survival time within and between classes. The log-rank test is used to test whether there are quantitatively significant differences in survival time between different categories. Furthermore, risk score survival analysis plots are plotted to observe changes in survival time between samples from different categories.

[0179] Analysis of significant differences in prognosis between categories. To verify the reliability of the prognostic categories found, a series of statistical methods based on sample distribution differences (such as Welch two-sample t test, one-way analysis of variance, Hotelling two-sample t test, etc.) were used. 2 Test, etc.) to achieve class comparison and evaluate whether there are significant differences in risk scores of different categories; on this basis, try to use training samples to build different types of classifiers, and use the prognostic survival results of test samples to verify the accuracy and effectiveness of the classification results.

[0180] The present invention discovered that pairs of microRNAs significantly correlated with survival time come from different characteristic clusters. This serves as the most direct experimental verification of the "cross-cluster hypothesis." Furthermore, the applicants have independently developed five practical, original differential expression and prognostic survival analysis tools (including: JSDSA, JCD-DEA, ECFS-DEA, AFS-DEA, and IOFS-SA). For details on previous and ongoing research work, please refer to the Research Foundation section.

[0181] Research foundation;

[0182] Since December 2012, the applicant has been engaged in research on differential expression analysis and prognostic survival analysis of cancer multi-omics data. In recent years, the applicant has conducted a large amount of preliminary research on this topic, including: (1) Research on covariate selection methods related to cancer prognosis and survival; (2) Research on clustering methods based on automatic cluster center search based on density descending order; (3) Research on re-clustering methods for prognostic survival category discovery; (4) Development of related differential expression analysis and prognostic survival analysis tools. The above research has laid the foundation for the development of this project.

[0183] Study on covariate methods related to cancer prognosis and survival

[0184] This research work discovered a pair of microRNAs that lead to significant differences in the prognosis of malignant brain gliomas through bottom-up multivariate joint enumeration. Further experimental results verified the "cross-cluster hypothesis." The applicant, together with relevant medical teams, conducted statistical experiments and pathway validation on the microRNA expression profiles of the TCGA public dataset, designed a roadmap based on covariate selection, and discovered a pair of key microRNAs that lead to differences in the prognosis and survival of malignant brain gliomas. While verifying the idea of bottom-up multivariate joint enumeration, this research work discovered that the two microRNAs significantly correlated with survival time may come from different gene clusters. Figure 6 It was revealed that the selected two significant microRNAs (miR-222 and miR-10b) came from different gene clusters. Figure 7 The solid black boxes represent gene clusters, and the dashed black boxes represent pairwise significant genes. This indicates that genes within a gene cluster are significantly positively correlated, and that pairwise significant genes cross clusters. These results serve as the most direct experimental verification of the "cross-cluster hypothesis."

[0185] The applicant has also independently developed five practical and original differential expression and prognostic survival analysis tools, including: JSDSA, JCD-DEA, ECFS-DEA, AFS-DEA and IOFS-SA. JSDSA is a bottom-up prognostic survival analysis tool based on Cox regression. The development of this tool provides strong technical support for the realization of pairwise feature joint enumeration based on Cox regression and bottom-up "cross-cluster enumeration" feature selection. JCD-DEA is a bottom-up differential expression analysis tool that combines sample distribution and classification differences. ECFS-DEA is a top-down differential expression analysis tool based on integrated classification. AFS-DEA is its network version. The development of the above tools can provide strong technical support for the analysis of significant prognostic differences between classes. IOFS-SA is a special tool for the construction of prognostic survival analysis indicators and the discovery of prognostic categories. It realizes interactive re-clustering of samples based on the constructed prognostic survival analysis indicators.

[0186] The above is only a preferred embodiment of a method for cancer prognosis and survival analysis based on a collaborative co-expression network. The scope of protection of a method for cancer prognosis and survival analysis based on a collaborative co-expression network is not limited to the above embodiment. All technical solutions based on this concept fall within the scope of protection of the present invention. It should be noted that for those skilled in the art, several improvements and variations without departing from the principles of the present invention should also be considered as the scope of protection of the present invention.

Claims

1. A cancer prognosis survival analysis method based on a collaborative co-expression network, characterized by: The following steps are involved: Step 1: Collect multi-omics data on cancer based on public databases and preprocess the data; Step 2: Based on the processed data, construct a collaborative co-expression network for cancer multi-omics data; Cox regression based on pairwise feature enumeration is used to fit the prognosis survival time, and Pearson correlation based on pairwise feature enumeration is used to measure the correlation between features. The two metrics, Cox regression based on pairwise feature enumeration and Pearson correlation based on pairwise feature enumeration, are unified to mark the weights of any two nodes in a fully connected network. An automatic cluster search algorithm based on density descending is used to delete edges with lower weights in the network, thereby achieving network construction and feature cluster mining. The automatic cluster search algorithm based on density descending is used to cluster the weights of all edges in the network in descending density, deleting edges with lower weight density from the network until disconnected feature clusters appear. The collaborative co-expression network uses individual features as network nodes and the relationships between two features as edges. The weights of the edges between nodes are calculated by comprehensively analyzing the significant relationship between the features and survival time, as well as the correlation between the features. The following metric is constructed to measure the weights of the edges between nodes: in, and Represents the joint pairwise features and The resulting binary Cox regression reordering p value, and Represents a single feature and The resulting univariate Cox regression reordering p value, represents rounding up; The first item is to measure the characteristics and The second term of the "synergistic" contribution to survival regression comes from the reordering obtained by joint enumeration of pairwise features based on Pearson correlation p value, corresponding to the "co-expression" part; To measure the nodes and The relationship between nodes: the larger the value, the more it indicates that the node and nodes may originate from the same feature cluster, and vice versa, when When the value of is less than zero, its value is set to the maximum value of the weight in the network, and the weighted undirected graph constructed is called a collaboration-co-expression network; Step 3: Extract prognostic biomarkers based on the collaborative co-expression network; Step 4: Based on the extracted prognostic biomarkers, sample clustering for survival analysis is performed, and prognostic survival analysis is performed.

2. The method for cancer prognosis and survival analysis based on a collaborative co-expression network according to claim 1, wherein: The step 1 is specifically as follows: Step 1.1: For the target tumor, collect high-quality RNA-seq data of the FPKM expression type from the TCGA public database, convert the expression value to log2 TPM+1, and perform normalization to ensure that the overall expression level of the data conforms to the normal distribution; Step 1.2: Collect DNA methylation data of target tumor samples from the TCGA database, preferably high-throughput methylation sequencing data, and use Minfi and ChAMPR packages for background correction, normalization, probe filtering, and batch effect correction; Step 1.3: For the target tumor, collect high-quality CNV data of the SegmentMean type from the TCGA public database through the Genomic Data Commons and convert the data into the GISTIC2 format. Perform data cleaning, including outlier removal, error correction, redundant information removal, and data standardization to eliminate the influence of experimental batch effects and sample processing differences. Z-score standardization is used to standardize the CNV data. Step 1.3: Collect FASTQ-formatted ChIP-Seq data from the TCGA database for the target tumor samples. Use FastQC to perform quality control on the raw sequencing data to identify and process low-quality reads and contamination. Use Trimmomatic or Cutadapt to remove adapter sequences and low-quality reads. Use alignment tools to align the cleaned reads to the reference genome. Use SAMtools to filter out reads with low alignment quality, multiple alignments, and nonspecific alignments. Use Picard Tools or SAMtools to remove PCR duplicates to ensure data accuracy.

3. The method for cancer prognosis and survival analysis based on a collaborative co-expression network according to claim 2, wherein: The step 2 is specifically as follows: Step 2.1: Based on the Cox regression, the Cox proportional hazard regression model is used to fit the survival time by enumerating the pairwise features, so as to reveal the significant relationship between any pair of features and prognosis survival. The regression coefficient corresponding to any pair of features is obtained by performing maximum likelihood estimation on the logarithmic partial likelihood function of the hazard function. , when it is assumed that the features are uncorrelated, then any component of the regression coefficient obeys the Gaussian distribution, and the Wald test is used to determine whether the regression coefficient component deviates significantly from zero; By reordering the samples and their corresponding survival results, the sample size can be further expanded, and the reordering corresponding to the Wald statistic can be generated on each feature. p Value, enumerate all two-way features to get paired p The value is used to select significant feature pairs closely related to survival from all pairwise feature combinations. To deal with overfitting, a penalty constraint is added to the logarithmic partial likelihood function corresponding to the hazard function. Considering that the hazard ratio may change with survival time, a weighted Cox regression model is used instead. Step 2.2: Pearson correlation-based joint enumeration of two features. To measure the correlation between two features, the Pearson correlation coefficient is used. Calculate the correlation between any two features and use it to build of samples Statistics, statistics are expressed as: , using the sample reordering method, disrupting the order of sample values of one feature in any pair of features, and repeatedly calculating the above Statistics, corresponding reordering p The value is expressed as: in, Characterized by and Correlation coefficient The corresponding statistics, Represents the statistic generated by randomly reordering the values of any feature. is the number of random reorderings; Step 2.3: Determine the association relationship of network nodes; Step 2.4: Based on the feature cluster mining of the collaborative-coexpression network, a density-decreasing cluster center automatic search clustering algorithm is introduced to measure the node association relationship of the weight of the collaborative coexpression network edge. The measurement threshold of the edge in the network is limited by the optimal automatic estimation of the number of categories and cluster boundaries to determine whether there is an associated edge between any two nodes in the network, thereby degenerating the network from a weighted undirected graph to an undirected graph.

4. The method for cancer prognosis and survival analysis based on a collaborative co-expression network according to claim 3, wherein: The step 3 is specifically as follows: Step 3.1: Bottom-up "cross-cluster enumeration" biomarker selection, based on the "cross-cluster hypothesis", enumerate single variables within each feature cluster, and The single variables of different feature groups are combined into multivariate variables, and their regression coefficients on the Cox survival hazard ratio function are calculated; Step 3.2: Top-down "cross-cluster random" biomarker selection, introducing a sample resampling voting strategy, randomly selecting a single variable within each feature cluster, and The single variables of different feature groups are combined into multivariate, and their regression coefficients for the Cox survival hazard ratio function are calculated; assuming that the features are uncorrelated, analogous to the pairwise feature joint enumeration method based on Cox regression, each group of "cross-group" multivariate can obtain the corresponding Reorder p Value, reorder each round p The multivariate vote count values are all significant; After rounds of voting, the features with the highest votes are selected from each feature cluster to form a significant prognostic feature set; Step 3.3: Unsupervised feature selection of joint multivariates; single variables need to be extracted from different feature clusters to form joint multivariates for regression analysis. For bottom-up "cross-cluster enumeration" biomarker selection, record the corresponding features of each enumerated multivariate. Reorder p For top-down "cross-group random" biomarker selection, multiple groups corresponding to multiple randomly selected variables can be recorded by setting thresholds. dimension p Value vector, in In the dimensional feature space, all enumerations or reorderings that meet the threshold p The value vector is used to automatically search for cluster centers based on density descending order, and the cluster closest to the origin is selected. dimensional feature set, and calculate the Mahalanobis distance from the scattered points to the origin. When using the bottom-up "cross-cluster enumeration" biomarker selection, the multidimensional variable with the closest distance to the origin is selected as the significant prognostic feature set; if the top-down "cross-cluster random" biomarker selection is used, it is still necessary to determine whether the multidimensional variable with the closest distance to the origin corresponds to the feature combination with the highest cumulative votes; Step 3.4: Construct prognostic survival analysis indicators using the Cox proportional hazards regression model. Due to the difficulty in estimating the baseline hazard function or baseline survival function, the linear part of the risk score is used to construct the prognostic survival indicator. The risk score of the sample is a linear combination of the corresponding sample values of the prognostic characteristics, and the combination coefficient is the regression coefficient.

5. The method for cancer prognosis and survival analysis based on a collaborative co-expression network according to claim 4, wherein: The step 4 is specifically as follows: Step 4.1: Sample clustering for prognostic survival analysis: To discover prognostic survival categories, key biomarkers from different omics or prognostic survival analysis indicators built from different omics data are combined as input to the convolutional neural network. The output is regarded as a new feature subspace, and the samples are automatically clustered based on density-decreasing cluster centroids to discover the prognostic category to which the samples belong. Step 4.2: Draw Kaplan-Meier survival analysis curves for samples from different categories to observe the changes in survival time within the category and the differences in survival time between categories. Use the log-rank test to measure whether there are significant quantitative differences in survival time between different categories. Draw a risk score survival analysis chart to observe the changes in survival time between samples from different categories. Step 4.3: Analysis of significant differences in prognosis between classes: Use statistics based on sample distribution differences to achieve class comparison and evaluate whether there are significant differences in risk scores between different categories. On this basis, use the training samples to construct different types of classifiers, and use the prognostic survival results of the test samples to verify the accuracy and effectiveness of the classification results.

6. The method for cancer prognosis and survival analysis based on a collaborative co-expression network according to claim 5, wherein: Synergy represents the correlation between the feature pair composed of two biomarkers and prognostic survival time; co-expression represents the correlation between two biomarkers.

7. The method for cancer prognosis and survival analysis based on a collaborative co-expression network according to claim 4, wherein: The cross-group assumptions are as follows: The significant feature set closely related to the prognosis of survival time is not entirely composed of a single significant variable related to the prognosis of survival time, and there is no obvious correlation between the variables in the significant feature set.

8. A cancer prognosis survival analysis system based on a collaborative co-expression network, characterized by: The system comprises: A data acquisition module, which collects multi-omics data of cancer based on public databases and pre-processes the data; A network construction module, wherein the network construction module constructs a collaborative co-expression network for cancer multi-omics data based on the processed data; Cox regression based on pairwise feature enumeration is used to fit the prognosis survival time, and Pearson correlation based on pairwise feature enumeration is used to measure the correlation between features. The two metrics, Cox regression based on pairwise feature enumeration and Pearson correlation based on pairwise feature enumeration, are unified to mark the weights of any two nodes in a fully connected network. An automatic cluster search algorithm based on density descending is used to delete edges with lower weights in the network, thereby achieving network construction and feature cluster mining. The automatic cluster search algorithm based on density descending is used to cluster the weights of all edges in the network in descending density, deleting edges with lower weight density from the network until disconnected feature clusters appear. The collaborative co-expression network uses individual features as network nodes and the relationships between two features as edges. The weights of the edges between nodes are calculated by comprehensively analyzing the significant relationship between the features and survival time, as well as the correlation between the features. The following metric is constructed to measure the weights of the edges between nodes: in, and Represents the joint pairwise features and The resulting binary Cox regression reordering p value, and Represents a single feature and The resulting univariate Cox regression reordering p value, represents rounding up; The first item is to measure the characteristics and The second term of the "synergistic" contribution to survival regression comes from the reordering obtained by joint enumeration of pairwise features based on Pearson correlation p value, corresponding to the "co-expression" part; To measure the nodes and The relationship between nodes: the larger the value, the more it indicates that the node and nodes may originate from the same feature cluster, and vice versa, when When the value of is less than zero, its value is set to the maximum value of the weight in the network, and the weighted undirected graph constructed is called a collaboration-co-expression network; An extraction module, which extracts prognostic biomarkers based on a collaborative co-expression network; A clustering module is used to cluster samples for survival analysis based on extracted prognostic biomarkers, and to perform prognostic survival analysis.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement a cancer prognosis survival analysis method based on a collaborative co-expression network as claimed in any one of claims 1 to 7.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method for cancer prognosis survival analysis based on a collaborative co-expression network according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method for screening potential biomarkers of gastric cancer based on weighted gene co-expression network analysis, and application thereof

    CN109872776A

  • Construction method of prostatic cancer prognosis significant correlation ceRNA regulation and control network

    CN112837744A