Hepatocellular carcinoma data processing method and system
By using the CIBERSORT tool and weighted gene co-expression network analysis to screen hepatocellular carcinoma-related genes, combined with differential expression and prognosis correlation analysis, and using a comprehensive machine learning algorithm to generate an optimal prognosis prediction model, the problem of strong subjectivity in the prediction results of existing models was solved, and a more accurate prognosis prediction of hepatocellular carcinoma was achieved.
Patent Information
- Application Number
- CN202511113874.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-11
AI Technical Summary
The existing hepatocellular carcinoma prognosis prediction model is based on a predetermined algorithm, which leads to highly subjective prediction results and fails to fully utilize the model's predictive effectiveness.
The CIBERSORT tool combined with weighted gene co-expression network analysis was used to screen gene sets related to M2 macrophage infiltration. Target genes were identified through differential expression analysis and prognostic correlation analysis. A comprehensive machine learning algorithm was used to generate a prognostic prediction model, and the optimal model was selected using C-index and complexity.
The objectivity and accuracy of the hepatocellular carcinoma prognosis prediction model are improved, and the generated prediction results are more realistic and reliable.
Smart Images

Figure CN120600124A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biological data analysis, and in particular relates to a hepatocellular carcinoma data processing method and system. Background Art
[0002] Hepatocellular carcinoma (HCC) is one of the most common malignant tumors. The current epidemiological status of HCC is not optimistic, and new breakthroughs are needed in all aspects of HCC prevention, diagnosis, treatment, and prognosis assessment.
[0003] It should be noted that most prognostic prediction models built based on machine learning theory use predetermined algorithms, which makes the prediction results obtained full of subjectivity and often fails to maximize the predictive effectiveness of the model. Summary of the Invention
[0004] Based on this, an embodiment of the present invention provides a hepatocellular carcinoma data processing method and system, aiming to establish a prognosis model and improve prediction accuracy.
[0005] A first aspect of an embodiment of the present invention provides a method for processing hepatocellular carcinoma data, the method comprising: The CIBERSORT tool combined with weighted gene co-expression network analysis was used to screen gene sets positively correlated with M2 macrophage infiltration; Determining target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis; Obtaining a data set containing the target gene, and using a comprehensive machine learning algorithm to generate a corresponding number of prognostic prediction models in each data set; Calculate the C-index of each prognostic prediction model under different data sets, calculate the average value of the C-index, and determine the prognostic prediction model corresponding to the maximum average value of the C-index; Determine whether the prognostic prediction model corresponding to the maximum average value of C-index is unique; If so, the prognostic prediction model corresponding to the maximum average value of the C-index is determined as the target prognostic prediction model; If not, the complexity of the prognostic prediction model corresponding to the maximum average value of the C-index is calculated according to the C-index of each prognostic prediction model under different data sets, and the prognostic prediction model corresponding to the minimum complexity is determined as the target prognostic prediction model.
[0006] Furthermore, the dataset is an RNA-seq dataset for hepatocellular carcinoma users, including the TCGA dataset, ICGC dataset, GSE14520 dataset, and GSE76427 dataset.
[0007] Furthermore, the step of using the CIBERSORT tool in combination with weighted gene co-expression network analysis to screen a gene set positively correlated with M2 macrophage infiltration includes: The whole RNA sequencing data in the TCGA dataset was fed into the CIBERSORT tool to generate an immune cell infiltration expression matrix for each HCC sample in the corresponding dataset; Weighted gene co-expression network analysis was used to analyze the immune cell infiltration expression matrix of each HCC sample in the dataset; The neighborhood relationships were converted into a topological overlap matrix, clustered into a chain hierarchy according to the average values of different topological overlap matrix metrics, and the gene sets positively correlated with M2 macrophage infiltration were identified.
[0008] Furthermore, in the step of analyzing the immune cell infiltration expression matrix of each hepatocellular carcinoma sample in the dataset using weighted gene co-expression network analysis, the independence β value was set to 0.9, and hepatocellular carcinoma samples with an average expression level > 0.5 were selected for analysis.
[0009] Furthermore, the step of determining a target gene with prognostic value from the gene set by differential expression analysis and prognostic correlation analysis includes: determining the first differentially expressed gene between hepatocellular carcinoma tissue and normal liver tissue in the TCGA dataset by transcriptome expression difference analysis with a threshold of |LogFC|≥1.5, wherein |LogFC| is the absolute value of the logarithmic fold change; searching for the first differentially expressed gene and the same member of M2 macrophages, and determining the second differentially expressed gene between hepatocellular carcinoma tissue and normal liver tissue; Cox regression analysis was used to compare the correlation between the expression of the second differentially expressed gene and the overall survival time of hepatocellular carcinoma patients to determine the target gene.
[0010] Furthermore, the comprehensive machine learning algorithms include 8 single algorithms and 22 joint algorithms, including SurvivalSVM, SuperPC, Enet, StepCox, Ridge regression, plsRcox, RSF and GBM.
[0011] Furthermore, the step of calculating the complexity of the prognosis prediction model corresponding to the maximum average value of the C-index of each prognosis prediction model under different data sets includes: Determining a likelihood function of a prognostic prediction model corresponding to a maximum average value of the C-index, wherein the likelihood function includes a parameter vector of the prognostic prediction model; Take the logarithm of the likelihood function to obtain the log-likelihood function, then take the derivative of the parameter vector and set the derivative to zero, solve the parameter estimate that maximizes the log-likelihood function, and obtain the maximum likelihood estimate; Determining the number of parameters of the prognostic prediction model corresponding to the maximum average value of the C-index, and calculating the complexity based on the number of parameters and the maximum likelihood estimate; Rank the C-index of each prognostic prediction model under different data sets according to the value from large to small. According to the ranking results and the importance of the data set, determine the adjustment factor corresponding to the C-index of each prognostic prediction model under different data sets; A target complexity is obtained by calculation according to the adjustment factor and the complexity.
[0012] A second aspect of the embodiments of the present invention provides a hepatocellular carcinoma data processing system, configured to implement the hepatocellular carcinoma data processing method provided in the first aspect, the system comprising: A screening module is used to screen gene sets positively correlated with M2 macrophage infiltration using the CIBERSORT tool combined with weighted gene co-expression network analysis; A first determination module is used to determine target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis; A generation module is used to obtain a data set containing the target gene, and use a comprehensive machine learning algorithm to generate a corresponding number of prognosis prediction models in each data set; A calculation module is used to calculate the C-index of each prognostic prediction model under different data sets, calculate the average value of the C-index, and determine the prognostic prediction model corresponding to the maximum average value of the C-index; A judgment module is used to judge whether the prognosis prediction model corresponding to the maximum average value of C-index is unique; A second determining module is configured to determine the prognosis prediction model corresponding to the maximum average value of the C-index as the target prognosis prediction model if it is determined that the prognosis prediction model corresponding to the maximum average value of the C-index is unique; The third determination module is used to calculate the complexity of the prognosis prediction model corresponding to the maximum average value of the C-index if it is determined that the prognosis prediction model corresponding to the maximum average value of the C-index is not unique, and determine the prognosis prediction model corresponding to the minimum complexity as the target prognosis prediction model.
[0013] A third aspect of the embodiments of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the hepatocellular carcinoma data processing method provided in the first aspect.
[0014] A fourth aspect of an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the hepatocellular carcinoma data processing method provided in the first aspect is implemented.
[0015] A method and system for processing hepatocellular carcinoma data provided in an embodiment of the present invention screens a gene set that is positively correlated with M2 macrophage infiltration by utilizing the CIBERSORT tool in combination with weighted gene co-expression network analysis; determines target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis; obtains a data set containing the target gene, and uses a comprehensive machine learning algorithm to generate corresponding prognostic prediction models in each data set; calculates the C-index of the data set in the corresponding prognostic prediction model, and screens the target prognostic prediction model based on the C-index and the complexity of the prognostic prediction model. Compared with the traditional modeling method in which the algorithm is selected at the beginning of the study, the target prognostic prediction model finally screened reflects more objective and authentic prediction results, and at the same time, the prediction accuracy of the target prognostic prediction model is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flowchart of a method for processing hepatocellular carcinoma data provided in Example 1 of the present invention; Figure 2 This is a structural block diagram of a hepatocellular carcinoma data processing system provided in Example 2 of the present invention; Figure 3 This is a structural block diagram of an electronic device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0017] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.
[0018] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0020] Example 1 See also Figure 1 , Figure 1 The flowchart of the implementation of a hepatocellular carcinoma data processing method provided in the first embodiment of the present invention is shown. The hepatocellular carcinoma data processing method specifically includes steps S01 to S07.
[0021] Step S01: Using the CIBERSORT tool combined with weighted gene co-expression network analysis, a gene set positively correlated with M2 macrophage infiltration was screened.
[0022] It should be noted that HCC samples were collected first. The inclusion criteria for samples were as follows: (1) patients diagnosed with HCC by histopathology; (2) patients with a follow-up period of more than 3 months and complete follow-up information. The exclusion criteria were as follows: (1) HCC samples from repeated patients; (2) samples from patients with metastatic HCC; (3) samples from patients who received anti-HCC treatment before surgery; (4) samples with missing clinical data; and (5) samples with missing RNA sequencing data.
[0023] In this example, in order to establish the target prognostic prediction model, four RNA-seq datasets containing large sample sizes of HCC patients were collected, including the TCGA dataset (n = 329), the ICGC dataset (n = 228), the GSE14520 dataset (n = 218), and the GSE76427 dataset (n = 90), where n represents the number of samples.
[0024] The step of using the CIBERSORT tool in combination with weighted gene co-expression network analysis to screen a gene set positively correlated with M2 macrophage infiltration includes: The whole RNA sequencing data in the TCGA dataset was fed into the CIBERSORT tool to generate an immune cell infiltration expression matrix for each HCC sample in the corresponding dataset; Weighted gene co-expression network analysis was used to analyze the immune cell infiltration expression matrix of each HCC sample in the dataset; The neighborhood relationships were converted into a topological overlap matrix, clustered into a chain hierarchy according to the average values of different topological overlap matrix metrics, and the gene sets positively correlated with M2 macrophage infiltration were identified.
[0025] Specifically, CIBERSORT is a tool for deconvolution of the expression matrix of human immune cell subtypes based on the principle of linear support vector regression. It provides a set of gene expression features of 22 immune cell subtypes by default, including plasma cells, naive B cells, CD8+T cells, M0 macrophages, M2 macrophages, dormant natural killer cells, activated natural killer cells, activated CD4+T cells, dormant CD4+T cells, M1 macrophages, memory B cells, activated mast cells, dormant mast cells, neutrophils, naive CD4+T cells, activated dendritic cells, dormant dendritic cells, gamma-delta T cells, regulatory T cells, eosinophils, monocytes, and follicular helper T cells.
[0026] Weighted correlation network analysis (WGCNA) is a systematic statistical method used to describe the gene associations between different samples. It is often used to identify gene sets with highly coordinated changes and can screen potential disease biomarkers and therapeutic targets based on the internal connections of gene sets and the associations between gene sets and selected phenotypes.
[0027] To more accurately describe the correlation between gene sets and phenotypes, a beta value of 0.9 was set during network construction, and samples with an average expression level >0.5 were selected for analysis. Neighborhood relationships were then converted into a topological overlap matrix (TOM), and clustered into a chain hierarchy based on the average values of different TOM metrics. Genes in the module that were significantly positively correlated with M2 macrophages were defined as M2 macrophage-associated genes.
[0028] Step S02: determining target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis.
[0029] Specifically, transcriptome expression differential analysis was performed to determine the first differentially expressed genes between HCC tissues and normal liver tissues in the TCGA dataset, using a threshold of |LogFC| ≥ 1.5, where |LogFC| is the absolute value of the logarithmic fold change and FC is the fold change, calculated by dividing the average expression level of the gene in the experimental group by the average expression level of the gene in the control group. searching for the first differentially expressed gene and the same member of M2 macrophages, and determining the second differentially expressed gene between hepatocellular carcinoma tissue and normal liver tissue; Cox regression analysis was used to compare the correlation between the expression of the second differentially expressed genes and the overall survival time of hepatocellular carcinoma users, and the target genes were determined. The target genes were VASP, TRMU, TRIM28, TIMM50, TBCB, SUPT5H, SNRPA, SNAI2, SMYD3, REEP4, RECQL4, RANBP1, PTK7, PSMD8, PRMT1, PPAN, PAF1, LAPTM4B, HMGA1, GRWD1, ERCC2, EIF4EBP1, CREB3L1, CPNE1, CD320, BCAT1, AKR1B1 and ACTN1.
[0030] Step S03: obtaining a data set containing the target gene, and using a comprehensive machine learning algorithm to generate corresponding prognosis prediction models in each data set.
[0031] The target genes and their corresponding patient survival data from the TCGA, ICGC, GSE14520, and GSE76427 datasets were incorporated into the "Comprehensive Machine Learning Algorithm" module. Specifically, 30 prognostic prediction models were generated for each of the four datasets using the "Comprehensive Machine Learning Algorithm" module. For example, the prediction model corresponding to the StepCox[backward]+SuperPC method had the highest average C-index among the four datasets, at 0.64. This indicates that the target prognostic prediction model could be the prediction model corresponding to the StepCox[backward]+SuperPC method.
[0032] It can be understood that the comprehensive machine learning algorithms include 8 single algorithms including SurvivalSVM, SuperPC, Enet, StepCox, Ridge regression, plsRcox, RSF and GBM, and 22 joint algorithms, specifically: StepCox[backward]+SuperPC, StepCox[both]+survivalSVM, StepCox[backward]+survivalSVM, survivalSVM, SuperPC, StepCox[both]+SuperPC, RSF+survivalSVM, StepCox[backward]+Enet[alpha=0.1], Enet[alpha=0.1], StepCox[both]+Ridge, StepCox[backward]+Ridge, RSF+StepCox[forward] , RSF+plsRcox, RSF+SuperPC, Ridge, RSF+Ridge, StepCox[backward]+Enet[alpha=0.1], StepCox[both], StepCox[backward], RSF, StepCox[forward], StepCox[backward]+RSF, StepCox[both]+plsRcox, StepCox[backward]+plsRcox, RSF+GBM, GBM, StepCox[both]+RSF, StepCox[both]+GBM, plsRcox, StepCox[backward]+GBM, Among them, RSF and StepCox algorithms can be used for feature selection (that is, selecting meaningful and helpful features from all features to avoid having to import all features into the model for training). Therefore, the above-mentioned joint algorithms all include at least one of the RSF and StepCox algorithms.
[0033] Based on the screened target prognostic prediction model, a risk score can be assigned to each HCC user in the four data sets. It should be noted that among the 28 M2 macrophage-related genes, only 12 of them were involved in the construction of the model risk score. The formula for calculating the risk score of each HCC sample is: Risk score = 0.5671879 × SNRPA expression + (-0.4282688) × RECQL4 expression + 0.6410663 × TRIM28 expression + (-0.3376426) × ERCC2 expression + 0.3590617 × GRWD1 expression + 0.4491523 × REEP4 expression + (-1.2190593) × PRMT1 expression + (-0.3341503) ×PAF1 expression + 0.2864963 ×LAPTM4B expression + 0.3805174 ×RANBP1 expression + (-0.3052321) ×BCAT1 expression + 0.4239179 ×PTK7 expression. Next, the optimal cutoff value in the survminer package of R software was used as the cutoff point to divide HCC samples into high-risk and low-risk groups.
[0034] Step S04 , calculating the C-index of each prognosis prediction model under different data sets, and calculating the average value of the C-index, and determining the prognosis prediction model corresponding to the maximum average value of the C-index.
[0035] For example, the C-index calculated by StepCox[backward]+SuperPC in the TCGA dataset is 0.683, the C-index calculated in the ICGC dataset is 0.692, the C-index calculated in the GSE14520 dataset is 0.530, and the C-index calculated in the GSE76427 dataset is 0.644; the C-index calculated by StepCox[both]+survivalSVM in the TCGA dataset is 0.671, and the C-index calculated in the ICGC dataset is 0.692. ex is 0.694, the C-index calculated in the GSE14520 dataset is 0.544, and the C-index calculated in the GSE76427 dataset is 0.639; the C-index calculated by StepCox[backward]+survivalSVM in the TCGA dataset is 0.671, the C-index calculated in the ICGC dataset is 0.694, the C-index calculated in the GSE14520 dataset is 0.544, and the C-index calculated in the GSE76427 dataset is 0.639.
[0036] It can be found that the average C-index of the StepCox[backward]+SuperPC combined algorithm, the StepCox[both]+survivalSVM combined algorithm, and the StepCox[backward]+survivalSVM combined algorithm are all 0.64, which is the prognostic prediction model corresponding to the maximum average C-index.
[0037] Step S05 , determining whether the prognosis prediction model corresponding to the maximum average value of the C-index is unique, if so, executing step S06 , if not, executing step S07 .
[0038] In step S06 , the prognosis prediction model corresponding to the maximum average value of the C-index is determined as the target prognosis prediction model.
[0039] In step S07 , the complexity of the prognosis prediction model corresponding to the maximum average value of the C-index is calculated according to the C-index of each prognosis prediction model under different data sets, and the prognosis prediction model corresponding to the minimum complexity is determined as the target prognosis prediction model.
[0040] Specifically, the likelihood function of the prognostic prediction model corresponding to the maximum average value of the C-index is determined. The likelihood function includes the parameter vector of the prognostic prediction model. It can be understood that different models have different likelihood function forms. For example, for a linear regression model, assuming that the error obeys a normal distribution, its likelihood function can be expressed as: ; Where θ is the parameter vector of the model, y=(y 1 ,y 2 ,...,y n ) is the observed response variable vector, X is the independent variable matrix, The first i observations, σ 2 is the variance of the error, n is the sample size, y i For the i For other models, such as logistic regression and survival analysis models, there are also corresponding likelihood function expressions, which can usually be derived based on the probability distribution assumptions of the model.
[0041] Take the logarithm of the likelihood function to obtain the log-likelihood function, then take the derivative of the parameter vector and set the derivative to zero, solve the parameter estimate that maximizes the log-likelihood function, and obtain the maximum likelihood estimate. For example, the log-likelihood function for the above linear regression model is: ; Solve to get parameter estimates Then, substitute it into the log-likelihood function to obtain the maximum log-likelihood value .
[0042] Determine the number of parameters of the prognostic prediction model corresponding to the maximum average value of C-index, and calculate the complexity based on the number of parameters and the maximum likelihood estimate. It should be noted that the number of parameters is recorded as k, It is understandable that the number of parameters is related to the data set that needs to be acted on. The larger the data set, the more corresponding parameters there are. For example, in the linear regression model In the number of parameters k = p +1, is the pth independent variable, is the regression coefficient corresponding to the pth independent variable, is a random error term. For more complex models, such as those containing interaction terms, polynomial terms, or models with hierarchical structures, the number of independent parameters needs to be carefully calculated. Furthermore, taking the linear regression model as an example, according to the formula Computational complexity, where 2 k is a penalty for model complexity, It reflects the degree of fit of the model to the data. AIC indicates complexity. The smaller the complexity, the better the balance between goodness of fit and complexity, and the better the model performance. Rank the C-index of each prognostic prediction model under different data sets according to the numerical value from large to small. According to the ranking results and the importance of the data set, determine the adjustment factor corresponding to the C-index of each prognostic prediction model under different data sets. Specifically, a mapping relationship between the ranking, the importance of the data set, and the adjustment factor is established in advance, that is, the combination of ranking A and importance B corresponds to the adjustment factor C. It should be noted that the importance of the data set can be determined by expert evaluation; According to the adjustment factor and the complexity, the target complexity is calculated, where there are four data sets, that is, there are four C-index rankings, and there are four corresponding adjustment factors. The adjustment factors are summed and multiplied by the complexity to obtain the target complexity. Finally, the prognosis prediction model corresponding to the minimum complexity value is determined as the target prognosis prediction model.
[0043] In summary, the hepatocellular carcinoma data processing method in the above embodiment of the present invention uses the CIBERSORT tool combined with weighted gene co-expression network analysis to screen a gene set that is positively correlated with M2 macrophage infiltration; through differential expression analysis and prognostic correlation analysis, target genes with prognostic value are determined from the gene set; a data set containing the target gene is obtained, and a comprehensive machine learning algorithm is used to generate corresponding prognostic prediction models in each data set; the C-index of the data set in the corresponding prognostic prediction model is calculated, and the target prognostic prediction model is screened based on the C-index and the complexity of the prognostic prediction model. The target prognostic prediction model finally screened is compared with the traditional modeling of the algorithm selected at the beginning of the study, and the prediction results reflected by the target prognostic prediction model are more objective and true. At the same time, the prediction accuracy of the target prognostic prediction model is improved.
[0044] Example 2 See also Figure 2 , Figure 2 2 is a block diagram of a hepatocellular carcinoma data processing system provided in a second embodiment of the present invention. The hepatocellular carcinoma data processing system 200 includes: a screening module 21, a first determination module 22, a generation module 23, a calculation module 24, a judgment module 25, a second determination module 26, and a third determination module 27, wherein: Screening module 21, for screening gene sets positively correlated with M2 macrophage infiltration using the CIBERSORT tool combined with weighted gene co-expression network analysis; A first determination module 22 is configured to determine target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis; a generation module 23 for obtaining a dataset containing the target gene, and using a comprehensive machine learning algorithm to generate corresponding prognostic prediction models in each dataset, wherein the dataset is an RNA-seq dataset of a hepatocellular carcinoma user, including a TCGA dataset, an ICGC dataset, a GSE14520 dataset, and a GSE76427 dataset. The comprehensive machine learning algorithm includes 8 single algorithms and 22 combined algorithms, including SurvivalSVM, SuperPC, Enet, StepCox, Ridge regression, plsRcox, RSF, and GBM; The calculation module 24 is used to calculate the C-index of each prognostic prediction model under different data sets, calculate the average value of the C-index, and determine the prognostic prediction model corresponding to the maximum average value of the C-index; A judgment module 25 is used to judge whether the prognosis prediction model corresponding to the maximum average value of the C-index is unique; A second determining module 26 is configured to determine the prognosis prediction model corresponding to the maximum average value of the C-index as a target prognosis prediction model if it is determined that the prognosis prediction model corresponding to the maximum average value of the C-index is unique; The third determination module 27 is used to calculate the complexity of the prognosis prediction model corresponding to the maximum average value of the C-index according to the C-index of each prognosis prediction model under different data sets if it is determined that the prognosis prediction model corresponding to the maximum average value of the C-index is not unique, and determine the prognosis prediction model corresponding to the minimum complexity as the target prognosis prediction model.
[0045] Furthermore, in other embodiments of the present invention, the screening module 21 includes: A generation unit is used to bring the overall RNA sequencing data in the TCGA dataset into the CIBERSORT tool to generate the immune cell infiltration expression matrix for each hepatocellular carcinoma sample in the corresponding dataset; An analysis unit is used to analyze the immune cell infiltration expression matrix of each HCC sample in the dataset using weighted gene co-expression network analysis, wherein the independence β value is set to 0.9, and HCC samples with an average expression level > 0.5 are selected for analysis; The conversion unit is used to convert the neighborhood relationship into a topological overlap matrix, cluster it into a chain hierarchy according to the average value of different topological overlap matrix metrics, and determine the gene set that is positively correlated with M2 macrophage infiltration.
[0046] Furthermore, in other embodiments of the present invention, the first determining module 22 includes: A first determining unit is configured to determine the first differentially expressed gene between hepatocellular carcinoma tissue and normal liver tissue in the TCGA dataset by transcriptome expression difference analysis, with |LogFC| ≥ 1.5 as a threshold, wherein |LogFC| is the absolute value of the logarithmic fold change; a second determining unit, configured to search for the first differentially expressed gene and the same member of M2 macrophages, and determine the second differentially expressed gene between the hepatocellular carcinoma tissue and the normal liver tissue; The third determining unit is configured to use a Cox regression analysis method to compare the correlation between the expression of the second differentially expressed gene and the overall survival time of hepatocellular carcinoma patients, and determine the target gene.
[0047] Furthermore, in other embodiments of the present invention, the third determining module 27 includes: A first determining unit is configured to determine a likelihood function of a prognosis prediction model corresponding to a maximum average value of the C-index, wherein the likelihood function includes a parameter vector of the prognosis prediction model; A first computing unit is configured to take the logarithm of the likelihood function to obtain a log-likelihood function, then derive the parameter vector and set the derivative to zero, and solve for a parameter estimate that maximizes the log-likelihood function to obtain a maximum likelihood estimate; A second calculation unit is used to determine the number of parameters of the prognosis prediction model corresponding to the maximum average value of the C-index, and calculate the complexity according to the number of parameters and the maximum likelihood estimate; The second determining unit is used to rank the C-index of each prognostic prediction model under different data sets according to the numerical value from large to small, and determine the adjustment factor corresponding to the C-index of each prognostic prediction model under different data sets according to the ranking result and the importance of the data set; The third calculation unit is configured to calculate a target complexity according to the adjustment factor and the complexity.
[0048] Example 3 Another aspect of the present invention provides an electronic device, see Figure 3 , shown is an electronic device in embodiment 3 of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, the hepatocellular carcinoma data processing method as described above is implemented.
[0049] In some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing access restriction programs.
[0050] The memory 20 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, and the like. In some embodiments, the memory 20 may be an internal storage unit of the electronic device, such as the hard disk of the electronic device. In other embodiments, the memory 20 may also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, and the like. Furthermore, the memory 20 may include both an internal storage unit and an external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or is about to be output.
[0051] It should be pointed out that Figure 3 The structure shown does not constitute a limitation to the electronic device. In other embodiments, the electronic device may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0052] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the hepatocellular carcinoma data processing method as described above is implemented.
[0053] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device.
[0054] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.
[0055] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0056] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0057] The above embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for processing hepatocellular carcinoma data, characterized in that: The method comprises: The CIBERSORT tool combined with weighted gene co-expression network analysis was used to screen gene sets positively correlated with M2 macrophage infiltration; Determining target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis; Obtaining a data set containing the target gene, and using a comprehensive machine learning algorithm to generate a corresponding number of prognostic prediction models in each data set; Calculate the C-index of each prognostic prediction model under different data sets, calculate the average value of the C-index, and determine the prognostic prediction model corresponding to the maximum average value of the C-index; Determine whether the prognostic prediction model corresponding to the maximum average value of C-index is unique; If so, the prognostic prediction model corresponding to the maximum average value of the C-index is determined as the target prognostic prediction model; If not, the complexity of the prognostic prediction model corresponding to the maximum average value of the C-index is calculated according to the C-index of each prognostic prediction model under different data sets, and the prognostic prediction model corresponding to the minimum complexity is determined as the target prognostic prediction model.
2. The hepatocellular carcinoma data processing method according to claim 1, characterized in that: The dataset is an RNA-seq dataset for hepatocellular carcinoma users, including the TCGA dataset, ICGC dataset, GSE14520 dataset, and GSE76427 dataset.
3. The hepatocellular carcinoma data processing method according to claim 2, characterized in that: The step of using the CIBERSORT tool in combination with weighted gene co-expression network analysis to screen a gene set positively correlated with M2 macrophage infiltration includes: The whole RNA sequencing data in the TCGA dataset was fed into the CIBERSORT tool to generate an immune cell infiltration expression matrix for each HCC sample in the corresponding dataset; Weighted gene co-expression network analysis was used to analyze the immune cell infiltration expression matrix of each HCC sample in the dataset; The neighborhood relationships were converted into a topological overlap matrix, clustered into a chain hierarchy according to the average values of different topological overlap matrix metrics, and the gene sets positively correlated with M2 macrophage infiltration were identified.
4. The hepatocellular carcinoma data processing method according to claim 3, characterized in that: In the step of analyzing the immune cell infiltration expression matrix of each hepatocellular carcinoma sample in the dataset using weighted gene co-expression network analysis, the independence β value was set to 0.9, and hepatocellular carcinoma samples with an average expression level > 0.5 were selected for analysis.
5. The hepatocellular carcinoma data processing method according to claim 4, characterized in that: The step of determining target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis includes: The first differentially expressed genes between HCC tissues and normal liver tissues in the TCGA dataset were identified by transcriptome expression differential analysis, with a threshold of |LogFC| ≥ 1.5, where |LogFC| is the absolute value of the logarithmic fold change; searching for the first differentially expressed gene and the same member of M2 macrophages, and determining the second differentially expressed gene between hepatocellular carcinoma tissue and normal liver tissue; Cox regression analysis was used to compare the correlation between the expression of the second differentially expressed gene and the overall survival time of hepatocellular carcinoma patients to determine the target gene.
6. The hepatocellular carcinoma data processing method according to claim 5, characterized in that: The comprehensive machine learning algorithms include 8 single algorithms and 22 joint algorithms, including SurvivalSVM, SuperPC, Enet, StepCox, Ridge regression, plsRcox, RSF and GBM.
7. The hepatocellular carcinoma data processing method according to claim 6, characterized in that: The step of calculating the complexity of the prognosis prediction model corresponding to the maximum average value of the C-index of each prognosis prediction model under different data sets includes: Determining a likelihood function of a prognostic prediction model corresponding to a maximum average value of the C-index, wherein the likelihood function includes a parameter vector of the prognostic prediction model; Take the logarithm of the likelihood function to obtain the log-likelihood function, then take the derivative of the parameter vector and set the derivative to zero, solve for the parameter estimate that maximizes the log-likelihood function, and obtain the maximum likelihood estimate; Determining the number of parameters of the prognostic prediction model corresponding to the maximum average value of the C-index, and calculating the complexity based on the number of parameters and the maximum likelihood estimate; Rank the C-index of each prognostic prediction model under different data sets according to the value from large to small. According to the ranking results and the importance of the data set, determine the adjustment factor corresponding to the C-index of each prognostic prediction model under different data sets; A target complexity is obtained by calculation according to the adjustment factor and the complexity.
8. A hepatocellular carcinoma data processing system, characterized in that: For implementing the hepatocellular carcinoma data processing method according to any one of claims 1 to 7, the system comprises: A screening module is used to screen gene sets positively correlated with M2 macrophage infiltration using the CIBERSORT tool combined with weighted gene co-expression network analysis; A first determination module is used to determine target genes with prognostic value from the gene set through differential expression analysis and prognostic correlation analysis; A generation module is used to obtain a data set containing the target gene, and use a comprehensive machine learning algorithm to generate a corresponding number of prognosis prediction models in each data set; A calculation module is used to calculate the C-index of each prognostic prediction model under different data sets, calculate the average value of the C-index, and determine the prognostic prediction model corresponding to the maximum average value of the C-index; A judgment module is used to judge whether the prognosis prediction model corresponding to the maximum average value of C-index is unique; A second determination module is configured to determine the prognosis prediction model corresponding to the maximum average value of the C-index as a target prognosis prediction model if it is determined that the prognosis prediction model corresponding to the maximum average value of the C-index is unique; The third determination module is used to calculate the complexity of the prognosis prediction model corresponding to the maximum average value of the C-index if it is determined that the prognosis prediction model corresponding to the maximum average value of the C-index is not unique, and determine the prognosis prediction model corresponding to the minimum complexity as the target prognosis prediction model.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the hepatocellular carcinoma data processing method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the method for processing hepatocellular carcinoma data according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Construction method of idiopathic pulmonary fibrosis plasma cell characteristic gene prognosis model
CN117497062A
Hepatocellular carcinoma prognosis scoring model construction method and device, equipment and storage medium
CN117936111A
Method for constructing pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS) related lncRNA
CN120072067A
Lung adenocarcinoma prognosis risk prediction model and construction method and application thereof
CN120280151A
Machine learning-based systems and methods for predicting liver cancer recurrence in liver transplant patients
US20240404707A1