Prognosis gene data processing method and system based on pseudouridine modified gene

By obtaining a dataset of pseudouridine-modified genes, performing standardization and visualization, selecting significantly prognostic-related genes using Cox regression analysis, and constructing a genome merging analysis model, the problem of insufficient prediction accuracy for renal clear cell carcinoma was solved, achieving a more accurate prognosis judgment.

CN120656549APending Publication Date: 2025-09-16THE UNIVERSITY OF HONG KONG SHENZHEN HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510778530.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, the prognosis prediction of renal clear cell carcinoma (KIRC) relies on clinical characteristics and molecular markers, which are limited by heterogeneity and lead to insufficient prediction accuracy.

Method used

By obtaining a pseudouridine-modified gene dataset, standardizing and visualizing it, Cox regression analysis was used to select significant prognosis-related genes, and a genome merging analysis model was constructed. After verification, it was used as the target model for prognosis classification.

Benefits of technology

The accuracy of prognosis prediction for renal clear cell carcinoma has been improved. The target model established through pseudouridine-related genes can effectively judge the prognosis and improve the independence and accuracy of the prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656549A_ABST
    Figure CN120656549A_ABST
Patent Text Reader

Abstract

The invention discloses a pseudouridine modified gene-based prognosis gene data processing method and system, and the method comprises the steps: obtaining a pseudouridine modified gene data set, carrying out standardization processing, and carrying out the visualization of the expression quantity of the processed pseudouridine modified synthase to obtain a target visualization map, according to single-factor Cox regression analysis, selecting a significant prognosis related gene from the visualization graph; combining all the significant prognosis related genes, carrying out correlation calculation, and selecting a target gene combination and a regression coefficient from all the significant prognosis related genes; constructing an analysis model according to the regression coefficient of the target gene combination and verifying the analysis model, and taking the analysis model as a target model when the verification is passed; and when to-be-processed target pseudouridine modified gene data is obtained, processing according to the target model to obtain a classification result, and outputting the classification result. According to the method, the target model is constructed, so that the pseudouridine modified gene can be processed, and corresponding information can be conveniently obtained from the pseudouridine modified gene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a method, system, terminal and computer-readable storage medium for processing prognostic gene data based on pseudouridine-modified genes. Background Art

[0002] Renal clear cell carcinoma (KIRC) is the most common subtype of renal cancer, with complex pathological features and molecular mechanisms and a poor prognosis. Currently, prognosis prediction for KIRC mainly relies on clinical features and some molecular markers.

[0003] However, due to the heterogeneity of KIRC, these traditional methods have limited predictive accuracy and cannot obtain accurate prognosis.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a prognostic gene data processing method, system, terminal and computer-readable storage medium based on pseudouridine modified genes, aiming to solve the problem in the prior art that due to the heterogeneity of KIRC, these traditional methods have limited prediction accuracy and cannot obtain accurate prognosis.

[0006] To achieve the above object, the present invention provides a method for processing prognostic gene data based on pseudouridine modified genes, the method comprising the following steps:

[0007] Obtaining a pseudouridine modification gene dataset, normalizing the pseudouridine modification gene dataset, and visualizing the expression of the processed pseudouridine modification synthase based on a target tool to obtain a target visualization graph, and selecting significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis;

[0008] All the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and each gene combination is associated with a Cox regression analysis, and a target gene combination and a regression coefficient of each gene in the target gene are selected according to the significance;

[0009] Constructing an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination, and verifying the analysis model. When the verification passes, the analysis model is used as the target model;

[0010] The target pseudouridine modified gene data to be processed is obtained and input into the target model. The target pseudouridine modified gene data to be processed is processed by the target model to obtain a classification result for determining the prognosis.

[0011] Optionally, the step of obtaining a pseudouridine modification gene dataset, normalizing the pseudouridine modification gene dataset, and visualizing the expression of pseudouridine modification synthase based on a target tool to obtain a target visualization graph, and selecting significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis, specifically includes:

[0012] obtaining a pseudouridine modification gene dataset, performing standardization processing on the pseudouridine modification gene dataset, and obtaining a pseudouridine modification synthase expression profile;

[0013] Based on the target tool, the expression profile of pseudouridine modification synthase is visualized to obtain a target visualization map;

[0014] According to the univariate Cox regression analysis, pseudouridine modification synthases with a p-value less than a first threshold are selected in the target visualization graph as the significant prognosis-related genes.

[0015] Optionally, all the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and association calculations are performed on each gene combination according to multivariate Cox regression analysis, and the target gene combination and the regression coefficient of each gene in the target gene are selected according to significance, specifically including:

[0016] According to a preset combination method, all the significant prognosis-related genes are randomly combined to obtain multiple gene combinations;

[0017] The association of prognosis with each gene combination was calculated based on multivariate Cox regression analysis, and the p value was obtained;

[0018] The gene combination with the smallest p-value was selected as the target gene combination, and the regression coefficient of each gene in it was obtained accordingly.

[0019] Optionally, constructing an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination specifically includes:

[0020] constructing an analysis model according to the target gene combination and the regression coefficient of each gene in the target gene combination;

[0021] Wherein, the analysis model is expressed as:

[0022] Risk score = 0.018*DKC1 gene expression - 0.037*RPUSD4 gene expression - 0.075*TRUB2 gene expression + 0.096*PUS1 gene expression;

[0023] Risk score represents the analysis model score, and the DKC1 gene expression level, RPUSD4 gene expression level, TRUB2 gene expression level, and PUS1 gene expression level are genes in the target gene combination.

[0024] Optionally, the verifying the analysis model and using the analysis model as a target model when the verification passes specifically includes:

[0025] obtaining a validation pseudouridine modification gene dataset, and validating the analysis model based on the validation pseudouridine modification gene dataset;

[0026] When the verification is passed, the analysis model is used as the target model.

[0027] In addition, to achieve the above-mentioned object, the present invention further provides a prognostic gene data processing system based on pseudouridine modified genes, wherein the prognostic gene data processing system based on pseudouridine modified genes comprises:

[0028] A significant prognosis-related gene selection module is used to obtain a pseudouridine modification gene dataset, standardize the pseudouridine modification gene dataset, and visualize the expression level of the processed pseudouridine modification synthase based on the target tool to obtain a target visualization graph, and select significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis;

[0029] A target gene combination selection module is used to combine all the significant prognosis-related genes according to a preset combination method to obtain multiple gene combinations, perform association calculations on each gene combination according to multivariate Cox regression analysis, and select target gene combinations and regression coefficients of each gene in the target genes according to significance;

[0030] A target model generation module is used to construct an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination, and to verify the analysis model. When the verification passes, the analysis model is used as the target model;

[0031] The application module is used to obtain the target pseudouridine modified gene data to be processed and input it into the target model, and the target pseudouridine modified gene data to be processed is processed by the target model to obtain a classification result for judging the prognosis.

[0032] Optionally, the significant prognosis-related gene selection module includes:

[0033] a normalization processing unit, configured to obtain a pseudouridine modification gene dataset, perform normalization processing on the pseudouridine modification gene dataset, and obtain a pseudouridine modification synthase expression profile;

[0034] A visualization unit, used for visualizing the expression profile of pseudouridine modification synthase based on a target tool to obtain a target visualization graph;

[0035] A selection unit is used to select pseudouridine modification synthases with a p-value less than a first threshold in the target visualization graph according to a univariate Cox regression analysis as the significant prognosis-related genes.

[0036] Optionally, the target gene combination selection module includes:

[0037] a gene combination generating unit, configured to randomly combine all the significant prognosis-related genes according to a preset combination method to obtain a plurality of gene combinations;

[0038] A calculation unit is used to calculate the association of prognosis for each gene combination based on multivariate Cox regression analysis and obtain a p value;

[0039] The target selection unit is used to select the gene combination with the smallest p-value as the target gene combination and obtain the regression coefficient of each gene in it.

[0040] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a prognostic gene data processing program based on pseudouridine modified genes stored on the memory and runnable on the processor, wherein the prognostic gene data processing program based on pseudouridine modified genes, when executed by the processor, implements the steps of the prognostic gene data processing method based on pseudouridine modified genes as described above.

[0041] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a prognostic gene data processing program based on pseudouridine modified genes, and when the prognostic gene data processing program based on pseudouridine modified genes is executed by a processor, the steps of the prognostic gene data processing method based on pseudouridine modified genes as described above are implemented.

[0042] In the present invention, a pseudouridine modification gene dataset is obtained, the pseudouridine modification gene dataset is standardized, and the expression of the processed pseudouridine modification synthase is visualized based on the target tool to obtain a target visualization graph, and significant prognosis-related genes are selected from the visualization graph according to univariate Cox regression analysis; all the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and each gene combination is associated with each gene combination according to multivariate Cox regression analysis, and the target gene combination and the regression coefficient of each gene in the target gene are selected according to significance; according to the target gene combination and the regression coefficient of each gene in the target gene combination, an analysis model is constructed, and the analysis model is verified. When the verification passes, the analysis model is used as the target model; the target pseudouridine modification gene data to be processed is obtained and input into the target model, and the target pseudouridine modification gene data to be processed is processed by the target model to obtain a classification result for judging the prognosis. The present invention performs prognostic treatment on renal clear cell carcinoma by pseudouridine-related genes, establishes a corresponding target model, and thus a classification result for judging the prognosis can be obtained by the target model, thereby improving the accuracy of the prognosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Flowchart of a preferred embodiment of the method for processing prognostic gene data based on pseudouridine modified genes of the present invention;

[0044] Figure 2 It is a visualization diagram of the expression level of pseudouridine modified genes in tumors and adjacent areas of cancer in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0045] Figure 3 The pseudouridine modified gene is obtained by single factor Cox regression and is associated with prognosis in the prognostic gene data processing method based on the pseudouridine modified gene of the present invention;

[0046] Figure 4 Schematic diagram of the relationship between each combination in the prognostic gene data processing method based on pseudouridine modified genes of the present invention and survival calculated through multivariate Cox regression analysis;

[0047] Figure 5 Schematic diagram of the results of univariate and multivariate Cox regression analysis of pseudouridine-related gene prognostic models and clinical factors in the prognostic gene data processing method based on pseudouridine-modified genes of the present invention;

[0048] Figure 6 Schematic diagram of survival analysis of the pseudouridine-related gene prognostic model in the prognostic gene data processing method based on pseudouridine-modified genes of the present invention;

[0049] Figure 7 Schematic diagram of the effect of the target model in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0050] Figure 8 This is a schematic diagram of the survival analysis of the target model in the prognostic gene data processing method based on pseudouridine modified genes of the present invention on another independent validation set;

[0051] Figure 9 Schematic diagram of the effect of the target model displayed by the ROC curve on another independent validation set in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0052] Figure 10 This is a schematic diagram of calculating differentially expressed genes between two groups with high and low risk scores using DESeq2 in the TC GA-KIRC dataset in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0053] Figure 11 This is a schematic diagram of using cluster profiler to perform GO pathway analysis on highly expressed genes in a group with high risk scores among differentially expressed genes in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0054] Figure 12 This is a schematic diagram of using cluster profiler to perform GO pathway analysis on highly expressed genes in a group with low risk scores among differentially expressed genes in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0055] Figure 13 This is a schematic diagram of using gsekegg to perform KEGG pathway analysis on highly expressed genes in a group with high risk scores among differentially expressed genes in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0056] Figure 14 This is a schematic diagram of using gsekegg to perform KEGG pathway analysis on highly expressed genes in a group with low risk scores among differentially expressed genes in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0057] Figure 15 This is a schematic diagram of comparing immune infiltration scores between high and low risk score groups using the R package IOBR in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0058] Figure 16This is a schematic diagram of the present invention's method for processing prognostic gene data based on pseudouridine-modified genes using the R package IOBR to compare the abundance of different types of immune cells in high and low risk score groups;

[0059] Figure 17 This is a schematic diagram of the comparison of different types of immune cell infiltration between high and low risk score groups using the R package IOBR in the prognostic gene data processing method based on pseudouridine modified genes of the present invention;

[0060] Figure 18 1 is a structural diagram of a preferred embodiment of the prognostic gene data processing system based on pseudouridine modified genes of the present invention;

[0061] Figure 19 1 is a structural diagram of a preferred embodiment of a module for selecting significant prognosis-related genes based on pseudouridine-modified genes according to the present invention;

[0062] Figure 20 1 is a structural diagram of a preferred embodiment of a target gene combination selection module based on pseudouridine modified genes of the present invention;

[0063] Figure 21 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0065] RNA modification, a crucial component of epitranscriptomics, has recently been shown to play a key regulatory role in a variety of biological processes. With the advancement of molecular biology techniques, genes associated with pseudouridine (Ψ) have been shown to play a crucial role in tumor invasion and metastasis. Pseudouridine is the earliest and most abundant modification discovered to date across various RNA types, and is catalyzed by the pseudouridine enzyme family. In mammals, Ψ modification is widely present in rRNA, tRNA, snRNA, and mRNA, and participates in various pathological processes by regulating RNA stability, splicing, translation efficiency, and structural conformation. Clear cell renal cell carcinoma (KIRC) is the most common subtype of renal cancer, characterized by complex pathological features and molecular mechanisms and a poor prognosis. Currently, prognostic prediction for KIRC relies primarily on clinical features and a few molecular markers. However, due to the heterogeneity of KIRC, the predictive accuracy of these traditional methods is limited. Currently, pseudouridine-related genes have not been applied to the prognosis of clear cell renal cell carcinoma, resulting in the inability to effectively utilize pseudouridine-related genes to provide prognostic information.

[0066] In response to one or more of the above problems, the present invention obtains a pseudouridine modification gene dataset, standardizes the pseudouridine modification gene dataset, and visualizes the expression level of the processed pseudouridine modification synthase based on the target tool to obtain a target visualization graph, and selects significant prognosis-related genes from the visualization graph according to the univariate Cox regression analysis; all the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and each gene combination is associated with a multivariate Cox regression analysis. The regression coefficient of each gene in the target gene combination and the target gene is selected according to the significance; an analysis model is constructed based on the target gene combination and the regression coefficient of each gene in the target gene combination, and the analysis model is verified. When the verification passes, the analysis model is used as the target model; the target pseudouridine modification gene data to be processed is obtained and input into the target model, and the target pseudouridine modification gene data to be processed is processed by the target model to obtain a classification result for judging the prognosis.

[0067] The prognostic gene data processing method based on pseudouridine modified genes described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the prognostic gene data processing method based on pseudouridine modified genes includes the following steps:

[0068] Step S10: Obtain a pseudouridine modification gene dataset, standardize the pseudouridine modification gene dataset, and visualize the expression level of the processed pseudouridine modification synthase based on the target tool to obtain a target visualization graph, and select significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis.

[0069] Specifically, in the present invention, the pseudouridine modification gene dataset is a collection of expression matrix data of the KIRC dataset in the TCGA database. By processing the pseudouridine modification gene dataset, genes associated with renal clear cell carcinoma among 13 pseudouridine synthases are selected, among which the corresponding pseudouridine modification synthase can be obtained according to the pseudouridine modification gene dataset.

[0070] Furthermore, the pseudouridine modification gene dataset is obtained, the pseudouridine modification gene dataset is normalized, and the expression level of the pseudouridine modification synthase is visualized based on the target tool to obtain a target visualization graph, and significant prognosis-related genes are selected from the visualization graph according to univariate Cox regression analysis, specifically including:

[0071] obtaining a pseudouridine modification gene dataset, performing standardization processing on the pseudouridine modification gene dataset, and obtaining a pseudouridine modification synthase expression profile;

[0072] Based on the target tool, the expression profile of pseudouridine modification synthase is visualized to obtain a target visualization map;

[0073] According to the univariate Cox regression analysis, pseudouridine modification synthases with a p-value less than a first threshold are selected in the target visualization graph as the significant prognosis-related genes.

[0074] Specifically, after obtaining the pseudouridine modification gene dataset in the present invention, TPM is used to normalize the data to obtain the pseudouridine modification synthase expression profile, i.e., the expression matrix; then, the target tool is used to visualize the expression of the pseudouridine modification synthase expression profile, wherein the target tool is the R package ggplot2. The target visualization diagram obtained is as follows: Figure 2 As shown, the horizontal axis represents the 13 pseudouridine synthases and the vertical axis represents the corresponding gene expression levels. By comparing the mRNA expression of 13 pseudouridine modification synthases between TCGA-KIRC tumor samples and the control group, it was determined whether there was differential expression of pseudouridine-related genes between the tumor and control groups. The results showed that most genes were significantly different.

[0075] In the target visualization diagram, the association between 13 pseudouridine synthases and survival prognosis was analyzed using univariate Cox regression analysis to obtain the corresponding significant prognosis-related genes, which are expressed as the corresponding Figure 3 After using univariate Cox regression analysis, the p-value of the association between each pseudouridine synthase and survival prognosis was obtained, where the first threshold was set at 0.05, and p<0.05 was screened to obtain prognosis-related genes. A total of eight genes were obtained, namely, significantly prognosis-related genes, including: PUS3, DKC1, RPUSD4, TRUB1, RPUSD2, TRUB2, PUSL1, and PUS1.

[0076] Step S20: Combine all the significant prognosis-related genes according to a preset combination method to obtain multiple gene combinations, perform association calculation on each gene combination according to multivariate Cox regression analysis, and select the target gene combination and the regression coefficient of each gene in the target gene according to significance.

[0077] Specifically, in the present invention, different combinations of significant prognosis-related genes are performed to generate corresponding gene combinations, and the correlations with prognosis are calculated through the gene combinations, thereby selecting target gene combinations.

[0078] Furthermore, all the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and each gene combination is associated with a multivariate Cox regression analysis. The target gene combination and the regression coefficient of each gene in the target gene are selected according to the significance, specifically including:

[0079] According to a preset combination method, all the significant prognosis-related genes are randomly combined to obtain multiple gene combinations;

[0080] The association of prognosis with each gene combination was calculated based on multivariate Cox regression analysis, and the p value was obtained;

[0081] The gene combination with the smallest p-value was selected as the target gene combination, and the regression coefficient of each gene in it was obtained accordingly.

[0082] Specifically, in the present invention, according to a preset combination mode, all the significant prognosis-related genes are randomly combined to obtain multiple gene combinations, wherein the preset combination mode is a random arrangement of 2-8 gene combinations for eight genes, i.e., 8 random combinations of genes, which can be 2 genes, 3 genes, or 8 genes. According to each combination, the relationship between the gene expression and survival is calculated by multivariate Cox regression analysis. Each combination can obtain a p-value related to survival, and the combination with the smallest p-value, i.e., the combination with the most significant survival prognosis, is screened to obtain the corresponding target gene combination and obtain the corresponding regression coefficient.

[0083] like Figure 4 The figure shows the relationship between each combination and survival, as calculated using multivariate Cox regression analysis. The gene combination with the lowest p-value, DKC1, RP USD4, TRUB2, and PUS1, has a p-value of 1.33e-14. Cox regression analysis was used to calculate the regression coefficients of the prognostic target genes. The regression coefficients for each gene were: DKC1: 0.018, RP USD4: -0.037, TRUB2: -0.075, and PUS1: 0.096.

[0084] Step S30: construct an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination, and verify the analysis model. When the verification is passed, the analysis model is used as the target model.

[0085] After obtaining the corresponding target gene combination and the regression coefficient of each gene in the target gene combination, an analysis model can be constructed, and after verification, the analysis model can be applied as a target model.

[0086] Furthermore, constructing an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination specifically includes:

[0087] constructing an analysis model according to the target gene combination and the regression coefficient of each gene in the target gene combination;

[0088] Wherein, the analysis model is expressed as:

[0089] Risk score = 0.018*DKC1 gene expression - 0.037*RPUSD4 gene expression - 0.075*TRUB2 gene expression + 0.096*PUS1 gene expression;

[0090] Risk score represents the analysis model score, and the DKC1 gene expression level, RPUSD4 gene expression level, TRUB2 gene expression level, and PUS1 gene expression level are genes in the target gene combination.

[0091] The obtained analysis model can be used to calculate the analysis model score of each sample.

[0092] Furthermore, the verification of the analysis model, and taking the analysis model as the target model when the verification passes, specifically includes:

[0093] obtaining a validation pseudouridine modification gene dataset, and validating the analysis model based on the validation pseudouridine modification gene dataset;

[0094] When the verification is passed, the analysis model is used as the target model.

[0095] Specifically, in the present invention, after the analysis model is obtained, the analysis model can be verified using a corresponding validation pseudouridine modified gene dataset to determine its performance.

[0096] In the present invention, the source of the pseudouridine modification gene verification dataset and the pseudouridine modification gene dataset may be the same. In one embodiment of the present invention, data in the EBI database may be selected as the pseudouridine modification gene verification dataset.

[0097] The validation of the analysis model in the present invention includes verifying whether the Risk Score is an independent factor for prognosis, verifying the prognostic efficacy of the model after 1, 3 and 5 years, and verifying whether the prognostic scores of different risks are the same.

[0098] Validation was successful when the Risk Score was an independent prognostic factor, the model's prognostic efficacy met the requirements after 1, 3, and 5 years, and there was a significant difference between the high-risk group and the low-risk group.

[0099] like Figure 5The figure shows the use of the analysis model to score samples in the validation pseudouridine modification gene dataset. Univariate Cox regression and multivariate Cox regression were used to perform independence tests on the sample's Risk Score (RS), clinical pathological stage (T), and age. The results showed that the RS score was determined to be an independent factor for prognosis, where Pathologic_t is pathological stage T, Age is age, pvalue is p value, and Hazard ratio is hazard ratio.

[0100] like Figure 6 The horizontal axis represents survival time, and the vertical axis represents overall survival rate. The expression levels of the four genes were used to calculate a risk score for each sample in the TCGA-KIRC cohort. Samples were divided into high-risk and low-risk groups based on the mean risk score. Kaplan-Meier curves were drawn, revealing significant differences between the high-risk and low-risk groups; the high-risk group had a significantly worse prognosis than the low-risk group (p<0.0001).

[0101] like Figure 7 The horizontal axis represents the false positive rate, and the vertical axis represents the true positive rate. This represents the area under the curve (AUC) for the corresponding model. Larger values ​​indicate better model performance. The R packages survival and tim eROC were used to perform time-dependent ROC curves to verify the prognostic efficacy of the model after 1, 3, and 5 years. The AUC values ​​for 1, 3, and 5 years were 0.715, 0.708, and 0.711, respectively, indicating that the model has good predictive value.

[0102] like Figure 8 As shown, the EBI dataset E-MTAB-1980 expression data was used as the validation dataset for pseudouridine modification genes. The Risk score was calculated for each sample, and the samples were divided into high-risk and low-risk groups based on the median RS score. Survival analysis and survival curves were performed for both groups using the R packages survival and survminer. Survival differences between the two groups were analyzed using the log-rank test, which revealed a significant difference in survival between the high-risk and low-risk Risk score groups (p = 0.021).

[0103] like Figure 9 As shown in the figure, the R package survival and timeROC were used to verify the prognostic efficacy of the model after 1, 3 and 5 years using the time-dependent ROC curve, which is specifically reflected in Figure 9 .like Figure 10As shown in the figure, the present invention uses DESeq2 to calculate the differentially expressed genes between the two groups with high and low risk scores in the TCGA-KIRC dataset, where the horizontal axis is the log2 change fold, the vertical axis is the -log10p value, DOWN indicates down-regulation of expression, UP indicates up-regulation of expression, and NOT indicates not significant. Figure 11 As shown in the figure, clusterprofiler was used to perform GO pathway analysis on the results of differential analysis. The GO signal pathways enriched by highly expressed genes in the RS-high group indicate that highly expressed genes in the RS-high group are enriched in GO signal pathways such as ion transport. The horizontal axis represents the ratio of the number of genes, the vertical axis represents the name of the GO signal pathway, the color represents the corrected p-value, and the size of the circle represents the number of genes. Figure 12 As shown in the figure, the results of differential analysis are analyzed using clusterprofiler for GO pathway analysis. The GO signal pathways enriched by highly expressed genes in the RS-low group indicate that highly expressed genes in the RS-low group are enriched in GO signal pathways such as immune response. The horizontal axis represents the ratio of the number of genes, the vertical axis represents the name of the GO signal pathway, the color represents the corrected p-value, and the size of the circle represents the number of genes. Figure 13 As shown in the figure, the KEGG pathway analysis was performed using gsekegg for the results of differential analysis. This figure represents the KEGG signaling pathway corresponding to the genes highly expressed in the RS-high group. Figure 14 As shown in the figure, the KEGG pathway analysis was performed using gsekegg for the results of differential analysis. This figure represents the KEGG signaling pathways corresponding to the genes highly expressed in the RS-low group, indicating that the highly expressed genes in the RS-low group are enriched in KEGG signaling pathways such as IL-17. Figure 15 As shown in the figure, the immune infiltration scores of the RS-high and RS-low groups were compared using the R package IOBR. The results showed that the RiskScore high group had higher immune cell infiltration; the horizontal axis represents high or low Risk score, and the vertical axis compares Stromal Score (stromal in tumor tissue), Immune Score (immune score) and ESTIMATE Score (the sum of the above two scores). Figure 16 As shown in the figure, the abundance of different types of immune cells is compared between RS-high and RS-low using the R package IOBR. The comparison is based on the difference in the proportion of various cells, such as primitive B cells, memory B cells, etc. Figure 17 As shown, the R package IOBR is used to compare the infiltration of different types of immune cells in RS-high and RS-low, among which the comparison is of large cells, B cells, CD4T, CD8T, neutrophils, macrophages and DC cells.

[0104] Step S40: Obtain target pseudouridine modified gene data to be processed and input it into the target model; process the target pseudouridine modified gene data to be processed by the target model to obtain a classification result for determining the prognosis.

[0105] Specifically, when the target pseudouridine modified gene data to be processed is obtained, the corresponding score, i.e., the classification result, can be output through the target model, and the target pseudouridine modified gene data to be processed is the pseudouridine modified gene data of patients with renal clear cell carcinoma.

[0106] The present invention obtains a pseudouridine modification gene dataset, standardizes the pseudouridine modification gene dataset, and visualizes the expression of the processed pseudouridine modification synthase based on a target tool to obtain a target visualization graph, and selects significant prognosis-related genes from the visualization graph according to univariate Cox regression analysis; all the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and each gene combination is associated with each gene combination according to multivariate Cox regression analysis. The regression coefficient of each gene in the target gene combination and the target gene is selected according to significance; an analysis model is constructed based on the target gene combination and the regression coefficient of each gene in the target gene combination, and the analysis model is verified. When the verification passes, the analysis model is used as the target model; the target pseudouridine modification gene data to be processed is obtained and input into the target model, and the target pseudouridine modification gene data to be processed is processed by the target model to obtain a classification result for judging the prognosis. The present invention performs prognostic treatment on renal clear cell carcinoma by pseudouridine-related genes, establishes a corresponding target model, and thus a classification result for judging the prognosis can be obtained by the target model, thereby improving the accuracy of the prognosis.

[0107] Furthermore, if Figure 18 As shown, based on the above-mentioned prognostic gene data processing method based on pseudouridine modified genes, the present invention also provides a prognostic gene data processing system based on pseudouridine modified genes, wherein the prognostic gene data processing system based on pseudouridine modified genes includes:

[0108] A significant prognosis-related gene selection module 91 is used to obtain a pseudouridine modification gene dataset, perform standardization on the pseudouridine modification gene dataset, and visualize the expression level of the processed pseudouridine modification synthase based on a target tool to obtain a target visualization graph, and select significant prognosis-related genes from the visualization graph based on a univariate Cox regression analysis;

[0109] A target gene combination selection module 92 is configured to combine all of the significant prognosis-related genes according to a preset combination method to obtain multiple gene combinations, perform association calculations on each gene combination according to multivariate Cox regression analysis, and select target gene combinations and regression coefficients of each gene in the target genes according to significance;

[0110] A target model generation module 93 is used to construct an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination, and to verify the analysis model. When the verification passes, the analysis model is used as the target model;

[0111] The application module 94 is used to obtain the target pseudouridine modified gene data to be processed and input it into the target model, and process the target pseudouridine modified gene data to be processed by the target model to obtain a classification result for determining the prognosis.

[0112] like Figure 19 As shown, the significant prognosis-related gene selection module includes:

[0113] A normalization processing unit 911 is used to obtain a pseudouridine modification gene dataset, perform normalization processing on the pseudouridine modification gene dataset, and obtain a pseudouridine modification synthase expression profile;

[0114] A visualization unit 912 is used to visualize the expression profile of pseudouridine modification synthase based on the target tool to obtain a target visualization graph;

[0115] The selection unit 913 is used to select pseudouridine modification synthases with a P value less than a first threshold in the target visualization graph according to the univariate Cox regression analysis as the significant prognosis-related genes.

[0116] like Figure 20 As shown, the target gene combination selection module includes:

[0117] A gene combination generating unit 921 is configured to randomly combine all the significant prognosis-related genes according to a preset combination method to obtain a plurality of gene combinations;

[0118] a calculation unit 922 for performing a prognostic correlation calculation on each gene combination according to a multivariate Cox regression analysis to obtain a p-value;

[0119] The target selection unit 923 is used to select the gene combination with the smallest p-value as the target gene combination, and obtain the regression coefficient of each gene therein.

[0120] Furthermore, if Figure 21As shown, based on the above-mentioned prognostic gene data processing method and system based on pseudouridine modified genes, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 21 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0121] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is about to be output. In one embodiment, a prognostic gene data processing program 40 based on pseudouridine modified genes is stored on the memory 20, and the prognostic gene data processing program 40 based on pseudouridine modified genes can be executed by the processor 10, thereby realizing the prognostic gene data processing method based on pseudouridine modified genes in the present invention.

[0122] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes or process data stored in the memory 20, such as executing the prognostic gene data processing method based on pseudouridine modified genes.

[0123] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0124] In one embodiment, when the processor 10 executes the pseudouridine-modified gene-based prognostic gene data processing program 40 in the memory 20 , the steps of the above pseudouridine-modified gene-based prognostic gene data processing method are implemented.

[0125] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a prognostic gene data processing program based on pseudouridine modified genes, and when the prognostic gene data processing program based on pseudouridine modified genes is executed by a processor, the steps of the prognostic gene data processing method based on pseudouridine modified genes as described above are implemented.

[0126] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0127] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0128] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for processing prognostic gene data based on pseudouridine modified genes, characterized in that: The prognostic gene data processing method based on pseudouridine modified genes includes: Obtaining a pseudouridine modification gene dataset, normalizing the pseudouridine modification gene dataset, and visualizing the expression of the processed pseudouridine modification synthase based on a target tool to obtain a target visualization graph, and selecting significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis; All the significant prognosis-related genes are combined according to a preset combination method to obtain multiple gene combinations, and each gene combination is associated with a Cox regression analysis, and a target gene combination and a regression coefficient of each gene in the target gene are selected according to the significance; Constructing an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination, and verifying the analysis model. When the verification passes, the analysis model is used as the target model; The target pseudouridine modified gene data to be processed is obtained and input into the target model. The target pseudouridine modified gene data to be processed is processed by the target model to obtain a classification result for determining the prognosis.

2. The method for processing prognostic gene data based on pseudouridine modified genes according to claim 1, characterized in that: The method comprises obtaining a pseudouridine modification gene dataset, performing standardization on the pseudouridine modification gene dataset, and visualizing the expression of pseudouridine modification synthase based on a target tool to obtain a target visualization graph, and selecting significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis, specifically comprising: obtaining a pseudouridine modification gene dataset, performing standardization processing on the pseudouridine modification gene dataset, and obtaining a pseudouridine modification synthase expression profile; Based on the target tool, the expression profile of pseudouridine modification synthase is visualized to obtain a target visualization map; According to the univariate Cox regression analysis, pseudouridine modification synthases with a p-value less than a first threshold are selected in the target visualization graph as the significant prognosis-related genes.

3. The method for processing prognostic gene data based on pseudouridine modified genes according to claim 1, characterized in that: The method combines all the significant prognosis-related genes according to a preset combination method to obtain multiple gene combinations, performs association calculation on each gene combination according to multivariate Cox regression analysis, and selects the target gene combination and the regression coefficient of each gene in the target gene according to significance, specifically including: According to a preset combination method, all the significant prognosis-related genes are randomly combined to obtain multiple gene combinations; The association of prognosis with each gene combination was calculated based on multivariate Cox regression analysis, and the p value was obtained; The gene combination with the smallest p-value was selected as the target gene combination, and the regression coefficient of each gene in it was obtained accordingly.

4. The method for processing prognostic gene data based on pseudouridine modified genes according to claim 1, characterized in that: The analysis model is constructed based on the target gene combination and the regression coefficient of each gene in the target gene combination, specifically comprising: constructing an analysis model according to the target gene combination and the regression coefficient of each gene in the target gene combination; Wherein, the analysis model is expressed as: Risk score = 0.018*DKC1 gene expression - 0.037*RPUSD4 gene expression - 0.075*TRUB2 gene expression + 0.096*PUS1 gene expression; Risk score represents the analysis model score, and the DKC1 gene expression level, RPUSD4 gene expression level, TRUB2 gene expression level, and PUS1 gene expression level are genes in the target gene combination.

5. The method for processing prognostic gene data based on pseudouridine modified genes according to claim 1, characterized in that: The verification of the analysis model, and taking the analysis model as the target model when the verification passes, specifically includes: obtaining a validation pseudouridine modification gene dataset, and validating the analysis model based on the validation pseudouridine modification gene dataset; When the verification is passed, the analysis model is used as the target model.

6. A prognostic gene data processing system based on pseudouridine modified genes, characterized in that: The prognostic gene data processing system based on pseudouridine modified genes includes: A significant prognosis-related gene selection module is used to obtain a pseudouridine modification gene dataset, standardize the pseudouridine modification gene dataset, and visualize the expression level of the processed pseudouridine modification synthase based on the target tool to obtain a target visualization graph, and select significant prognosis-related genes from the visualization graph based on univariate Cox regression analysis; A target gene combination selection module is used to combine all the significant prognosis-related genes according to a preset combination method to obtain multiple gene combinations, perform association calculations on each gene combination according to multivariate Cox regression analysis, and select target gene combinations and regression coefficients of each gene in the target genes according to significance; A target model generation module is used to construct an analysis model based on the target gene combination and the regression coefficient of each gene in the target gene combination, and to verify the analysis model. When the verification passes, the analysis model is used as the target model; The application module is used to obtain the target pseudouridine modified gene data to be processed and input it into the target model, and the target pseudouridine modified gene data to be processed is processed by the target model to obtain a classification result for judging the prognosis.

7. The prognostic gene data processing system based on pseudouridine modified genes according to claim 6, characterized in that: The significant prognosis-related gene selection module includes: a normalization processing unit, configured to obtain a pseudouridine modification gene dataset, perform normalization processing on the pseudouridine modification gene dataset, and obtain a pseudouridine modification synthase expression profile; A visualization unit, used for visualizing the expression profile of pseudouridine modification synthase based on a target tool to obtain a target visualization graph; A selection unit is used to select pseudouridine modification synthases with a p-value less than a first threshold in the target visualization graph according to univariate Cox regression analysis as the significant prognosis-related genes.

8. The prognostic gene data processing system based on pseudouridine modified genes according to claim 6, characterized in that: The target gene combination selection module includes: a gene combination generating unit, configured to randomly combine all the significant prognosis-related genes according to a preset combination method to obtain a plurality of gene combinations; A calculation unit is used to calculate the association of prognosis for each gene combination based on multivariate Cox regression analysis to obtain a p value; The target selection unit is used to select the gene combination with the smallest p-value as the target gene combination and obtain the regression coefficient of each gene in it.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a prognostic gene data processing program based on pseudouridine modified genes stored in the memory and runnable on the processor. When the prognostic gene data processing program based on pseudouridine modified genes is executed by the processor, the steps of the prognostic gene data processing method based on pseudouridine modified genes as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a prognostic gene data processing program based on pseudouridine modified genes, and when the prognostic gene data processing program based on pseudouridine modified genes is executed by a processor, the steps of the prognostic gene data processing method based on pseudouridine modified genes as described in any one of claims 1 to 7 are implemented.