A Bladder Cancer Intelligent Prediction System and Biomarker Combination Based on Seven Gene Expression Characteristics
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-14
AI Technical Summary
[0007]本发明的目的在于提供一种基于七基因表达特征的膀胱癌预测系统及生物标志物组合,以解决现有技术中膀胱癌早期症状缺乏特异性导致最佳干预时机延误、单一标志物检测方法灵敏度和特异性有限、现有诊断模型对复杂生物学数据间潜在关联的挖掘能力不足、以及缺乏便捷智能化临床应用软件等问题
(1)通过基于单细胞转录组测序数据的系统性筛选策略,发现并验证了在膀胱癌中显著升高的七个关键基因标志物,相比单一标志物检测方法,七基因组合能够整合多维度生物学特征,显著提高膀胱癌诊断的灵敏度和特异性;在保留测试集验证中,灵敏度达到0.9012,特异性达到1.0000,AUC达到0.9835,诊断性能优异;
Smart Images

Figure CN122392651B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical and artificial intelligence-assisted diagnostic technology, and specifically relates to a bladder cancer intelligent prediction system, method, application, and biomarker combination based on the expression characteristics of seven genes. Background Technology
[0002] Bladder cancer is one of the most common malignant tumors worldwide, and its incidence has been rising steadily in recent years. It is characterized by high recurrence rates, high treatment costs, and significant differences in prognosis, seriously threatening human life and health and increasing the burden on healthcare systems. Because early clinical symptoms of bladder cancer lack specificity, patients are often diagnosed only when the disease has progressed, leading to delays in optimal intervention. Therefore, a thorough understanding of the key molecular mechanisms underlying the development and progression of bladder cancer is crucial for early screening, precision prevention, and personalized treatment.
[0003] Current research indicates that the development of bladder cancer is the result of multiple factors, including genetic factors, environmental exposure, and abnormal tumor microenvironment, involving complex biological processes such as abnormal cell proliferation, inflammatory response, immune escape, metabolic reprogramming, and epigenetic regulation. However, the core driving events and upstream regulatory networks of bladder cancer development are not yet fully elucidated, and significant molecular heterogeneity exists among different patients. Traditional diagnostic methods for bladder cancer mainly rely on cystoscopy and urine cytology. However, cystoscopy is an invasive procedure with poor patient compliance; while urine cytology is non-invasive, its sensitivity is low, especially for early or low-grade bladder cancer.
[0004] In recent years, biomarker-based methods for bladder cancer detection have gained increasing attention. Current technologies sometimes employ single molecular markers for bladder cancer diagnosis, such as detecting the expression levels of specific proteins or nucleic acids in urine. However, the sensitivity, specificity, and stability of single-marker detection methods are often limited, making it difficult to meet the needs of precise clinical diagnosis. Furthermore, due to the high heterogeneity of bladder cancer, a single marker cannot comprehensively reflect the biological characteristics of the tumor, resulting in low diagnostic accuracy.
[0005] With the development of bioinformatics and machine learning technologies, constructing diagnostic models by combining multiple biomarkers has shown promising application prospects in the early screening and auxiliary diagnosis of various tumors. This strategy can integrate multi-dimensional biological characteristics and improve the sensitivity, specificity, and stability of traditional single-marker detection. However, most existing bladder cancer diagnostic models are based on traditional statistical methods or simple machine learning algorithms, which have limited ability to uncover potential associations between complex biological data and lack systematic biomarker screening strategies. The generalization ability and stability of these models still need to be improved.
[0006] Furthermore, most existing bladder cancer diagnostic tools remain in the research stage, lacking convenient and intelligent clinical application software. Current diagnostic methods often rely on specialized laboratory equipment and complex data analysis processes, resulting in high operational barriers and limiting their widespread application in clinical practice. Therefore, there is an urgent need to develop an intelligent bladder cancer auxiliary diagnostic system based on multi-marker joint detection, combined with advanced machine learning algorithms, and featuring a user-friendly interface, to achieve early screening, accurate diagnosis, and risk assessment of bladder cancer. Summary of the Invention
[0007] The purpose of this invention is to provide a bladder cancer prediction system and biomarker combination based on the expression characteristics of seven genes, in order to solve the problems in the prior art, such as the lack of specificity of early symptoms of bladder cancer leading to delays in the best intervention time, the limited sensitivity and specificity of single biomarker detection methods, the insufficient ability of existing diagnostic models to explore potential associations between complex biological data, and the lack of convenient and intelligent clinical application software.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: A smart bladder cancer prediction system includes a processor and a memory, and further includes a data acquisition module, a preprocessing module, a model inference module, and a result output module. The data acquisition module acquires the expression values of seven genes from a sample to be tested, including TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1. The preprocessing module is communicatively connected to the data acquisition module and is used to standardize the expression values of the seven genes using a Z-score standardization method. The model inference module is communicatively connected to the preprocessing module and includes a pre-trained multilayer perceptron model used to calculate the probability of bladder cancer occurrence based on the standardized expression values of the seven genes. The result output module is communicatively connected to the model inference module and is used to determine the bladder cancer classification result based on the bladder cancer occurrence probability and a preset threshold obtained through cross-validation.
[0009] Preferably, the data acquisition module obtains the expression values of the seven genes using quantitative real-time PCR or RNA sequencing technology.
[0010] Preferably, the multilayer perceptron model includes a multilayer perceptron (MLP) classifier comprising 7 input nodes, 7 tanh-activated hidden neurons, and 1 sigmoid output node; the training parameters are alpha=0.0001, solver=adam, and max_iter=5000.
[0011] Preferably, the result output module further includes a graphical user interface, which includes an input box for inputting the expression values of the seven genes, a prediction button for triggering prediction, and an output area for displaying the prediction probability and classification results.
[0012] Preferably, the preset threshold is determined by hierarchical K-fold cross-validation combined with an out-of-fold threshold calibration strategy, so that the accuracy, sensitivity and specificity of the model are all greater than 0.85.
[0013] A combination of biomarkers capable of diagnosing or predicting bladder cancer, the combination of biomarkers comprising the following genes: TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1.
[0014] Application of TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR and E2F1 genes in the preparation of kits for bladder cancer diagnosis.
[0015] The beneficial effects of this invention are as follows: (1) Through a systematic screening strategy based on single-cell transcriptome sequencing data, seven key gene biomarkers that were significantly elevated in bladder cancer were discovered and validated. Compared with single biomarker detection methods, the combination of seven genes can integrate multi-dimensional biological characteristics and significantly improve the sensitivity and specificity of bladder cancer diagnosis. In the validation of the retained test set, the sensitivity reached 0.9012, the specificity reached 1.0000, and the AUC reached 0.9835, demonstrating excellent diagnostic performance. (2) A multilayer perceptron model combined with standardization is adopted. The nonlinear feature transformation is achieved by activating hidden neurons with 7 tanh, which effectively explores the complex relationship between the expression patterns of seven genes and significantly improves the prediction accuracy compared with the traditional linear model. The model stability and generalization ability are optimized by hierarchical K-fold cross-validation and OOF threshold calibration strategy. The accuracy is 0.8993 and the AUC reaches 0.9554 in repeated nested cross-validation. After OOF threshold calibration, the accuracy is 0.9034 and the AUC is 0.9550. The model shows high stability under different validation strategies. (3) The trained prediction model is packaged into a graphical prediction software based on Tkinter, BC_7genes_Pre.exe. Users only need to input the expression values of seven genes to automatically obtain the probability and classification results of bladder cancer. It is easy to operate and does not rely on professional laboratory equipment and complex data analysis processes, which significantly reduces the threshold for clinical application and facilitates its application in early screening, accurate diagnosis and risk assessment of bladder cancer. (4) The prediction model constructed in this invention has a score greater than 0.85 in all three core indicators of accuracy, sensitivity and specificity, which meets the strict requirements of clinical diagnosis. It provides a new technical solution for early screening and accurate diagnosis of bladder cancer and has good clinical application value and promotion prospects. Attached Figure Description
[0016] Figure 1 Annotation results of different sample sources and cell types in single-cell sequencing data of normal tissue and bladder cancer tissue.
[0017] Figure 2 Further subgroup annotation of epithelial cell subsets and their three-dimensional distribution maps constructed based on the first three principal components (PC1–PC3) of principal component analysis.
[0018] Figure 3 Analysis of the composition and fold change of specific epithelial cell subsets in normal and bladder cancer tissues, and correlation analysis between the subset score and tumor copy number variation.
[0019] Figure 4 Bubble plot of the top 10 genes significantly upregulated in epithelial cell subset 4, and the specificity and activity of transcriptional regulatory factors in different epithelial cell subsets obtained based on pySCENIC analysis.
[0020] Figure 5 In normal tissue and bladder cancer tissue, the effect of Figure 4 The expression of key genes identified in the screening was validated at the mRNA and protein levels. Immunohistochemical data were obtained from the HPA database, and mRNA data were obtained from the TCGA-BLCA and GTEx databases.
[0021] Figure 6 Area under the receiver operating characteristic curve (AUC) for different gene combinations.
[0022] Figure 7 Overall development flowchart of the model.
[0023] Figure 8 The performance evaluation results of this model include evaluation indicators such as accuracy, sensitivity, specificity, and AUC.
[0024] Figure 9 A diagram illustrating the user interface of a software tool.
[0025] Figure 10 The detection efficacy of existing biomarkers with different gene combination schemes.
[0026] Figure 11 ROC curve analysis of existing detection technology biomarkers for single genes and gene combination models. Detailed Implementation
[0027] Specific embodiments are given below with reference to the accompanying drawings. These specific embodiments are only used to describe the technical solution of the present invention in detail, and are not intended to limit the scope of protection of this application.
[0028] Example 1: Screening of seven gene biomarkers based on single-cell transcriptome sequencing This embodiment discloses a method for screening bladder cancer diagnostic biomarkers based on single-cell transcriptome sequencing data, and supports the application of TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR and E2F1 genes in the preparation of kits for bladder cancer diagnosis.
[0029] 1.1 Data Sources and Preprocessing This study collected single-cell transcriptome sequencing data from normal bladder tissue and bladder cancer tissue, including 25 normal bladder tissue samples and 75 bladder cancer tissue samples, sourced from the GEO database. 10× Genomics expression matrices were read from both bladder cancer and normal samples, and Seurat objects were constructed using CreateSeuratObject with parameters min.cells = 3 and min.features = 200.
[0030] The proportions of mitochondrial genes (percent.mt) and erythrocyte-related genes (percent.rbc) were calculated, and cell quality control was performed according to the criteria of nFeature_RNA < 4500, percent.mt < 5, and percent.rbc < 1. After quality control, NormalizeData was used for normalization, FindVariableFeatures(selection.method="vst", nfeatures=3000) was used to screen for hypervariable genes, ScaleData was used for standardization, RunPCA was used for dimensionality reduction, and the RunHarmony function was used to correct for batch effects.
[0031] 1.2 Cell Grouping and Annotation like Figure 1As shown, proximity graphs, clustering, and UMAP visualization were performed based on Harmony correction results, primarily using dims = 1:10 and resolution = 0.1. For cell annotation, classic lineage marker genes were used for manual annotation. Markers used in the first round of annotation included epithelial cell marker genes EPCAM, KRT8, KRT18, and KRT19; immune cell marker genes PTPRC, TRAC, LST1, and MS4A1; stromal cell marker genes COL11 and DCN; and endothelial cell marker genes PECAM1 and VWF. After observing marker expression using DotPlot, FeaturePlot, and VlnPlot, Seurat cluster numbers were mapped to cell types, successfully identifying major cell types such as epithelial cells, immune cells, stromal cells, and endothelial cells.
[0032] 1.3 Analysis of Epithelial Cell Subpopulations like Figure 2 As shown, epithelial cell populations were extracted, and then NormalizeData, FindVariableFeatures, ScaleData, RunPCA, RunHarmony, FindNeighbors, FindClusters, and RunUMAP analyses were performed again. Further re-clustering of epithelial cells was then performed using dims = 1:15 and resolution = 0.2. Multiple epithelial cell clusters were initially obtained, then low-quality clusters or those not included in subsequent analyses were removed, and the remaining clusters were finally annotated into nine epithelial cell subpopulations (epithelial cell subpopulation 1 to epithelial cell subpopulation 9).
[0033] The FindAllMarkers function was used to screen for epithelial cell subset marker genes, with parameters set to only.pos=TRUE, min.pct=0.25, and logfc.threshold=0.25. Based on the principal component analysis (PCA) results, the coordinates of three principal components (PC1, PC2, and PC3) were extracted, and three-dimensional principal component analysis was performed to analyze the distribution characteristics of different epithelial cell subsets in normal tissues and bladder cancer tissues.
[0034] like Figure 3As shown, the analysis revealed that epithelial cell subsets 4, 5, and 7 were significantly enriched in bladder cancer tissues, and their principal component characteristics differed significantly from those of normal tissues. To further evaluate the association between these three subsets and malignant transformation of bladder cancer, epithelial cell subset scores were calculated based on the TCGA expression matrix. Specifically, the gene expression matrix was standardized using gene-wise Z-scores, and then the average Z-score of a specified gene set was calculated as the scores for epithelial cell subsets 4, 5, and 7. Samples were assigned to the corresponding epithelial subgroup based on the maximum score among the three scores, with a maximum score of 0.5 used for screening. The TCGA bladder cancer CNV file was read, and the global CNV load for each sample was calculated, i.e., CNV_burden=colMeans(abs(cnv2), na.rm=TRUE). The correlation between epithelial cell subset scores and CNV load was analyzed using Spearman correlation coefficient. The results showed that only the characteristic score of epithelial cell subset 4 was significantly positively correlated with CNV load. r =0.29, P =0.00011), indicating that the higher the proportion of this subpopulation, the more obvious the genomic instability of the patient, suggesting that epithelial cell subpopulation 4 may play a key role in the malignant transformation of bladder epithelial cells and tumor progression.
[0035] 1.4 Screening for differentially expressed genes like Figure 4 As shown, a detailed analysis of differentially expressed genes in epithelial cell subset 4 was conducted. The FindAllMarkers function was used to screen for significantly elevated genes in this subset, with parameters set to only.pos=TRUE, min.pct=0.25, and logfc.threshold=0.25. The results showed that the top 10 significantly elevated genes in this subset included CDK1 and UBE2C.
[0036] 1.5 Transcriptional Regulatory Network Analysis The pySCENIC workflow was used to analyze transcription factor regulatory activity in single-cell transcriptome data. Epithelial cell Seurat objects and .loom files generated by SCENIC were read, and expression matrices, regulator information, and corresponding regulator activity AUC matrices were extracted. Regulators were converted into corresponding target gene sets using the `regulonsToGeneLists` method, and the AUCell method was used to calculate the activity score of each regulator in different cell types.
[0037] Combining cell type annotation information, the calcRSS method was used to calculate the cell-specific regulatory activity (RSS) of different transcription factors in each epithelial cell subset, and plotRSS was used to draw a heatmap of transcription factor-specific activity. Figure 4 As shown, the network regulated by transcription factor E2F1 exhibits high specificity and activity in epithelial cell subset 4, suggesting that it may be involved in regulating biological processes related to this subset.
[0038] 1.6 Validation of Seven Gene Markers Based on the above analysis, eleven candidate biomarkers were selected: E2F1, TPX2, DLGAP5, CENPF, ASPM, TOP2A, MKI67, CDK1, UBE2C, NUSAP1, and HMMR. Further experimental validation of these eleven genes was then conducted.
[0039] Transcriptional level data of seven genes in bladder cancer and normal tissues were obtained from the TCGA-BLCA and GTEx databases. The bladder cancer transcriptome feature matrix file was read, and the expression data was loaded using the pandas.read_csv function. Expression values for the eleven genes were extracted, with bladder cancer samples marked as 1 and normal bladder tissue samples marked as 0. The Mann–Whitney U test was used to compare the statistical differences in gene expression levels between the bladder cancer group and the normal group.
[0040] like Figure 5 As shown, TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1 all exhibited significantly elevated mRNA expression levels in bladder cancer tissues. Protein level validation data were obtained from immunohistochemical (IHC) staining images from the HPA database, showing that the protein expression levels of these seven genes were also significantly elevated in bladder cancer tissues, consistent with the mRNA expression trend. These biomarkers can stably distinguish between normal tissues and bladder cancer tissues, demonstrating good potential for auxiliary diagnosis and early screening of bladder cancer.
[0041] 1.7 Performance Analysis of Biomarker Prediction Model To verify the necessity of the seven-gene combination, prior analysis was conducted based on the predictive performance of different gene combinations. After assessing the importance of each gene through regression analysis, combination models containing the top six, top five, top three, and top two genes were constructed. The results showed that the area under the receiver operating characteristic (AUC) of the seven-gene combination was significantly higher than that of the combinations of the top six, top five, top three, and top two genes. Figure 6 As shown, the seven-gene combined model has better diagnostic efficacy for bladder cancer.
[0042] Example 2: Construction of a Bladder Cancer Prediction Model Based on Seven Gene Expression Characteristics This embodiment discloses a method for constructing a bladder cancer prediction model based on multilayer perceptron (MLP). 2.1 Data Preparation like Figure 7 As shown, the model input features consist of the expression values of seven genes, including TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1. The transcriptome expression data used in the model were obtained from the TCGA-BLCA and GTEx databases, where sample labels were encoded in a binary classification format, with bladder cancer samples labeled as 1 and normal bladder tissue samples labeled as 0. Based on the constructed dataset, 407 bladder cancer samples and 28 normal samples were ultimately included for model development. In the input data, each row represents a sample, and each column represents the expression value of the corresponding gene, forming a seven-gene expression matrix.
[0043] 2.2 Dataset Partitioning The seven-gene expression matrix was divided into training and test sets. The hold-out method was used for data splitting, using the `train_test_split` function with parameters set to `test_size=0.2` and `random_state=42`. This means that 80% of the samples (348 cases) were used for model training, and 20% of the samples (87 cases) were used to retain the test set for evaluation.
[0044] 2.3 Data Standardization To reduce the impact of differences in gene expression levels at different scales on model training, StandardScaler was used to perform Z-score normalization on the input expression values. The normalization parameters were calculated from the training set and saved and retrieved in subsequent prediction stages to ensure consistency in data processing between the training and actual prediction phases.
[0045] 2.4 MLP Model Construction A multilayer perceptron classifier is used as the core prediction model. For example... Figure 7 As shown, the model's input layer contains 7 nodes, corresponding to the standardized expression values of seven genes: TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1. The hidden layer has 7 neurons and uses the tanh activation function for non-linear feature transformation. The output layer has 1 node and uses the sigmoid function to output a predicted probability between 0 and 1. The model training parameters are set to alpha=0.0001, solver=adam, and max_iter=5000. After training, the model can output the probability value of a sample belonging to bladder cancer based on the expression features of the seven genes.
[0046] 2.5 Classification Threshold Optimization After model training, further optimization of the classification threshold is performed. Instead of simply using the default threshold of 0.5 for bladder cancer diagnosis, the model filters different thresholds based on the model's predicted probabilities, selecting the optimal diagnostic threshold that simultaneously meets the requirements of accuracy, sensitivity, and specificity. The threshold optimization requirement is that accuracy, sensitivity, and specificity are all greater than 0.85. The final determined classification threshold is saved along with the model parameters for automatic retrieval during subsequent software predictions.
[0047] 2.6 Model Validation and Performance Evaluation like Figure 7 As shown, in the model validation and screening phase, a combination of nested cross-validation and OOF (Out-of-Fold) evaluation was used to assess model stability. The inner layer used StratifiedKFold for threshold calibration with parameters n_splits=4 and shuffle=True; the outer layer used StratifiedKFold for full-data OOF evaluation with parameters n_splits=5. Simultaneously, a 5-fold × 5 repeated nested cross-validation approach was employed, recalibrating the threshold in each round of outer layer validation to reduce the risk of overfitting and improve generalization ability.
[0048] like Figure 8 As shown, to evaluate the predictive performance of BC_7genes_Pre for bladder cancer, three methods were used to evaluate the model: repeated nested cross-validation, validation with the test set retained, and OOF threshold calibration. The results show: (1) Repeated nested cross-validation: The model accuracy was 0.8993, sensitivity was 0.9021, specificity was 0.8533, precision was 0.9896, F1 score was 0.9429, and the area under the ROC curve (AUC) reached 0.9554.
[0049] (2) Validation on the test set: The model accuracy was 0.9080, the sensitivity was 0.9012, the specificity was 1.0000, the precision was 1.0000, the F1 score was 0.9481, and the AUC was 0.9835.
[0050] (3) OOF threshold calibration: model accuracy is 0.9034, sensitivity is 0.9042, specificity is 0.8929, precision is 0.9919, F1 score is 0.9460, and AUC is 0.9550. The ROC curve results further demonstrate that BC_7genes_Pre has a high ability to identify bladder cancer. The AUC of the retained test set ROC analysis was 0.9835 (95% CI: 0.9444–1.0000). P<0.0001), the AUC after OOF threshold calibration was 0.9550 (95% CI: 0.9127–0.9882, P The value <0.0001 indicates that the model has high stability and generalization performance under different verification strategies.
[0051] 2.7 Saving Model Files After model training and validation are complete, the core training outputs are retained, including standardized parameters, neural network model parameters, classification threshold parameters, and related result files. The model file is saved in .joblib format, while the threshold and other configuration parameters are saved in .json format. These files together constitute the core parameter files required for BC_7genes_Pre to run.
[0052] 2.8 Performance of the BC_7genes_Pre model This embodiment also evaluated the effectiveness of the BC_7genes_Pre model, comparing its performance with commonly used bladder cancer auxiliary detection biomarkers in the prior art (covering TP53, GATA3, TP63, VEGFA, KRT20, FGFR3, and UCA1 genes). Single-gene and multi-gene (3, 5, 6, and 7 genes) combination analyses were performed on the aforementioned existing genes. Figure 10 As shown, both single-gene and combinations of existing genes exhibit significantly lower diagnostic efficacy than the BC_7genes_Pre model constructed in this invention. Specifically, even when all seven existing genes are combined for model construction, their accuracy, sensitivity, F1 score, and AUC value remain significantly inferior to the model of this invention. Figure 11 As shown, the predictive ability of existing single genes is limited, with the highest AUC (TP53) being only 0.7181; while the AUC of existing 7-gene combinations is only 0.6646, and the AUC of other combinations is all below 0.7, resulting in low overall diagnostic efficacy.
[0053] Example 3: Development and Application of BC_7genes_Pre Prediction Software This embodiment discloses a method for developing and applying BC_7genes_Pre, a bladder cancer prediction software based on Tkinter. 3.1 Software Interface Design like Figure 9 As shown, a BC_7genes_Pre prediction software tool based on the Tkinter interface was constructed. The software interface includes the following functional modules: (1) Input module: It contains 7 text boxes, which are used to input the expression values of seven genes: TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR and E2F1. The corresponding gene name is marked next to each text box for easy identification by the user.
[0054] (2) Function button module: It contains three buttons, namely "Predict", "Clear" and "Fill Sample". The "Predict" button is used to trigger prediction calculation; the "Clear" button is used to clear the values in all input boxes; the "Fill Sample" button is used to automatically fill sample data, which makes it easy for users to quickly test the software functions.
[0055] (3) Output module: contains a text display area for displaying prediction probability, classification results and threshold information.
[0056] 3.2 Software Packaging like Figure 7 As shown, the trained model is packaged into a standalone software tool. A .ps1 packaging script calls PyInstaller for single-file packaging, using the `--windowed` mode and specifying the program name as BC_7genes_Pre. A separate parameter file is added using the `--add-data` parameter, ultimately generating the executable program BC_7genes_Pre.exe, which automates the output of bladder cancer incidence probability and classification results.
[0057] 3.3 Software Application Examples Users can input the expression values of seven genes in the interface. During program execution, the system automatically loads standardized parameters, neural network weights, and threshold files, sequentially performing Z-score standardization, inference calculations for the seven tanh hidden neurons, and sigmoid probability output. The system then compares the output probability with a preset threshold. If the predicted probability is higher than the threshold, it is determined to be bladder cancer; if the predicted probability is lower than the threshold, it is determined to be normal.
[0058] like Figure 9 As shown in the example, in this case, the user enters the expression values of TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1 in the seven text boxes on the interface, and then clicks the "Predict" button. The system automatically completes the standardization process and model inference, and outputs a bladder cancer probability of 0.9862, with a predicted classification result of "bladder cancer" and a corresponding threshold of 0.92. Since the predicted probability of 0.9862 is higher than the diagnostic threshold of 0.92, this sample is identified as a bladder cancer sample.
[0059] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the inventive concept of this application, and these all fall within the protection scope of this application.
Claims
1. A smart bladder cancer prediction system based on the expression characteristics of seven genes, comprising a processor and a memory, characterized in that, The system includes: Data acquisition module: used to acquire the expression values of seven genes in the sample to be tested, including TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR and E2F1; Preprocessing module: Communicatively connected to the data acquisition module, used to standardize the expression values of the seven genes using the Z-score standardization method; Model inference module: Communicatively connected to the preprocessing module, containing a pre-trained multilayer perceptron model for calculating the probability of bladder cancer based on the standardized expression values of seven genes; The result output module is connected in communication with the model inference module and is used to determine the bladder cancer classification result based on the bladder cancer occurrence probability and a preset threshold obtained by optimization through cross-validation.
2. The intelligent bladder cancer prediction system according to claim 1, characterized in that, The data acquisition module obtains the expression values of the seven genes using quantitative real-time PCR or RNA sequencing technology.
3. The intelligent bladder cancer prediction system according to claim 1, characterized in that, The multilayer perceptron model includes a multilayer perceptron (MLP) classifier comprising 7 input nodes, 7 tanh-activated hidden neurons, and 1 sigmoid output node.
4. The intelligent bladder cancer prediction system according to claim 1, characterized in that, The result output module also includes a graphical user interface, which includes an input box for inputting the expression values of the seven genes, a prediction button for triggering prediction, and an output area for displaying the prediction probability and classification results.
5. The intelligent bladder cancer prediction system according to claim 1, characterized in that, The preset threshold of the result output module is determined by optimization through cross-validation, so that the accuracy, sensitivity and specificity of the model are all greater than 0.
85.
6. A combination of biomarkers capable of diagnosing or predicting bladder cancer, characterized in that, The biomarker combination includes TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR, and E2F1.
7. Application of TOP2A, MKI67, CDK1, UBE2C, NUSAP1, HMMR and E2F1 genes in the preparation of kits for bladder cancer diagnosis.
Citation Information
Patent Citations
Bladder cancer combined treatment method based on CTSE enhancement
CN121796591A
Tissue and blood-based mirna biomarkers for the diagnosis, prognosis and metastasis-predictive potential in colorectal cancer
WO2014145612A1